Stanford Real-Time EXPO-FT makes slow-but-smart VLAs real-time on robots: 42% to 97% with just 10 minutes of online data

New Stanford research (arXiv 2609.18207, Perry Dong, Kuo-Han Hung, Dorsa Sadigh, Chelsea Finn): large VLA models have high inference latency, so the world has moved on by the time their actions execute — and real robots do not pause for inference. Real-Time EXPO-FT splits control into two timescales: a slow asynchronous tier where the pretrained VLA proposes candidate action chunks using its strong behavior prior, and a fast synchronous tier where a lightweight edit policy applies bounded edits to actions at execution time using the latest observation, with a Q-value stage picking the best candidate chunk on the fly. The edit policy is trained with RL (built on EXPO/EXPO-FT; the Q signal touches only the edit, never backpropagating into the VLA backbone), with online robot data capped at 10 minutes and no human intervention. Results: average success 42% to 97% across four dynamic real-world tasks (object passing, ball balancing, table soccer kicking, dynamic object picking); best among delayed and non-delayed methods in 10/10 Kinetix environments. The inversion is notable: approaches like Astra have the slow smart model review the fast policy — here the fast lightweight policy edits the slow smart model. As LeoKharon put it: a big model's value may be its prior, not its control.





