
RLT: Physical Intelligence masters the last millimeter with RL Tokens
Physical Intelligence (Sergey Levine's team, Mar 19, 2026) present RL Tokens (RLT): freeze a pretrained VLA, attach an encoder-decoder that compresses its internal embeddings into a bottleneck RL token, then run online RL with a ~1M-parameter actor-critic on the real robot, refining only the critical phase (editing VLA action chunks rather than generating from scratch, anchored by a BC regularizer plus reference-action dropout). Across four sub-millimeter tasks the critical phase speeds up by up to 3x, screw success goes 20% to 65%, and half of Ethernet insertion episodes beat every human teleoperation demo - with just 15 minutes of real robot data.
Physical Intelligence's RLT: "RL Tokens" Let VLAs Master the Last Millimeter with Online RL
Source: Physical Intelligence research blog, "Precise Manipulation with Efficient Online RL" (March 19, 2026), by Charles Xu, Jost Tobias Springenberg, Michael Equi, Ali Amin, Adnan Esmail, Sergey Levine, Liyiming Ke.
The Problem: VLAs Get 90% of the Way There
Pretrained VLA models generalize impressively, but precision tasks in real deployments — driving an M3 screw into a robot arm, threading a zip tie through its narrow slot, inserting an Ethernet cable into a recessed port — all stall at the "contact-rich critical phase," where tiny errors in position, orientation or timing cause failure. Prior robot RL methods, including pi's own Recap, target broad improvement over long-horizon tasks with large-scale data collection. RLT goes the opposite way, targeting fine-grained manipulation with only minutes to hours of real-world experience.
Method: Freeze the VLA, Train a ~1M-Parameter Head
The core insight: don't fine-tune the multi-billion-parameter VLA with RL — adapt the VLA to be RL-friendly instead. The pipeline has two stages:
Stage 1 — Training the RL token (offline, ~1 hour). Attach a lightweight encoder-decoder transformer to the frozen VLA: the encoder compresses the VLA's internal token-embedding sequence into a special "RL token," and a decoder autoregressively reconstructs the original embeddings from it (with stop-gradient on the originals so the VLA is untouched). The reconstruction loss forces the RL token to act as an information bottleneck — keep everything task-relevant, drop the rest. Optionally, if out-of-the-box VLA performance is poor, the VLA can be jointly supervised-fine-tuned with the RL token module on demos. After this stage, both are permanently frozen.
Stage 2 — Online RL refinement of action chunks (on-robot, 15 min to 5 hours). Small actor and critic networks (~1M parameters total, one three-thousandth of the VLA) take the RL token plus proprioceptive state as input:
- Critic: standard off-policy temporal-difference learning (TD3-style target network); the reward is sparse +1 success / 0 failure, no reward engineering.
- Actor: outputs a Gaussian over action chunks (aligned with the VLA's output structure rather than single control steps), and receives the VLA's predicted action as input — it learns to edit the VLA's action rather than generate from scratch; a BC regularization term anchors the policy near the reference action, deviating only when the critic identifies something better.
- Reference-action dropout: 50% of transitions in each batch zero out the reference action, preventing the actor from degenerating into copying the VLA and keeping an independent action-generation pathway.
- The replay buffer stores VLA, actor and human-intervention transitions alike; 5 updates per environment step and hundreds of parameter updates per second make training responsive enough to improve after every attempt.
RL applies only to the human-marked critical phase of each task (5-20 s insertion/fastening/rotation segment); grasping and transport remain with the base VLA, with a post-hoc fine-tuned switch triggered automatically at test time.
Experiments: Four Sub-Millimeter Tasks
Tasks: electric-screwdriver M3 screw driving, zip-tie fastening, Ethernet insertion, charger insertion (USB-C into a power strip). Demonstration data: 1-10 hours per task; online RL data: 400-1,000 episodes (15 min to 5 hours of actual execution time).
| Task | Challenge | Critical-phase speedup | Success-rate change |
|---|---|---|---|
| M3 screw installation | sub-mm alignment + screw position tolerance | ~3× | 20% → 65% (+45pp) |
| Zip-tie fastening | bimanual deformable threading | ~2× | +40pp (+60% full task) |
| Ethernet insertion | pose alignment + decisive force | ~2× | significant gain |
| Charger insertion | limited prong visibility | ~1.5× | significant gain |
The headline result: on Ethernet insertion, RLT's median critical phase takes 66 steps (at 50 Hz) versus 146 for expert teleoperation demos and 228 for the base VLA — half of the RL episodes are faster than every human teleoperation demonstration, meaning the robot discovered execution strategies absent from the demo data (one fluid insertion motion, with small compliance wiggles on retry, replacing the VLA's probing-retreat-readjust behavior). Total training took 2 hours, of which only 15 minutes was actual robot data, the rest resets and overhead. Against baselines: single-step RL methods (HIL-SERL, PLD) fail to learn — sparse-reward credit assignment over a long effective horizon is impractical; DAgger reaches high success but is speed-capped by human demonstrations; DSRL's latent-space steering limits how much speed it can add. RLT wins on the combined success × speed × throughput picture.
Ablations: replacing the RL token with an ImageNet-pretrained ResNet-10 halves throughput; removing action chunks can't even match the base VLA; removing the BC regularizer causes the largest single drop (exploration drift in the full high-dimensional action space); removing reference-action pass-through learns significantly slower with more mid-training failures. Full RLT beats every ablation variant after just 5 minutes of critical-phase data.
Why It Matters
- The "frozen big model + tiny editing head" division of labor is likely to be widely copied: generalization comes from the VLA's perception and language priors; sample efficiency comes from squeezing RL's learning burden down to 1M parameters. It complements Recap (offline RL for overall behavior) — a production system could use both.
- RL is finally fast enough to deploy online: hundreds of updates per second and improvement after every attempt mean robots can, for the first time, practice on the job.
- 1X CEO Bernt Børnich's comment marks the boundary: real-world RL is the industry's "end game," but scaling it requires safe, compliant hardware to survive the fail-and-adapt phase; pi is betting on the software-first path (fresh off a $600M round at a $5.6B valuation), positioning the RL token as a modular interface for third-party embodiments.
- Honest limitations: humans still provide sparse reward labels, safety interventions and phase marking (future: learned reward models and progress predictors); RL covers only the critical phase; the RL token is trained per task — generalization is next.
Resources
- Blog: https://www.pi.website/research/rlt
- Paper PDF: https://www.pi.website/download/rlt.pdf
- Prior work: Recap (pi*0.6, offline RL for long-horizon tasks)
- Keywords: VLA, online RL, RLPD, action chunks, sub-millimeter precision, real-world RL
Source:Physical Intelligencehttps://www.pi.website/research/rlt


