PAPER DEEP DIVE
EgoLAP: Learning from Egocentric Human Data through Language-Action Reasoning
Egocentric human data offer a path to scaling robot learning beyond costly robot demonstrations, yet the embodiment gap makes raw human trajectories a poor supervisory target for control. Our key insight is that, although low-level actions are embodiment-specific, their underlying motion intent can capture task-relevant structure that transfers across humans and robots. We introduce EgoLAP, a VLA pre-training framework that jointly learns from human and robot trajectories through a shared language-based action chain-of-thought. EgoLAP expresses motion intent as structured, temporally abstracted language actions and pairs them with motion-level reasoning grounded in scene geometry, physics, and object affordances. Across extensive real-world and simulated experiments, EgoLAP transfers human experience to robot control more effectively than alternative action representations and reaches 80.1% mean real-world task progress, a 2.3x performance gain over alternative action representations. Motion-level reasoning also outperforms a composite reasoning format that combines subtask, object-box, and visual-trace reasoning.
Paper Metadata
Title: EgoLAP: Learning from Egocentric Human Data through Language-Action Reasoning
Authors: Lihan Zha, Shresth Grover, Tenny Yin, Samuel M. Bateman, Hengkai Pan, Mengchao Zhang, Aykut Onol, Allen Z. Ren, Dhruv Shah, Anirudha Majumdar (Princeton University, Toyota Research Institute, Physical Intelligence)
Links: arXiv:2610.08726v1 [cs.RO], submitted 2026-10-06; project page https://ego-lap.github.io/
Code status: The authors have not released EgoLAP training code; only the project page is public. The direct predecessor LAP (Language-Action Pre-training) is open source at github.com/lihzha/lap, and its language-action serialization and parsing code corresponds line by line to the method described here, so it serves as a practical reference for reproduction.
One-Sentence Summary
EgoLAP writes both human hand motion and robot arm motion into a single shared textual chain of thought composed of a natural-language action plus motion-level physical reasoning, letting egocentric human video and robot trajectories supervise one another in the same semantic space, and reaching 80.1% mean task progress on a bimanual real robot, 2.3x higher than any competing action representation.
Background and Motivation
Progress in robot foundation models has long been coupled to the scale of robot teleoperation data. That coupling is expensive in a very concrete way: every additional hour of demonstration requires a robot, an operator, a scene, and a task definition, and the resulting distribution stays narrow because the collection process itself is narrow. Egocentric human video has the opposite profile. It is cheap, it grows independently of any robot platform, it is recorded in real homes and workplaces rather than in a lab cell, and it covers long-horizon manipulation of the kind people actually perform. Turning that corpus into supervisory signal for control is therefore one of the central open problems in embodied AI.
Most prior work bridges the two embodiments through a hand-designed intermediate representation: interaction hotspots and affordances, hand pose or image-space trajectories, latent plans inferred from video, or embodiment-transfer edits that rewrite a human clip so that it looks like a robot performing it. Each of these captures some genuine facet of human behaviour, and each inherits the inductive bias and the failure modes of its own extraction pipeline. Hand-pose estimation produces anatomically impossible wrists under occlusion; point trackers drift across long horizons; embodiment transfer introduces artefacts that a policy then learns to imitate. None of these representations scales cleanly, because the quality of the supervision is capped by the quality of the extractor rather than by the amount of data.
The deeper obstacle is the embodiment gap itself. A human hand and a robot arm differ in morphology, in kinematic chain, in workspace, and in control frequency. Supervising robot control directly with human low-level actions is not merely noisy, it is category-mismatched: the same instruction executed by a person and by a YAM arm does not produce the same joint trajectory, the same end-effector path, or even the same gripper timing. Any representation defined at the level of raw actions is therefore untransferable by construction.
The key insight of this paper is that humans and robots do not need to share low-level actions. What they need to share is motion intent: what motion should be produced next, and why that motion is appropriate given the current physical state of the scene. The first half is expressed as a language action, a structured and temporally abstracted natural-language description of the pose and gripper change over one second, for example "move forward 3 cm, rotate clockwise 14 degrees, close gripper". The second half is expressed as motion-level reasoning, a local physical justification that ties contact evidence, geometry, and object affordances to the expected effect of the motion that is about to be commanded.
Language actions work because they reuse machinery the backbone already has. A vision-language model was pretrained on text describing motion, space, and contact; writing the control target into that vocabulary places the supervision inside a semantic space the trunk already understands, and makes it compositionally generalizable. Non-linguistic action tokens such as FAST or OAT are learned codebooks optimized for compression or token structure; they carry no pretrained semantics, so the trunk must learn their meaning from co-occurrence alone, and nothing in that learning transfers between embodiments. But "what" is not enough on its own: the same task on two different embodiments requires two different concrete motions, so a policy that only knows the target motion has no way to adapt it. Motion-level reasoning fills exactly that gap. It sits lower than subtask planning and lower than visual traces, it connects directly to the low-level action, and because it is stated in terms of contact, geometry, and affordance rather than in terms of joints, it survives the change of embodiment.
Put together, the two give a physically grounded action chain of thought: first reason about why this motion is appropriate, then emit the motion to execute. EgoLAP is the VLA pre-training framework built around that chain. It uses textual targets to align human and robot motion semantics during training, and at inference it rolls out only the continuous action expert without decoding a single token of text, so real-time control is preserved.
Preliminaries
VLA models and action representations. A vision-language-action model typically supervises a VLM trunk with a discrete action target, sometimes paired with a continuous action expert that handles real-time control. The choice of discrete target matters disproportionately, especially under knowledge insulation, where gradients from the action expert do not flow back into the trunk and the trunk is shaped entirely by its textual objective. Learned tokenizers such as FAST and OAT optimize action compression or token structure. The LAP family instead describes motion in natural language so that the target lands in the VLM's pretrained semantic space.
Flow matching. A continuous action chunk $a_{t:t+K}$ is learned by a velocity-field network $v_{\phi}$ through interpolation and denoising. During training the ground-truth action and Gaussian noise are mixed according to an interpolation time $\tau$, and the network is regressed onto the velocity pointing from the action toward the noise. At inference the model starts from pure noise and integrates the learned velocity field back to an action. Sampling is substantially cheaper than diffusion and suits closed-loop control.
Knowledge insulation. A randomly initialized continuous expert produces large, poorly conditioned gradients early in training. If those gradients reach the trunk they erode the pretrained language and visual representations that the whole method depends on. Insulation stops the gradient at the trunk-expert boundary: the trunk is optimized only by text objectives, and the expert learns to consume whatever representation the trunk provides.
Egocentric canonicalization. Because raw trajectories live in different coordinate conventions, a shared representation requires a shared frame. Canonicalization defines a per-sample action frame anchored on a camera, so that "forward", "left", and "down" mean roughly the same thing for a head-mounted human camera, a robot wrist camera, and a robot base camera.
Method
3.1 Problem Formulation and the Shared Discrete Target
Let the robot dataset $\mathcal{D}_{R}$ contain samples $(o_{t},\ell,s_{t},a_{t:t+K})$, where $o_{t}$ is one or more RGB observations, $\ell$ is the task instruction, $s_{t}$ is proprioceptive state, and $a_{t:t+K}$ is an action chunk of horizon $K$. The human dataset $\mathcal{D}_{H}$ contains egocentric samples $(o_{t},\ell,q_{t:t+K})$, where $q$ is one or two hand poses in the head-camera frame; the current pose $q_{t}$ plays the role of the state input $s_{t}$, and the relative motion between consecutive poses defines the continuous action chunk $a_{t:t+K}$. The two datasets are thus brought into the same tuple shape before any modelling happens, which is what makes a single shared objective possible.
Each sample supplies at most three training signals. The first is produced by a deterministic canonicalizer $G_{e}$ that depends only on embodiment metadata $e$ (single-arm robot, bimanual robot, human) and yields a shared discrete action-token target $m_{t}$:
$$m_{t}=G_{e}\!\left(q_{t:t+K}\right)\quad\text{or}\quad m_{t}=G_{e}\!\left(a_{t:t+K},s_{t}\right)\tag{1}$$
In EgoLAP, $m_{t}$ is a language action summarizing one second of hand or end-effector motion in the egocentric frame. Controlled baselines replace this discrete target with FAST or OAT tokens while holding every other part of the recipe fixed. The second signal is the motion rationale $r_{t}$, which grounds $m_{t}$ in visible spatial relations, contact evidence, object affordances, and the expected physical effect of the motion. The third is the continuous action chunk $a_{t:t+K}$ itself, which supervises real-time control directly. This three-signal decomposition is what lets the paper run clean ablations: swapping $m_{t}$ changes the representation, removing $r_{t}$ changes the reasoning, and neither touches the control target.
3.2 Architecture: Structured Attention and Structured Gradients
EgoLAP uses a Mixture-of-Transformers (MoT) architecture that pairs a PaliGemma-2B vision-language trunk with a 300M-parameter action expert initialized from scratch. The VLM receives up to three RGB images (one base camera plus two wrist cameras), the task instruction, and the proprioceptive state; missing camera views are zero-padded rather than dropped, so the token layout is stable across embodiments with different sensor suites. The state is normalized and discretized into 256 uniform bins and spliced into the prompt as text, which means state enters through the same embedding table as language.
The attention mask is the core of the design. Image, instruction, and state tokens form a bidirectional prefix. Reasoning tokens and language-action tokens attend to that prefix and remain causal among themselves. The action expert attends only to its own tokens and to the observation-instruction-state prefix, and never attends to the reasoning or language-action tokens. Consequently no text needs to be generated at inference: the expert's conditioning path does not include the text target at all. Knowledge insulation then truncates gradients coming from the action expert at the VLM boundary, so the trunk is optimized exclusively by textual objectives.
Figure 1: Attention and gradient structure. Orange marks permitted attention, grey marks blocked attention, and blue marks attention whose gradient is stopped at the VLM boundary. Action-expert tokens never attend to the textual targets, so control requires no text decoding.
The detailed block layout confirms how thin the expert really is. The action expert receives $a_{t:t+30}^{\tau}\in\mathbb{R}^{B\times 30\times 14}$ and projects it to $B\times 30\times 1024$; a sinusoidal embedding of $\tau$ produces one $B\times 1024$ conditioning vector consumed by adaptive RMS normalization in all 18 expert blocks; the expert uses FFN width 4096, eight query heads and a single key-value head; and its velocity head returns $v_{\phi}\in\mathbb{R}^{B\times 30\times 14}$, exactly matching the target chunk shape. The tied VLM head produces logits of shape $B\times L\times 257{,}152$, and only reasoning and action-target positions receive text loss.
Figure 2: The full MoT architecture with tensor shapes. The 300M action expert is 18 blocks of width 1024 and shares nothing but the prefix with the 2B trunk.
With $D$ the action dimension, the expert predicts $a_{t:t+K}\in\mathbb{R}^{K\times D}$ in delta end-effector pose space under a flow-matching loss $\mathcal{L}_{\mathrm{FM}}$; $\mathcal{L}_{\mathrm{text}}$ is the expected masked token-level objective over language actions and motion-level reasoning; and $\lambda_{\mathrm{text}}$ trades the two off. The complete objective is:
$$\mathcal{L}=\mathbb{E}_{(x_{t},a_{t:t+K},m_{t},r_{t})\sim p_{\mathrm{mix}}}\left[\lambda_{\mathrm{text}}\mathcal{L}_{\mathrm{text}}+\mathcal{L}_{\mathrm{FM}}\right]\tag{2}$$
where $p_{\mathrm{mix}}$ is the fixed human-robot sampling distribution. The paper uses $K{=}30$ and $\lambda_{\mathrm{text}}{=}0.8$, deliberately downweighting the autoregressive objective because the discrete token loss converges faster than the flow-matching expert and would otherwise overfit the trunk early. At inference the expert integrates the learned velocity field from $\tau{=}1$ Gaussian noise down to $\tau{=}0$ in ten Euler steps, conditioned only on the current observation, instruction, and state.
3.3 Language Actions: Egocentric Canonicalization and a Deterministic Grammar
Raw trajectories cannot share a representation until their coordinate conventions agree. For every sample an egocentric action frame $F_{t}=(R_{t}^{F},p_{t}^{F})$ is defined at the start of the prediction window: single-arm robots with a wrist camera use the end-effector frame, bimanual robots and single-arm robots without a wrist camera use the base frame, and egocentric human data uses the head-camera frame. These camera-centric frames follow broadly similar viewing conventions, so motion directions are far more consistent across embodiments than they would be in each robot's own world frame. That single design decision is what makes a shared vocabulary of "forward", "left", and "down" meaningful at all.
For an end-effector or hand pose $(R_{\tau},p_{\tau})$, the representation in that frame is:
$$\bar{p}_{\tau}=(R_{t}^{F})^{\top}(p_{\tau}-p_{t}^{F}),\qquad\bar{R}_{\tau}=(R_{t}^{F})^{\top}R_{\tau}\tag{3}$$
Over a one-second window the net translation $\Delta p=\bar{p}_{t+1\mathrm{s}}-\bar{p}_{t}$ and relative rotation $\Delta R=\bar{R}_{t}^{\top}\bar{R}_{t+1\mathrm{s}}$ are computed; bimanual samples run the identical procedure on the left and right trajectories independently and preserve arm identity; rotations are converted to extrinsic XYZ Euler angles before serialization. LAP's structured grammar then turns $(\Delta p,\Delta R,\Delta c)$ into text:
$$m_{t}^{\mathrm{lang}}=\text{``}\langle\text{verb}\rangle\;\langle\text{direction}\rangle\;\langle\text{magnitude}\rangle\;\langle\text{unit}\rangle\text{''}\tag{4}$$
Here move denotes translation, tilt denotes roll or pitch, and rotate denotes yaw. Translation magnitudes are rounded to whole centimetres, rotation magnitudes to multiples of five degrees, and components that round to zero are omitted. A gripper phrase (open, close, or maintain) is appended, and bimanual motion uses explicit left-arm and right-arm clauses. A representative output reads: "Left arm: move forward 5 cm, move left 2 cm, tilt left 10 degrees, rotate clockwise 10 degrees, close gripper. Right arm: move down 10 cm, tilt forward 15 degrees, open gripper."
Two properties of this target deserve emphasis. First, it is lossy by design: quantizing to centimetres and five-degree bins discards exactly the sub-centimetre dynamics that a raw action token would preserve. The paper's bet is that this loss is worth paying, because what remains is the transferable part. Second, because $G_{e}$ is deterministic and the grammar is machine-parseable, a language action is a verifiable target. Any sampled completion can be parsed back into deltas and scored against the demonstrated motion with no learned reward model. That property is what makes the GRPO post-training in Section 3.6 possible at all, and it is not available to FAST or OAT without training a separate decoder and reward.
3.4 Motion-Level Reasoning: the Physically Grounded "Why"
A language action says what motion happened; it does not say why that motion was appropriate in the current physical state. The authors therefore annotate a motion rationale $r_{t}$ for a subset of both human and robot trajectories. Given the task instruction, an ordered set of trajectory frames, a compact text summary of recent action history, and the ground-truth language action for the current frame, the annotator produces a rationale that names the specific end-effector-to-object relation, the relevant contact or affordance evidence, and the immediate physical effect of the labelled motion. For a dough-kneading segment the rationale reads roughly: the dough sits at the bottom of a metal bowl, both palms are in direct contact with its top and side, kneading requires pressing the dough against the bowl surface to deform it, and the hands therefore drive a backward translation with raised wrists and spread fingers to pull the compliant dough toward the body before the next press.
Motion-level reasoning is deliberately local and physics-focused, which is precisely how it differs from the high-level task or subtask reasoning used in earlier work. It sits at a lower level, it connects directly to the low-level action, and it describes a motion intent that is shared across embodiments rather than a plan that is specific to one. The annotation pipeline is careful about tense: the prompt distinguishes the state visible at the anchor from the command to be executed over the following segment, which prevents a future grasp, release, or contact event from being described as already complete. For DROID, anchors are formed from adaptive action segments with a nominal stride of 15 filtered steps (about one second), and a segment can terminate early at a persistent gripper-state change or a large change in translational direction. Gemini returns a trajectory summary, a subtask plan, and per anchor the completed, current and next subtasks, an observation-action grounding, a main rationale, a visual-relation sentence, a motion-effect sentence, and an alternative-motion diagnostic, but only the main rationale is used as a training target.
Because rationales are only available for part of the corpus, the training scheme must tolerate their absence. Two forms of dropout handle this: reasoning dropout removes the rationale from the target entirely, and reasoning loss dropout keeps it as context without supervising it. Both are formalized in Section 3.5.
3.5 Flow-Matching Objective and Stochastic Supervision
For a continuous action chunk $a_{t:t+K}$, Gaussian noise $\epsilon\sim\mathcal{N}(0,I)$, and interpolation time $\tau=0.001+0.999z$ with $z\sim\operatorname{Beta}(1.5,1)$, define the interpolated state and the target velocity:
$$a_{t:t+K}^{\tau}=(1-\tau)a_{t:t+K}+\tau\epsilon,\qquad u_{t}=\epsilon-a_{t:t+K}\tag{5}$$
The Beta(1.5,1) schedule biases sampling toward larger $\tau$, i.e. toward the noisier end of the path, which is where the velocity field is hardest to estimate. The action expert $v_{\phi}$ is trained to recover that velocity:
$$\mathcal{L}_{\mathrm{FM}}=\mathbb{E}_{\tau,\epsilon}\left[\frac{1}{KD}\left\|v_{\phi}\!\left(a_{t:t+K}^{\tau},\tau,\operatorname{sg}[h_{\theta}(x_{t})]\right)-u_{t}\right\|_{F}^{2}\right]\tag{6}$$
where $x_{t}=(o_{t},\ell,s_{t})$ is the multimodal prefix, $h_{\theta}$ is the trunk prefix representation, and $\operatorname{sg}$ is the stop-gradient that realizes the knowledge-insulation boundary. The $1/KD$ normalization makes the loss scale invariant to chunk length and action dimension, which matters when the same recipe is reused across embodiments with different $D$. At inference a fixed-step ODE solver integrates from $\tau{=}1$ to $\tau{=}0$.
To stop the VLM from overfitting to rationales, the authors randomize both whether reasoning appears in the target and whether it receives loss. For the rationale tokens $r_{t}$ and the discrete action tokens $m_{t}$, two independent variables are drawn:
$$B\sim\operatorname{Bernoulli}(p_{\mathrm{ctx}}),\qquad C\sim\operatorname{Bernoulli}(p_{\mathrm{sup}})\tag{7}$$
If $B{=}1$, the rationale is inserted as a teacher-forced causal prefix before $m_{t}$; otherwise the VLM predicts $m_{t}$ directly. Given that the rationale is present, $C$ decides whether its tokens receive language-modelling loss. The masked single-sample objective is $\ell_{\mathrm{text}}(B,C)$ and the text term is $\mathcal{L}_{\mathrm{text}}=\mathbb{E}_{B,C}[\ell_{\mathrm{text}}(B,C)]$. The paper sets $p_{\mathrm{ctx}}=p_{\mathrm{sup}}=0.5$; samples without a rationale get $B=C=0$. Among annotated samples, then, half train direct language-action prediction, a quarter use reasoning as causal context only, and a quarter both condition on reasoning and supervise it. The discrete action tokens receive loss in every case, so the control-relevant target is never dropped.
3.6 Verifiable Targets Enable GRPO Post-Training
Because a sampled language action can be parsed and compared against the demonstrated motion, the authors can run reinforcement learning with a purely programmatic reward. For arm $b\in\{\mathrm{left},\mathrm{right}\}$, let $p_{b}$ be the parsed translation in centimetres, $R_{b}$ the rotation reconstructed from the parsed XYZ Euler angles, and $g_{b}$ the binary gripper command, with $\star$ marking the target. Writing the target-relative rotation vector as $\omega_{b}=\operatorname{Log}\!\big((R_{b}^{\star})^{-1}R_{b}\big)^{\vee}\in\mathbb{R}^{3}$ in radians, the cost and reward are:
$$C=\sum_{b}\left[\frac{\lVert p_{b}-p_{b}^{\star}\rVert_{2}^{2}}{(2\,\mathrm{cm})^{2}}+\frac{\lVert\omega_{b}\rVert_{2}^{2}}{(\pi/18)^{2}}+\mathbf{1}\!\left[g_{b}\neq g_{b}^{\star}\right]\right],\qquad r=-C\tag{8}$$
Terms are summed over arms rather than averaged, so a 2 cm single-axis translation error, a 10-degree relative orientation error, and one wrong gripper command each cost one unit. The rotation error is the true $SO(3)$ residual rather than a difference of Euler angles, and the gripper term is a discrete mismatch penalty rather than a cross-entropy. Valid rewards are non-positive and uncapped, and an exact match scores zero. A completion counts as valid only if the parser recovers exactly one clause per arm with finite motion values; an invalid completion receives $r_{\mathrm{invalid}}=-\max(1001,\,1+\max_{j\,\mathrm{valid}}C_{j})$, ranking it below every valid completion in its group. There is no reward term for rationale quality, response length, entropy, or simulator success, which keeps the signal purely about motion accuracy.
The advantage is group-normalized and clipped:
$$A_{ij}=\operatorname{clip}\!\left(\frac{r_{ij}-\bar{r}_{i}}{\max(\sigma_{i},10^{-4})},\,-4,\,4\right)\tag{9}$$
Each update samples 32 observations and 64 completions per observation, and all 2,048 completions enter the gradient; the group is used for relative advantage estimation, not for selecting a single best response to imitate. Only the VLM backbone excluding the vision encoder is updated, the vision encoder and the action expert stay frozen, and there is no auxiliary flow-matching, SFT, rationale, or entropy term. The PPO clip range is $\varepsilon=0.2$ with a non-negative $k_{3}$ reference penalty of weight $\beta=0.02$ against the frozen SFT checkpoint.
Figure 3: Language-action GRPO post-training. Reasoning pre-training supplies a stronger initialization, and sampling only the language action during RL beats sampling reasoning as well.
3.7 Pipeline Overview
flowchart TD
subgraph SRC["Data sources about 10k hours"]
H1["MECKA egocentric human hands"]
H2["Scale and Aria egocentric"]
R1["ABC YAM bimanual 30 percent"]
R2["AgiBot and Galaxea and MolmoAct2 25 percent"]
R3["OXE single arm including DROID 20 percent"]
end
H1 --> CANON
H2 --> CANON
R1 --> CANON
R2 --> CANON
R3 --> CANON
CANON["Egocentric canonicalization G_e shared action frame"]
CANON --> LA["Language action m_t structured natural language"]
GEM["Gemini offline labels only MECKA plus DROID 32 percent"] --> MR["Motion-level reasoning r_t contact geometry affordance effect"]
OBS["Images plus instruction plus state prefix"] --> VLM
OBS --> EXP
LA --> VLM
MR --> VLM
VLM["VLM trunk PaliGemma 2B text loss L_text"]
VLM -- "stop gradient boundary" --> EXP
EXP["Action expert 300M flow matching L_FM 30 step chunk"]
EXP --> ACT["Inference rolls out expert only 10 Euler steps no text decoded"]
Figure 4: EgoLAP training and inference. Human and robot trajectories pass through egocentric canonicalization to yield shared language actions; Gemini-annotated motion-level reasoning covers only MECKA and DROID; textual objectives supervise the VLM trunk alone, and the continuous expert learned by flow matching rolls out independently at inference.
3.8 Correspondence with Open-Source Code
EgoLAP itself is not released, but the LAP repository (github.com/lihzha/lap, same first author) implements the language-action pipeline reused here, and the correspondence is exact enough to read as a reference implementation. In src/lap/policies/transforms/action_text.py, the function summarize_numeric_actions() is the deterministic serialization of Eq. (4): translation is converted to centimetres via round(abs(dx_m * 100.0), decimals), rotation is snapped to five-degree multiples via _round_to_nearest_n(..., rotation_precision), components are concatenated in the fixed order forward-back, up-down, left-right, roll, pitch, yaw, gripper, and any component that rounds to zero is dropped. This matches the vocabulary table in the paper exactly.
In the same file, summarize_bimanual_numeric_actions() returns f"Left arm: {left_summary}. Right arm: {right_summary}", which is the bimanual template described in Appendix A.2. Human hand trajectories reuse this same bimanual template rather than a separate vocabulary, so human and robot motion supervise the identical set of VLM tokens, which is a non-obvious but important implementation detail for cross-embodiment alignment. The inverse mapping is provided by parse_language_to_deltas() in src/lap/policies/lang_action_formats.py, which uses regular expressions to recover translation, rotation, and gripper deltas from strings such as "move", "tilt", "rotate", and "open gripper". That parser is the engineering foundation for the reward in Eq. (8): verifiability without a learned reward model is not a claim about language in general, it is a claim about this specific grammar being invertible.
Experiments
4.1 Experimental Design
Real-robot evaluation runs on a custom bimanual YAM platform whose arm spacing and camera configuration differ from every embodiment in the training mixture, so nothing about the test robot was seen during pre-training. Five behaviours are tested, spanning rigid, articulated, and deformable manipulation: Pick Place (Simple), Fold Pants, Handover and Place, Wipe Object, and Pick Place (Complex). Each task gets 10 trials (12 for Complex) and the reported metric is mean task progress, which credits partial completion rather than collapsing a long-horizon task to a binary outcome.
Simulation evaluation builds a bimanual benchmark on ABC with 14 tasks (3 pick, 8 place, 2 colour-selection, 1 cloth-stacking), 100 episodes per model per task for a total of 1400 episodes. Because no model is trained on simulation data, this is a zero-shot simulation evaluation, and it is the harder of the two tests: the policies have never seen the simulator, its physics, or its rendering.
Baselines form three policy families that differ only in action representation: EgoLAP (language actions plus motion-level reasoning), EgoFAST (FAST tokens plus reasoning), and EgoOAT (OAT tokens plus reasoning). Unless stated otherwise, all variants share everything except the action representation: same backbone, same architecture, same data mixture, same preprocessing including egocentric canonicalization, same optimization budget, and same evaluation protocol. All representations correspond to the same one-second action chunk. Every variant trains for 40k steps with checkpoint selection by the same validation procedure, the sole exception being EgoLAP, which converges faster and is trained to 30k steps. The paper also compares against LAP without reasoning and without human data, and against ABC-VLA trained directly on ABC real-robot data.
The OAT baseline is trained rather than reused: an OAT codec is trained on normalized $30\times 14$ chunks from the same mixture for 100k updates, with hidden width 256, eight heads, six blocks, dropout 0.1, and finite scalar quantization at levels $(8,8,6,5)$ giving 1,920 codes per token, and is then frozen during policy training. FAST uses the released tokenizer, also frozen. This level of care in the baselines is what makes the representation comparison credible.
4.2 Language Actions Transfer Human Experience Most Effectively
Holding the representation fixed and varying only whether human data is included: adding egocentric human data lifts simulation success from 7.2% to 16.4%, a factor of 2.3. The simulation suite is the harder transfer test since neither policy has seen the simulator, and human data pays off precisely there. Holding human data out and varying only the action representation without reasoning, EgoLAP reaches 16.4% in simulation and 44.5% on the real robot, against 1.6% / 35.2% for EgoFAST and 2.7% / 18.0% for EgoOAT. The gap is widest exactly where the domain shift is largest: in the unseen simulator, language actions beat FAST and OAT by an order of magnitude.
The full EgoLAP also beats ABC-VLA, which trains directly on in-domain ABC real-robot data, by 16.9 points in simulation and 54.5 points on the real robot. That is the strongest single result in the paper: a model that never saw the evaluation data source outperforms one that did, by a margin large enough that it cannot be explained by checkpoint selection noise.
Figure 5: Benefit of human data. With the language-action policy held constant, adding egocentric human data more than doubles simulation success.
A useful control appears in the appendix. With motion-level reasoning, adding human data improves language actions from 18.4% to 38.3%, a 19.9-point gain, whereas FAST improves from 1.4% to 8.7%, a 7.3-point gain. Language actions already provide the stronger robot-only baseline, yet their advantage widens under human-robot co-training: the gap over FAST grows from 17.0 to 29.6 points. That 12.6-point difference in human-data gains is what rules out the alternative explanation that language actions simply start from a better robot-only policy.
4.3 Motion-Level Reasoning and Language Actions Reinforce Each Other
Comparing each EgoX policy without reasoning to its reasoning-enabled counterpart: motion-level reasoning lifts EgoLAP's real-robot mean progress from 44.5% to 80.1%, a gain of 35.6 points that holds on all five tasks, and simulation success from 16.4% to 38.3%, a gain of 21.9 points that holds across every task family. The same supervision gives EgoFAST only 7.1 points in simulation and actually degrades it on the real robot, from 35.2% down to 26.1%. For EgoOAT the gains are 0.7 points in simulation and 7.6 on the real robot.
This asymmetry is the paper's central mechanistic claim, and it is worth stating precisely. Reasoning is not a generic auxiliary loss that helps any representation; it helps language actions dramatically, helps OAT marginally, and hurts FAST on the real robot. The proposed explanation is that language actions and motion-level reasoning form a single textual action chain of thought, moving from a physical justification of why a motion is appropriate to the specification of what to execute. The two are semantically continuous, so the trunk can use its pretrained knowledge to connect them. Non-linguistic tokens have no such semantic counterpart, so the trunk can only infer correlation from co-occurrence, and the added tokens act as capacity drain rather than as knowledge.
Figure 6: Zero-shot evaluation on the bimanual YAM. EgoLAP reaches 80.1% mean real-robot progress (2.3x other action representations) and 38.3% simulation success (1.8x ABC-VLA); paired with language actions, motion-level reasoning contributes 35.6 points, far more than with any other representation.
4.4 Human Data Amplifies Reasoning, and Motion-Level Reasoning Beats Existing Formats
To test whether the gain comes only from robot-side rationale supervision, the authors compare language-action policies with and without human data on the 14-task simulation suite. Reasoning lifts the robot-only policy from 7.2% to 18.4%, a gain of 11.2 points, while with human data it lifts from 16.4% to 38.3%, a gain of 21.9. Human data contributes 9.2 points without reasoning and 19.9 points with it. Robot-side rationale supervision alone therefore cannot account for the full effect: the two data sources and the reasoning objective interact.
The paper further compares against an ECoT-style composite target that predicts the current subtask, visible object boxes, and a visual trace before the language action. Motion-level reasoning helps every target it is paired with (+7.1 for FAST, +9.1 for the ECoT composite, +21.9 for plain language actions), but the composite target does not help on its own: it reaches only 5.3%, well below the 16.4% of language actions without reasoning, and stacking it before motion-level reasoning drags the policy down to 14.4%, less than half of the 38.3% that motion-level reasoning achieves alone. Because all variants train on identical trajectories with identical action-conditioned rationales, and because the ECoT control retains the same language-action target, the larger gain cannot be attributed to privileged action information or to the mere presence of extra auxiliary supervision.
The composite baseline is implemented carefully enough to be a fair test. Subtask context is reused from the same offline Gemini annotations; object boxes come from a separate grounding pass with gemini-3.5-flash at temperature 0, grounded every second anchor, mapped through the same aspect-preserving resize and symmetric padding, quantized into PaliGemma's 1,024 location bins and serialized in $(y_{\min},x_{\min},y_{\max},x_{\max})$ order. Visual traces are computed from trajectories and camera parameters rather than generated: recorded end-effector positions are projected through calibrated intrinsics and extrinsics for DROID, and left and right palm positions through the synchronized head pose for MECKA, with five future image-space positions sampled at 5 Hz over the next second. The resulting target deliberately excludes the motion-level rationale, which is what isolates ECoT-style "what and where" supervision from this paper's physical "why".
Figure 7: Motion-level reasoning carries the gain. It outperforms ECoT-style reasoning and substantially strengthens it when stacked, while the ECoT composite on its own underperforms having no reasoning at all.
4.5 Downstream Adaptation and Cross-Embodiment Skill Alignment
Post-training with GRPO on language actions, with the vision encoder and action expert frozen, improves simulation success in all three arms. Two findings stand out. First, reasoning pre-training supplies a stronger initialization for the post-RL policy. Second, generating reasoning during RL sampling hurts relative to sampling language actions only (+12.7% versus +19.4%): ABC rationales were never seen in training, so sampled rationales are out of distribution and degrade the subsequent language action. The authors are explicit that they leave proper exploitation of motion-level reasoning for RL to future work, which is an honest reading of their own negative result.
For cross-embodiment skill retrieval, the authors take paired checkpoints that differ only in reasoning supervision, mean-pool the last-layer features, subtract each embodiment's mean, and run cosine retrieval. Robot-to-human R@1 rises from 51% to 76%, a gain of 25 points, and human-to-robot R@1 rises from 42% to 63%, a gain of 21. Since the two embodiments remain linearly separable, reasoning does not collapse the domain identity; it gives both domains the same internal organization, so that folding sits near folding and placing sits near placing regardless of which body performed it. This is the most direct representational evidence in the paper that the shared textual target is doing what the introduction claims.
4.6 Training Recipe and Data Mixture
All policy variants share one recipe: batch size 2048 on 64 TPU v6e devices, 40k steps, 2500-step linear warmup then constant schedule, peak learning rate $5\times10^{-5}$, Adam with $\beta_{1}{=}0.9$, $\beta_{2}{=}0.95$, $\epsilon{=}10^{-8}$, weight decay $1\times10^{-4}$, gradient clipping at norm 1.0, EMA decay 0.999, language-action loss weight 0.8, flow-matching loss weight 1.0, action chunk length 30, action dimension 14, state dimension 20, image resolution $224\times224$, bfloat16, per-category 1st-99th percentile normalization for actions and states, flow-time distribution $0.001+0.999\operatorname{Beta}(1.5,1)$, ten inference integration steps, wrist-image dropout 0.1, random camera masking 0.2, and validation every 1000 steps. Images are resized with aspect-ratio-preserving symmetric padding, then augmented with a random crop of $\lfloor0.95W\rfloor\times\lfloor0.95H\rfloor$, a uniform in-plane rotation in $[-5^{\circ},5^{\circ}]$, and Augmax colour jitter at strength 0.2 for brightness, contrast, and saturation.
The data mixture spans four groups: ABC robot data at 30% from the ABC YAM platform, egocentric human data from EgoVerse at 25% drawn from the MECKA, Scale, and Aria subsets, additional bimanual robot data at 25% including AgiBot, Galaxea R1 Lite, and MolmoAct2 YAM, and single-arm OXE datasets at 20% of which DROID contributes 12%. Only MECKA (20%) and DROID (12%) carry motion-level reasoning labels, so 32% of each batch has reasoning supervision; the remainder is supervised with language actions and continuous actions alone. Total raw data is around 10k hours. The robot-only ablation removes EgoVerse while preserving the relative weights of every robot dataset and renormalizing, and is otherwise identical in architecture, hyperparameters, and compute.
4.7 Key Numbers
| Action representation / configuration | Sim 14-task success, no reasoning | Sim with motion-level reasoning | Gain |
|---|---|---|---|
| Language action (EgoLAP) | 16.4% | 38.3% | +21.9 |
| FAST (EgoFAST) | 1.6% | 8.7% | +7.1 |
| OAT (EgoOAT) | 2.7% | 3.3% | +0.7 |
| ECoT composite (subtask + boxes + trace) | 5.3% | 14.4% (stacked before motion-level reasoning) | +9.1 |
| Language action, robot data only | 7.2% | 18.4% | +11.2 |
| FAST, robot data only | 1.4% | 8.7% | +7.3 |
| Method | Real mean task progress, no reasoning | Real with reasoning | Delta |
|---|---|---|---|
| EgoLAP (language action) | 44.5% | 80.1% | +35.6 |
| EgoFAST | 35.2% | 26.1% | -9.1 |
| EgoOAT | 18.0% | 25.6% | +7.6 |
| ABC-VLA (in-domain real data) | - | 25.6% (implied by the 54.5-point gap) | - |
| Cross-embodiment retrieval | No reasoning | EgoLAP | Delta |
|---|---|---|---|
| Robot to human, 3 skills, R@1 | 51 | 76 | +25 |
| Human to robot, 3 skills, R@1 | 42 | 63 | +21 |
| Robot to human, 2 skills, R@1 | 69 | 92 | +23 |
| Robot to human, 3 skills, MRR | 61 | 80 | +19 |
| Training mixture, share of each batch | Motion-level reasoning labels |
|---|---|
| ABC robot data, 30% | No |
| EgoVerse human (MECKA / Scale / Aria), 25% | MECKA only (20%) |
| Other bimanual robots (AgiBot / Galaxea R1 Lite / MolmoAct2), 25% | No |
| Single-arm OXE (DROID 12%), 20% | DROID only (12%) |
Limitations
Stated by the authors: language actions currently summarize fixed one-second windows with quantized translation and rotation, which may miss the dynamics or precision required for fast, contact-rich manipulation; the evaluation remains focused on bimanual tabletop tasks; and the reinforcement-learning study is confined to simulation and optimizes a motion-matching proxy rather than task reward.
Our assessment: First, motion-level reasoning labels are produced offline by Gemini (gemini-3.5-flash for DROID, gemini-3.7-flash for MECKA), so label quality is bounded by the annotating model, and only 32% of each batch carries them. The marginal value of reasoning supervision is therefore entangled with annotation coverage, and the paper does not report a coverage sweep, so we cannot tell whether 32% is near saturation or still on the rising part of the curve. Second, real-robot evaluation covers 5 tasks with 10-12 trials each; the error bars are correspondingly wide, and although the 35.6-point gap over the no-reasoning variant is large, the absolute sample size behind 80.1% is small. Third, the "zero-shot simulation" ABC benchmark shares a platform and object distribution with the ABC real-robot data that constitutes 30% of the training mixture. No simulation trajectories are used, but that platform-level overlap may inflate the transfer numbers relative to a genuinely held-out embodiment. Fourth, the ECoT comparison shows that a poorly chosen auxiliary textual target can actively hurt, which suggests the design space of textual action targets is not yet well understood and that the specific form of the rationale matters more than the paper's ablations can establish.
Conclusion and Outlook
EgoLAP converts human video into supervisory signal for robot control with a single clean claim: low-level actions are embodiment-specific, but motion intent transfers. Writing "what to do" as a language action and "why" as motion-level reasoning yields one textual chain of thought shared by both data sources. The experiments show that this chain not only improves policy performance (80.1% real-robot mean progress, 38.3% zero-shot simulation success, beating an in-domain ABC-VLA by 54.5 and 16.9 points respectively) but also reshapes the representation geometry, improving cross-embodiment skill retrieval by up to 25 points. The synergy is specific: reasoning pairs strongly with language actions and barely at all with non-linguistic tokens, which is the most convincing evidence that the mechanism is semantic rather than a generic auxiliary-loss effect. Because textual targets exist only during training and inference rolls out the continuous expert alone, none of this costs real-time control.
Looking forward, the authors argue that the linguistic form of the target naturally supports adaptive precision, combining coarse long-horizon descriptions with fine high-frequency corrections, and tolerating approximate labels from pose tracking, UMI, or internet video, which would widen the data funnel well beyond carefully annotated corpora. On the RL side, the obvious next step is to jointly post-train motion-level reasoning and the action expert against task-level feedback rather than a motion-matching proxy. The negative result on sampling reasoning during GRPO is the interesting obstacle here: the rationales a reasoning-pretrained model generates for a new embodiment are out of distribution, so making reasoning useful at RL time likely requires reasoning supervision over the target embodiment, not merely transfer of a reasoning-capable trunk.
Golden Quotes
"Although low-level actions are embodiment-specific, their underlying motion intent can capture task-relevant structure that transfers across humans and robots."
"Language actions and motion-level reasoning form a single textual action chain of thought, moving from a physical justification of why a motion is appropriate to the specification of what motion to execute."
"Reasoning helps language actions dramatically, helps OAT marginally, and hurts FAST on the real robot; the synergy is with the representation, not with supervision in general."



