PAPER DEEP DIVE
Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models
Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are harder to scale than web-scale image-text data because embodied collection is costly and sparsely covers the physical world. This makes representation quality a central bottleneck: under a fixed robot-data budget, continued pre-training must turn limited trajectories into transferable visual-action knowledge rather than merely fit actions. We propose VLAct, a VLA-oriented VLM backbone trained on broad, heterogeneous, multi-embodiment robot data before task-specific fine-tuning. VLAct preserves the broad VLM prior and encourages shared action semantics across embodiments through VLM-prior preservation, multi-head continuous action co-supervision, and a partially unified cross-embodiment action layout, while allowing task-specific action heads during fine-tuning. Across simulation, real-world, and unseen-embodiment transfer, VLAct consistently improves downstream performance under fixed fine-tuning protocols. On LIBERO-Plus and RoboTwin 2.0, VLAct surpasses industrial VLA systems including ABot-M0 and LingBot-VLA, achieving success rates of 82.6% and 92.5%. On RoboDojo, VLAct ranks sixth among all policies by success rate and outperforms all explicitly designated world-action model (WAM) entries on both metrics. Most notably, on RoboCasa-GR1, an unseen humanoid embodiment, VLAct using only 20% of downstream trajectories outperforms the full-data GR00T-N1.6 baseline. These results are obtained using fully open-source data and only a 16-GPU training setup, showing that representation-centric continued pre-training can deliver highly competitive performance under a modest compute budget and is an important independent axis of VLA progress beyond data scaling.
Paper Metadata
Title: Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models
Authors: Senqiao Yang, Chengyao Wang, Yuxin Chen, Zixuan Wang, Longxiang Tang, Haokun Gui, Jinhui Ye, Changsheng Lu, Xiaoyang Wu, Mingkang Zhu, Pengguang Chen, Shu Liu, Zhuotao Tian, Hengshuang Zhao, Bei Yu, Jiaya Jia
Links: arXiv:2608.27550 (v2) · project page starvla.github.io/VLAct · code github.com/starVLA/VLAct (open source, MIT license; training scripts and checkpoints promised for release)
Pre-training corpus: four fully open robot datasets (DROID, InternA1, RoboCoin, MolmoAct) mixed with captioning data as a representation anchor; backbone Qwen3-VL-4B; training codebase StarVLA.
Compute: 16 GPUs for continued pre-training; 8x H800 and 50k steps for real-robot fine-tuning. All comparisons use a fixed downstream fine-tuning protocol.
One-Sentence Summary
VLAct reframes VLA continued pre-training from fitting actions on the pre-training distribution to shaping a transferable visual-action representation: shallow-layer protection and caption mixing preserve the VLM prior, multi-head co-supervision prevents collapse toward one action head, and a partially unified action layout with a wrap-aware loss aligns action semantics across embodiments, yielding a backbone that beats several industrial VLA systems with 16 GPUs and fully open data.
Background and Motivation
Progress in vision and language foundation models has long been driven by three scaling curves: data, parameters, and compute. The paradigm works not merely because web corpora are large, but because they cover an extremely broad range of visual and semantic variation. VLA models naturally aspire to the same paradigm: more robot trajectories should, in principle, support more generalist robot policies.
Robot trajectories, however, cannot be scraped from the web. They must be produced through embodied execution, typically under human teleoperation or carefully designed collection protocols, and the space a policy must generalize over is combinatorial and continuous, spanning scenes, objects, task goals, embodiments, and contact-rich dynamics. Even the largest robot datasets remain sparse samples of the physical interaction space, with partial and uneven coverage.
This does not diminish the value of scale; it changes the role scale should play. If continued pre-training cannot count on exhaustive coverage, its success depends not only on how many trajectories are collected but on how effectively those trajectories induce transferable visual-action representations. The paper therefore poses a fixed-budget question: given a fixed robot-data budget, how can continued pre-training learn representations that generalize beyond the trajectories they were trained on?
The answer is a representation-centric view: treat continued pre-training as distilling robot trajectories into reusable visual-action knowledge inside the backbone. A useful VLA backbone should not only predict the actions present in the pre-training data, but internalize transferable priors about objects, affordances, spatial relations, and action-conditioned physical interactions that downstream policies can reuse across new tasks, scenes, and embodiments.
One terminological clarification. "Continued pre-training" here means starting from an already pretrained VLM and training it on broad, multi-embodiment robot trajectories before any task-specific fine-tuning, rather than pre-training a foundation model from scratch. Representative generalist VLA systems such as pi-0, pi-0.5, and GR00T N1/N1.5 adopt exactly this setting and commonly call it VLA pre-training; the paper uses the more precise term.
Preliminaries: Four Action-Head Families
Almost every design decision in the paper revolves around the action-head axis, so the four representative heads are fixed first. FAST preserves the VLM's autoregressive interface by encoding an action chunk into discrete action tokens $z_{1:M}=\mathrm{Tok}_{\mathrm{FAST}}(A_t)$ and training with the standard next-token objective:
$$\mathcal{L}_{\mathrm{FAST}}=-\sum_{m=1}^{M}\log p_\theta(z_m\mid H_t,\,z_{<m})$$
At inference the tokens are generated autoregressively and mapped back to continuous actions by the inverse tokenizer. OFT removes the tokenizer and appends a fixed set of action query tokens to the backbone sequence, regressing the whole action chunk in parallel with a small MLP:
$$\mathcal{L}_{\mathrm{OFT}}=\frac{1}{K\,d_a}\sum_{k=0}^{K-1}\left\|g_\phi(h^{\mathrm{act}}_{t,k})-a_{t+k}\right\|_1$$
PI and GR00T generate continuous action chunks with flow matching: sample noise $\epsilon\sim\mathcal{N}(0,I)$ and flow time $\tau\in[0,1]$, follow the linear probability path $A_t^\tau=(1-\tau)\,\epsilon+\tau\,A_t$ with target velocity $u_t=A_t-\epsilon$, and train
$$\mathcal{L}_{\mathrm{PI}}=\mathbb{E}_{A_t,\epsilon,\tau}\left[\left\|v_\phi(A_t^\tau,\tau;H_t)-(A_t-\epsilon)\right\|_2^2\right]$$
GR00T uses the same flow-matching objective but inside an explicitly separated dual-system architecture: the VLM acts as a slow semantic reasoner while a DiT-style action module acts as a fast motor generator, conditioned on robot state $s_t$ and an embodiment identifier $e$. All four families supervise the same object, a future action chunk; they differ in decoding geometry: discrete autoregressive, one-shot regression, iterative flow matching, and dual-system flow matching. That difference is exactly where the representation-collapse problem comes from.
Method in Detail
Figure 1: Overview of VLAct. Naive VLA pre-training overfits the pre-training scenario and fails on new tasks, environments, embodiments, and action heads; representation-centric pre-training improves simulation benchmarks and real-robot experiments simultaneously.
Motivating probe: how action supervision reshapes the backbone
The paper opens with a controlled probe (Figure 2): fix the Qwen3-VL-4B backbone and vary only the action head used during pre-training and fine-tuning, then observe how the backbone representation is reshaped. Two findings follow. First, discrete supervision transfers but loses information: a FAST-pretrained backbone paired with a continuous GR00T head slightly beats from-scratch GR00T fine-tuning, yet keeping the FAST head itself is far worse than continuous-head fine-tuning, indicating that discrete tokens teach coarse action structure while losing the fine-grained temporal and amplitude information manipulation needs.
Figure 2: Action supervision reshapes VLA backbone representations. FAST pre-training transfers to a GR00T head but is limited by discretization loss; OFT pre-training helps same-head fine-tuning yet degrades sharply with PI or GR00T heads, revealing head-specific representation collapse.
Second, single-head continuous supervision induces head-specific representation collapse: OFT pre-training helps a lot when fine-tuning also uses OFT, but the same backbone performs poorly with PI or GR00T heads. The action information is not absent; it is organized in a geometry that only the pre-training head can decode. Strong same-head performance therefore overstates backbone reusability. The appendix names this tendency decoder lock-in.
These probes yield three failure modes of naive continued pre-training: full end-to-end updating erodes broadly useful vision-language features; single-head supervision over-specializes the backbone to that head's decoding geometry; and embodiment-specific output spaces isolate physically comparable actions, such as gripper open/close, inside separate heads. The three components of VLAct address these three modes respectively.
Component 1: shallow-layer protection and caption mixing preserve the VLM prior
Figure 6: Layer-wise attention visualization. Lower layers attend broadly to visual and spatial information while deeper layers focus on semantic and task-relevant regions, motivating freezing of the vision encoder and the lower half of the LLM.
Robot trajectories are far narrower in visual diversity than web corpora: fixed camera viewpoints, repeated manipulation scenes, embodiment-specific action correlations. Updating every layer end-to-end lets gradients from this narrow distribution shift the backbone away from its general VLM representation before robust action-conditioned features emerge. The layer-wise attention visualization (Figure 6) shows lower layers carrying broad visual and spatial information and deeper layers carrying semantic and task-relevant regions, so the paper freezes the entire vision encoder and the lower half of the LLM layers during pre-training, updating only the upper layers and the action heads; the full model is unfrozen during downstream fine-tuning. This simple strategy adds 3.7 points on LIBERO-Plus and 3.4 points on Agilex (Appendix C: full-backbone update 78.9, freeze vision encoder only 81.3, freeze vision encoder plus lower half of the LLM 82.6).
Caption mixing is the second anchor. Comparing five auxiliary sources (captions, BBox-QA, Point-QA, Spatial-QA, pure language instructions), captioning gives the strongest single-source gain: LIBERO-Plus rises from 75.0 for robot-only training to 82.6. The reason is intuitive: dense captions over objects, attributes, spatial relations, and scene context sit closest to the original VLM pre-training distribution. Every pre-training minibatch contains both robot samples and auxiliary samples, optimized with
$$\mathcal{L}_{\mathrm{total}}=\mathcal{L}_{\mathrm{action}}+0.5\,\mathcal{L}_{\mathrm{VLM\text{-}CE}}$$
where 0.5 is the fixed weight of the auxiliary cross-entropy. Notably, pure language instruction data also helps even though it is text-only and unrelated to robot perception or control; such a gain can only be explained at the representation level, as preserving the backbone's foundation-model operating regime and diversifying feature updates, rather than as task transfer.
Component 2: multi-head co-supervision turns head diversity into a supervision signal
To counter head-specific collapse, VLAct attaches three continuous heads, OFT, PI, and GR00T, to the same backbone in parallel. Given the same vision-language input, the backbone produces a shared latent $z$; each head reads the same $z$ and predicts the same ground-truth action chunk, and the objective is the sum of the three head losses:
$$\mathcal{L}_{\mathrm{action}}=\mathcal{L}_{\mathrm{OFT}}+\mathcal{L}_{\mathrm{PI}}+\mathcal{L}_{\mathrm{GR00T}}$$
The design is deliberately simple: no new action head and no alignment module, only head diversity itself as supervision. Because the three heads impose different objectives and decoder biases on the same action-prediction problem, the backbone cannot rely on features that help only one head; to reduce all three losses simultaneously it must encode action information in a form accessible to several parameterizations. Since all heads share a single backbone forward pass, multi-head supervision adds only lightweight head-specific computation.
Co-supervision therefore plays two roles: it makes the backbone more head-agnostic, improving transfer to alternative downstream heads, and it regularizes the representation, yielding stronger action features even when the downstream head matches a pre-training head. Appendix E separates the two. Taking PI as the downstream fine-tuning head: no pre-training gives 60.5; OFT-only pre-training drops it to 55.1 (lock-in); OFT+GR00T pre-training, with PI excluded, raises it to 63.1; including all three heads gives 77.0. Adding a second pre-training head moves the unseen-head result from below the from-scratch baseline to above it, which is direct evidence that diversity reduces decoder lock-in.
Same-head adaptation does not suffer either: for OFT, PI, and GR00T downstream heads, head-diverse pre-training improves over the matched single-head pre-training by 1.7, 1.6, and 4.3 points respectively (the chains 61.7 to 78.8 to 80.5, 60.5 to 75.4 to 77.0, and 51.2 to 71.7 to 76.0). The gains are moderate, but they show that multi-head supervision does not tax same-head performance.
Component 3: partially unified action space and wrap-aware loss
Figure 4: Three cross-embodiment output designs. Embodiment-specific heads isolate supervision; a fully unified space naively aligns coordinates with different physical meanings; VLAct shares only physically comparable coordinates such as gripper open/close and masks the rest.
The common cross-embodiment practice is one head or output projector per robot, which is flexible but hides shared structure inside embodiment-specific heads; the opposite extreme pads every embodiment to a common dimension, naively aligning coordinates with different physical meanings. VLAct takes the middle path: one shared action head emitting a fixed 20-dimensional vector per step, $\hat{a}_{t+k}\in\mathbb{R}^{20}$, with dimensions assigned by embodiment role: dimensions 1-6 are the left 6-DoF arm of bimanual embodiments such as AgileX in absolute joint angles, dimensions 7-12 the right arm likewise, dimensions 13-18 the 6-DoF arm of single-arm embodiments such as Franka in delta end-effector pose, dimension 19 the shared gripper coordinate used by Franka's single gripper and AgileX's left gripper, and dimension 20 the right gripper of bimanual embodiments.
During training each sample contributes loss only on the active dimensions of its embodiment; inactive dimensions are masked. The layout preserves each corpus's native low-level convention (delta end-effector for Franka, absolute joint angles for AgileX) and does not force all embodiments into one controller parameterization; it aligns output coordinates only at the level of embodiment role, with no embodiment adapter, routing module, or embodiment-conditioned decoder. The ablation (Appendix I.2) reports separate heads 78.5/81.1 (RoboTwin/LIBERO-Plus), a unified head without space alignment 79.5/81.4, and the full unified action representation 80.5/82.6.
The last detail is angular wrapping for periodic joints. Absolute joint angles live on a periodic space: $\pi+\epsilon$ and $-\pi+\epsilon$ describe nearly identical configurations while their raw values differ by almost $2\pi$. The paper first wraps all absolute joint-angle dimensions into the canonical range on the data side:
$$a_{\mathrm{wrap}}=(a+\pi)\bmod 2\pi-\pi$$
It then corrects the distance measure at the wrap boundary on the loss side: when the target is near $-\pi$ and the prediction near $+\pi$, the naive residual is close to $2\pi$ although the angles are physically close. Defining the wrapped residual
$$\delta_{\mathrm{wrap}}=\big((\hat{a}-a)+\pi\big)\bmod 2\pi-\pi,\qquad \mathcal{L}_{\mathrm{wrap}}=\left|\delta_{\mathrm{wrap}}\right|$$
and adding it to each head's original objective gives the per-head total
$$\mathcal{L}^{(h)}_{\mathrm{total}}=\mathcal{L}^{(h)}_{\mathrm{action}}+\mathcal{L}^{(h)}_{\mathrm{wrap}}$$
For OFT the term applies to the direct regression output; for GR00T and PI it applies to the final generated action sample after denoising or flow integration, not to intermediate noise or velocity predictions. It is applied only to absolute joint-angle dimensions, excluding gripper commands and delta end-effector translations. Under the RoboTwin base setting the ablation reads: raw joint-angle regression 75.5, adding data-side unified joint space 78.6, adding the wrap-aware loss 80.5.
Overall recipe and code correspondence
flowchart TB
subgraph PT["VLAct continued pre-training 16 GPUs"]
A["open robot data DROID InternA1 RoboCoin MolmoAct"] --> B["Qwen3-VL-4B shared backbone"]
C["caption mixing LLaVA-ReCap-CC3M OneVision"] --> B
B --> D["shallow-layer protection freeze vision encoder and lower LLM half"]
B --> E["multi-head co-supervision OFT PI GR00T on one shared latent"]
B --> F["partially unified action space 20-dim layout masked inactive dims"]
E --> G["wrap-aware loss modulo residual on periodic joints"]
F --> G
D --> H["L_action plus 0.5 times L_VLM-CE"]
E --> H
G --> H
end
subgraph FT["downstream fine-tuning fixed protocol"]
I["discard pre-training heads and caption stream"] --> J["freshly initialize task-specific action head"]
J --> K["LIBERO-Plus VLA-Arena RoboTwin2.0 DOMINO"]
J --> L["RoboCasa-GR1 RoboDojo real Franka"]
end
H --> I
The open repository starVLA/VLAct maps each component to concrete code. The wrap-aware loss lives in starVLA/model/modules/action_model/flow_matching_loss.py: blend_angular_wrap_abs_error computes the wrapped residual with torch.remainder(raw_diff + math.pi, 2*math.pi) - math.pi, blends it into the raw absolute error with 0.5/0.5 weights, and gates it by angular_mask so only joint-angle dimensions are affected; the OFT framework calls it at QwenOFT.py:309, while the PI and GR00T heads route through flow_matching_loss_with_endpoint_wrap. The partially unified layout lives in starVLA/model/framework/QwenHybrid_xrobot_padding.py: DISJOINT_ACTION_LAYOUT_DIM = 20, DISJOINT_DUAL_ARM_JOINT_SLICE = slice(0, 12), DISJOINT_SINGLE_ARM_EEF_SLICE = slice(12, 18), together with MULTI_ROBOT_ACTION_SPECS that validates each corpus's action dimension by robot tag (franka and oxe_droid as 7-dim delta end-effector, interna1_split_aloha and robocoin as 14-dim joints plus grippers), with all three heads sharing one encoder pass. Shallow-layer protection is not new code but configuration-driven: freeze_backbones at training/trainer_utils/trainer_tools.py:151 freezes submodules listed by relative path in config.trainer.freeze_modules.
Experiments
The evaluation protocol is the source of the paper's argumentative force: across every VLAct comparison the only thing that changes is the VLM backbone weight. The action head and its fresh initialization, the downstream data, the optimizer, and the fine-tuning budget are identical, so gains of 7.6 to 21.4 points are attributable to the backbone representation alone. Pre-training uses fully open data and 16 GPUs.
Single-arm Franka. On LIBERO-Plus, where models train on standard LIBERO and are evaluated on perturbed test sets, VLAct reaches 82.6 overall, beating the same-backbone same-head Qwen3VL-OFT (75.0) by 7.6 points and the large-scale industrial system Abot-M0 (80.5) by 2.1 points; the largest gains fall on the Camera, Robot, Noise, and Layout perturbation axes, indicating a more robust visual-spatial representation. On VLA-Arena, VLAct is best on all four behavioral axes (Safety, Distractor, Extrap., Long-H.) with a 54.8 average, 21.4 points above the same-backbone baseline and 10.5 points above the strongest public baseline pi-0.5, with margins of 21.0 and 11.7 points on Long-Horizon and Safety.
| Method | Camera | Robot | Lang. | Light | Bg. | Noise | Layout | Total |
|---|---|---|---|---|---|---|---|---|
| OpenVLA-OFT | 56.4 | 31.9 | 79.5 | 88.7 | 93.3 | 75.8 | 74.2 | 69.6 |
| pi-0 | 13.8 | 6.0 | 58.8 | 85.0 | 81.4 | 79.0 | 68.8 | 53.6 |
| pi-0-FAST | 65.1 | 21.6 | 61.0 | 73.2 | 73.2 | 74.4 | 68.8 | 61.6 |
| Abot-M0 | 60.4 | 67.9 | 86.4 | 96.2 | 91.6 | 86.4 | 82.6 | 80.5 |
| Qwen3VL-OFT (same-backbone baseline) | 47.0 | 60.1 | 87.0 | 96.3 | 95.3 | 73.1 | 79.2 | 75.0 |
| VLAct | 73.9 | 68.4 | 81.5 | 96.7 | 96.7 | 86.0 | 83.3 | 82.6 |
Table 1: Per-dimension success rate (%) on LIBERO-Plus. All models train on standard LIBERO and are evaluated on the perturbed test set.
Bimanual AgileX. On RoboTwin 2.0, the Base setting with only 50 clean trajectories per task gives VLAct-OFT 80.5 versus 61.7 for the same-backbone baseline; under Data Scaling it reaches 92.5 Clean and 90.8 Random, above InternVLA-A1, Being-H0.7, Motus, LingBot-VLA, Abot-M0, and pi-0.5, and just below HoloBrain-0 and Fast-WAM. The head-agnostic claim shows up here too: the same VLAct backbone scores 76.0/89.6/87.4 with a GR00T head and 77.0/93.0/88.8 with a PI head, the PI head even taking the best Clean score, with all three heads within 3.4 points of each other. What transfers is the backbone, not a head. On the dynamic manipulation benchmark DOMINO, VLAct-OFT leads both metrics with 18.50 SR and 34.20 MS, 7.64 and 3.71 above the same-backbone baseline.
| Method | Family | Base (Clean) | Data Scaling (Clean) | Data Scaling (Random) |
|---|---|---|---|---|
| pi-0 | Flow | 46.4 | 65.9 | 58.4 |
| pi-0.5 | Flow | 60.2 | 82.7 | 76.8 |
| LingBot-VLA | Flow | - | 88.6 | 86.7 |
| Abot-M0 | AML | - | 86.1 | 85.1 |
| InternVLA-A1 | Flow | - | 89.4 | 89.6 |
| Being-H0.7 | Flow | - | 90.2 | 89.6 |
| Fast-WAM | WAM | - | 91.9 | 91.8 |
| HoloBrain-0-QW | Diff. | - | 91.9 | 92.3 |
| Qwen3VL-OFT (same-backbone baseline) | OFT | 61.7 | 88.2 | 88.3 |
| VLAct (OFT head) | OFT | 80.5 | 92.5 | 90.8 |
| VLAct (GR00T head) | Flow | 76.0 | 89.6 | 87.4 |
| VLAct (PI head) | Flow | 77.0 | 93.0 | 88.8 |
Table 2: Task success rate (%) on RoboTwin 2.0. Base uses 50 clean trajectories per task; Data Scaling adds 500 domain-randomized expert trajectories per task.
Real robots and cross-embodiment transfer. The physical platform is a Franka Research 3 7-DoF arm with an external RealSense D435 and a wrist-mounted D405, all images resized to 224 pixels. On in-domain single-arm short-horizon tasks VLAct averages 92.5 versus 77.5 for the non-pretrained baseline; on novel-object out-of-domain tasks 90.0 versus 73.3 and 65.0; on long-horizon table cleaning and bean scooping 86.6/80.0 versus 73.3/33.3; on harder out-of-domain long-horizon variants (extended sequence, full object substitution) 82.5/83.3 versus 47.5/46.6; and although pre-trained only on single-arm data, VLAct transfers to dual-arm coordination with 72.0 versus 44.0, the largest gains on deformable and synchronization-heavy tasks such as pants folding (70 versus 40) and towel folding (50 versus 20).
Figure 5: Real-machine evaluation (a-c) and RoboCasa-GR1 cross-embodiment transfer (d). VLAct beats the non-pretrained baseline across single-arm short-horizon, long-horizon, and dual-arm settings, and surpasses full-data GR00T-N1.6 with only 20% downstream data.
Cross-embodiment transfer is the hardest test: the GR-1 humanoid and the ARX X5 bimanual platform are held out entirely during continued pre-training. On RoboCasa-GR1 with full fine-tuning data VLAct reaches 54.0, above full-data Qwen3VL-OFT (48.8), GR00T-N1.6 (47.6), and pi-0.5 (37.0); with only 20% of downstream trajectories it already reaches 49.5, matching or exceeding full-data Qwen3VL-OFT, then 51.0 at 50% and 54.0 at 100%. On the official RoboDojo leaderboard snapshot of 2026-08-24 with 35 policies, VLAct ranks eighth by average score and sixth by success rate (10.66 / 7.60%), outperforming every explicitly designated world-action-model entry including the strongest, X-WAM (7.69 / 3.83), as well as industrial systems such as Xiaomi-Robotics-0, GalaxeaVLA G0, LingBot-VLA, and Abot-M0; against the same-backbone StarVLA-alpha entry it improves the average score by 4.26 points and success by 4.36 points.
| Ablated component | Settings | Result |
|---|---|---|
| Shallow-layer protection | full update / freeze vision encoder / freeze vision encoder plus lower LLM half | LIBERO-Plus 78.9 / 81.3 / 82.6; RoboTwin 77.1 / 79.3 / 80.5 |
| Auxiliary co-training | robot-only / caption-only / mixed auxiliary | LIBERO-Plus 75.0 / 82.6 / 82.5 |
| Head diversity (downstream PI head) | none / OFT only / OFT+GR00T / all three | PI fine-tuning 60.5 / 55.1 / 63.1 / 77.0 |
| Unified action representation | separate heads / unified head / unified action representation | RoboTwin 78.5 / 79.5 / 80.5; LIBERO-Plus 81.1 / 81.4 / 82.6 |
| Wrap-aware loss | baseline / plus unified joint space / plus wrap loss | RoboTwin 75.5 / 78.6 / 80.5 |
| Heterogeneous UMI data | original mixture / plus 20k RealOmin trajectories | LIBERO-Plus 82.6 / 83.7 |
Table 3: Component ablation summary (numbers from Appendices C, D, E, F, G, and I). Each component is validated with all other settings fixed.
One final openness check: adding 20k RealOmin trajectories collected with a UMI-style handheld gripper, from a different embodiment and collection paradigm and converted to delta end-effector actions to fit the layout, raises LIBERO-Plus from 82.6 to 83.7. The recipe absorbs heterogeneous data rather than being disrupted by it, which is a useful practical signal for how far open-data mixtures can be pushed.
Limitations
First, the authors state that memory remains a clear limitation. Among the six RoboDojo capability axes, VLAct scores only 0.66 / 0.56% on Memory while the top entry DM0.5 scores 47.74 / 47.44 on that axis; the gains concentrate on Precision and Long-Horizon. Tasks that require maintaining and recalling online history are exactly what the current representation-centric recipe does not cover.
Second, the authors also acknowledge that mixed auxiliary co-training (82.5) falls slightly below caption-only co-training (82.6): with the total pre-training budget fixed, mixing several auxiliary sources reduces the sampling frequency of captions, the strongest single source. How much to mix and in what proportion remains a budget-allocation question, and the paper offers only a fixed-ratio snapshot.
Third, two caveats follow from the experimental design itself. The extra compute of multi-head co-supervision is recouped only indirectly through the backbone, since the pre-training heads are discarded at fine-tuning time; whether that cost pays off depends on downstream users actually switching heads, while the main and real-robot results still default to the OFT head. In addition, the backbone scale is fixed at 4B, real-robot evaluation uses only 10 trials per task, and the RoboDojo leaderboard does not normalize training compute, so the recipe's behavior at larger scale and under stricter statistics remains open.
Conclusion and Outlook
VLAct's contribution is not a new module but a redefinition of the objective of continued pre-training: not fitting actions on the pre-training distribution, but shaping a backbone representation that transfers across tasks, embodiments, environments, and action heads. Shallow-layer protection and caption mixing answer "what must not be damaged", multi-head co-supervision answers "what must not be over-specialized", and the partially unified action space with the wrap-aware loss answers "where should supervision be shared". All three answers are backed by controlled probes and validated under fixed downstream protocols, so the gains can be attributed cleanly to the backbone.
The direct implication for practitioners is that the VLM backbone should be treated as a first-order design variable of VLA systems rather than a fixed component inherited from general vision-language pre-training. Under limited data and compute budgets, spending the budget on representation design may yield more than doubling trajectories. The paper's open stance, fully open data, 16 GPUs, and promised release of scripts and checkpoints, makes this axis independently reproducible and extensible.
Golden Quote
Effective VLA continued pre-training requires designing how trajectory data shapes the backbone representation, rather than simply optimizing action prediction on the pre-training distribution.



