PAPER DEEP DIVE
MobileVISTA: Generative Data Augmentation for Pose Generalization in Mobile Manipulation
Mobile manipulators such as humanoid robots are increasingly deployed in dynamic, unstructured environments to perform dexterous manipulation tasks. However, end-to-end manipulation policies trained to imitate demonstration data collected from a single robot pose are brittle: even centimeter-scale deviations in robot pose at deployment can drive ego-centric observations and end-effector trajectories out of the training distribution, leading to sharp drops in performance. We introduce MobileVISTA, a data generation framework that transforms demonstrations captured at canonical poses into diverse, pose-perturbed training data by jointly (1) augmenting egocentric visual observations and (2) retargeting actions to compensate for base pose changes. Unlike prior methods, which assume a camera rigidly mounted off the actuated chain or non-trivial articulated robot geometry largely out of frame, MobileVISTA targets compatibility with egocentric platforms (e.g., humanoids) where the camera is both influenced by and must observe the robot's kinematic chain as it moves. We study MobileVISTA in simulated tasks spanning humanoid and bimanual embodiments, and on a real Galaxea R1 Pro. We find policies trained on MobileVISTA-augmented data demonstrate improved robustness to previously out-of-distribution poses encountered at test time, without additional demonstration collection or a trained generative model. Additionally, we find MobileVISTA's benefit is largest on tested humanoids, where the camera rides the actuated chain and the robot fills much of the frame. Additional videos and appendix can be found on our website: https://mobilevista.github.io
Paper Info
MobileVISTA: Generative Data Augmentation for Pose Generalization in Mobile Manipulation
Authors: Suzannah Wistreich, Stephen Tian, Isabella Huang, Vitor Campagnolo Guizilini, Sergey Zakharov, Katherine Liu, Jiajun Wu. Affiliations: Stanford University and Toyota Research Institute (TRI). arXiv ID 2610.07511, v1 submitted October 5, 2026; the full HTML version is at arxiv.org/html/2610.07511v1.
Project page: sdwistreich.github.io/mobilevista, with method animations, real-robot deployment videos and the side-by-side comparison against GEN3C.
Code status: as of v1 the authors released only the project page; the same-named GitHub repository mobileVISTA/mobileVISTA.github.io is merely the hosting repo for that page and contains no method implementation. Every building block of the pipeline is open source, though: mask rendering rides on robosuite and MuJoCo, point-cloud deprojection and reprojection use Open3D, hole filling uses the official E2FGVI weights E2FGVI-HQ-CVPR22.pth, whole-body IK uses mink (a MuJoCo-based Python solver), and the policy-side baselines use the official implementations of real-stanford/diffusion_policy, YanjieZe/3D-Diffusion-Policy and Physical-Intelligence/openpi. What is missing for a reproduction is the glue code that chains these together, plus the robot URDFs and the camera calibrations.
One-Paragraph Summary
From demonstrations collected at a single canonical base pose, MobileVISTA jointly synthesizes perturbed egocentric observations and the matching retargeted actions, using geometric reprojection, an off-the-shelf video inpainting model and inverse-perturbation action retargeting. The same Diffusion Policy then works from any stance inside a 15 cm disk on a real Galaxea R1 Pro, while three strong baselines, including a fine-tuned π0.5, collapse under the same randomized stances. The whole method trains no generative model and needs not one extra demonstration.
Why "the robot stood 15 cm off" is a real problem
The usual division of labour in mobile manipulation is that the base navigates to somewhere near the target and the arm does the work. The catch is that navigation has error: wheel slip, uneven floors, localization drift. A few centimetres of offset is the norm, not the accident. Imitation-learning demonstrations, however, are typically collected at exactly one stance, so the policy learns a conditional distribution conditioned on that stance, a condition that was never written into the training objective.
The paper narrows the setting very specifically, and that specificity is what separates it from earlier augmentation work: the policy input is egocentric RGB-D, the camera sits on the kinematic chain (the humanoid's head, which moves with the waist joint), and the robot's own arms occupy a large part of the frame. Translate the base and two things change in the image at once: the viewpoint on the scene, and the position of the robot's own geometry in the frame.
A comparison table (paper Table I) shows that prior work satisfies at most one of these two properties: VISTA, RoVi-Aug, ROPA and RoboSplat all lack a mobile base; EgoDemoGen has a mobile base but its egocentric camera is rigidly mounted on a head fixed to the base, and it synthesizes with a video generation model fine-tuned per domain, with no public implementation at submission time; 1001 Demos puts the camera on the chain but the robot body never appears in frame; R2RGen has a mobile base but synthesizes no images at all. MobileVISTA is the only row in that table with all three boxes ticked, and the only one validated on a humanoid.
"Then why not just collect demonstrations at more stances?" In simulation the paper does exactly that as an upper bound, called Simulator (Oracle): move the base to the perturbed pose, retarget the actions with the same rules, record the observations again. But that upper bound has no real-robot counterpart: re-recording expert demonstrations per stance does not scale. Augmented data therefore has to be generated from a single demonstration, which is what "generative" in the title means: what gets generated is data, not a model.
"Then why not learned novel-view synthesis (NVS)?" Because NVS extrapolates the robot as part of a static scene, while under an egocentric camera the camera motion and the robot joint motion are coupled. The paper does not stop at argument: it runs GEN3C, a mainstream camera-controllable 3D-consistent video generation model, as a same-condition control, comparing both PSNR and the policies trained on its synthetic data, detailed in the experiments section.
Prerequisites
The four comparison methods each represent one "change the representation or the prior, not the data" idea: Canonical-2D trains Diffusion Policy directly on 2D images from the single stance; DP3 swaps the input for an explicit 3D representation such as point clouds, testing whether geometric inputs carry pose robustness on their own; Fine-tuned π0.5 starts from the pi05_base checkpoint and applies LoRA fine-tuning, testing whether large-scale pretraining priors suffice; Simulator (Oracle) supplies the perfect-synthesis upper bound. The key point is that MobileVISTA is a data generation framework, not a policy architecture: MobileVISTA, Canonical-2D and every ablation variant share exactly the same Diffusion Policy architecture, hyperparameters and training pipeline, and differ only in the data fed to them.
A base perturbation is written $\Delta p=(\Delta b,\ \Delta\theta)$, with translation $\Delta b\in\mathbb{R}^{2}$ and yaw $\Delta\theta$. To sample uniformly by area inside the disk (rather than clustering at the centre) the paper uses the standard transformation:
$$ (r,\theta)\sim\left(\sqrt{u}\cdot r_{\text{train}},\ 2\pi v\right),\qquad u,v\sim\mathcal{U}(0,1),\qquad \Delta b=(r\cos\theta,\ r\sin\theta,\ 0) $$
No vertical offset is added; when yaw jitter is enabled the paper additionally draws $\Delta\theta\sim\mathcal{U}(-5^{\circ},\ 5^{\circ})$. Training radius and evaluation radius take the same value throughout, $r_{\text{train}}=r_{\text{eval}}=15\ \text{cm}$, because the authors argue this range covers realistic post-navigation base error. Evaluation then sweeps the radius over $r_{eval}\in\{0,1,3,5,7,10,15\}\ \text{cm}$ to watch how the policy degrades at increasingly absurd stances.
Method in Detail
One demonstration contains time-aligned egocentric and wrist RGB-D images, proprioceptive state, and joint-space actions including gripper or finger commands. MobileVISTA samples one $\Delta p$ per demonstration and then changes two things at once: it synthesizes the egocentric observation at the perturbed stance, and it retargets the actions to compensate for that stance change. The wrist views are left untouched.
Visual branch, step 1: cut the robot out of the frame
The whole pipeline assumes you hold a simulation model of the robot (URDF or MuJoCo model), and it is used in three places: segmenting the robot in the original image, re-rendering the robot at the retargeted joint angles, and recovering the egocentric camera trajectory along the base-to-camera kinematic chain.
The simulation segmentation is deliberately engineering-flavoured: repaint the scene skybox in a conspicuous colour (red for the GR-1 humanoid, yellow for the dual-arm Franka, yellow chosen so it cannot collide with red objects in the scene), hide all non-robot geometry, render once, and threshold on the background colour to get a binary mask. On the real robot the route is different: render the R1 Pro model from its URDF in MuJoCo, take the depth buffer directly as the silhouette (threshold 0.995 to exclude background), then dilate the mask by 5 pixels to swallow edge fringes. Simulation renders at 512x512; the real robot renders at the dataset's native resolution and is upsampled 6x when the point cloud is reprojected.
Visual branch, step 2: deproject, lift to world, reproject
Applying the mask to the depth map yields a robot-free depth map, which a pinhole model deprojects into a coloured point cloud: $[X,Y,Z]^{\top}=Z\,K^{-1}[u,v,1]^{\top}$. Simulation intrinsics are derived from MuJoCo's vertical field of view cam_fovy; the real robot uses the calibrated ZED2 intrinsics $f_x=58.4,\ f_y=58.5,\ c_x=74.7,\ c_y=42.9$ (scaled proportionally with image size). The point cloud is built with Open3D's RGBDImage.create_from_color_and_depth, depth scale 1000 mm, depth truncation 5 m in simulation and 10 m on the real robot.
The next step is the most easily overlooked and most important design in the paper: the point cloud is not translated in the camera frame. It is first lifted into the world frame using the original per-frame camera extrinsics (MuJoCo's cam_xpos / cam_xmat), the base perturbation $\Delta p$ is applied in the world frame, and only then is it reprojected from the new egocentric camera extrinsics. And the new extrinsics are not picked arbitrarily: they come from feeding the retargeted actions into the robot model and reading them off the kinematic chain. The paper's statement can be written as one transformation chain:
$$ P_w = T^{\text{orig}}_{c\to w}\,P_c,\qquad P'_w = T_{\Delta p}\,P_w,\qquad \hat{p} = \pi\!\left(K,\ T^{\text{new}}_{w\to c}\,P'_w\right) $$
In other words, camera pose and actions are not two independent degrees of freedom. Under the same perturbed base, unconstrained intermediate joints would in principle admit many legal camera poses; MobileVISTA takes only the one resolved by IK from the retargeted actions. This constraint is the essential difference from augmentation methods that "just render the scene from another viewpoint", and it is the precondition that makes the egocentric setting work at all.
Visual branch, step 3: inpaint the holes with E2FGVI
The reprojected frame necessarily has holes: one region left by the robot that was just segmented away, one region of disocclusion newly exposed by the viewpoint change, plus missing pixels from incomplete depth. The authors choose video inpainting over nearest-neighbour fill because disocclusions expose content near boundaries that was never observed, not repetitive texture that can be copy-pasted. Concretely, all frames of a trajectory and their per-frame hole masks are fed as a temporal batch into the official E2FGVI pipeline; before inpainting, the hole masks get a 3x3 morphological closing (3 iterations) to seal the thin gaps caused by point-cloud sparsity; the weights are E2FGVI-HQ-CVPR22.pth.
Visual branch, step 4: re-render the robot at retargeted joints and composite
After inpainting you hold a clean background, but the frame still lacks the robot itself. So the robot is rendered once more in MuJoCo at the retargeted joint angles (again hiding all non-robot geometry) and composited onto the background. Simulation uses the colour mask (excluding skybox-coloured pixels) to handle background bleed-through; the real robot uses the re-render's depth buffer to decide robot pixels, again dilated by 5 pixels. The real-robot output is finally downsampled back to the dataset's native resolution. The entire chain is automatic, with no manual curation at any step.
Action branch: inverse perturbation in end-effector space
The retargeting objective is a single sentence: keep the world-space trajectory of the end effector as unchanged as possible, because that trajectory is known to complete the task. If the source actions are joint-space, forward kinematics first converts them into base-frame end-effector pose targets, and each frame then receives the inverse of the perturbation. The position part is a 4x4 homogeneous transform (inverse yaw rotation plus translation):
$$ \mathbf{p}'_t = T^{-1}_{\Delta p}\cdot[\mathbf{p}_t;\ 1] $$
The orientation part, represented as a rotation vector, is composed with the inverse yaw rotation and converted back to a rotation vector:
$$ R'_t = R_{\text{inv yaw}}\cdot R(\mathbf{r}_t) $$
Gripper and finger commands pass through untouched. For the dual-arm Franka rig, both base poses and the shared camera pose are cached; the yaw rotation is applied about the midpoint of the two base positions to preserve their relative configuration, so each arm's compensation contains both a rotation and a translation derived from the change of that arm's base pose.
Action branch: whole-body IK on the real robot
Simulation demonstrations already record end-effector poses, so retargeting there is closed-form; the real robot's JoyLo teleoperation records joint-space commands, which must go through IK. The authors use mink for whole-body inverse kinematics with a configuration detailed enough to copy: a FrameTask on each end-effector site (left_ee_center / right_ee_center) with position cost 1.0 and orientation cost 1.0; a PostureTask regularizing toward the previous frame's configuration with cost $10^{-5}$; ConfigurationLimit on joint positions; per-step velocity caps of 1.0 on arm joints and torso joint limits of 0.02 on Book Stacking versus 0.04 on Laundry Sort, a deliberate difference that prefers arm motion over torso motion to stabilize the viewpoint while letting the torso move more on the long-horizon Laundry Sort; and a CollisionAvoidanceLimit between end effectors and torso with minimum collision distance 0.04 m and detection distance 0.18 m. At most 20 solver iterations per frame, $\Delta t=0.01$, quadprog as the underlying solver, damping $10^{-3}$, convergence position tolerance 0.01 m. As an optimization problem it reads roughly:
$$ \min_{q_t}\ \sum_{i\in\{L,R\}}\Big(\lVert \mathrm{FK}_i(q_t)-\mathbf{p}'_{t,i}\rVert^{2}+\lVert \mathrm{orient}_i(q_t)-R'_{t,i}\rVert^{2}\Big)+10^{-5}\lVert q_t-q_{t-1}\rVert^{2} $$
subject to joint limits, per-step velocity caps and end-effector-to-torso collision distances. The IK target for frame 0 is the original end-effector pose transformed by $\Delta p$ about the rig pivot; later frames simply let IK chase the original end-effector trajectory, allowing the arms to reach the same world-space target with different elbow configurations from the new base.
The authors' attitude toward IK failure is pragmatic: within the sampled 15 cm range every trajectory converged, so no filtering is applied; and once a perturbation is large enough that IK cannot solve it, that case largely coincides with "the task was unreachable from that stance anyway", which is a repositioning problem, not a data-augmentation problem.
Why the wrist camera stays untouched
The paper augments only the egocentric camera and keeps wrist views as recorded. Two reasons: the wrist field of view is too local and its close-range interaction depth is too sparse for geometric reprojection to be reliable on such views; and since the retargeting objective is to keep the end-effector trajectory unchanged, what the wrist camera sees is largely self-consistent under base perturbation. This trade is also one of the limitations the authors list themselves.
Data ratios
In simulation each source demonstration yields 1 augmented trajectory: 200 source demonstrations become 200 augmented trajectories (1:1). On the real robot each source demonstration yields 10: 50 source demonstrations become 500 augmented trajectories (1:10). Each real task collects 50 expert demonstrations (about 150 timesteps each for Book Stacking), teleoperated with the JoyLo framework (from BEHAVIOR Robot Suite) under ROS 2 at 10 Hz, recording synchronized RGB-D, proprioception and joint actions.
The flowchart below is drawn from the paper's actual implementation (Section III and Appendix A). Note that the visual branch's new camera extrinsics come from the action branch's solution: the two branches meet at exactly that point.
flowchart TD D["Demonstration at one canonical stance
RGB-D + proprioception + joint actions"] --> S["Sample perturbation delta_p
disk r_train = 15 cm, yaw +/-5 deg"] S --> V1["Render the robot in MuJoCo / robosuite
to obtain a 2D mask"] V1 --> V2["Mask the robot out of the depth map
real: URDF depth buffer, thresh 0.995, dilate 5 px"] V2 --> V3["Pinhole deprojection to a coloured point cloud
Open3D, depth truncation 5 m sim / 10 m real"] V3 --> V4["Lift to the world frame with original extrinsics
cam_xpos / cam_xmat"] V4 --> V5["Apply delta_p in the world frame"] S --> A1["Joint actions -> forward kinematics -> base-frame EE targets"] A1 --> A2["Apply the inverse perturbation
position and rotation compensated, gripper passed through"] A2 --> A3{"Where do the actions come from"} A3 -->|"sim: demonstrations are EE poses"| A4["Closed-form retargeting"] A3 -->|"real: JoyLo joint-space commands"| A5["mink whole-body IK
up to 20 steps per frame, dt 0.01, tol 0.01 m
torso limits 0.02 / 0.04, collision 0.04 m"] A4 --> V6["Forward kinematics along the chain
new egocentric camera extrinsics"] A5 --> V6 V5 --> V7["Reproject from the new extrinsics
OffscreenRenderer outputs RGB + hole mask"] V6 --> V7 V7 --> V8["E2FGVI video inpainting
3x3 closing x3, weights E2FGVI-HQ-CVPR22"] V8 --> V9["Re-render the robot at retargeted joints
composite onto the inpainted background"] V9 --> O1["Perturbed egocentric observation"] A4 --> O2["Retargeted actions"] A5 --> O2 O1 --> T["Augmented dataset
sim 1:1 (200 to 200), real 1:10 (50 to 500)"] O2 --> T T --> P["Diffusion Policy training
architecture and hyperparameters identical to Canonical-2D"] W["Wrist camera views"] -.->|"kept as-is, never reprojected"| T

Figure 1 (paper Fig. 1): left, demonstrations collected at a single canonical stance; middle, the jointly synthesized egocentric observations and retargeted actions; right, the trained policy completing the task across a whole field of unseen stances.

Figure 2 (paper Fig. 2): framework overview. The visual branch removes the robot, reprojects the scene, then renders the robot back in at its retargeted configuration; the action branch retargets actions with the inverse perturbation in end-effector space. Only the two branches together make one usable augmented datum.
Experiments
Simulation uses robomimic plus MuJoCo with four tasks from DexMimicGen: Can Sort and Pouring on the Fourier GR-1 humanoid, Box and Drawer on dual-arm Franka. The GR-1 head camera field of view is widened from DexMimicGen's default 50 degrees to 90 degrees so that the entire manipulation workspace stays in frame. Every result is the mean over 3 random seeds with 50 rollouts per seed. Can Sort balances training demonstrations evenly by can colour; the other tasks train on the first 200 demonstrations and evaluate on DexMimicGen's environment variants.
Q1: who still stands up outside the training stance distribution
The table below is an excerpt of the paper's Appendix Table III (success rate, mean plus/minus std, 3 seeds x 50 rollouts), showing four representative evaluation radii. Italics mark the Simulator (Oracle) upper bound.
| Task | Method | r=0 | r=5 cm | r=10 cm | r=15 cm |
|---|---|---|---|---|---|
| Can Sort (GR-1 humanoid) | Canonical-2D | 0.480 ± 0.02 | 0.440 ± 0.05 | 0.240 ± 0.00 | 0.120 ± 0.06 |
| DP3 | 0.733 ± 0.08 | 0.407 ± 0.05 | 0.153 ± 0.04 | 0.060 ± 0.02 | |
| Fine-tuned π0.5 | 0.926 ± 0.05 | 0.827 ± 0.10 | 0.607 ± 0.09 | 0.320 ± 0.03 | |
| MobileVISTA | 0.953 ± 0.03 | 0.887 ± 0.05 | 0.793 ± 0.06 | 0.527 ± 0.04 | |
| Simulator (Oracle) | 0.913 ± 0.02 | 0.927 ± 0.03 | 0.727 ± 0.01 | 0.527 ± 0.08 | |
| Pouring (GR-1 humanoid) | Canonical-2D | 0.547 ± 0.03 | 0.353 ± 0.03 | 0.080 ± 0.02 | 0.073 ± 0.06 |
| DP3 | 0.287 ± 0.08 | 0.173 ± 0.06 | 0.080 ± 0.06 | 0.013 ± 0.01 | |
| Fine-tuned π0.5 | 0.467 ± 0.03 | 0.333 ± 0.05 | 0.093 ± 0.03 | 0.013 ± 0.01 | |
| MobileVISTA | 0.673 ± 0.02 | 0.613 ± 0.04 | 0.547 ± 0.12 | 0.460 ± 0.02 | |
| Simulator (Oracle) | 0.773 ± 0.03 | 0.727 ± 0.03 | 0.700 ± 0.07 | 0.560 ± 0.11 | |
| Box (dual-arm Franka) | Canonical-2D | 0.307 ± 0.11 | 0.147 ± 0.10 | 0.047 ± 0.02 | 0.000 ± 0.00 |
| DP3 | 0.407 ± 0.01 | 0.333 ± 0.02 | 0.213 ± 0.03 | 0.087 ± 0.05 | |
| Fine-tuned π0.5 | 0.187 ± 0.10 | 0.147 ± 0.01 | 0.060 ± 0.03 | 0.020 ± 0.00 | |
| MobileVISTA | 0.520 ± 0.14 | 0.440 ± 0.07 | 0.393 ± 0.07 | 0.247 ± 0.05 | |
| Simulator (Oracle) | 0.313 ± 0.09 | 0.407 ± 0.06 | 0.373 ± 0.06 | 0.333 ± 0.01 | |
| Drawer (dual-arm Franka) | Canonical-2D | 0.287 ± 0.10 | 0.100 ± 0.04 | 0.040 ± 0.02 | 0.013 ± 0.01 |
| DP3 | 0.260 ± 0.00 | 0.180 ± 0.03 | 0.067 ± 0.06 | 0.040 ± 0.02 | |
| Fine-tuned π0.5 | 0.580 ± 0.09 | 0.373 ± 0.11 | 0.167 ± 0.08 | 0.060 ± 0.02 | |
| MobileVISTA | 0.653 ± 0.04 | 0.593 ± 0.06 | 0.487 ± 0.09 | 0.413 ± 0.01 | |
| Simulator (Oracle) | 0.720 ± 0.03 | 0.713 ± 0.07 | 0.640 ± 0.09 | 0.393 ± 0.03 |
Three readings. First, the degradation rate: Canonical-2D collapses quickly as the radius grows, averaging under 45 percent across the four tasks at r=15 cm; DP3 and fine-tuned π0.5 match or slightly beat Canonical-2D on most tasks but fail to generalize as the perturbation grows. Second, the exception: on Can Sort, fine-tuned π0.5 already approaches MobileVISTA and the Oracle for $r_{eval}$ below 5 cm, showing that a large pretraining prior does patch small perturbations, but not out to 15 cm. Third, and most interesting: the Oracle is not a strict upper bound. On Box at r=0 the Oracle scores only 0.313 while MobileVISTA scores 0.520; on Drawer at r=15 cm MobileVISTA's 0.413 overtakes the Oracle's 0.393. The distributional diversity that synthetic data brings is, in some cells, worth more than observations that are perfect but single-sourced.
Q2: both modalities must be augmented together, and the robot must be re-rendered
The ablation design is clean: Aug. Actions Only pairs the original canonical observations with retargeted actions, Aug. Images Only pairs augmented observations with the original actions, and both strictly preserve demonstration order, so exactly one modality is augmented relative to the source demonstration. Fixed Robot Geom. removes the robot segmentation and re-render-composite stages from the visual branch entirely, reprojecting the whole scene including the robot from depth into the target view and treating the robot as a static object in the environment.
The result: augmenting either modality alone consistently hurts, and Aug. Actions Only hurts worst (near 0.00 ± 0.00 on Pouring, only 0.07 to 0.11 on Can Sort), because images and actions that disagree teach the policy to look at one thing and do another. The cost of dropping the robot re-render is strongly task-dependent: averaged over radii, MobileVISTA's lead over Fixed Robot Geom. is +0.47 on Pouring, +0.28 and +0.24 on Can Sort and Box, and about 0.00 on Drawer. The authors' hypothesis is that two properties decide whether re-rendering pays: how much of the egocentric frame the robot geometry occupies, and whether the camera sits on the kinematic chain. Pouring uses the GR-1 with its active waist joint and scores on both, hence the largest gap; on the two Franka tasks the camera is on a fixed mount and sees little of the robot body, so the effect weakens. The authors are candid that this is only a hypothesis: four tasks cannot disentangle the two properties, nor explain the difference between Box and Drawer.
Q3: against learned NVS, geometric reprojection wins on fidelity and speed
For each task the paper samples 200 perturbed configurations inside the 15 cm disk, renders the oracle view from the simulator as ground truth, and synthesizes the same camera perturbation with MobileVISTA and with GEN3C, then computes PSNR. Below are the complete numbers from the paper's Appendix Table V (in dB).
| Method | Can Sort 0-5 / 5-10 / 10-15 cm | Pouring 0-5 / 5-10 / 10-15 cm | Box 0-5 / 5-10 / 10-15 cm | Drawer 0-5 / 5-10 / 10-15 cm |
|---|---|---|---|---|
| GEN3C | 16.29 / 14.42 / 13.08 | 17.74 / 15.09 / 13.28 | 16.37 / 13.75 / 11.78 | 16.32 / 15.76 / 14.35 |
| MobileVISTA | 25.02 / 24.05 / 23.28 | 23.71 / 23.25 / 22.18 | 16.47 / 16.01 / 14.72 | 18.67 / 18.93 / 18.29 |
| Delta | +8.73 / +9.63 / +10.20 | +5.97 / +8.16 / +8.90 | +0.10 / +2.26 / +2.94 | +2.35 / +3.17 / +3.94 |
The gap is not uniform across embodiments: in the 10 to 15 cm band MobileVISTA leads by 10.2 dB and 8.9 dB on the two GR-1 humanoid tasks but only 2.9 dB and 3.9 dB on the two Franka tasks. The humanoid tasks are exactly where the robot body fills more of the frame and the camera moves with the waist, which is where geometric reprojection's advantage over learned NVS is largest. On speed, mean synthesis time per trajectory is 14.93 / 13.65 / 14.93 / 16.85 minutes for GEN3C (Can Sort / Pouring / Box / Drawer) versus 3.81 / 4.31 / 2.39 / 2.99 minutes for MobileVISTA, a 4x to 7x speedup; and MobileVISTA's timing was measured on a single L40S while GEN3C ran on a single, stronger H200, so the ratio is conservative. The authors attribute the gap to GEN3C's iterative diffusion sampling per frame, versus MobileVISTA's one geometric reprojection pass plus one inpainting pass.

Figure 3 (paper Fig. 6): each dot is one synthesized trajectory's PSNR against its sampled base offset (top view, in cm). MobileVISTA sits higher overall and separates most cleanly on the humanoid tasks; on tasks where the two methods' PSNR is close, the downstream policy success gap is also smaller.
More important is the downstream effect: the authors train policies on GEN3C-synthesized observations paired with the same action retargeting and the same Diffusion Policy configuration, and those policies lose to the MobileVISTA-trained ones on all four tasks and at every evaluation radius, with gap sizes oriented the same way as the PSNR gaps. PSNR here is not a decorative metric: it actually predicts whether the policy can learn.

Figure 4 (paper Fig. 8): qualitative comparison at the same perturbed base pose, showing for each task the simulator oracle view, the MobileVISTA synthesis and the GEN3C synthesis. MobileVISTA keeps crisp object boundaries and consistent scene geometry; GEN3C tends to smear object edges and warp surfaces as the viewpoint change grows.
Q4: real-robot results on the Galaxea R1 Pro
The hardware is one Galaxea R1 Pro humanoid with a head-mounted ZED2 RGB-D providing the shared egocentric view for all policies, plus one Intel RealSense RGB per wrist. The evaluation protocol deserves its own paragraph because it rules out luck thoroughly: a floor grid is taped by hand, centred on the canonical stance and extending 15 cm in every direction; each trial samples an offset uniformly in the disk and a human drives the robot to that position; yaw is not controlled separately because training already covered the full yaw range; from the same offset, all four policies (MobileVISTA, Canonical-2D, DP3, Fine-tuned π0.5) run back to back without repositioning, in a randomized order unknown to the scorer. Inference runs on a single RTX 4090. Each task gets 25 canonical-stance trials plus 25 randomized-stance trials.
Below are the complete numbers from the paper's Table II (each cell is successes out of 25 trials; randomized stances within a 15 cm radius).
| Task / stance | Method | Metric | Successes |
|---|---|---|---|
| Book Stacking randomized stance | Canonical-2D | grasped | 1/25 |
| placed on stack | 1/25 | ||
| DP3 | grasped | 2/25 | |
| placed on stack | 1/25 | ||
| Fine-tuned π0.5 | grasped | 2/25 | |
| placed on stack | 2/25 | ||
| MobileVISTA | grasped | 17/25 | |
| placed on stack | 13/25 | ||
| Book Stacking canonical stance | Canonical-2D | grasped | 25/25 |
| placed on stack | 24/25 | ||
| DP3 | grasped | 20/25 | |
| placed on stack | 8/25 | ||
| Fine-tuned π0.5 | grasped | 25/25 | |
| placed on stack | 25/25 | ||
| MobileVISTA | grasped | 19/25 | |
| placed on stack | 18/25 | ||
| Laundry Sort randomized stance | Canonical-2D | at least 1 / 2 / 3 / 4 sorted | 6 / 3 / 2 / 1 (each /25) |
| DP3 | at least 1 / 2 / 3 / 4 sorted | 3 / 2 / 2 / 2 (each /25) | |
| Fine-tuned π0.5 | at least 1 / 2 / 3 / 4 sorted | 14 / 2 / 1 / 1 (each /25) | |
| MobileVISTA | at least 1 / 2 / 3 / 4 sorted | 18 / 12 / 6 / 2 (each /25) | |
| Laundry Sort canonical stance | Canonical-2D | at least 1 / 2 / 3 / 4 sorted | 21 / 15 / 5 / 3 (each /25) |
| DP3 | at least 1 / 2 / 3 / 4 sorted | 19 / 11 / 8 / 3 (each /25) | |
| Fine-tuned π0.5 | at least 1 / 2 / 3 / 4 sorted | 21 / 8 / 7 / 5 (each /25) | |
| MobileVISTA | at least 1 / 2 / 3 / 4 sorted | 24 / 15 / 9 / 4 (each /25) |
The randomized-stance comparison is lopsided: on Book Stacking MobileVISTA grasps 17/25 and places 13/25 while all three baselines stay at or below 2/25 in both stages. On Laundry Sort MobileVISTA sorts at least one item in 18/25 trials and at least two in 12/25, against the strongest baseline, fine-tuned π0.5, at 14/25 for at-least-one but only 2/25 for at-least-two, and Canonical-2D at 6/25 and 3/25. The cost is stated openly: on canonical-stance Book Stacking MobileVISTA grasps 19/25 and places 18/25 while Canonical-2D reaches 25/25 and 24/25 and fine-tuned π0.5 reaches 25/25 and 25/25. Augmentation bought robustness and shaved a little peak precision. Interestingly that precision cost does not appear on Laundry Sort, where MobileVISTA is best or tied-best on three of the four thresholds.
The failure-mode analysis says more than the success rates. MobileVISTA generalizes especially well to leftward and forward offsets; its remaining failures are mostly joint-limit violations when reaching for books, which the authors attribute to retargeted trajectories that require the arms to cross the torso. The baselines, by contrast, replay the demonstration trajectory as-is: forward offsets drive them into the shelf, backward offsets leave the gripper short of the book; fine-tuned π0.5 additionally emits actions beyond joint limits even though its qualitative failure modes resemble Canonical-2D and DP3. On Laundry Sort, Canonical-2D and DP3 fail to compensate backward translation and stop early during grasping, and Canonical-2D still performs the two-hand transfer after an empty grasp; neither pattern appears in any MobileVISTA rollout. MobileVISTA does hit the table and baskets occasionally, but on backward offsets, manifesting as overshooting items rather than failing to reach them.
Limitations
The three the authors state themselves: first, synthesis quality depends on camera field of view, since a narrower field leaves less visible context and can degrade both generated-view consistency and policy performance, a knob that was in fact already turned in their own setup, where the GR-1 head camera was widened from the default 50 degrees to 90 degrees to keep the workspace visible. Second, only the egocentric camera is augmented; wrist cameras are not, and the robustness gains multi-view augmentation might bring remain unexplored. Third, every conclusion is bounded by the sampled range (15 cm translation, ±5 degrees yaw); behaviour outside it was never measured.
Four additional judgments of my own:
First, the pipeline hard-depends on holding an accurate robot model and camera calibration. The real-robot chain needs the URDF, the calibrated ZED2 intrinsics, and a MuJoCo model that can render a matching depth buffer. Many real deployments lack exactly these, especially a URDF and hand-eye calibration that agree with the physical robot.
Second, the scene is treated as a static point cloud and the objects themselves have no model. The authors state explicitly that, as in standard behaviour cloning, task-relevant contact and object dynamics are assumed roughly invariant under the sampled perturbations. That is reasonable for rigid cans, books and drawers, and less so for deformable cloth (the socks and towels of Laundry Sort); and E2FGVI inpaints content that was never observed, so an inpainting error teaches the policy something wrong. The paper never quantifies semantic correctness of inpainted regions, proxying it with whole-frame PSNR.
Third, statistical power is limited. Each real-robot cell holds only 25 trials, and a gap like 17/25 versus 14/25 is not significant under a binomial model; the paper runs no significance test. The truly uncontested comparison is the 17/25 versus at-most-2/25 tier.
Fourth, the learned-NVS comparison uses a single model, GEN3C, and never fine-tunes it for robot scenes. The authors themselves note in the Table I discussion that methods like EgoDemoGen use video generation models fine-tuned per domain, meaning the opponent in this comparison did not get its best configuration.
Takeaways and Outlook
MobileVISTA's contribution is not a new generative model but a purely geometric pipeline that solves a setting nobody had handled all at once: egocentric views, camera on the kinematic chain, robot body in frame, and it solves it more faithfully and faster than learned NVS. Three transferable conclusions are worth remembering: actions and visuals must be augmented coupled, since augmenting one side alone is worse than augmenting neither; the robot body must be re-rendered at its retargeted configuration, or platforms with chain-mounted cameras lose a large chunk of performance; and camera extrinsics must be solved from the actions, never sampled independently. All three apply directly to any team doing embodied data augmentation, regardless of whether diffusion models are involved.
What it leaves unsolved is equally clear: the canonical-stance precision loss on the real robot, the missing wrist views, and the dependence on robot models and calibration. Of the two directions the authors name, combining MobileVISTA with large pretrained policies and extending augmentation to wrist cameras, the former deserves special attention: the tables already show fine-tuned π0.5 approaching MobileVISTA at small perturbations, so feeding MobileVISTA-generated data into fine-tuning π0.5 instead of training a from-scratch Diffusion Policy could in principle stack the two routes' strengths.
Golden Quotes
"Within the sampled 15 cm range all trajectories converged, so we apply no filtering. Under significant perturbations, IK failure would largely coincide with the task being unreachable from that base pose, a case for repositioning rather than augmentation."
On why the pipeline needs no IK-failure filter: augmentation has a jurisdiction, and unreachable stances are outside it.
"Egocentric camera poses are not chosen independently of actions: a perturbed base may admit multiple valid camera poses, but the camera extrinsics are inherited from the joint configuration resolved by the retargeted actions via IK."
The sentence that separates this method from "render the scene from another angle": the camera is a consequence of the action, not a free variable.



