RoboRender: Turning Simulated Trajectories into Photorealistic Video for Policy Training

Loading video
Loading videoRoboRender tackles the visual sim-to-real gap by converting simulated trajectories into photorealistic RGB video for policy learning. A robot-oriented video generation model is conditioned on simulated depth, language instructions and robot RGB masks, preserving simulator geometry, robot motion and action labels while synthesizing realistic textures, backgrounds and distractors. Policies trained on the generated video reach a 71% average zero-shot real-world success rate across pick-and-place, articulated-object and mobile manipulation, about 7.1x raw simulation rendering and 3.6x visual domain randomization. From Stanford, Nvidia and Microsoft Research.
Category: research
Author: @RavenHuang4
Date: 2026-10-08T00:00:00
Duration: 122.299s
Reference: https://arxiv.org/abs/2610.09254





