VLK: Learning Humanoid Loco-Manipulation from Synthetic Interactions in Reconstructed Scenes
VLK synthesizes paired vision-language-kinematics supervision inside 3DGS-reconstructed real scenes: it generates navigation and object-interaction trajectories with privileged scene info, renders egocentric views after the fact, and produces 48,000 paired trajectories to train a policy predicting Unitree G1 whole-body motion, enabling sim-to-real perception-based humanoid loco-manipulation.
Yen-Jen Wang, Jiaman Li, Sirui Chen·Jun 29, 2026
Humanoidloco-manipulationSynthetic DataJun 29, 2026