Robocurve raised $10M to build independent physical-AI benchmarks on real robots.
Wheeled humanoid for factory automation, cited at 300 Nm joint torque and a 10 kg single-arm payload.
Three LLM controllers, one MuJoCo apple-plate task: Jev 1.13 and GPT-6 Astra place it; GPT-4.1 mini hits the 160-cycle limit. Jev cost $0.0188 vs $5.9336.
WholeBodyWAM introduces a transferable predictive motion prior for humanoid world-action modeling: instead of treating the vast motion data from humans and other robots as direct robot actions, it first learns the regularities of whole-body coordination. Trained on UniMotion-4K, a corpus of more than 4000 hours of human, 3D and humanoid motion, the model learns to predict future whole-body movement before seeing any target-robot action data; this motion expert is then combined with video and action models to generate embodiment-specific behavior, improving motion prediction, task performance, real-world manipulation and data efficiency when robot demonstrations are limited. The 29.8-second real-robot capture shows a black-and-white humanoid in an indoor room picking and relocating a grey storage bin beside a white shelving table, while an output future motion window overlays the predicted future whole-body skeleton in coloured lines, the motion prior made visible. Paper arXiv:2609.18197 with a project page and further demos.
A lab member introduced booster_mjlab, an mjlab (MuJoCo lab framework) training setup for the Booster K1 humanoid that ships a velocity tracking task, motion tracking, an AMP (Adversarial Motion Priors) implementation and more; the project page hosts additional interactive demos. The 11-second capture runs in a MuJoCo indoor football pitch with green turf, a centre circle and windowed walls: the K1 runs across the field, chases the ball and swings a foot to kick it, while a grey reference humanoid motion sequence plays alongside in a second viewport, the reference-versus-execution view of the motion-tracking task. The repo plugs the K1 model and training tasks into mjlab so velocity control and motion-imitation experiments can be reproduced in simulation.
Jev (typesafe.ai) never sees the sim camera: MuJoCo hands it the environment state as JSON and it answers with bounded typed actions (hover, descend, grasp, lift, place) plus a confidence score. Put the soup can in the yellow bin takes 9 decisions at ~150 ms each, and unmappable tasks are refused rather than guessed.
LimX Dynamics TRON 2 runs a temporally constrained VLA sequence: identify the drawer, pull it open, locate the toy, then retrieve it precisely. One wrong step fails the task.
Harvard SceneAgent turns image, video, LiDAR or 3DGS captures into interactive sim scenes: per-Gaussian semantics, baked predictive physics, part decomposition, articulation and digital sisters, exported as USDZ for Isaac Lab, MuJoCo and Unreal.
XPENG released XPACE, a unified embodied world model for IRON: one shared video backbone learns both the action and the post-action visual future, trained on 5,000 hours of embodied video spanning human egocentric, human video-action, bridge and teleoperation data. Its world simulator perturbs expert trajectories, simulates deviation and recovery, filters the generated examples and feeds recoveries back into policy training.
RealMan demonstrated an autonomous restroom-cleaning robot that plans paths, switches cleaning end-effectors, and completes multi-step workflows.
ugo Nova is a 22-DoF wheeled semi-humanoid with a 10 kg dual-arm payload, Jetson Thor onboard, and 100 Hz teleoperation that records synchronized video, joint, torque, and command data.
Odyssey-3 renders interactive generated worlds driven in real time by keyboard and mouse: a rabbit in a meadow, first-person cycling, a warehouse picker among shelves. If a world model can simulate infinite, intelligent worlds, it becomes the place where agents themselves are trained.