BLOG
Light Origins' Real2Sim2Real engine turns 2,000+ internet-sourced scenes into 4,000+ hours of aligned VLA experience. LightNav-0 tops 10 simulated settings and transfers zero-shot to humanoid, quadruped, aerial, and wheeled robots.
Official Jetson AI Lab tutorial: fine-tune models directly on Jetson with Unsloth and JetPack 7.2 using memory-efficient QLoRA, export to GGUF, and run locally with llama.cpp. Two hands-on examples: Qwen3.5-4B vision-language model on Jetson Orin Nano (LaTeX OCR fine-tuning), and NVIDIA Nemotron 3.5 Lightning 30B-A3B on Jetson AGX Thor (3 training steps in 44.8s, 66.1 tokens/sec with Q4_K_M quantization). No cloud required.
Tsinghua AIR and Z-Trans AI present Zetta ζ, a closed-loop embodied harness that improves a frozen VLA policy without a single gradient update. Instead of fine-tuning, it evolves the execution harness around the policy: code-based runtime critics watch every action, recovery skills take over on deviation, and a validation gate admits only skills that generalize. The frozen baseline scores 31.0% on LIBERO-Pro; the same policy under Zetta ζ reaches 92.5% (+56.3 absolute points), with a +20-point gain to 93.6% across 18 RoboCasa tasks. Z-Infra scales valid rollout throughput 20.6× and speeds inference 11.1×. Skills transfer zero-shot, and clear robotic "aha moments" emerge.
1X announces 25-DOF tendon-driven hands for NEO. Low gear ratios enable force transparency and backdrivability, combined with high-resolution tactile skin, turning every grasp into an experiment — a perception stack, not just an actuator.
1X introduces 1XWM, a video-pretrained world model integrated into NEO as a robot policy. Unlike VLAs, 1XWM derives robot actions from text-conditioned video generation, leveraging world dynamics from internet video for zero-shot generalization without large-scale teleoperation data.
Ludo Robotics presents Ludi₀.₁, their first agentic system integrating perception, navigation, and manipulation with interactive speech, dialogue, memory, and social reasoning — powered by a fine-tuned Qwen3.5 VLM, GR00T VLA policies, and KISS-ICP navigation.
Across ICRA REAL-I and CRAIC 2026, 200+ student teams used Leju's open-source LeTools chain to go from algorithm to real-robot deployment in as little as one day. Hardware-free dry-run of a 37-node behavior tree, a 1,000-hour production-grade LET-Base dataset, a one-line simulation-to-real switch in deploy.yaml, and 10+ architectures behind a unified Adapter layer show embodied-AI competition shifting toward low-barrier dev environments covering data, models, and deployment.
DYNA Robotics introduces Dyna-2, a world-action model pre-trained on 1M+ hours of egocentric human video, demonstrating the first human-to-robot transfer scaling law. With just hours of fine-tuning, it performs tasks across embodiments, achieving 87% zero-shot pass rate at customer sites.
Riemann Dynamics releases Riemann-1.0: a fully causal autoregressive World Action Model that unifies executable robot policy and action-conditioned world simulation in one architecture. Through progressive pretraining on 232K+ hours of heterogeneous embodied experience, it achieves SOTA on LIBERO (99.0%), RoboTwin 2.0 (94.3%), RoboCasa365 (62.6%), and real-world manipulation (85.0% SR).
Official 6B base weights, RoboTwin post-training weights, training code, open-loop eval, simulation eval, and deployment entry are all public. Labs can run inference, RoboTwin closed-loop evaluation, or custom-data post-training, but the full 60K-hour raw corpus and training budget are missing, making equivalent base pre-training reproduction infeasible.
Published in Nature Machine Intelligence, ergoCub is a humanoid robot designed via a shared embodied intelligence architecture that jointly optimizes hardware and control for human ergonomic metrics. L5-S1 torque drops ~50% during collaborative lifting, and walking step length increases 25% over its predecessor iCub3.
Reka releases RekaDaily-10k: 10,312 hours of unscripted first-person household life video, recorded by a global paid collector network in real homes, with ~1,670 hours in native 4K. Released on Hugging Face under Apache 2.0 in both raw and processed+captioned tiers.