0:57
0:57
1:05
0:11
0:08
0:22
0:23
0:59
2:34
Large-scale vision-language-action (VLA) models provide powerful priors for robot manipulation, yet adapting them to a specific deployment remains challenging. Supervised fine-tuning (SFT) on task-specific demonstrations provides a step toward deployment, but faces two persistent limitations: static data provide limited coverage of out-of-distribution states, and standard imitation objectives do not distinguish progressing behavior from less useful data. Interactive post-training can address these limitations, but typically requires repeated policy execution and human intervention on a physical robot. We introduce HIL-UMI, a policy-guided Universal Manipulation Interface (UMI) framework for robot-free human-in-the-loop VLA post-training. During handheld UMI demonstrations, HIL-UMI queries the current policy on the same observation stream without executing its predictions. The Energy Score compares the human action trajectory with policy inference and triggers collection when their discrepancy indicates an out-of-distribution region. In a separate feedback loop, low online advantage predictions identify essential segments for refining a progress-based advantage estimator. The updated estimator then guides advantage-conditioned behavioral cloning using a balanced mixture of base demonstrations and new policy data. This design preserves the iterative and policy-aware nature of human-in-the-loop learning while decoupling data collection from robot deployment. Experiments on four real-world tasks spanning long-horizon and precise manipulation show that HIL-UMI achieves consistent improvement over SFT and benefits from both targeted collection and advantage refinement. Moreover, HIL-UMI outperforms HG-DAgger on Clean Up Table with lower per-frame collection time, suggesting a scalable path for VLA post-training across operators and locations.

Humanoid locomotion across complex terrain demands forward-looking exteroception to anticipate obstacles, yet this signal is unreliable in real-world deployment, failing partially and intermittently. Existing perceptive policies often assume that depth observations remain clean and in-distribution, while recent attempts to unify perceptive and blind control typically route or switch between separate sub-policies, leaving recoverable information in partially corrupted depth unexploited. We instead propose CAP, a single-stage humanoid locomotion policy that recovers this signal with a perceptive world-model encoder trained as a learned denoiser to reconstruct clean depth from a corrupted input, together with a co-active proprioceptive variational encoder that supplies depth-free body-state information. A coupled training recipe pairs a depth-noise curriculum on the world-model input with world-model feature dropout on the policy-facing latent, exposing the policy to failures across the entire perception-quality spectrum. In simulation, CAP matches or improves upon perceptive baselines when depth remains informative, and degrades more smoothly than a binary-switching baseline as perception worsens. On the Unitree G1, controlled trials and indoor-outdoor deployments demonstrate perception-robust locomotion under intermittent occlusion, real-sensor corruption, and outdoor depth artifacts.

World Action Models (WAMs) offer a promising approach to general-purpose robot manipulation by jointly modeling visual dynamics and actions. However, most WAM studies focus on tabletop or arm-centric manipulation, while humanoid loco-manipulation remains less explored. To address this gap, we introduce WholeBodyWAM, which jointly predicts future visual dynamics, manipulation actions, and whole-body control intents for generalizable humanoid loco-manipulation. It preserves pre-trained world-action priors while grounding heterogeneous whole-body controller (WBC) semantics and coordinating whole-body behavior. Extensive experiments show that WholeBodyWAM achieves an overall simulation task success rate of 91.9%, with a 0.23 improvement in real-world out-of-distribution task progress and a 70% reduction in success-rate variance across WBCs relative to the respective baselines. These results suggest a path toward scalable humanoid whole-body intelligence by extending pre-trained world-action priors through structured WBC grounding and coordination, rather than relearning whole-body behavior from scratch. Project page: https://wholebodywam.github.io/.

JEPA-Anything introduces orthogonal predictive factorization for domain-agnostic world modeling, splitting latent targets into complementary factors with dedicated predictive pathways. Across seven domains, it improves matched dynamics tasks and supports intervention, OOD generalization, and long-horizon forecasting.

As coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands into long trajectories of reasoning, tool use, and feedback. SoL-Pi scales recursive auto-research across executable environments to discover reusable harness mechanisms. Four retained mechanisms improve action execution, context compaction, observation handling, and delegated reading, reducing recorded token traffic by 44.7-49.0% and API cost by about one third at comparable EdgeBench performance.

ModularRSI is a benchmark-disjoint, contrastive, and modular framework for generalizable agent-harness evolution. It diagnoses recurring failures from paired successful and failed trajectories, evolves five functional modules independently, and integrates validated changes into a unified harness. On Terminal-Bench 2.0 and SWE-Bench Verified, the evolved harness improves unseen in-domain and cross-domain tasks and transfers across foundation models.

A deep read of TypeSafe's first System One Model: three question primitives, the economics of parallel calls, confidence-gated routing, eval caveats, eight jagged edges, and the OpenJev local repro.

The run book behind PaperRoute: mechanics before art, engine and look as separate threads, Blender driven by headless Python, Meshy for faces, review renders driving iteration, 39 tracked hours.

Meta distills compliance expertise into 200+ structured files, splits what the agent knows from how it reasons via recipes, and compiles expert fixes into regression-tested edits, no model retraining.
Hardware★11855%Next generation open-source KVM over IP for $69
JetKVM
Gadgets★6145%A lightweight EDC with an M390 blade that can cut through almost anything...
Hacksmith Industries
3D Printing★20614%4 Toolheads | 5s Toolhead Swap | Multi-Material | Low Waste | 500 mm/s Speed | Smart Calibration | Auto Filament System | App Control
Snapmaker
Hardware★10290%No noise to disturb sleep with Active Noise Cancellation and Snore Masking System. Sleep better with AI brainwave audio. Ultra comfort.
soundcore

A 100 km 3D city that opens in seconds — the digital twin of the AI agents that build this site: crawling, blogs, papers, industry and investment analysis.
100km city · opens in seconds · Live task stream · Agent-maintained

An SO-101 6-DoF arm running MuJoCo physics in your browser: joint teleoperation, IK end-effector dragging and gripper pick-and-place into a basket. No install.
MuJoCo WASM physics · SO-101 · 6 DoF · Contacts / telemetry HUD

Take a humanoid joint module apart layer by layer: brushless motor, magnetic encoder, planetary / harmonic / cycloidal drive and output flange. Switch architectures on one page — exploded view plus analytic kinematics.
Three gearbox types · Planetary / harmonic / cycloidal · Analytic kinematics · explode

Matcha-TTS mixed zh/en speech synthesis: server-side synthesis with sentence-streamed playback and full-article read-aloud for papers and blogs. Sign-in required.
Matcha-TTS · server-side · Mixed zh / en · Full-article read-aloud

A caring coding sprite on your desktop and a cockpit that never stops: the work resumes itself after a crash or a reboot. Every coding terminal, ssh session and local shell in one tree, every start, finish and permission request reported in time. Open source, one-line install.
Your caring coding desktop sprite · Resumes after crash & reboot · First-rate terminal / ssh / shell

The first Unreal-adapted physics-AI kernel for the browser: a deterministic ECS on a fixed timestep, a WebGPU scale layer driving 100k particles and 20k soft-body nodes, a policy-gradient learner training live in the page, and UE 5.5 levels playing smoothly in three.js.
UE 5.5 levels, read live · 100k particles · WebGPU · in-page RL training