Skip to content
←Back to Papers

PAPER DEEP DIVE

World ModelsNavigationEmbodied AI

SparseVideoNav: Video Generation Models Crack Beyond-the-View Navigation — First Nighttime BVN, 2.5× SOTA Success

SparseVideoNav (HKU & OpenDriveLab) replaces LLMs with a Video Generation Model as the carrier of long-horizon foresight: predicting only sparse keyframes (interval 3, spanning 20 seconds) lifts average BVN success to 25.0%, 2.5× the strongest LLM baseline; PCM distillation plus sparsification yields a 27× end-to-end speedup (inference 21.6s→0.8s). Built on a 140-hour real-world navigation dataset, it demonstrates nighttime BVN capability for the first time. An in-depth read of the sparse design, the four-stage training pipeline, and 240 real-robot trials.

Overview & Motivation: Why "Beyond the View" Is the Real Problem

If the last two years of progress in Vision-Language Navigation (VLN) have mostly come from plugging large language models into robots, this paper from the University of Hong Kong and OpenDriveLab asks a sharper question: why does navigation have to be tied to detailed, verbose language instructions at all? In a real home or campus, humans say things like "find the desk and stop next to it" or "go look for the trash can up ahead" — not a step-by-step script of every turn and corridor. The authors formalize this setting, where only high-level intent is given and the goal lies far outside the field of view, as Beyond-the-View Navigation (BVN), and explicitly distinguish it from traditional Instruction-Following Navigation (IFN).

The paper's core diagnosis is that existing LLM-based navigation methods are systematically "myopic" on BVN. They are trained under short-horizon supervision — typically action sequences of only 4 to 8 steps — forcing the model to infer long-horizon intent from an extremely short causal chain. At deployment this produces two characteristic failure modes: first, when the target is invisible over long distances, uncertainty explodes and the robot spins aimlessly; second, at dead ends it misjudges the situation and gets stuck. Crucially, the authors argue the intuitive fix — simply stretching the supervision horizon — is not viable, because extending the horizon destabilizes LLM training.

This motivates the paper's paradigm-level shift: use a Video Generation Model (VGM). The key insight is that VGMs are inherently pretrained to predict long-horizon futures conditioned on language; their representations are naturally aligned with the temporal evolution of visual scenes. In contrast, the language token space of an LLM carries no such spatiotemporal prior. The authors therefore position the VGM as the natural interface for the "long-horizon foresight" that BVN demands.

Core Contribution: Generate a Sparse Future, Not a Full Video

But the paper does not simply adopt the standard video generation recipe. The authors ask a more interesting question: does navigation really need continuous, high-frame-rate future video? Their answer is no. The high-frequency temporal detail that continuity demands is redundant for deciding where to go. This observation leads to the paper's most important technical claim — sparse video generation: instead of predicting every frame, the model predicts only carefully chosen key timesteps, and those sparse frames serve as the direct supervision signal for the VGM.

This yields a double win: the prediction horizon is significantly extended (a fixed number of predicted frames now covers a much longer time span), while training and inference overhead actually decrease. Through ablations, the authors identify a sparse interval of 3 as the optimum, balancing prediction horizon against visual fidelity. An interval of 1 (dense prediction) preserves image quality but the horizon is too short for the model to "see" distant goals; an interval of 5 extends the horizon further but causes pronounced fidelity degradation, and the resulting blur itself harms goal recognition.

To keep action prediction precise, the authors retain continuous generation for the first two observation chunks (covering 8 timesteps). The final sparse generation schedule is [T+1, T+2, T+5, T+8, T+11, T+14, T+17, T+20], spanning 20 seconds at 4 FPS, where T is the current time. This is precisely what enables SparseVideoNav to complete trajectory reasoning in sub-second time.

flowchart TD A["Current observation I_t
History H"] --> B["History compression
Q-Former temporal + Video-Former spatial"] L["Language instruction l
umT5 encoder"] --> C["VGM backbone
Wan2.1-1.3B"] B --> C A --> C C --> D["Sparse future latents
T+1,T+2,T+5,...,T+20
20s @ 4 FPS"] D --> E["DiT action head
cross-attention instruction injection"] E --> F["Continuous action
a_0"] G["DA3 relabeling of generated frames"] --> E

A Four-Stage Training Pipeline: From Text-to-Video to Action Learning

Turning "sparse foresight" into a deployable navigation system faces two engineering challenges. First, injecting full history plus the denoising steps needed for dynamic scene generation brings inference latency no real robot can afford. Second, unlike LLM methods that can bridge the sim-to-real gap by co-training on heterogeneous real data, VGMs lack such a direct mechanism. The authors answer the former with a four-stage training pipeline and the latter with a purpose-built real-world navigation dataset.

Stage 1: T2V → I2V. The backbone is Wan2.1-1.3B, a text-to-video (T2V) model, which compresses the spatiotemporal dimensions with a 3D causal VAE: the input video is encoded into latent chunks, each of shape [H/8, W/8, 16], and all subsequent training operates at the chunk level. Since T2V generates the future primarily from language rather than visual input, the first stage adapts it into an image-to-video (I2V) model to ensure generated futures stay consistent with the initial observation. Fine-tuning follows Wan's original flow matching objective: given the current chunk's latent, sparse target frames x₁, random noise x₀, and a timestep t sampled from a logit-normal distribution, the intermediate latent is constructed as x_t = t·x₁ + (1−t)·x₀ with ground-truth velocity v_t = x₁ − x₀; the loss is the mean squared error between predicted and true velocity.

Stage 2: History injection. A navigation foundation model — unlike general vision-language-action models — must consume full observation history, and VGMs, unlike LLMs, have no built-in ability to process long sequences of image tokens. Borrowing from the CDiT architecture, the authors insert an extra cross-attention block into every transformer block of the Wan backbone to explicitly inject history; to preserve the generative prior of the fine-tuned I2V model, the final linear layer of each new block is zero-initialized. Because raw history is too long and too high-dimensional, it is compressed in two steps: a Q-Former first processes temporal features, then a Video-Former handles spatial features, yielding a history embedding h_T that is merged into the training objective.

Stage 3: Diffusion distillation. Navigation differs from robotic manipulation: manipulation scenes change little, so few denoising steps suffice for faithful reconstruction, whereas navigation involves drastic scene transitions — generating high-fidelity future frames in few steps is hard in itself and is exactly the bottleneck for real deployment (existing autonomous-driving video generation methods can take tens of seconds to minutes). The authors adapt PCM to the flow matching paradigm: the history-injected I2V model serves as the teacher; a structurally identical student is cloned from the same weights; the noise schedule is divided into 4 phases, and the student learns to predict each phase's solution point along the teacher's probability-flow ODE trajectory by minimizing a consistency loss between adjacent timesteps — reducing inference steps from N = 50 to M = 4.

Stage 4: Action learning. The authors freeze the distilled I2V model and predict continuous actions via an inverse dynamics paradigm: the generated sparse futures and the language instruction are injected through cross-attention into a DiT-based action head. One notably honest engineering detail: the authors observe a clear visual gap between generated frames and the raw ground-truth future frames, misaligning "synthetic dynamics" with the original action labels. To eliminate this inconsistency they relabel the generated frames with DA3 (Depth Anything 3), ensuring the action supervision is precisely aligned. Training uses DDIM, learning a denoising function that approximates the noise.

Data Construction: 140 Hours of Real-World Navigation Video

Rather than sidestepping the data problem, the paper treats it as a systematic contribution. The authors state plainly that they cannot reuse the LLM playbook of co-training on simulated navigation data plus real VQA data: simulation-only data causes mode collapse due to excessive domain gap, and existing real navigation datasets suffer from severe fisheye distortion and limited scale, making them unfit for fine-tuning VGMs.

The team therefore built its own data pipeline: human operators collected diverse videos with a handheld camera. To suppress hand shake — which would corrupt the consistent dynamics a VGM needs to learn — they explicitly used a DJI Osmo Action 4 with RockSteady+ stabilization. In total they collected 140 hours of real-world navigation video, processed via uniform temporal sampling into roughly 13,000 trajectories averaging 140 frames @ 4 FPS. Camera poses were estimated with Depth Anything 3 (DA3) to extract continuous action labels, and language instructions were hand-annotated by human experts. The authors call this the largest real-world VLN dataset to date and commit to open-sourcing it.

Experimental Setup: Six Scenes, 240 Trials, Success Defined by Distance

Evaluation is conducted in the real world across six unseen scenes spanning three categories: indoor (rooms, an academic building), outdoor (courtyards, parks), and nighttime (plazas, hilly terrain), probing zero-shot generalization. Each scene hosts four distinct navigation tasks — two standard IFN and two challenging BVN. For statistical reliability, every model is tested 10 times per task, 240 trials in total.

Three strong LLM baselines are compared: Uni-NaVid, a video LLM unifying multiple navigation tasks; StreamVLN, a streaming framework accelerating inference via KV-cache; and InternVLA-N1, the first dual-system LLM model for VLN. A particularly commendable evaluation detail: the authors note that LLM baselines often stop sideways next to the goal while their model typically stops facing it. To keep the comparison fair, they define success purely by distance — the robot succeeds if it stops within 1.5 meters of the goal. Furthermore, all comparative tests are run within the same time window to minimize environmental differences from lighting and weather.

Deployment hardware is a Unitree Go2 quadruped carrying an upward-facing DJI Osmo Action 4 for stabilized RGB observation; since InternVLA-N1 requires depth input, an Intel RealSense D455 (pitched down 15°) is added per its original configuration. The camera is mounted on the Go2's back via a custom 3D-printed bracket at about 1 meter height, kept uniform across all methods. Compute runs on a remote workstation with an RTX 4090: the Go2 continuously streams visual observations to the server, which returns navigation commands for the robot to execute.

Main Results: 2.5× SOTA Success on BVN, First-Ever Nighttime BVN

The results are unambiguous: SparseVideoNav achieves consistent state-of-the-art zero-shot performance across all real-world scenes on both IFN and BVN. Against the strongest baseline, StreamVLN, average success improves by +15.0% on IFN and +15.0% on BVN. More telling are the absolute numbers: 25.0% average success on BVN versus 10.0% — a 2.5× gap (the paper's headline number).

MethodIndoor IFNIndoor BVNOutdoor IFNOutdoor BVNNight IFNNight BVNAvg IFNAvg BVN
Uni-NaVid15.02.515.05.00.00.010.02.5
StreamVLN42.512.540.017.522.50.035.010.0
InternVLA-N120.02.532.522.50.00.017.58.3
SparseVideoNav55.027.557.530.037.517.550.025.0

The nighttime setting is the most persuasive evidence. The authors note that reduced visibility further amplifies the "myopia" defect, causing all existing baselines to fail systematically on BVN; powered by robust long-horizon guidance, SparseVideoNav is the only method able to reach distant targets in such extreme environments, achieving a 17.5% success rate and demonstrating nighttime BVN capability for the first time. Qualitative results (Figure 4) further show it traversing dead ends, narrow passable ramps, and steep slopes.

The paper also analyzes why it works (Figure 5): LLM baselines, constrained by short-horizon supervision, exhibit unexpected turns under long-distance uncertainty and get trapped prematurely at dead ends, whereas SparseVideoNav combines sparse foresight with closed-loop feedback to mitigate both.

The Efficiency Ledger: Where the 27× End-to-End Speedup Comes From

The 27× headline figure is not a single optimization but the compound effect of sparsification plus distillation. The teaser contrast: inference drops from 21.6 s to 0.8 s (27×), and Stage 1+2 training time from 357 hours to 49 hours (7.1×). The ablations break the ledger down:

Design axisComparisonResult
Sparse designContinuous vs interval-3 sparse generationInference 1.35 s → 0.79 s, a 1.7× speedup
Diffusion distillation50 steps vs 4 stepsInference 7.56 s → 0.79 s, roughly 10×, with 4-step quality comparable to 50
History compressionWith / without Former (history length N=45)Without Former, a +54.9% latency penalty; with it, latency is decoupled from history length and stays flat
Data scale8h / 50h / 140hFVD 2534 → 1755 → 1390, monotonically improving
Pretraining orderProgressive adaptation vs direct Stage-2 training32h vs 64h to converge, a 2× training speedup (32× H200)

The paper also runs three variants to validate the sparse design. Variant (a), "4-step distillation + 2 continuous chunks," deliberately mimics the short-horizon behavior of LLM baselines and scores only 15.8 / 2.5; variant (b) extends continuous chunks to 10, improving to 36.7 / 11.7 — longer horizons help but are not sufficient; variant (c), "no distillation (50 steps) + 20 continuous chunks," can be read as an accuracy oracle at 62.5 / 35.8 — better than SparseVideoNav's 50.0 / 25.0, but at 1.7× slower inference and 1.4× longer training convergence, which is exactly the paper's argued trade-off between efficiency and effectiveness. Variant (d), without the Former, drops to 45.0 / 22.5, degrading both performance and latency.

Two "Emergent" Behaviors and Honest Limitations

Section 4 of the paper discusses two illuminating phenomena. First, dynamic pedestrian avoidance. Because DA3 cannot produce reliable action estimates in scenes with oncoming pedestrians, the authors filtered such trajectories out during data construction; yet at deployment SparseVideoNav emergently avoids oncoming pedestrians — it successfully swerves around them and reaches the doorway. The authors take this as evidence of strong generalization. Second, camera-height insensitivity. LLM-based navigation is quite sensitive to camera-height changes, while the video-generation paradigm stays robust: although training data was captured at roughly 1 meter height, SparseVideoNav still navigates with the camera fixed at 50 centimeters.

The limitations section is equally candid. The authors list two points: first, their 140-hour dataset is still modest compared to web-scale data, and scaling it is identified as the key improvement direction; second, despite extensive optimization for real deployment, inference remains slightly slower than existing LLM-based navigation paradigms, and VGM-specific acceleration distillation and quantization are flagged as future work. The paper likewise concedes that a single clean detection pass is evidence, not proof — true visual quality still requires manual inspection across viewports.

Why This Paper Matters: Three Paradigm-Level Reminders

First, swap the pretrained prior, not just the model size. The most elegant move here is not scaling up but identifying a mismatch that has gone unnoticed: BVN demands long-horizon visual foresight, LLM pretraining contains no such prior, and VGM pretraining has it natively. For the same task, changing the modality that carries foresight beats stretching the supervision horizon within the same modality — which destabilizes training anyway.

Second, sparsification is not just a compute saver; it is a capability source. Sparsity is usually framed as an efficiency trick; here it is simultaneously the means to extend the prediction horizon. With a fixed prediction-frame budget, the frame interval is the knob that trades horizon length against resolution, and interval 3 is the optimal compromise between "seeing far enough" and "seeing clearly enough." This observation transfers readily to any system that supports long-horizon decisions with a fixed-length future prediction.

Third, engineering details decide success — and deserve to be in the paper. Zero-initialized cross-attention preserves the generative prior; relabeling generated frames with DA3 resolves the misalignment between synthetic dynamics and original action labels; the Former decouples latency from history length; and even the evaluation bias of "baselines stop sideways while we stop facing the goal" is explicitly discussed and neutralized by distance-based success. These are the costs of taking a beautiful idea all the way to a robot that actually runs.

Placed in a broader context, the paper converses with two recent threads. One is "world models for decision-making": from the Dreamer line to Genie, learning a future-predicting representation and distilling actions from it has been a long-held dream; SparseVideoNav's contribution is showing that generic video generation priors are already strong enough — no world model needs to be learned from scratch, only "aimed" at the robot's viewpoint and decision needs with modest real navigation data. The other is "video pretraining transfers to robotics": much recent work distills manipulation skills from internet video, and this paper shows the same transfer logic holds on navigation — arguably more naturally, since a navigation future is literally the evolution of an egocentric video. For engineering teams there is also a pragmatic signal: every acceleration technique in the paper (sparse supervision, flow-matching consistency distillation, feature-compression latency decoupling) is backbone-agnostic and applies directly to other embodied systems built on video generation.

SparseVideoNav is by the University of Hong Kong and OpenDriveLab, with code and data pledged to github.com/OpenDriveLab/SparseVideoNav. The paper appears as arXiv:2602.05827 (cs.CV, submitted February 5, 2026), by Hai Zhang, Siqi Liang, Li Chen, Yuxian Li, Yukuan Xu, Yichao Zhong, Fu Zhang, and Hongyang Li (the first two contributed equally), supported by the National Natural Science Foundation of China (62206172) and the JC STEM Lab of Autonomous Intelligent Systems funded by the Hong Kong Jockey Club Charities Trust.

Related Papers

SparkDiffusion: Mitigating the High-Sparsity Trap — A Unified Framework for up to 265× Single-GPU Acceleration of Visual Generation

SparkDiffusion: Mitigating the High-Sparsity Trap — A Unified Framework for up to 265× Single-GPU Acceleration of Visual Generation

SparkDiffusion (arXiv:2609.23153) from Peking University, Tsinghua and Alibaba diagnoses the high-sparsity trap: at 97% attention sparsity, step-local training loss keeps falling while terminal video quality stagnates or degrades — the root cause is supervision, since flow-matching-style step-local losses cannot constrain terminal errors (an oracle probe shows fixing the five highest-noise steps removes most terminal error). The three-stage recipe: a short sparse warm-up (compensated sparse attention, RoLA), few-step trajectory-mixed distillation (CrossDistill: high-noise PCM consistency + low-noise DMD distribution matching, 3-step CFG-free student), and FP8 W8A8 quantization with fused kernels. It sustains 97% sparsity on long 720P generation, achieving a measured 265x end-to-end speedup on Wan2.1-T2V-14B-720P on a single RTX 5090 (220x on H100) and 1.3 s for 1.3B-480P videos, with seed-level diversity closest to the dense reference (VBench 83.15 vs 83.69). Code open-sourced.

Yuxi Liu, Haoyu Li, Zekun ZhangSep 19, 2026
Video GenerationDiffusion modelsDistillationSep 19, 2026
CAP: Continuously Adaptive Perception-Blind Humanoid Locomotion via Learned Denoising

CAP: Continuously Adaptive Perception-Blind Humanoid Locomotion via Learned Denoising

Humanoid locomotion across complex terrain demands forward-looking exteroception to anticipate obstacles, yet this signal is unreliable in real-world deployment, failing partially and intermittently. Existing perceptive policies often assume that depth observations remain clean and in-distribution, while recent attempts to unify perceptive and blind control typically route or switch between separate sub-policies, leaving recoverable information in partially corrupted depth unexploited. We instead propose CAP, a single-stage humanoid locomotion policy that recovers this signal with a perceptive world-model encoder trained as a learned denoiser to reconstruct clean depth from a corrupted input, together with a co-active proprioceptive variational encoder that supplies depth-free body-state information. A coupled training recipe pairs a depth-noise curriculum on the world-model input with world-model feature dropout on the policy-facing latent, exposing the policy to failures across the entire perception-quality spectrum. In simulation, CAP matches or improves upon perceptive baselines when depth remains informative, and degrades more smoothly than a binary-switching baseline as perception worsens. On the Unitree G1, controlled trials and indoor-outdoor deployments demonstrate perception-robust locomotion under intermittent occlusion, real-sensor corruption, and outdoor depth artifacts.

Chen, Hongjin, Xu, Zijun, Ma, ShihaoSep 10, 2026
Humanoid LocomotionWorld ModelsDepth DenoisingSep 10, 2026
JEPA-Anything: Learning Predictive Models across Different Worlds

JEPA-Anything: Learning Predictive Models across Different Worlds

JEPA-Anything introduces orthogonal predictive factorization for domain-agnostic world modeling, splitting latent targets into complementary factors with dedicated predictive pathways. Across seven domains, it improves matched dynamics tasks and supports intervention, OOD generalization, and long-horizon forecasting.

Cui, Taoyong, Wang, Zhongyao, Xu, XinyueSep 17, 2026
World ModelsJEPAPredictive Representation LearningSep 17, 2026
Dream-RSI: Recursive Self-Improvement through Evolving Worlds

Dream-RSI: Recursive Self-Improvement through Evolving Worlds

Recursive self-improvement is becoming increasingly vital for autonomous AI agents, where progress hinges on discovering high-value solutions across complex domains. The driver of this process is effective exploration, however, managing and improving exploration strategies remains a major bottleneck. Current systems face a fundamental dilemma: fixed strategies fail to adapt as search spaces scale, while online policy optimization requires navigating vast meta-search spaces under delayed and expensive feedback over long-horizon rollouts. We introduce {Dream-RSI}, a framework for scalable and recursively self-improving exploration. A lightweight orchestration layer makes exploration explicit and programmable while leaving the underlying coding agent unchanged. Our key insight is that accumulated discovery history can serve as a replay simulator over the realized search space. By performing dreaming in the replay simulator constructed from historical discovery trees, {Dream-RSI} secures immediate, low-cost off-policy feedback to evaluate and refine exploration policies without invoking repetitive, expensive online evaluations. The improved policy is subsequently redeployed online to drive further discovery, continuously expanding the simulator pool in a self-improving loop. Across algorithm engineering, mathematical optimization, and GPU kernel engineering, {Dream-RSI} achieves competitive or improved discovery quality while substantially reducing discovery cost in several settings.

Tong Zheng, Xidong Wu, Zheng ZhangSep 14, 2026
Recursive self-improvementWorld ModelsLLM AgentSep 14, 2026

As an Amazon Associate, we earn from qualifying purchases.