PAPER DEEP DIVE
NavVerse: a physics-enabled benchmark for indoor-to-outdoor embodied navigation — best zero-shot VLA reaches 11.6% and halves across the boundary
UMich-CURLY (CoRL 2026) introduces NavVerse: the first physics-enabled benchmark for continuous indoor-to-outdoor embodied navigation. It spans 200 scenes (100 indoor + 50 outdoor urban + 50 connected indoor-to-outdoor), 10,000 episodes and three tasks (ObjNav, VLN, and the new place-level PlaceNav). Built on Isaac Sim for both photorealistic visuals and executable robot dynamics; connected scenes use door-to-facade assembly so the robot walks from corridor to street in one physical world. Evaluation covers three orthogonal axes: success (SR/SPL), efficiency (CE coverage efficiency) and safety (CR collision rate, ADO distance to obstacles, NSR navigable surface ratio). All four zero-shot baselines fall short: the strongest VLA (UniNaVid) reaches 11.62% ObjNav, 11.38% PlaceNav, 10.67% VLN; the modular method is safest yet barely explores; the VLA-RL policy keeps the largest clearance but most often stops at wrong goals. The sharpest finding is the transition gap: for UniNaVid, PlaceNav success drops from 17.65% outdoors to 3.64% on connected episodes (a 14.01-point absolute fall), with 25-48% of episodes never even reaching the outdoors, and every method's coverage efficiency declines post-exit (manifesting as in-place spinning). Diagnostics also show that on the same oracle trajectories the legged Spot completes 100% under both friction settings while a wheeled Turtlebot manages only 30.75-47.50% — benchmarks ignoring embodiment dynamics systematically overestimate executability. Prior work evaluates indoor and outdoor separately and abstracts away execution; NavVerse supplies the three long-ignored stages: exit finding, boundary crossing, and post-exit re-adaptation.
One paper, one plain question: can today's robots step out the door and actually get the job done?
Imagine telling a robot: "Find a place to get burgers and fries." For a human this is trivial — walk out, cross the street, recognize which storefront is a restaurant, get there. For today's navigation agents, every link in that chain is a pit: leave the building, cross into an outdoor world with completely different lighting and structure, understand a place-level semantic goal like "restaurant," and physically walk there under kinodynamic constraints. NavVerse, published by the UMich-CURLY lab at CoRL 2026, is the first benchmark to put this entire chain under systematic evaluation in physics-based simulation.
(corridors/shelves)"] --> B["Exit finding"] B --> C["Boundary crossing
doorway → street"] C --> D["Outdoor adaptation
scale/lighting/terrain shift"] D --> E["Place-level search
PlaceNav: find restaurant/bank"] B -.-> F["Not Reached Outdoor
~25-48% stuck here"] C -.-> G["Post-exit efficiency drop
CE 1.07 to 0.84"]
Benchmark composition: 200 scenes, 10,000 episodes, three tasks
NavVerse is built on NVIDIA Isaac Sim, providing both photorealistic rendering and physics-aware robot execution — the fundamental difference from the Habitat family (vision-heavy, physics-light) and the MuJoCo/Gazebo family (physics-heavy, visually sparse). Scenes come in three types: 100 indoor scenes (from GRUtopia's GRScenes: apartments, supermarkets, hospitals, offices); 50 outdoor urban scenes (from Virtual Community's Google 3D Tiles, each roughly 1km×1km with on average 546 buildings and 36.4km of roads); and 50 indoor-to-outdoor connected scenes — assembled via a "door-to-facade" procedure that embeds an indoor layout into a road-facing building: cut an opening in the facade, place a compatible indoor layout behind it, align entrance height, remove blocking meshes, so the robot walks seamlessly from corridor to street in one physical world with no teleportation and no viewpoint switching.
Three tasks: ObjNav (locate and approach an object category), VLN (follow natural-language instructions), and the newly proposed PlaceNav — the goal is a functional place such as a restaurant, cafe or bank rather than a concrete object, requiring agents to reason over road topology, building frontages and storefront semantics in long-horizon search. 10,000 episodes total: 4,027 ObjNav, 2,973 PlaceNav, 3,000 VLN. Episode generation is a three-step loop: offline NavMesh sampling of start-to-goal paths → online filtering via simulated robot execution → VLM-generated VLN instructions with human web-based quality control. To support PlaceNav, the authors generated 101 storefront textures (restaurants/cafes/banks) with Gemini 3 Pro Image, projected onto building facades via ray-casting at OSM store locations, and populated scenes with objects from 121 Objaverse categories — up to 200 cars and 1,000 sidewalk objects per outdoor scene, plus 30 "contextual object groups" (bus stops get shelters, briefcases and umbrellas; benches get notebooks and coffee) for semantic realism.
Evaluation protocol: turning "safety" into a score
All baselines run through a unified executable robot interface: the default embodiment is a simulated Boston Dynamics Spot (legged, RGB-D input); model outputs are converted to waypoints and executed by a shared PID waypoint follower on top of an Isaac Lab locomotion policy trained on flat ground, rough ground and stairs. Metrics cover three orthogonal axes: task completion (SR, SPL); exploration efficiency (CE — newly covered space per unit travel distance, especially critical in large outdoor scenes); and safety (CR collision rate, ADO average distance to obstacles, NSR navigable surface ratio — whether the robot stays on terrain-annotated valid surfaces). Episodes terminate on success, timeout, fall, or stopping at a wrong goal. This is an evaluation that prices "physical viability" into the score: the classic trick of sliding along obstacles on indoor benchmarks gets billed here through collision rate and falls.
Results: the best zero-shot VLA reaches only 17% — and collapses the moment it goes outside
Four baselines — modular semantic navigation SGImagineNav, indoor RL policy PoliFormer, VLA navigation model UniNaVid, and VLA-RL policy LongNav-R1 — all evaluated zero-shot. The conclusion is clear and brutal: UniNaVid achieves the best zero-shot task completion (11.62% ObjNav, 11.38% PlaceNav, 10.67% VLN) but the absolute numbers remain low; the modular method is the safest (lowest CR at 0.05, highest NSR); LongNav-R1 keeps the largest obstacle clearance (highest ADO) yet most often calls stop at wrong goals (87.1% of its PlaceNav failures). Each dimension has a different "champion" — no method balances success, safety and efficiency simultaneously, which is precisely the non-negotiable baseline for real deployment.
The most informative finding is the transition gap. Splitting the strongest baseline UniNaVid's results by scene type: going from pure outdoor to indoor-to-outdoor connected scenes, ObjNav drops 17.33% → 8.42% (−8.91), VLN 14.00% → 7.33% (−6.67), and PlaceNav plunges from 17.65% to 3.64% — an absolute drop of 14.01 points, nearly 80% in relative terms. Two reasons: VLN leans on instruction following, which transfers across domains relatively stably; PlaceNav requires topology-aware place search, and indoor and outdoor topologies follow fundamentally different logics, making the adaptation bottleneck steepest.
Failure-mode analysis pinpoints the problem further. Exit finding is a hidden level: invisible in pure indoor or pure outdoor evaluation — in indoor-to-outdoor episodes, UniNaVid and LongNav-R1 fail to even reach the outdoors in 48.42% and 47.37% of ObjNav episodes respectively; even SGImagineNav, with the best reach-outdoor rate, gets stuck there 25.26% (ObjNav) / 30.91% (PlaceNav) of the time. Systematic post-exit efficiency drop: every method's coverage efficiency is lower after the exit than before (UniNaVid: 1.07 → 0.84 on ObjNav, 1.26 → 0.81 on PlaceNav, 1.14 → 0.80 on VLN) — the authors observe agents spinning in place and pacing back and forth after exiting, unable to adjust exploration strategy to changes in scene scale and topology. Indoor and outdoor stress different failure modes: indoors, local execution dominates (falls cause 31.3% of PoliFormer and 43.0% of SGImagineNav indoor ObjNav failures); outdoors, exploration and grounding take over (95.1% of PoliFormer's outdoor ObjNav failures are timeouts; 95.0% of LongNav-R1's are wrong goals).
Diagnostics: longer routes compound errors, and wheels refuse to cooperate
Three diagnostic experiments stand out. Difficulty splits: grouping episodes into path-length tertiles (Easy/Medium/Hard), the Hard split consistently drags down success across all tasks — long-horizon navigation is limited not just by goal recognition but by compounding errors in exploration, memory, terrain selection, instruction grounding and closed-loop control. Embodiment dynamics: on the same oracle reference trajectories, the legged Spot completes them 100% under both friction settings, while the wheeled Turtlebot manages only 30.75%–47.50% at 0.36–0.58 m/s — a navigation benchmark that ignores embodiment dynamics systematically overestimates kinodynamic executability, which is exactly what physics simulation is irreplaceable for. Instruction format: on PlaceNav, switching from store-category prompts to route-style VLN instructions raises SPL from 3.30 to 4.84 (mainly outdoors), but indoor-to-outdoor SR stays frozen at 3.64% — route descriptions do not solve exit finding; adding an intent-to-place grounding step (Intention Driven) drops SR further to 4.07%, with indoor-to-outdoor at zero.
Where it sits in the evaluation landscape
In the related-work coordinate system: Talk The Walk and TouchDown do outdoor VLN over street-view panoramas, but non-interactive and bodiless; R2R/RxR/HM3D-ObjNav dig deep indoors without crossing the threshold; VLN-PE (2025) moves to continuous control in Isaac Sim but stays indoors. NavVerse is the first benchmark to fit indoor + outdoor + connected transition scenes, three tasks, an executable robot interface and safety metrics into one physics-based simulation platform. Its two honestly stated limitations are worth recording: most environments are currently single-floor meshes (multi-floor layouts and skybridges are future work); dynamics are limited (pedestrians, vehicles, and interactive doors are on the roadmap), and the reverse outdoor-to-indoor transition is unexplored.
How the scenes are built: from OSM road networks to "feeding the cat"
The scene pipeline's engineering details deserve their own section, because they define exactly what makes cross-context navigation hard. On the outdoor side, NavVerse starts from 50 Virtual Community cities whose original meshes carry only sparse urban geometry, then adds four layers of enhancement. Real-to-sim terrain: road centerlines and widths are extracted from OSM (tagged widths when available, otherwise derived from lane count and road type), converted into footprint polygons imprinted onto the terrain mesh with a uniform depression inside road regions — so roads really do sit a bit lower than sidewalks and wheels really do feel the curb, turning "trajectory feasibility" from a semantics question into a physics question. Surface semantics and materials: terrain is segmented into five classes, each with distinct appearance and physics materials, so policies experience different friction and contact behavior on sidewalks versus grass. Contextual object placement: no random scattering — objects are sampled conditioned on each map region's semantics (bike racks get bicycles; bus stops get shelters, briefcases and umbrellas; there is even a feeding_cat group of cat, rice bowl and nursing bottle), with placement constrained away from existing groups to avoid clutter, and objects stackable by size and support surface. The storefront pipeline: Gemini 3 Pro Image generates 101 storefront textures with relative depth maps; each is manually inspected, depth-extruded into a 3D mesh in Blender, ray-cast onto the nearest building facade, oriented toward the closest road, scaled, terrain-aligned, and cut into the facade — accepted only if the cut succeeds, the storefront faces the road and stays unoccluded. The result: PlaceNav goals are not textures but physical, correctly-placed, approachable places.
Indoor scenes get equal care: raw GRScenes meshes carry 21.77 million faces; an automated Blender pipeline decimates, dissolves and remeshes them down 89.91% while keeping textures — cutting NavMesh build time from 705ms to 107ms. NavMesh generation uses C++ RecastNavigation with finer voxels indoors (0.05m cells) to capture narrow passages and coarser ones outdoors (0.10m) for efficiency; max climb and max slope (45°) parameters directly encode "physically walkable" constraints, so reference paths reflect true traversability. Online episode verification requires the oracle rollout to succeed under full physics within duration and distance bounds without triggering any failure termination — failing episodes are discarded.
Three action interfaces and an honest runtime bill
NavVerse does not force agents to accommodate the evaluator: it offers three action interfaces — world-frame waypoint paths (most general), discrete primitives (forward/left/right with different step sizes indoors vs outdoors, two rotation magnitudes), and body-frame velocity commands that drive the locomotion policy directly, bypassing the PID controller. All baselines share one low-level stack running at 60Hz: lookahead path tracking, PID velocity computation with clipping (linear, lateral, angular limits), and a learned locomotion policy identical in structure to Isaac Lab's default. This cleanly separates "good policy, poor execution" from "good execution, poor policy" — LongNav-R1's frequent wrong-goal stops are a policy problem, while PoliFormer's post-exit timeouts reflect an exploration strategy mismatched to scene scale. The appendix also provides a rarely-seen honest runtime bill: Isaac Sim 5.1.0 + Isaac Lab 2.3.0, 0.005s physics steps (200Hz), ~20 FPS rendering on an RTX 4090, up to three parallel simulation instances per L40S (48GB), peak VRAM of ~9GB indoors and ~13GB outdoors, with large-scale evaluation on four L40S GPUs. For teams planning to reproduce or train on this benchmark, that is the entry ticket in plain numbers.
Why "transition" is harder than indoor or outdoor: a structural view
lighting/texture/scale shift"] Q --> P2["Spatial re-anchoring
exit finding + boundary crossing"] Q --> P3["Policy re-adaptation
exploration switches with topology"] P1 --> R["The old benchmarks' blind spot:
indoor and outdoor scored separately"] P2 --> R P3 --> R R --> S["NavVerse's answer:
continuous episodes in one physical world +
success/efficiency/safety axes"]
The failure data fits this frame neatly: P2 is the steepest wall (25–48% of episodes never reach the outdoors), P3 the most insidious pit (every method's CE drops post-exit, manifesting as in-place spinning — inefficiency that "looks like work"), and P1 shows up mostly in PlaceNav's topology-level semantic search — finding a reception desk inside a hospital and finding a bank on the street draw on entirely different spatial priors. VLN's instruction-following transfers most stably precisely because explicit linguistic structure offloads part of the P1/P3 burden. The path forward is explicit: rather than piling more single-context data, model the transition itself — exit detection, boundary states, and cross-context transfer of exploration strategies.
Three statistics plots that read like the exam syllabus
The paper's three appendix treemaps are more revealing than the main tables because they expose the benchmark's "syllabus bias." Object categories: shopping trolley (4.0%), couch (3.8%), table (3.3%), shelf (2.7%) lead, with garbage trucks, ambulances, police vans, seagulls and puddles also on the board and the longest tail at 0.5% each — ObjNav's goal distribution deliberately favors everyday semantics over rare objects, testing whether agents "understand daily life" rather than memorize long tails. POI distribution (PlaceNav's syllabus): restaurants dominate at 20.6%, convenience stores 16.1%, fast food 9.9%, cafes and toilets 7.6% each — food and daily retail make up roughly 60%, banks/ATMs/bureaux de change ~12.5%, while cultural venues (schools, universities, libraries, theatres, community centres) total under 10%. PlaceNav mostly tests the highest-frequency urban errands, matching the tweet's "burgers and fries" framing perfectly. Instruction verbs (VLN's language syllabus): continue (9.7%), move and keep (8.9% each), arrive (8.6%), pass (7.8%) lead, with the top 19 verbs covering ~83% of occurrences — a strongly route-style instruction register, consistent with Table 8's finding that route instructions beat category prompts on SPL. For teams chasing the leaderboard, these plots publish the examiner's preferences; for critics, they mark the ready-made weak spots: cultural venues and rare objects are underrepresented, and the winning strategy may simply be solid semantic grounding of food and retail.
What to take away from this paper
For researchers, NavVerse delivers a clearly partitioned case file: is your method like UniNaVid — winning on completion but losing on safety (higher CR), or like SGImagineNav — stable but barely able to explore (single-digit SR), or like LongNav-R1 — keeping its distance yet stopping rashly at wrong goals? Single-axis scores can no longer masquerade as overall capability. For engineers, the lesson is more direct: if your robot works between buildings and streets, exit finding and post-exit re-adaptation are two modules that must be tested separately — indoor-only tuning will not save them. For the community, the paper's most important contribution may be a "negative-results infrastructure": it measures the VLA, RL and modular lines against the same physical ruler, and the gap it reveals (best-in-class 17%, halved again across the transition) is enough to feed years of method research. And the plain wish of "stepping out the door and getting something done" finally has an honest scale.
References
- Paper: NavVerse: Benchmarking Indoor-to-Outdoor Embodied Navigation in Continuous Robot Simulation — arXiv:2607.19695 (CoRL 2026)
- Project page: umich-curly.github.io/NavVerse-Benchmark
- Code: github.com/UMich-CURLY/NavVerse-Benchmark
- Video: @GhaffariMaani demo (CoRL 2026)
SOURCE LINKS
