PAPER DEEP DIVE
UniLM-Nav: A Unified Framework for Zero-Shot Last-Mile Navigation
Mobile manipulation requires a robot to navigate to a target object or receptacle and then perform intended manipulation. However, reaching the vicinity of the target does not guarantee a manipulation-ready base pose, a problem known as last-mile navigation. Prior methods for last-mile navigation either rely on manual pose annotation or task-specific training, limiting their scalability to open-vocabulary settings with fine-grained spatial constraints. We propose UniLM-Nav, a unified framework for zero-shot open-vocabulary last-mile navigation. UniLM-Nav decomposes last-mile navigation into view selection, task-conditioned affordance grounding, and geometry-aware base-pose reasoning, all resolved with a shared multimodal large language model (MLLM) backend. Specifically, UniLM-Nav first selects a reference view that best captures the target object or receptacle from recently collected observations. It then grounds task-relevant affordance point in the selected view and lifts the result into the robot-centric coordinate frame. Finally, conditioned on the grounded affordance, task context, and robot geometry, it infers a manipulation-ready base pose for the robot. We evaluate UniLM-Nav on the OVMM benchmark, where it outperforms the previous state-of-the-art method, MoTo, by 3.13 percentage points. Analyses show that the components of our method are crucial to final performance, and that the choice of MLLM also has a substantial effect. We further deploy UniLM-Nav on a Unitree B2 quadruped robot with a 6-DoF Unitree Z1 manipulator, validating its applicability to real-world mobile manipulation tasks.
Paper Metadata
Title: UniLM-Nav: A Unified Framework for Zero-Shot Last-Mile Navigation
Authors and affiliations: Zhuofan Zhang, Tianxu Wang, Guoxi Zhang, Yixiong Lin, Xilin Wang, Hongming Xu, Qing Li, Song-Chun Zhu, Lifeng Fan (Tsinghua University, Beijing Institute for General Artificial Intelligence (BIGAI), Harbin Institute of Technology, Peking University).
Links: arXiv:2607.06537 (v2, 2026-07-11) · project page unilm-nav.github.io
Code status: at the time of this close reading (2026-10-08) the Code button on the project page is still a placeholder (aria-disabled), so the implementation is pending release. GitHub hosts only the project-page repository; no third-party reproduction exists yet.
In One Sentence
UniLM-Nav takes the most neglected stretch of mobile manipulation, the segment between "the robot can see the target" and "the robot can act on the target", and decomposes it into three MLLM sub-problems (choose a reference view, ground a task-conditioned interaction point, reason about where the base should stand), while keeping every metric quantity outside the model through depth back-projection and a closed-form heading. The result is a fully zero-shot policy that reaches 23.77% Overall SR on the HomeRobot OVMM validation split, 3.13 points above the previous strongest method MoTo, and that closes the loop on real hardware: a Unitree B2 quadruped carrying a 6-DoF Unitree Z1 arm.
Background and Motivation
Mobile manipulation is a relay race. The robot first navigates toward an object or a container, then performs a pick or a place. The handover between those two legs decides whether the whole chain succeeds, and it is exactly where current systems are weakest. An object-navigation policy typically guarantees only that the robot ends up somewhere in a one-to-two meter neighborhood of the target. That granularity is enough for perception, and nowhere near enough for manipulation: the robot may stop at an angle from which the arm cannot reach the table, or it may stop facing a chair that blocks any approach path.
The paper makes this concrete with its teaser scene (Figure 1). The instruction is to place a water bottle on the desk in front of a monitor. After object navigation terminates, the robot is indeed near the desk, but its final pose is either too far from the tabletop for the arm to span the gap, or wedged between a chair and the desk so that a placement motion is physically impossible. The authors name this segment last-mile navigation: everything that happens after object navigation ends and before the base reaches a manipulable pose. Solving it requires answering two questions at once. Where in the scene is the interaction region relevant to this task, and given that region, where should the base stand and which way should it face?
Figure 1: the real-robot teaser. The task is to put a water bottle on the desk in front of the monitor. Pose 1 is where object navigation terminates, too far from the tabletop for the arm; pose 2 has its approach path blocked by a chair; pose 3 is the manipulation-ready pose reached after last-mile navigation. The three panels on the right are the first-person observations at those poses.
Existing routes each carry a hard limitation. One family relies on manually annotated, task-specific base poses; change the task or the environment and the annotations must be redone. A second family handles the handover implicitly with imitation learning or reinforcement learning, or learns an explicit "manipulation-conditioned pose preference", but both are data hungry and hard to extend to open-vocabulary instructions. A third, zero-shot family (MoTo being the representative) leans on visual foundation models and MLLMs for object-level reasoning, and object-level cues simply cannot answer fine-grained spatial relations: "in front of the monitor" is not a request to localize the monitor, it is a request to find a container region that satisfies a spatial constraint.
Meanwhile two capabilities have matured in parallel. Visual foundation models turned open-vocabulary perception into a plug-and-play component, and MLLMs keep improving on spatial reasoning benchmarks. Wiring both into the handover of a mobile manipulation stack, so that an off-the-shelf MLLM reasons out a manipulable base pose with zero task-specific training, is the natural next step and is not yet solved well. The difficulty is that the three sub-decisions involved (which frame to look at, which pixel to point at, which position to stand at) are highly heterogeneous; thrown into a single call, current models tend to trade one off against another.
UniLM-Nav answers with a unified framework and an explicit decomposition. Three stages share one MLLM backend, and each stage asks the model for exactly one thing that matches its competence: a choice among labeled candidates, a pointing answer in image space, or a position in a local metric frame. Everything the model is bad at, metric distance from a single image, heading arithmetic, reachability checks, is handed to deterministic code. The design keeps zero-shot and open-vocabulary flexibility while removing metric estimation from the reasoning loop.
Problem Setting: Inputs and Outputs
The paper works in the open-vocabulary mobile manipulation (OVMM) setting: the robot receives one natural-language instruction and must find an object, pick it, find a receptacle, and place it, all inside open-vocabulary indoor scenes. Evaluation follows the HomeRobot OVMM benchmark and reports the conditional stage success rates FindObj, Pick and FindRec, plus Overall SR for the full chain and Average SR across stages.
The system assumes the object-navigation policy and a pre-built scene map have already brought the robot into the target neighborhood, meaning the target appears in recent observations at roughly one to two meters. What the last-mile stage may consume is standard equipment for any mobile manipulation stack and requires no extra annotation: egocentric RGB-D observations, proprioceptive state, odometry, the task instruction, and an obstacle map estimated from onboard sensors.
The output is a manipulable base pose $\mathbf{b}=(x,y,\theta)\in\mathbb{R}^{3}$, where $(x,y)$ is the target base position and $\theta$ is the desired heading, subject to being collision-free, reachable, and suitable for the manipulation that follows. The central question of the paper is how to reason that pose out of an off-the-shelf MLLM with zero shot examples.
Method
UniLM-Nav splits last-mile navigation into three stages: view selection, task-conditioned affordance grounding, and geometry-aware base-pose reasoning. All three query the same MLLM backend (Figure 1). The description below follows the data flow.
Figure 2: UniLM-Nav method overview. The MLLM first selects a reference view from short-term memory, then grounds the task-relevant affordance inside that view and lifts it to 3D with depth, and finally predicts a manipulable base pose conditioned on the 3D affordance, the robot configuration and the task instruction.
Stage 1: short-term observation memory and view selection. The frame captured at the termination step $t$ of object navigation is an accidental viewpoint at the end of a trajectory. The target surface may be occluded, and the surrounding layout may not offer enough spatial context. UniLM-Nav therefore keeps a short-term memory buffer during the last steps of object navigation, storing observations together with the corresponding robot states:
$$\mathcal{M}_{t}=\{(o_{t-k},s_{t-k})\}_{k=0}^{K-1}$$
Each record holds the egocentric observation $o_{t-k}$ and the state $s_{t-k}$ that gives the robot and camera pose, with $K=5$ by default. So that the model can name a candidate unambiguously, the system overlays an ID from 0 to 4 in the lower-right corner of each frame and asks the MLLM to return exactly one reference view $o_{\mathrm{sel}}$ under two criteria: whether the target object or container surface is clearly visible, and whether a feasible approach path exists that lets the robot get close enough to manipulate. The answer comes back as a single ID in JSON. The qualitative case in the appendix (Figure 3) shows why this matters: the approach path in the termination frame is blocked by a chair, while the frame from $t-4$ exposes free floor space at the lower left of the desk and is selected as the reference view.
Figure 3: view selection in practice. The short-term memory $\mathcal{M}_{t}$ holds the five most recent frames. The red box is the object-navigation termination frame, whose approach path is blocked by a chair; the green box is the $t-4$ frame selected by the MLLM, which gives a far more reliable reference for placement.
Stage 2: task-conditioned affordance grounding. Once the reference view is fixed, the MLLM is asked to ground the instruction into a single interaction point in image space, not into an object box and not into an object centroid:
$$(u,v)=\mathrm{MLLM}(o_{\mathrm{sel}},\mathcal{T})$$
For a pick task this point should fall on a graspable region of the target object. For a place task it should be a safe, free spot on the container surface away from the edges. The prompt states three hard constraints for placement points: the surface must be flat and stable, the point must avoid the container rim so the object does not fall off, and there must be clear space around it so that base and arm can both reach. To decouple the output from image resolution, the model predicts normalized coordinates strictly inside the open unit interval, $(u,v)\in(0,1)^{2}$, which are then mapped back to pixels using the image width and height; for backends that ship their own grounding convention, such as Qwen3-VL and RoboBrain-2.5, the native 0-to-1000 coordinate space is used and rescaled instead. This step forces the model to reason jointly about task semantics, object semantics, local spatial layout and manipulation constraints, and it is what makes the pipeline task-aware rather than merely object-aware.
From 2D to 3D: depth back-projection. Given $(u,v)$, the system lifts it into a 3D affordance point $\mathbf{p}_{a}$ expressed in the frame of the selected robot pose, using the aligned depth image and the camera intrinsics. This is the standard pinhole back-projection $\mathbf{p}_{a}=D(u,v)\,\mathbf{K}^{-1}(u,v,1)^{\top}$, where $D(u,v)$ is the depth of that pixel and $\mathbf{K}$ is the intrinsic matrix. The point is also drawn onto the selected observation as a red dot, producing the visually prompted observation $\tilde{o}_{\mathrm{sel}}$. This is the hinge of the whole design: it converts visual evidence into an explicit metric quantity so that the next stage no longer has to guess distances and reachability implicitly from a single 2D image.
Stage 3: geometry-aware base-pose reasoning. The most direct alternative is to ask the MLLM to point at a floor pixel and call that the base position. That would require the model to infer metric distance, arm-reach feasibility and local traversability from one 2D image, which is precisely the current weakness of MLLMs. UniLM-Nav instead feeds explicit geometry into the prompt:
$$(x,y)=\mathrm{MLLM}(\tilde{o}_{\mathrm{sel}},\mathbf{p}_{a},\mathcal{R},\mathcal{T})$$
The inputs are the red-dot observation, the robot configuration $\mathcal{R}$, the task instruction and the 3D target coordinates; the output is the target base position $(x,y)$ in the local frame whose origin is the robot pose of the selected observation. The prompt also encodes the engineering constraints: the target should end up between 70% and 80% of the maximum arm reach, the robot should face the target, and both the pose and the arm setting must stay clear of obstacles. Besides the position, the model predicts an arm extension arm_reach and a lift height arm_lift, and those two scalars form the simple manipulation policy of UniLM-Nav. For picks, the prompt gives the arithmetic for the lift height, target height plus half the object height plus a 0.2 m gripper margin:
$$z_{\mathrm{lift}}=z_{\mathrm{target}}+\tfrac{1}{2}h_{\mathrm{obj}}+0.2\ \mathrm{m}$$
The heading is never asked of the model. Instead of letting the MLLM predict the heading $\theta$, the paper computes it geometrically by pointing the base at the lifted affordance, in the local frame of the selected pose, $\theta=\mathrm{atan2}(y_{a},x_{a})$. The ablation in Table 3 shows how much weight that decision carries: on the 20% subset the geometric heading yields 25.42% Overall SR, while letting the MLLM predict the heading directly gives only 17.08%. With FindRec roughly comparable, the overall gap exceeds eight points. Heading is a degree of freedom that geometry can lock down, so there is no reason to spend model reasoning budget on it.
Execution loop. The local base pose is transformed into the global frame by the robot pose of the selected observation and handed to a low-level navigation policy that drives there while avoiding obstacles; on the real robot this is the ROS 2 Navigation2 stack. Real-robot deployment adds one refinement step. After arriving at the predicted pose, odometry residuals, depth noise and calibration error can shift the projected point in the final manipulation view, so the system queries the MLLM once more with the current wrist-camera image and revises the manipulation target, whether a grasp point or a placement point.
Why the decomposition has to be explicit. The merging ablation in the appendix answers this. Fusing view selection and affordance grounding into a single call (five images in, one ID plus one coordinate pair out) costs 3.95, 2.91 and 3.33 points of Overall SR on the Gemini-3-Flash-Preview, Qwen3-VL-8B and Qwen3-VL-32B backends respectively. A human can do both in a single glance, but today's MLLMs cannot select a frame and pinpoint a pixel well inside one inference. Dropping base-pose reasoning and pointing directly at the floor degrades results just as clearly, and the visualization in the paper shows predicted positions hugging obstacles or off to the side of the target rather than in front of it. The decomposition is not engineering fastidiousness; it is drawing responsibility boundaries along the model's actual competence.
flowchart TD A["Object navigation ends: target inside the 1-2 m neighborhood"] --> B["Short-term memory M_t: last K=5 observations with robot states"] B --> C["Stage 1 view selection: MLLM returns one reference view ID on visibility plus approachability"] C --> D["Stage 2 affordance grounding: MLLM outputs a normalized interaction point (u,v)"] D --> E["Depth back-projection: lift (u,v) to a 3D point p_a in the robot frame and mark it red on the image"] E --> F["Stage 3 base-pose reasoning: MLLM conditioned on p_a, R, T returns (x,y) plus arm_reach and arm_lift"] F --> G["Geometric heading: base faces p_a, theta computed in closed form instead of predicted"] G --> H["Transform to the global frame: low-level navigation drives there while avoiding obstacles"] H --> I["Manipulate: execute the pick or place with the predicted reach and lift"]
Experiments
Main OVMM results. On the HomeRobot OVMM validation split, UniLM-Nav with a Gemini-3-Flash-Preview backend reaches 23.77% Overall SR, ahead of the previously strongest zero-shot method MoTo at 20.64%, and above the trained MoManipVLA (15.80%) as well as the RL (14.80%) and heuristic (7.30%) HomeRobot baselines. Its Average SR of 53.50% is also the best in the table. The FindObj column is worth a second look: UniLM-Nav scores 69.47%, higher than every baseline, which says that a last-mile policy makes the approach itself more reliable, since view selection and standing-position reasoning improve the situation right after navigation terminates. Swapping in a 4B-class backend, RoboBrain-2.5-4B, still gives 19.19% Overall SR, better than all baselines except MoTo, which is a practical deployment-versus-performance trade.
| Method | FindObj | Pick | FindRec | Overall SR | Average SR |
|---|---|---|---|---|---|
| HomeRobot (RL) | 66.60% | 61.10% | 50.90% | 14.80% | 48.30% |
| HomeRobot (Heuristic) | 65.40% | 54.80% | 43.70% | 7.30% | 42.80% |
| MoManipVLA | 66.10% | 62.60% | 53.10% | 15.80% | 49.40% |
| UniTeam | 66.13% | 62.65% | 54.69% | 17.96% | 50.36% |
| MoTo | 66.67% | 60.95% | 49.87% | 20.64% | 49.53% |
| UniLM-Nav (Gemini-3-Flash-Preview) | 69.47% | 66.22% | 54.55% | 23.77% | 53.50% |
| UniLM-Nav (RoboBrain-2.5-4B) | 68.97% | 63.05% | 52.38% | 19.19% | 50.90% |
Backend choice: embodied data beats parameter count. Comparing 13 backends on a scene-stratified 20% subset, almost all of the spread sits downstream of FindRec, in the placement stage. FindObj stays between 66% and 69% for every backend, while Overall SR runs from 3.75% for Qwen3-VL-4B to 25.42% for Gemini-3-Flash-Preview. Inside the Qwen3-VL family the scaling trend is clean (4B at 3.75%, 8B at 9.58%, 32B at 15.83%), yet RoboBrain-2.5-4B, fine-tuned from the same architecture on robot-oriented embodied spatial reasoning data, reaches 20.50%. That is not only far above its own base size, it is above Qwen3-VL-235B-A22B-Instruct at 17.50%, a model two orders of magnitude larger. For a robot stack this is a much cheaper upgrade path than swapping in a bigger model.
| MLLM backend (20% subset) | FindObj | Pick | FindRec | Overall SR | Average SR |
|---|---|---|---|---|---|
| GPT-5.4 | 68.33% | 66.25% | 55.00% | 19.17% | 52.19% |
| Gemini-3-Flash-Preview | 67.91% | 65.83% | 54.58% | 25.42% | 53.43% |
| Qwen3-VL-4B-Instruct | 68.33% | 61.25% | 50.42% | 3.75% | 45.93% |
| Qwen3-VL-32B-Instruct | 67.92% | 62.08% | 52.08% | 15.83% | 49.48% |
| Qwen3-VL-235B-A22B-Instruct | 67.92% | 59.58% | 48.33% | 17.50% | 48.33% |
| RoboBrain-2.5-4B | 67.36% | 60.67% | 51.05% | 20.50% | 49.90% |
Component ablations: every stage is carrying load. Removing last-mile navigation entirely (turn to face the affordance and manipulate right away) drops Overall SR below 5%, which means that even with the target visible, standing in the right place is not a cosmetic refinement. Removing view selection moves the score from 25.42% to 20.42%. Replacing base-pose reasoning with direct floor pointing significantly hurts both Pick and Overall. The thinking-model ablation adds a sharper insight about stage heterogeneity: swapping Qwen3-VL-8B-Instruct for its Thinking variant only at the base-pose stage lifts Overall SR from 9.58% to 20.83%, while swapping it in at the affordance-grounding stage pushes the score down to 5.83%, and adding an explicit chain-of-thought prompt to the instruct model gives only 9.17%. Reasoning budget belongs to the stage that requires spatial deduction; on perceptual stages it interferes.
| Heading strategy (20% subset, Gemini backend) | FindObj | Pick | FindRec | Overall SR | Average SR |
|---|---|---|---|---|---|
| Geometric heading (face the lifted affordance) | 67.91% | 65.83% | 54.58% | 25.42% | 53.43% |
| MLLM predicts the heading directly | 67.50% | 63.33% | 52.92% | 17.08% | 50.21% |
| Per-stage model configuration (view / grounding / pose) | FindObj | Pick | FindRec | Overall SR | Average SR |
|---|---|---|---|---|---|
| Instruct / Instruct / Instruct | 68.75% | 60.00% | 47.92% | 9.58% | 46.56% |
| Thinking / Instruct / Instruct | 67.92% | 61.67% | 50.00% | 10.00% | 47.40% |
| Instruct / Thinking / Instruct | 67.92% | 55.83% | 45.83% | 5.83% | 43.85% |
| Instruct / Instruct / Thinking | 69.17% | 62.92% | 52.50% | 20.83% | 51.35% |
| Instruct / Instruct / Instruct with CoT prompt | 69.17% | 60.00% | 50.00% | 9.17% | 47.08% |
Error analysis: the bottleneck is outside the last mile and inside it. Manual categorization of failed episodes on the 20% subset attributes 61.0% of failures to navigation (object not found, receptacle not found, or navigating to a receptacle that is semantically correct but physically unreachable), 18.1% to manipulation (release height too high, or residual motion after placement failing the benchmark's stability threshold), and 20.9% to last-mile navigation. Inside that last group the most common error is affordance grounding (the point lands on the container rim so the placement is unstable), followed by view selection (the chosen view has an approach path blocked by a chair) and base-pose errors (the pose hugs a side wall and limits the arm workspace). These numbers say two things at once: object navigation is still the biggest bottleneck in OVMM, and each of the three stages has independent headroom.
Figure 4: left, manual categorization of failed episodes (navigation 61.0%, last-mile 20.9%, manipulation 18.1%); right, the three last-mile failure categories, with green marking the desired view, region or pose and red marking the model's erroneous prediction.
Real-robot experiments: closing the loop on a quadruped with a 6-DoF arm. The authors deploy UniLM-Nav on a Unitree B2 quadruped carrying a 6-DoF Unitree Z1 arm. The stock two-finger gripper is replaced by a custom 3D-printed parallel gripper with an integrated Orbbec Gemini 335 RGB-D camera in an eye-in-hand mount, an onboard LiDAR runs LIO-SAM for the occupancy map and odometry, and Navigation2 executes the last-mile displacement. Across four tasks repeated ten times each: taking a cup off a table 7/10, putting a cake onto a plate 6/10, placing a water bottle in front of a monitor 4/10, and "imagine you are facing the monitor, put the cake at the lower-left corner of the table" 4/10, for an overall success rate of 52.5%. The first two tasks confirm the handover chain works in the real world; the last two, which require resolving fine-grained spatial relations such as "in front of the monitor" or "the lower-left corner of the table", succeed markedly less often, matching the simulation finding that placement is the weak stage.
| Take a cup off the table | Cake onto a plate | Bottle in front of monitor | Cake at lower-left corner | Overall |
|---|---|---|---|---|
| 7/10 | 6/10 | 4/10 | 4/10 | 52.5% |
Figure 5: an OVMM success case. When object navigation ends (middle) the robot is wedged between a chair and a sofa; after last-mile navigation (right) the base has moved to a manipulable pose facing the desk. The red dot is the grounded affordance and the green dot is the predicted base position.
Limitations
The authors state three limitations. First, the framework assumes object navigation delivers the robot into a near-target state where the goal appears in recent observations; it has no active local exploration, and relaxing that assumption is listed as future work. Second, view selection, affordance grounding and base-pose reasoning still make decision errors of the kind shown in Figure 4, which the authors plan to reduce with stronger prompts and verification mechanisms. Third, evaluation is concentrated on OVMM plus one office environment, so generalization needs more benchmarks.
Our own reading adds four points. First, 23.77% Overall SR is state of the art but still low in absolute terms: placement reliability is the ceiling of the whole chain, and last-mile navigation improves the handover rather than the manipulation itself. Second, each episode costs at least three MLLM calls, plus a fourth for wrist-view refinement on the real robot, which in a closed control loop means second-scale latency and per-token cost; the paper reports no end-to-end timing, so anyone deploying this has to measure it themselves. Third, the accuracy of the 3D lift inherits depth noise and calibration error, and the fact that real hardware needs an extra refinement round shows how sensitive that step is to sensor quality. Fourth, a manipulation policy made of two scalars, arm_reach and arm_lift, only covers how far to extend and how high to lift; it is not enough for operations that need pose planning, so the benefit of the framework stops at the base handover.
Takeaways and Outlook
The contribution of UniLM-Nav is not a new model but a clean division of labor: semantics and spatial relations go to the MLLM, metric geometry goes to deterministic code, and every sub-decision is reshaped into a form the model can actually answer. View selection settles which eye to look through, task-conditioned grounding settles where to look, and depth back-projection plus a closed-form heading settle which quantities should never be left for the model to guess. The zero-shot state of the art on OVMM, the side-by-side comparison of 13 backends, and the closed-loop run on a quadruped together support the claim that an MLLM can serve as the visuospatial reasoner for the last mile, while RoboBrain-2.5-4B beating a 235B general model points the next step toward fine-tuning on embodied spatial data rather than merely scaling up.
Looking forward, the authors list active exploration upstream of the last mile, verification and self-correction across the three decisions, and evaluation on more benchmarks with longer-horizon tasks. From a systems angle, the more durable legacy may be the explicit geometric interface this framework exposes: a 3D affordance point plus a closed-form heading as a standard intermediate representation for mobile manipulation stacks, letting different MLLM backends be swapped in and out without touching the robot side.
Quote
Reaching the vicinity of the target is only the ticket into manipulation; the last meter that actually decides success or failure is hidden inside the base pose.
SOURCE LINKS



