PAPER DEEP DIVE
Wheels, Legs and Hierarchical RL: The Engineering Behind 8.3 km of Autonomous Urban Delivery at ETH Zurich
ETH Zurich's Hutter lab (Science Robotics 2024) unifies wheeled-legged urban delivery with three RL layers (LLC locomotion + HLC navigation + global planning): 8.3 km with only 3 human interventions, 1.68 m/s average speed, mechanical COT 0.16 (3x faster and 53% lower COT than legged DARPA SubT), 0.34 ms navigation inference — three orders of magnitude faster than sampling planners with zero collisions; gaits fully emerge (crawl-drive, active suspension, knee climbing), with honest boundaries on perception range, missing semantics and manual mapping.
Wheels + Legs + Hierarchical RL: the engineering behind 8.3 km of autonomous delivery on Zurich streets
The last parcel delivered to your door may one day be carried by a robot with four wheels on four legs. This work from Marco Hutter's Robotic Systems Lab at ETH Zurich, published in Science Robotics (Vol. 9, Issue 89, 2024), turns the question "how does a wheeled-legged robot actually work autonomously in a city?" into a complete system: adaptive locomotion control (LLC) + mobility-aware local navigation (HLC) + city-scale global path planning, all three layers driven by reinforcement learning, validated with kilometer-scale autonomous navigation in Zurich and Seville.
Why hybrid wheels-and-legs? The paper opens with the accounting: the purely legged ANYmal averaged roughly 2.2 km/h during the DARPA Subterranean Challenge — about half a human walking speed — with at most one hour of endurance; a purely wheeled robot cannot climb stairs. Urban delivery needs both ends of that curve: high speed and efficiency over large flat expanses, obstacle negotiation for stairs, curbs, gravel and lawns. A wheeled-legged robot puts a driven wheel at the end of each leg, rolling on flat ground and stepping on complex terrain — the natural answer to that demand curve, provided someone can solve the hybrid-gait question of when to roll and when to step.
(pre-scanned point cloud + graph, human picks goals)"] --> WP["Two waypoints WP1/WP2
(pure-pursuit look-ahead interpolation)"] WP --> HLC["HLC high-level navigation controller (10 Hz)
inputs: LLC belief state + terrain height +
memory of positions visited over last 10 m"] HLC -->|"velocity target"| LLC["LLC low-level locomotion controller (50 Hz)
RNN policy + privileged learning
autonomously selects gait: walk/drive/hybrid"] LLC -->|"joint positions + wheel speeds"| R["ANYmal wheeled-legged robot"] R -->|"IMU / encoders / height map"| HLC
LLC: gait selection handed entirely to learning, with no human intuition
The low-level locomotion controller (LLC) is an RNN policy built on Miki et al.'s perceptive locomotion framework with two key changes: an improved observation/action space, and the complete removal of engineered motion primitives (CPGs). Training uses privileged learning: during training the agent sees true velocities/accelerations, terrain properties and noise-free exteroceptive measurements, while at deployment the policy relies only on raw IMU, joint encoders and onboard height-map readings — no conventional state estimator at all, using raw sensor readings directly to remove heuristic filtering as a failure point. The result is that gaits emerge entirely on their own: an asymmetric crawl-and-drive gait on large discrete obstacles, a return to the classic trot on stairs and steep slopes, and pure rolling over undulating terrain (height variation close to the wheel radius) — with all four legs finely adjusting extension as active suspension to keep wheels on the ground, and the body actively lowered on descents to prevent tipping. Peak flat-ground speed is 5.0 m/s (hardware limit 6.3 m/s). The two most dramatic moments: driving off a ~60 cm table, the front legs extend then tuck to keep the torso level, and the moment the front wheel lands the robot rolls forward back into balance; climbing a ~40 cm ledge with all four wheels in the air, it crawls on its knees until some wheel touches down again — behavior no model can design, and precisely the value of model-free RL.
HLC: replacing the classic "plan + track + communicate" trio with 0.34 milliseconds
The high-level navigation controller (HLC) is the paper's biggest technical contribution: it merges the three modules of a conventional navigation stack — path planning, path tracking and inter-module communication — into a single neural network, directly outputting velocity targets at 10 Hz (matching the onboard height-map update rate). It has three input streams: the LLC's belief state (RNN hidden state, encoding terrain properties and disturbances), terrain height values around the robot, and a distinctive position memory — the last 20 visited positions (one every 50 cm, covering 10 m, exactly a typical waypoint spacing), letting the HLC decide based on its own navigation history. Training uses the videogame concept of a Navigation Graph plus WFC procedural content generation: each episode generates a new obstacle-free path and randomly sampled waypoints, including detours, dynamic obstacles, rough terrain and narrow passages, with reward for following the shortest path to the goal — this controlled compositional training outperforms randomly scattering obstacles.
8.3 km, 30-minute-class tasks, three human interventions
The delivery demonstration in Zurich's Glattpark: a handheld laser scanner takes 90 minutes to scan a 245 m × 345 m urban area, producing a point cloud and a navigation graph, with goals placed by a human (the graph also encodes social preferences — avoiding green belts and private property). At deployment the robot localizes with LiDAR inside the pre-scanned point cloud (more stable than GPS among tall buildings), and after receiving a GPS goal it plans and drives fully autonomously, with the point cloud used only for localization, not for navigation. The robot accumulated 8.3 km of autonomous driving with only three human interventions: a child in the path (avoidable, but it stopped proactively for safety); tall grass that grew after the map was built (safe stop, human triggered global replanning); and localization drift in geometrically degenerate environments such as long corridors (kept driving safely on the local terrain map until localization recovered). The speed and energy comparison is the paper's brightest result: average 1.68 m/s and mechanical COT 0.16 — against ANYmal's DARPA SubT numbers, that is 3× the speed at 53% lower COT. The gain comes mainly from driving mode: with four legs sharing load evenly and joints nearly still, the joints' mechanical COT contribution approaches zero; the leg joints' energy consumption alone is 16% lower than ANYmal's despite being 12 kg heavier and faster — directly relevant to heat dissipation losses.
Four moments where local navigation looks genuinely smart
The paper's local navigation cases read most like "the robot has a mind." Exploratory detouring: when the path is blocked, the robot reverses and searches along walls until it finds stairs to the goal — position memory tells it where it has been. Narrow doorway traversal: two doors, a person standing between them, a gap exactly the width of the robot — passed without collision (and at that point human detection was not even enabled). Asymmetric obstacle cognition: facing a combined stair plus 0–50 cm ledge, the robot first reverses to explore, finds a feasible height of ~20 cm along the steps and climbs; data analysis further shows it can descend ledges higher than it can ascend — breaking the assumption of direction-independent symmetric estimates in conventional traversability cost maps, evidence that the HLC really decides from current terrain, its own state and the low-level controller's characteristics rather than consulting a static cost table. Dynamic person avoidance: camera-based human detection injects a height offset within a 50 cm radius around pedestrians, and the HLC overtakes them at a constant distance safely.
Head-to-head against a sampling-based planner: 3000× faster, zero collisions
The comparison reveals the essential advantage of learned tightly-coupled navigation over classical sampling planning. The baseline is a traversability-cost-map sampling planner (SOTA for legged robots, taking seconds per plan). Results: ① speed — HLC averages 0.34 ms from observation update to network inference, while the baseline exceeds one second on a desktop (Ryzen 9 3950X + RTX 2080), three orders of magnitude apart; ② collisions — only this method achieves collision-free trajectories, while the baseline repeatedly crashes into dynamic changes due to replanning latency plus an assumption of perfect tracking, and distant pose goals cause overshoot at high commanded speeds; ③ tracking error — this method averages 0.24 m/s with a uniform distribution, versus 0.45 m/s for the baseline with spikes at abrupt command changes (the LLC simply refusing unreasonable velocity commands); ④ exploration — in partially observable settings this method dynamically explores new regions with the highest success rate; ablation shows the policy without position memory falls into repetitive behavior and local minima with the highest failure rate. Together: when the robot moves fast enough that second-scale planning cannot react in time, navigation must be trained jointly with the actual capabilities of the motion controller (latency, error, gait characteristics), not optimized separately and stitched together.
Beyond the numbers: why the baseline keeps crashing
The failure-case analysis in the comparison is more informative than the success rates. The sampling planner's failures fall into two root causes. Occlusion handling: the traversability cost map derives from the height map, so occluded regions default to impassable and the planner gets stuck in front of them repeatedly; painting occluded regions as "passable" heuristically solves part of it, but replanning latency still leaves it flat-footed when the environment changes — whereas the HLC's position memory plus exploratory behavior lets it "go take a look" before deciding. Tracking-error assumptions: almost all classical methods assume the motion controller executes velocity commands perfectly, but a real LLC has latency, error and protection logic that refuses unreasonable commands; when the baseline emits distant pose goals, the resulting high-speed commands cause overshoot and the robot leaves the traversable area — the "smart planner, lagging actuator" mismatch that speed amplifies on fast robots. In contrast, the HLC and LLC here are jointly trained: every high-level velocity command has "experienced" the low-level response during training, so commands naturally fall inside the envelope of low-level capability. This confirms the DARPA SubT lesson from the introduction — systems assembled from individually optimal modules glued by heuristics exhibit stop-and-replan and zig-zag behaviors in the real world, while tightly-coupled learned systems swallow the engineering black hole of "inter-module communication" directly into network weights.
Honest boundaries
The authors are equally candid about limitations. First, semantics are almost absent — apart from injecting a height offset for pedestrians the system is purely geometric, so paved-surface detection and visual traversability estimation are explicit next steps. Second, perception range is limited — the HLC sees only 3 m ahead, inherent to the height map; the hardware can run 6.2 m/s, but autonomous deployment never dares approach that speed because of mapping latency, and "dropping the height map to consume raw sensor streams directly" is named the most promising direction. Third, mapping still requires humans — a 90-minute handheld scan is a real bottleneck for large-scale deployment. These admissions make the "kilometer-scale autonomous delivery" conclusion more credible: the system works in a complex real city, and its ceiling is clearly marked.
The engineering details that hold the whole system together
Kilometer-scale autonomy is not carried by two neural networks alone, and the system-integration description is equally dense. Payload includes three LiDARs, a front-facing RGB stereo camera, a delivery box, a 5G router and a GPS antenna — the RGB camera exists independently of the point-cloud terrain map because the map cannot capture dynamic obstacles: the camera runs human detection at high frequency, tracking pedestrians within 20 m in real time and injecting a buffer offset into the height map around detections. Global paths come from a shortest-path algorithm on the navigation graph with 2–20 m waypoint spacing, interpolated into two look-ahead waypoints WP1/WP2 between the robot's projected position and the next graph node, pure-pursuit style; goals are pushed over the mobile network and paths computed onboard. The localization choice is worth noting: LiDAR localization in the pre-scanned point cloud (LiDAR + IMU + joint encoders), which the authors found more stable than GPS among tall buildings — GPS multipath in urban canyons being a classic pitfall for delivery robots. The three interventions are recorded precisely by scenario: a child in the path (avoidable but stopped proactively), tall grass that grew between mapping and deployment (a real-world time-gap problem), and localization degradation in long corridors (geometric degeneracy being LiDAR localization's soft spot). These details together answer the question any reviewer must ask: did this system really run in the city, or is it another perfect curve in simulation?
A quantitative picture of hybrid gaits: why descending beats ascending
The paper's quantitative analysis of hybrid locomotion is revealing precisely because it exposes complexity that model-based planning struggles to capture. Step height vs. commanded speed: the maximum step height the robot clears varies with commanded speed, and descending is clearly stronger than ascending — consistent with the HLC's behavior of avoiding upward ledges it cannot clear, protecting its knees and ensuring safety; "can it pass" is not a static property of terrain and robot but a three-way function of terrain × robot state × motion direction. Slope × speed gait boundaries: climbing simulated slopes at fixed friction 0.7 and fixed linear speed (2 m counts as success), stepping gaits emerge only when the slope is steep and commanded speed exceeds 0.5 m/s — at low speed driving mode is more stable. Conventional model-based planning and path tracking therefore struggle to express coupled decisions like "which gait at what speed for this slope," even with a perfect traversability map. Comparison against a conventional MPC controller: the team's earlier MPC controller lacked robustness and could not even run on the real terrains in Fig. 6 (grass, sand, gravel, stairs, slopes), whereas the learned LLC produces adaptive gaits on all of them. Together these results point at the paper's deeper thesis: the design space of wheeled-legged systems is too large (gait × speed × terrain × payload) for hand-written rules and model optimization to cover, and handing that design space to RL with privileged learning is the pragmatic choice.
Why explicit hierarchy rather than end-to-end: a sample engineering decision
The methods section unusually publishes the reasoning behind the architecture choice, valuable reference for systems researchers. The team considered three options: ① a unified end-to-end policy handling both locomotion and navigation (as in Rudin et al.); ② decoupled high/low levels with learned latent subgoals (flexible but black-box); ③ two-level HRL with explicit subgoals — low level focused on locomotion, high level outputting velocity targets. They chose the third, for purely engineering reasons: the explicit split lets the two controllers be developed independently (two subteams in parallel without blocking each other), the low-level policy can be reused across high-level applications (standard practice in legged robotics), and the high level's velocity targets are physically interpretable so a human can read them while debugging. The team also tried having the high level command gait parameters directly (in the spirit of Tsounis et al.), with experiments in the supplementary material. This "chose the maintainable option, not the most elegant one" decision process, together with the honest record of three human interventions, makes up the engineering honesty a top-tier systems paper should have.
Why this paper belongs on the robot-learning must-read list
Its value operates on three levels. Systems: it is one of the few works that truly connects perception, navigation and locomotion end-to-end and accumulates kilometer-scale mileage in a real city rather than improving a single line of code — from handheld scan mapping, social-preference encoding in the navigation graph, goal delivery over the mobile network and LiDAR point-cloud localization to the hierarchical collaboration of two neural networks, every link is backed by engineering detail and failure records. Method: it replaces the seam that has always connected navigation planning and locomotion control through hand-crafted interfaces with joint training — the HLC consumes the LLC's RNN hidden state rather than an explicit state estimate, outputs velocity targets rather than pose goals, and uses position memory instead of explicit map search; 0.34 ms inference latency lets navigation decisions match 5 m/s-class locomotion for the first time. Conclusions: 3× speed and 53% lower mechanical COT (with lower leg-joint energy too) show that wheels-plus-legs is not merely "two modes glued together" but a path that directly pushes past the legged robot's real bottlenecks of one-hour endurance and 2.2 km/h in flat urban environments. For anyone in embodied AI, perhaps the most memorable lesson is the DARPA SubT one: gluing individually SOTA modules together with heuristics usually yields a system that runs but looks ugly, pausing and zig-zagging; true autonomy grows in the connections between modules.
References
- Paper: Learning Robust Autonomous Navigation and Locomotion for Wheeled-Legged Robots — arXiv:2405.01792; published in Science Robotics, Vol. 9, Issue 89 (2024)
- Authors: Joonho Lee, Marko Bjelonic, Alexander Reske, Lorenz Wellhausen, Takahiro Miki, Marco Hutter (ETH Zurich Robotic Systems Lab)
- Main result video: Movie 1; full mission: Movie S1; locomotion experiments: Movie S3
SOURCE LINKS



