PAPER DEEP DIVE
Agile and Generalized Legged Locomotion via Attention-Based Neural Map Encoding
Achieving agile and generalized legged locomotion across terrains requires tight integration of perception and control, especially under occlusions and sparse footholds. Existing methods have demonstrated agility on parkour courses but often rely on end-to-end sensorimotor models with limited generalization and interpretability. By contrast, methods targeting generalized locomotion typically exhibit limited agility and struggle with visual occlusions. We introduce a unified reinforcement learning (RL) framework for agile and generalized locomotion that incorporates a novel attention-based map encoder in the control policy. This encoder extracts local and global mapping features and uses attention mechanisms to focus on salient regions, producing an interpretable and generalized embedding for RL-based control. We further propose a learning-based mapping pipeline that provides fast, uncertainty-aware terrain representations robust to noise and occlusions, serving as policy inputs. It uses neural networks to convert depth observations into local elevations with uncertainties, and fuses them with odometry. The pipeline also integrates with parallel simulation so that we can train controllers with online mapping, aiding sim-to-real transfer. We validate our framework with the proposed mapping pipeline on a quadruped and a biped robot, and the resulting controllers demonstrate strong agility and generalization to unseen terrains in simulation and in real-world experiments.
Paper Metadata
Title: Agile and Generalized Legged Locomotion via Attention-Based Neural Map Encoding
Authors: Chong Zhang, Victor Klemm, Fan Yang, Marco Hutter (Robotic Systems Lab, ETH Zurich; corresponding: chong.zhang@ai.ethz.ch; also affiliated with the Secure, Reliable, and Intelligent Systems Lab and the ETH AI Center)
Links: arXiv:2601.08485v3 [cs.RO], updated 2026-09-07; full text https://arxiv.org/abs/2601.08485; project page https://sites.google.com/leggedrobotics.com/ame-2 ; result video https://www.youtube.com/watch?v=FNA1crvtBLs
Code status: a minimal reproduction repo is released at github.com/leggedrobotics/ame2_minimal (~78 stars / 3 forks, GPL-3.0, created 2026-10-05). It is an Isaac Lab extension focused on training a Unitree G1 teacher policy, and it also ships AME-1 and MoE baseline implementations, an optional TAGA-style active-gaze variant, and one G1 teacher checkpoint (modelzoo/g1_gaze_64000env_10h). The README explicitly states "half-archived as a reference to a finished work" and "No feature contribution"; the three pieces of the paper's ANYmal-D / TRON1 pipeline — Mid-360 lidar neural mapping in the loop, G1 student training, and the mapping model weights — are struck through and marked "unavailable until the next paper release". So the policy-side and reward/termination logic and the probabilistic winner-take-all fusion can be read line by line, while the learned mapping network weights are not yet available.
One-Sentence Summary
AME-2 adds a global terrain-context branch to an attention-based map encoder and uses it to modulate the local contact features, paired with an uncertainty-aware neural elevation-mapping pipeline that runs identically in simulation and on hardware; the same reward and training recipe, unchanged, drives a quadruped (ANYmal-D) and a biped (TRON1) to climb 1 m ledges zero-shot, run parkour at 2 m/s, cross a 19 cm balance beam, and spontaneously discover whole-body behaviors such as knee support and landing impact absorption.
Lead figure (Fig. 1): the same method runs on the quadruped ANYmal-D and the biped TRON1 with onboard sensing and computation only, traversing stairs, boxes, slopes, debris, and sparse terrain.
Background and Motivation
Legged robots need agility and generalization at the same time: climb high ledges, jump gaps, step on sparse footholds, and still not collapse on terrain never seen before. The hard part is that perception and control must be tightly coupled, while occlusions, sparse footholds, and sensor noise are the norm.
One mature route is classical model-predictive control: maintain a state estimator, build an explicit elevation map from lidar or a depth camera, and plan/control on top of that map. This is stable on structured terrain, but mapping often relies on hand-tuned filtering and probabilistic fusion, is slow and terrain-specific, and breaks down when occlusions punch holes into the elevation map — too slow to support agile motion.
A second route is end-to-end: a single network maps raw sensor streams and proprioception directly to joint commands. This family has reached agility on parkour courses, but generalization is limited and interpretability is poor: perception and action are welded into one black box that degrades when the test terrain distribution shifts.
The predecessor AME-1 (Science Robotics 2025, ref [15]) took a middle road: keep the modular structure of an explicit elevation map, but compress map features into a policy embedding with an attention map encoder, achieving strong generalization — at the cost of agility. Its way of encoding local footholds is essentially "model-based foothold scoring", looking only at local terrain features and robot state, and cannot express "what kind of terrain is this and which contact pattern should I use".
This paper aims to get both agility and generalization from two claims. First, the map encoder must carry both fine-grained local contact features and global terrain context, and the global context must modulate the local features. Second, the perception side needs a fast, lightweight, differentiable, explicitly uncertainty-aware neural mapping pipeline that runs across thousands of parallel simulation environments and then runs the very same code on hardware, closing the sim-to-real perception gap.
Preliminaries
POMDP and PPO. The control problem is cast as a partially observable Markov decision process and optimized with PPO in massively parallel simulation (Isaac Gym). The policy outputs joint PD targets at 50 Hz, tracked at 400 Hz on hardware.
Privileged teacher, distilled student. A privileged teacher is first trained with ground-truth map and noiseless proprioception; then a student, fed by neural mapping, is trained jointly on the PPO RL loss, an action-distillation loss to the teacher, and a representation-alignment loss between the two map embeddings.
Elevation maps and the AME-1 attention encoder. An egocentric elevation map writes each point as a 3D coordinate $(x,y,z)$; AME-1 extracts local features with a CNN, builds a query from proprioception, and applies multi-head attention over the local features to yield a map embedding. AME-2 adds global features on top of this.
Method
3.1 System Overview
The system has two levels: a teacher trained by RL with ground-truth mapping in simulation, and a student that replaces the ground-truth map with online neural mapping and is trained under the teacher's action distillation and representation alignment, then deployed to hardware. The controller only solves goal-reaching (to a position and heading), not velocity tracking, which gives the policy more freedom to discover agile motions.
Fig. 1: System overview (paper Fig. 2). Top: the privileged teacher is trained in simulation with ground-truth mapping; bottom: the deployable student uses online neural mapping, is supervised by the teacher, and drives the robot to position and heading goals.
3.2 Controller Inputs/Outputs and Map Representation
Proprioception includes base linear velocity $v_b$ (teacher only), base angular velocity $\omega_b$, projected gravity $g_b$, joint positions $q$, joint velocities $\dot{q}$, previous actions, and the goal command $c$. A ground-truth map writes each egocentric grid point as $(x,y,z)$; neural mapping augments this with an uncertainty channel $u$, giving a 4D representation $(x,y,z,u)$ — exactly the extra information the student has relative to the teacher.
3.3 The AME-2 Encoder: Global Context Modulates Local Contact Features
AME-2 is the core of the paper. It has two parallel branches that are then merged:
First, pointwise local features. The map goes through an MLP to produce positional embeddings (16-dim in the paper's figure), while the non-coordinate channels go through a two-layer CNN to produce local embeddings (48-dim); the two are concatenated and fused by another MLP into 96-dim pointwise local features $(\text{points}\times 96)$.
Second, global terrain context. The pointwise local features go through another MLP and are max-pooled over the point dimension into a 1×64 global feature that captures the "type" of the whole terrain.
The key step follows: the global feature is concatenated with the proprioception embedding and passed through an MLP to form a 96-dim query vector; with the pointwise local features as keys and values, multi-head attention (32 heads, dim 96) yields local features "weighted by the current proprioceptive state and the global context". Finally, the global feature and the weighted local features are concatenated into the map embedding and fed to the action-decoder MLP.
This is exactly the difference from AME-1: AME-1's query comes only from proprioception, so attention only distributes over local footholds; AME-2 first injects a global judgment ("is this terrain sparse, climbing, or obstacle-laden") into the query, and then lets attention pick contact regions. The authors' rationale: high-obstacle climbing may need whole-body contacts, and obstacle avoidance needs to avoid certain contacts — neither is expressible by "pick the single best foothold"; the policy must judge "where to step" and "whether/how to make contact" together.
Fig. 2: Policy and AME-2 encoder architecture (paper Fig. 3). Top-left: Actor overview. Blue box: the AME-2 encoder — positional embedding + CNN local embedding → pointwise local features; MLP + max pool → global features; the query comes from "proprioception embedding ⊕ global features", and MHA produces the weighted local features. Right: the proprioception encoder — an MLP for the teacher, LSIO over the past 20 steps for the student. Bottom: the critic is an MoE.
The proprioception encoder differs between teacher and student: the teacher uses a plain MLP over noiseless proprioception; the student stacks the past 20 steps of typed observations and uses LSIO (Long-Short I/O) for a temporal embedding, then feeds it with the command into an MLP. The reason is that on real hardware the environment is uncertain and history is important for judging robot state and environment dynamics.
3.4 Asymmetric Actor-Critic and Teacher-Student RL
The critic does not reuse the actor's attention structure, because the critic is neither deployed nor required to generalize beyond training terrain, while an MHA over $L\times W$ local features is expensive. Instead the critic uses a mixture-of-experts (MoE) design that is powerful for function fitting yet cheaper to optimize. Note one counter-intuitive ablation: an "AME-2 actor + MoE critic" trains a generalized teacher, but swapping in an MoE actor does not yield a generalized teacher.
During teacher training, the actor gets noiseless proprioception and the ground-truth map, while the critic additionally gets per-link contact states; contact states are withheld from the actor to keep an appropriate information gap and avoid hurting student training. Left-right symmetry augmentation is applied only to the critic.
The student objective is a linear combination of three parts: the PPO RL loss, an action-distillation loss to the teacher, and a representation loss between the teacher's and student's map embeddings (mean squared error). For the first few thousand iterations, the PPO surrogate loss is disabled with a large learning rate so distillation settles first; then RL is brought in.
3.5 Reward Design: No Foothold Reward, Let Whole-Body Contact Emerge
Rewards fall into three groups: task rewards (reaching goals), regularization/penalties, and simulation-fidelity terms (penalizing near-limit joint states). The quadruped and biped share one reward set. The position-tracking reward follows the predecessor:
$$r_{\rm position\_tracking}=\frac{1}{1+0.25\,d_{xy}^{2}}\cdot t_{\rm mask}(4)\tag{1}$$
where $d_{xy}$ is the horizontal distance to the goal and $t_{\rm mask}$ is a time-based mask that decays toward the end of the episode:
$$t_{\rm mask}(T)=\frac{1}{T}\cdot\mathbf{1}(t_{\rm left}<T)\tag{2}$$
The heading reward is analogous but active only near the goal:
$$r_{\rm heading\_tracking}=\frac{1}{1+d_{\rm yaw}^{2}}\cdot t_{\rm mask}(2)\cdot\mathbf{1}(d_{xy}<0.5)\tag{3}$$
After reaching the goal the robot must hold a stable stance, hence a standing reward that exponentially decays with the fraction of non-contacting feet $d_{\rm foot}$, the base tilt $d_{\rm g}$, the mean joint deviation from the standing reference $d_{\rm q}$, and the distance:
$$r_{\rm stand}=\mathbf{1}(d_{xy}<0.5\;\land\;d_{\rm yaw}<0.5)\cdot\exp\left(-\frac{d_{\rm foot}+d_{\rm g}+d_{\rm q}+d_{xy}}{4}\right)\tag{5}$$
A crucial choice: unlike prior work, this paper does not explicitly reward or penalize foot-contact positions; instead it defines penalties for all links, allowing whole-body contact to emerge. Foothold rewards are hard to define in a terrain-agnostic way — quadruped climbing benefits from active knee contacts and near-edge foot placements, whereas on sparse terrain near-edge placements are exactly what to avoid. The penalized events include: excessive yaw rate (>2.0 rad/s), leaping on flat terrain, non-foot contacts and their switches, stumbling (horizontal contact force larger than vertical), slippage (a link moving while in contact), and self-collision. All weights are integer powers of ten and one to two orders of magnitude smaller than task rewards, so training succeeds without extensive tuning.
3.6 Termination: Biokinetic Thresholds That Induce Impact-Absorption Behaviors
Early-termination conditions include: bad orientation (projected-gravity components out of range), base collision force exceeding the robot's total weight, high thigh acceleration, and stagnation (too little movement within 5 s). The most interesting is the thigh-acceleration threshold: the authors take biokinetic literature values from dogs ($60\ \mathrm{m/s^2}$) and humans ($100\ \mathrm{m/s^2}$) directly as thresholds to suppress impact during jumping. This seemingly pragmatic termination rule directly induces the hardware behaviors in Section 6 — ANYmal-D touches down gently with its knee when descending, and TRON1 retracts its leg before landing and then switches the support leg.
3.7 Terrain and Curriculum
Training terrain falls into three categories: dense (rough ground, stairs, boxes, obstacles), climbing (climbing out of a pit, down a platform, consecutive climbs), and sparse (gap jumps, pallets, stepping stones, beams). Half of the sparse terrains get a physical floor and the other half a "virtual floor" — visible in the map but without collision — to stop the robot from lazily walking on the floor. Curriculum-wise, roughness grows from ±0 m to ±0.2 m (ANYmal-D), box height from 0.05 m to 0.4 m, obstacle density from 0 to 0.5 per m², and slope from 5° to 45°.
Generalization is tested on four terrains never seen during training: Test 1 is randomly overlapping stepping stones (sparse generalization); Test 2 combines stepping stones, rotated pallets, and high-platform climbing (compositional generalization); Test 3 requires avoiding a large obstacle, then climbing, jumping a gap, avoiding another obstacle, and climbing down (obstacle avoidance + agile climbing + jumping); Test 4 is a realistic debris mesh requiring climbing over rubble without getting a foot stuck (realistic climbing generalization).
Fig. 3: Training and test terrains (paper Fig. 4). Top: three primitive training-terrain categories; bottom: four test terrains, all unseen or unseen combinations during training.
3.8 Neural Mapping Pipeline: From Depth Clouds to an Uncertainty-Aware Elevation Map
The mapping pipeline is the paper's second main contribution. Each frame, the depth point cloud is projected onto a local 2D height grid; if multiple points fall into the same cell, the maximum $z$ is kept (the most relevant for locomotion), and empty cells get a fixed minimum value. The resulting local elevations are noisy and incomplete under occlusions, so a lightweight CNN jointly predicts base-relative elevations and uncertainties (as log-variance). This uncertainty captures both measurement noise and occlusions, while the elevation prediction itself suppresses noise.
Odometry then fuses the local predictions into a global grid map $\mathcal{M}$ with two layers: elevation and uncertainty (variance). The global map is initialized as flat ground at the robot's standing height with a large uncertainty. Fusion does not use standard Bayesian fusion, because "repeated observations of the same occluded area should not reduce uncertainty merely due to consistent predictions". Instead the authors use a Probabilistic Winner-Take-All strategy:
First, compute an effective measurement variance, lower-bounded by the prior to prevent over-confidence:
$$\hat{\sigma}^{2}_{t}=\max\left(\sigma^{2}_{t},\,0.5\cdot\sigma^{2}_{\rm prior}\right)\tag{6}$$
An update is valid only if the effective variance is not significantly larger than the prior; for valid updates, the overwrite probability is set by relative precision:
$$p_{\rm win}=\frac{(\hat{\sigma}^{2}_{t})^{-1}}{(\hat{\sigma}^{2}_{t})^{-1}+(\sigma^{2}_{\rm prior})^{-1}}\tag{7}$$
Finally the map is updated stochastically: sample $\xi\sim\mathcal{U}[0,1]$ and let the new prediction take over the cell if the sample falls within the threshold:
$$(h_{\rm new},\sigma^{2}_{\rm new})\leftarrow\begin{cases}(h_{t},\hat{\sigma}^{2}_{t})&\text{if }\xi<p_{\rm win}\\ (h_{\rm prior},\sigma^{2}_{\rm prior})&\text{otherwise}\end{cases}\tag{8}$$
This strategy has four benefits: the uncertainty of an occluded point does not decrease from consistent predictions; over-confident but inconsistent predictions cannot take over; the map can update quickly when high-confidence measurements appear for dynamic terrain; and it is easy to integrate with parallel simulation yet fast enough to run on hardware.
Fig. 4: Neural mapping pipeline (paper Fig. 5). Depth clouds are projected to local grids, a neural network predicts elevations with uncertainties, odometry fuses them into a global map, and local maps are queried by robot pose as controller input. Green points are elevation estimates, blue lines are uncertainties.
3.9 Mapping Model Training and Architecture
The mapping model is trained with ray-traced sampling from procedurally generated terrain and random poses, without physics simulation. Using Warp ray tracing on a single RTX 4090, sampling reaches hundreds of thousands of frames per second. Training applies several augmentations: per-cell additive uniform noise, random cropping from the four borders, simulated occlusions with random sensor poses and fields of view, random elevation clipping ranges, and random missing points and outliers — synthesizing noisy, partially observable inputs whose labels are the ground-truth elevations.
The loss is a $\beta$-NLL ($\beta=0.5$):
$$L_{0.5}=\mathbb{E}_{X,Y}\left[\mathrm{sg}\!\left[\hat{\sigma}(X)\right]\left(\frac{\log\hat{\sigma}^{2}(X)}{2}+\frac{(Y-\hat{\mu}(X))^{2}}{2\hat{\sigma}^{2}(X)}\right)\right]\tag{9}$$
where $\mathrm{sg}[\cdot]$ is the stop-gradient operator. Compared with the standard NLL of classical Bayesian learning, this form curbs the tendency to "raise uncertainty on hard samples to trivially reduce the loss", encouraging high uncertainty where accurate prediction is impossible and low uncertainty with accurate predictions where it is possible.
Because terrain roughness varies widely within a batch, flat samples can dominate the loss. Samples are therefore reweighted by total variation (TV):
$$\mathrm{TV}(Y_{b})=\frac{1}{HW}\left(\|\nabla_{x}Y_{b}\|_{1}+\|\nabla_{y}Y_{b}\|_{1}\right),\qquad w_{b}=\frac{\mathrm{TV}(Y_{b})}{\sum_{b^{\prime}=1}^{B}\mathrm{TV}(Y_{b^{\prime}})+\varepsilon}\tag{10}$$
The network is a shallow U-Net with a gated residual design: the CNNs output uncertainty, a raw estimate $h'$, and a gating map $G$, and the final estimate is a gated combination $h=G\odot h'+(1-G)\odot h_{\rm in}$, preserving accuracy in clearly observed areas while selectively overwriting noisy or occluded ones. Each robot is trained on 54 million frames and each model converges in under an hour. ANYmal-D's local grid is 51×31 at 4 cm resolution; TRON1's is 31×31 at 4 cm.
Fig. 5: Mapping model architecture (paper Fig. 7). A gated-residual U-Net: input $h_{\rm in}$ passes through two CNN layers with a pool–upsample skip, splitting into an uncertainty head log σ², a raw-estimate head h′, and a gating head G, giving $h=G\odot h'+(1-G)\odot h_{\rm in}$.
3.10 Simulation Integration and Deployment
With 1000 parallel ANYmal-D environments, mapping inference is under 0.3 ms with ~3 GB of GPU memory (TRON1 is roughly 60% of that); storing 1000 8 m×8 m global maps takes ~0.3 GB, and the map can recenter around the robot as it nears the boundary. Depth clouds and other intermediate overheads add ~1.2 GB per 1000 environments. On hardware, the whole mapping pipeline takes ~5 ms per frame, with ~2.5 ms spent on ONNX Runtime inference, fast enough to keep up with the depth camera. ANYmal-D merges point clouds from two front cameras, TRON1 uses a single front camera; odometry uses CompSLAM + Graph-MSF on ANYmal-D and the lighter DLIO on the higher-acceleration TRON1. Because odometry velocity estimates are noisy and delayed (a 1 cm drift within 20 ms implies a 0.5 m/s velocity error), the authors remove the linear-velocity observation from the student.
3.11 Pipeline Overview
flowchart TD
subgraph MAP["Neural mapping pipeline 50Hz shared sim and hardware"]
DC["Depth cloud two front cameras merged"] --> LG["Local height grid 51x31 4cm keep max z"]
LG --> UN["Lightweight gated-residual U-Net predicts elevation and log-variance"]
UN --> PW["Probabilistic winner-take-all fusion variance floor and overwrite prob"]
ODO["Odometry CompSLAM+Graph-MSF or DLIO"] --> PW
PW --> GM["Global map 8m x 8m elevation layer and variance layer"]
GM --> Q["Query egocentric local map 4 channels x y z u"]
end
Q --> AME["AME-2 encoder"]
P["Proprioception stacked past 20 steps"] --> LSIO["LSIO temporal encoder"] --> PE["Proprioception embedding"]
PE --> AME
AME --> ME["Map embedding global features concat weighted local features"]
ME --> DEC["Action decoder MLP"]
DEC --> A["Joint PD targets 50Hz tracked at 400Hz"]
PE --> QRY["Query vector from proprio embedding concat global features"]
AME --> QRY
Fig. 6: AME-2 training and deployment flow. Left: the neural mapping pipeline (depth cloud → local grid → U-Net elevation and uncertainty → probabilistic winner-take-all fusion → global map → query local map). Right: the policy side, where the proprioception and map embeddings feed the action decoder that outputs joint targets.
3.12 Correspondence with the Open-Source Code
Although the repo targets G1 teacher training, its AME-2 encoder and probabilistic winner-take-all fusion map line by line to the paper. The encoder lives in rsl_rl/rsl_rl/models/AME2_models.py at AME2Model.get_latent(), where the local features, global features, and attention query are built exactly as the encoder design behind Eqs. (6)-(8):
# AME2_models.py :: AME2Model.get_latent()
fc_out = self.fc_part(map_4d) # positional embedding (B, L, W, 16)
cnn_out = self.cnn_part(map_4d[..., 2:].permute(0,3,1,2)) # local embedding (B, 48, L, W)
local_feat = self.local_proj(torch.cat([fc_out, cnn_out], -1)).view(B, L*W, 96)
global_pre_pool = self.global_pool_mlp(local_feat) # (B, L*W, 64)
global_feat, _ = torch.max(global_pre_pool, dim=1) # max pool -> (B, 64)
query = self.query_mlp(torch.cat([prop_embedded, global_feat], -1)).unsqueeze(1)
attn_out, _ = mha_math(self.mha, query, local_feat, local_feat) # K=V=local_feat
Here fc_part is Linear(dmap, 16), cnn_part processes only the non-coordinate channels via Conv2d(dmap-2, ...), and local_proj projects 16+48 dims to 96 — matching the "positional 16 / local 48 / pointwise local 96" dimensions of paper Fig. 3. The attention module mha_math (in the same repo, rsl_rl/rsl_rl/modules/ame2_modules.py) pins SDPA to the math kernel, with a comment explaining this avoids backward crashes of the fused kernel on Blackwell at head_dim=3 with batch×heads above ~1e5 — i.e., the paper's 96-dim, 32-head attention.
The probabilistic winner-take-all fusion in ame2/ame2/sensors/neural_mapping_utils.py at BatchedGlobalProbWinnerPipeline.step() is a line-by-line implementation of Eqs. (6)-(8):
# neural_mapping_utils.py :: BatchedGlobalProbWinnerPipeline.step()
R_eff = torch.maximum(R_meas, self.tau_min * P_prior) # Eq.(6) variance floor
better = (R_eff < P_prior * (1+self.tau_min)) | (R_eff < 0.04) # validity gate
prec_eff, prec_prior = 1.0/(R_eff+eps), 1.0/(P_prior+eps)
p_meas = prec_eff / (prec_eff + prec_prior + eps) # Eq.(7) overwrite prob
take = torch.rand_like(p_meas) < p_meas # Eq.(8) stochastic overwrite
M_new = torch.where(take, m_meas, M_prior)
P_new = torch.where(take, R_eff, P_prior)
tau_min=0.5 is exactly the 0.5 coefficient in Eq. (6); the validity gate matches the paper's "update only if the effective variance is not significantly larger than the prior"; and the repo also ships LivoxGlobalProbWinnerPipeline — so the struck-through "Mid-360 lidar mapping in the loop" is in fact present at the code level, with only the learned mapping network weights missing.
Experiments
4.1 Comparison with Prior Work
The paper compares against three lines of prior work: the most generalizable predecessor AME-1 [15], the quadruped generalist agile predecessor (end-to-end visual recurrent policy) [48], and a biped agile predecessor [17]. The comparable metric is the maximum climb-up/down height.
| Method | Quad climb up/down | Biped climb up | Biped climb down | Sparse terrain | Generalization |
|---|---|---|---|---|---|
| AME-1 [15] (generalization SOTA) | ~0.5 m | ~0.3 m | ~0.3 m | Yes | No |
| Quadruped generalist agile SOTA [48] | 1 m | — | — | Yes | Few-shot finetune |
| Biped agile SOTA [17] | — | 0.5 m (H1) | ≥0.5 m (H1) | No | No |
| Ours | 1 m | 0.48 m (TRON1) / 0.65 m (H1) | 0.88 m (TRON1) / 1 m (H1) | Yes | Zero-shot |
The paradigm-level difference is equally key: this work uses unified rewards across quadruped and biped, replaces privileged and classical mapping with neural mapping, and uses the same mapping pipeline for training and deployment (most predecessors use $\neq$), with only 2 stages/policies instead of a dozen. The generalist student in [48] generalizes to unseen environments via "few-shot finetuning", whereas this work is zero-shot. Since comparing TRON1 with the larger, stronger H1 may be unfair, the authors also apply the framework to H1 and report simulation results.
Fig. 7: The ANYmal-D controller zero-shots the hardest parkour and rubble terrains reported by prior work [32,48] (paper Fig. 8). The paper trains only on primitive terrain yet directly crosses terrain that prior work needed finetuning for.
4.2 Real-World Parkour and Sparse Terrain
On a parkour course included in neither the controller's nor the mapping model's training terrain, ANYmal-D traverses stably at up to 2 m/s, composing climbing and jumping. On sparse terrain both generations show their strengths: ANYmal-D crosses a 19 cm balance beam, two unfixed beams with gaps, and two rows of 20 cm stepping stones; TRON1 crosses two unfixed 19 cm floating blocks, a beam-then-gap, and diamond stepping stones — several of which are unseen in training.
Fig. 8: ANYmal-D and TRON1 traversing diverse sparse terrains (paper Fig. 10), including a balance beam, floating blocks, beams with gaps, and stepping stones, several unseen during training.
4.3 Biped Maneuvering and Robustness
To show omnidirectional perceptive locomotion, the authors command TRON1 through a 38 cm platform, a gap, a staircase, and rough terrain with smooth motions in all directions. On robustness, the robot climbs onto and balances on an unlocked tiltable platform cart; ANYmal-D also recovers from unfixed tilted stepping stones using its knees.
Fig. 9: TRON1 over a 38 cm platform, a gap, a staircase, and rough terrain (paper Fig. 11); orange arrows indicate approximate trajectories, showing omnidirectional perceptive locomotion.
4.4 Mapping Quality
After traversing complex terrain unseen during training, the estimated elevation grid faithfully reconstructs steps, boxes, and inclined surfaces, while regions never covered by cameras and occluded regions are marked with high uncertainty. In the visualizations, meshes are elevation estimates and thin lines are uncertainties; the red-blue colormap is only used to enhance height contrast.
Fig. 10: Maps obtained after ANYmal-D traverses terrains unseen in training (paper Fig. 13). Left column is the real terrain, right is the estimated elevation mesh, with never-covered regions excluded from coloring.
4.5 Emergent Behaviors and Interpretable Attention
Because foothold positions are no longer explicitly rewarded, and because biokinetic thresholds are used for termination, the controller spontaneously learns two classes of whole-body behavior: whole-body contact — ANYmal-D actively uses knee contacts to stabilize itself and to climb hard terrain; and impact absorption — descending a platform it touches down gently with its knee, or retracts its leg before landing to absorb the impact. The authors argue whole-body contact will matter increasingly for whole-body dexterity in complex environments, but caution that most legged robots are designed for foot contacts, so such motions stress hardware.
Fig. 11: Emergent impact-reduction behaviors (paper Fig. 14). (a) ANYmal-D supports the body with its knee and touches down gently when climbing down; (b) TRON1 extends a leg when jumping down and retracts it before landing, then switches support legs.
Attention visualization gives a direct clue to where the generalization comes from: local attention tends to concentrate on future contact regions, while the global context focuses on a sparse set of characteristic points that capture terrain-level structure. This pattern appears on both training and test terrain, implying transferable representations. On training terrain, attention varies with terrain type: during stair descent, local attention targets steppable regions and the global context attends to the next stair levels; during obstacle avoidance, local attention targets future footholds and the global context highlights obstacles and free space; during climbing, local attention covers the front-foot landing and hind-knee support regions while the global context indicates the upcoming climb. On test terrain these patterns are reused and recombined: on Test 2 the global points move to the platform top, and local attention shifts from future footholds to knee-support regions. The authors also admit that local attention becomes smooth on nearly flat contact regions, because the features are not supervised with contact labels.
Fig. 12: Feature patterns of the ANYmal-D controller (paper Fig. 15). Each panel compares local attention and global context; higher-intensity red means higher weight. Local attention provides fine-grained locomotion control patterns, while the global context focuses on sparse characteristic points of the terrain type.
4.6 Ablations and Benchmarking
All ablations are run on ANYmal-D in simulation. Teacher architectures are compared against AME-1 and MoE; student designs against 7 variants. The conclusions: all architectures scale well on training terrain, but MoE generalizes very poorly, AME-1 does fine on purely sparse terrain (Test 1) yet struggles on mixed terrain (Tests 2-4), which the authors attribute to AME-1's attention query depending only on proprioception and thus missing global context, whereas this paper's teacher generalizes reliably on all test terrains.
On the student side: direct teacher supervision (action distillation + representation alignment) already meets the bar on training terrain but does not generalize strongly; RL matters for test terrain and the representation loss adds more; RL without teacher supervision makes exploration hard. Against the "end-to-end visual recurrent student [48]": the latter is slightly better on Test 3 (pure obstacle parkour) but worse on every test terrain containing sparse regions — on Test 3, most of this paper's failures occur when the robot attempts to climb an obstacle that it misreads as a higher platform due to occlusion. Additionally, a student without uncertainty inputs performs worse, confirming that mapping uncertainty is useful for locomotion.
| Mapping method | Dense | Climbing | Sparse | Test1 | Test2 | Test3 | Test4 | Test avg. |
|---|---|---|---|---|---|---|---|---|
| Ours | -0.006 | -0.108 | 0.020 | 0.033 | 0.227 | 0.036 | -0.111 | 0.046 |
| Ours (Loco-Only) | 0.020 | -0.090 | 0.081 | 0.104 | 0.234 | 0.087 | -0.075 | 0.088 |
| Temporal Recurrent | -0.162 | -0.114 | -0.101 | 0.316 | 0.135 | 0.066 | -0.176 | 0.085 |
Table: $L_{0.5}$ loss of different neural mapping methods on training/test terrain (paper Table III), lower is better. Ours is best on the test average (0.046).
Two more mapping-side comparisons: first, training the mapping model only on locomotion terrain — test error grows markedly, showing that additional traversable and untraversable terrain diversity aids generalization; second, a temporal recurrent mapping model — it must be trained with the policy in the loop, converges ~5× slower, and reaches a Test 1 loss of 0.316; the authors observe that it "overfits training terrain, produces less meaningful uncertainty maps, and loses fine-grained detail".
Fig. 13: Qualitative comparison between the temporal recurrent model and this paper's mapping pipeline (paper Fig. 17). Ours estimates accurately when confident and assigns high uncertainty to occluded regions; the recurrent model produces less meaningful uncertainty maps, loses fine-grained detail, and performs worse on unseen terrain.
Robustness experiments further show: under missing points and artifacts beyond the training randomization range, the policy's success rate does not drop and is sometimes slightly higher (fewer points → more uncertainty → more conservative behavior); disabling the front upper camera still leaves most terrains working well, but interactive-perception terrains (Test 2 and Test 3 with high obstacles) degrade noticeably, showing these behaviors remain sensitive to the available viewpoints.
4.7 Key Numbers
| Item | Value |
|---|---|
| Control rate / joint tracking | Policy 50 Hz; joint PD 400 Hz |
| Training iterations / parallel envs | Teacher 80k, student 40k; 4800 envs |
| Training cost | ANYmal-D ~60 RTX-4090-days (8 GPUs); TRON1 ~30 RTX-4090-days (4 GPUs) |
| Mapping training frames / convergence | 54M frames per robot; converges in <1 hour |
| Mapping inference | Sim: 1000 envs <0.3 ms, ~3 GB; HW: whole pipeline ~5 ms, model ~2.5 ms |
| Policy inference (hardware) | ~2 ms on Intel Core i7-8850H |
| Local grid | ANYmal-D 51×31, TRON1 31×31, both 4 cm |
| Real-world limits | Parkour 2 m/s; ANYmal-D climbs 1 m; TRON1 climbs 0.88 m, 38 cm platform |
Limitations
Authors' three: (1) it uses a 2.5-D elevation map and does not address fully 3-D locomotion; future work could use multi-layer elevation maps or attention-based voxel representations; (2) the controllers are not designed for severely degraded perception (high grass, snow) and could be combined with robust controllers into a single scene-aware policy; (3) the mapping module can fail in highly dynamic, heavily occluded environments and should reason explicitly about moving elements.
Our assessment: (1) training cost is high (~60 GPU-days for the quadruped), and the authors acknowledge that attention over the whole map dominates compute with sparse attention only proposed, not measured. (2) The ablation shows an end-to-end visual policy is still slightly better on pure obstacle parkour (Test 3), and most of this paper's student failures occur at skill transitions (e.g., decelerating in a sparse region then climbing), so the ceiling of zero-shot generalization is not reached. (3) Whole-body contact is shown to help agility/robustness, but existing robots are mainly designed for foot contacts and knee/torso landings add hardware stress — the paper only points to the direction, without long-term durability data. (4) Generalization relies on training-terrain coverage plus heavy randomization; the four test terrains are still compositions of the same primitive categories, and there is no evidence for out-of-distribution terrains (e.g., soft deformable ground).
Summary and Outlook
The paper gives a clear path to both agility and generalization: do not push perception straight into a black-box policy; instead keep an interpretable, reusable, parallelizable elevation map; on the policy side, give the map encoder both local contact detail and global terrain context, and let the latter modulate the former; then use a unified reward and teacher-student recipe so the same method, unchanged, applies to a quadruped and a biped. The results are zero-shot climbing of 1 m ledges, parkour at 2 m/s, and spontaneously emerging knee support and landing absorption. Two counter-intuitive engineering details are worth remembering: removing foothold rewards and penalizing only all links instead induces whole-body contact; and using dog/human thigh-acceleration thresholds for termination forces impact-absorption behaviors to emerge. Looking ahead, the authors point to multi-layer elevation maps and attention voxels, combination with robust controllers, and explicit reasoning about dynamic occluders.
Quotes
"The policy must reason not only about where to step, but also about which contacts to make or avoid."
"Repeated observations of the same occluded or uncertain area should not reduce uncertainties simply due to consistent predictions."
"For generalized locomotion, explicitly modeling uncertainty is important: occluded regions are not completed by learned priors but kept uncertain, and newly observed geometry can be integrated into the map based on the predicted uncertainty once it becomes visible."

