PAPER DEEP DIVE
PredActor: Predictive Action Diffusion for Steerable Onboard Humanoid Control
Diffusion models provide a flexible framework for motion generation, but turning this flexibility into closed-loop humanoid control remains challenging. Hierarchical generator-tracker systems steer motion through reference trajectories, yet these references may exceed the capabilities of the downstream tracker, leaving physical feasibility and disturbance recovery largely to a separate control module. Action-only diffusion avoids this separation by directly generating executable actions, but provides no explicit future-state trajectory that can be steered toward test-time motion objectives. Joint state-action diffusion offers a natural alternative, but existing controllers often rely on privileged full-body states, while learned behavior selection and test-time motion steering remain only partially integrated. We present PredActor, a predictive action diffusion policy that unifies both steering modes in one directly executed policy using proprioception alone. Given proprioceptive history and optional task context, PredActor jointly predicts actions and an internal future-state trajectory that enables guidance: classifier-free guidance strengthens text-conditioned motion, while classifier guidance steers future states toward test-time objectives. Only actions are executed, requiring neither a motion-reference tracker nor privileged full-body states. In simulation, PredActor reaches 44 of 45 destination targets and achieves a text retrieval score of 0.539 versus 0.424 for conditional action diffusion, with similar disturbance survival. Rolling denoising and computation-preserving runtime optimizations reduce the complete callback to 16.790 ms median and 19.383 ms p95 on a Jetson Orin NX, within the 20 ms control period. Deployed on a Unitree G1, PredActor demonstrates text-conditioned motion, disturbance response, joystick control, and semantic interpolation in simulation and hardware.
PredActor: Predictive Action Diffusion for Steerable Onboard Humanoid Control
Lei Ye, Haibo Gao, Yitang Li, Peng Xu, Zetong Jing, Junhan Sun, Fanrong Dong, Ziqi Han, Xue Wang, Jianhua Sun, Cewu Lu, Hao Zhao, Liang Ding (Harbin Institute of Technology / Shanghai Innovation Institute / RoboParty Lab / Tsinghua University / Shanghai Jiao Tong University; corresponding authors Haibo Gao, Peng Xu, Jianhua Sun, Liang Ding)
arXiv:2609.24840 [cs.RO] (primary cs.RO, cross-listed cs.LG), v1 submitted 2026-09-21, revised to v3 on 2026-10-05 · arXiv:2609.24840 · project page · GitHub: MasterYip/PredActor
Code status: evaluation code is public (MIT). The repository currently ships a browser-based MuJoCo evaluation harness, the G1 assets, and the released PDP051 checkpoint (weights hosted on HuggingFace at MasterYip/PredActor_Artifacts; uv sync --locked then predactor-eval starts a local server, with no Isaac Sim and no training repo required). In the README release checklist, four items are still unchecked ([ ]): data collection and annotation, behavior-cloning training, DAgger-style interactive aggregation, and onboard deployment. In other words, neither the training recipe nor the real-robot half is third-party reproducible today.
In One Sentence
PredActor compresses "predict future states" and "emit actions directly" into a single proprioception-only diffusion policy: one interleaved state-action transformer denoises 20 future states and 20 action chunks in the same forward pass, and the predicted states never leave the policy. They exist purely as a differentiable steering bus, where classifier guidance pushes them toward test-time goals and classifier-free guidance sharpens the text condition, while the only thing sent to the 29 joints is the action. No motion-reference tracker, no privileged full-body state estimator. With rolling denoising plus a set of computation-preserving runtime optimizations, the complete control callback lands at 16.790 ms median and 19.383 ms p95 on a Jetson Orin NX, inside the 20 ms control period for 593 of 600 calls, which the authors present as the first fully onboard, 50 Hz joint state-action diffusion policy on a Unitree G1.
1. The Problem: Three Diffusion-Control Paradigms, Each Missing One Piece
Figure 1 (paper Figure 2): the paradigm spectrum of diffusion-based humanoid control. (a,b) Generator-tracker systems hand generated motion to a separate tracker for execution; (c) action-only diffusion predicts actions directly but has no future state to guide; (d) joint diffusion predicts states and actions but needs extra state estimation; (e) PredActor executes actions directly while its internal future-state prediction carries both CG and CFG steering. O denotes the observation.
What makes diffusion attractive for action generation is its multimodality: one observation window can admit several reasonable action sequences, and the sampling process delegates the choice to conditioning and guidance. Turning that flexibility into closed-loop humanoid control, however, does not fail at generation quality. It fails at the seam between generation and execution, specifically at the question of who guarantees physical feasibility once the generated motion is handed off. The paper sorts prior work along exactly this axis and shows that each family is missing something.
The first family is the hierarchical generator-tracker. A kinematic generator (diffusion or flow matching) produces a reference motion, and a physics-trained tracker executes it. Open-loop variants never see robot feedback; closed-loop variants such as CLoSD and ReactiveBFM feed executed states back into later planning windows. Two structural problems follow directly. References can land outside the tracker's support, which is to say the generator has no model of what the tracker can actually do. And when a disturbance arrives, the system's response is mostly the tracker pulling state back toward the reference rather than adapting the behavior itself. Running two clocks, one for planning and one for control, also complicates synchronization and replanning.
The second family is action-only diffusion, the Diffusion Policy and DiffuseLoco line, which maps observations and task conditions straight to actions. Dropping the reference-tracking layer buys better closed-loop behavior, but the price is that predictions contain no explicit future-state trajectory. A test-time objective like "send the robot to that point" then has nothing differentiable to attach to, because classifier guidance needs a state quantity to take gradients with respect to.
The third family is joint state-action diffusion, the idea Diffuser introduced and Diffuse-CLoC, BeyondMimic, SCDP, and SCRIPT implement in different ways. It restores the predictive representation, so state-level shaping can influence actions. The paper identifies two imbalances here: representative controllers tend to condition on privileged full-body states, which forces sim-to-real to carry an extra deployable state estimator; and the two guidance interfaces are unevenly supported, since CFG amplifies a learned behavior condition while CG pushes predicted motion toward a test-time objective, yet CFG is uncommon among representative joint controllers.
Rather than stating a benchmark target, the authors phrase the open problem as a functional question: can a directly executed policy, using only deployable observations, combine learned behavior selection with test-time motion steering while retaining closed-loop disturbance robustness? PredActor's answer is to keep the state prediction inside the policy as an interface and keep executing actions directly.
Table 1 in the paper lines up 14 systems against those three paradigms. Its persuasive force is not in any single row but in the last one, which is the only row in the table with all five capability columns filled:
| Method | Paradigm (Fig. 2) | Proprioception only | CFG | CG | Text | Real robot | Inference compute | Control rate |
|---|---|---|---|---|---|---|---|---|
| MotionBricks + SONIC | generator + tracker (a) | partial | no | no | no | yes | Orin (onboard) | 50 Hz tracker |
| ARDY + SONIC | open-loop tracking (a) | partial | partial | no | partial | no | RTX 4090 | - |
| TextOp | open-loop tracking (a) | partial | yes | no | yes | yes | RTX 4090 (offboard) | 50 Hz tracker |
| RoboGhost | latent generation + tracker (a) | yes | no | no | yes | yes | Orin NX (partial) | 50 Hz tracker |
| ReactiveBFM | closed-loop tracking (b) | partial | no | no | yes | yes | RTX 4090 (offboard) | 50 Hz |
| CLoSD | closed-loop tracking (b) | no | yes | no | yes | no | - | N/A |
| DiffuseLoco | action diffusion (c) | yes | no | no | no | yes | RTX 4060M (onboard) | 30 Hz target |
| SENTINEL | action flow + residual (c) | yes | yes | no | yes | yes | RTX 4090 (offboard) | 50 Hz |
| Diffuse-CLoC | joint diffusion (d) | no | no | yes | no | no | RTX 4060 (sim) | N/A |
| BeyondMimic | latent joint diffusion + decode (d) | partial | no | yes | no | yes | RTX 4060M (onboard) | 25 Hz |
| SCDP | joint diffusion (d) | yes | no | no | no | yes | RTX 5090 (offboard) | 50 Hz |
| SCRIPT | joint diffusion (d) | no | yes | no | yes | no | - | N/A |
| UniPhys | latent joint diffusion + decode (d) | no | yes | yes | yes | no | A100 (sim) | N/A |
| PredActor | joint diffusion (e) | yes | yes | yes | yes | yes | Orin NX (onboard) | 50 Hz |
Table 1 (condensed from the paper's capability matrix): PredActor is the only entry that is proprioception-only, supports CFG and CG and text conditioning, runs on real hardware, and does all of it on an onboard Orin NX at 50 Hz.
Reading the table honestly matters, because "the only row with everything checked" is a claim about interface coverage, not about winning every metric. Several entries marked N/A simply do not report the axis at all. The paper is careful to say that behavioral execution and control-period timing are evaluated separately from interface availability, which is the right way to keep the two claims from blurring into each other.
2. Preliminaries: Noise Levels, DDIM, and Two Guidance Channels
Two pieces of diffusion machinery carry the whole design. First, the noise level is indexed rather than scalar: each position in the horizon carries its own level $k\in\{0,\ldots,T\}$, so a single buffer can hold entries that are simultaneously half-denoised at the tail and nearly clean at the head. Second, DDIM gives a deterministic jump between levels, which is what makes it legal to carry a partially denoised trajectory from one control tick into the next.
On top of that sit two guidance interfaces with genuinely different jobs, and the paper keeps the distinction sharp. Classifier-free guidance (CFG) mixes conditional and null predictions to amplify a learned behavior condition such as a text prompt; it needs no external objective and no gradients through a cost. Classifier guidance (CG) takes the gradient of a test-time objective $\mathcal{J}(\hat{\mathbf{s}}^{0};\mathbf{g})$ parameterized by steering commands $\mathbf{g}=[g_{1},\ldots,g_{r}]$ and applies it to the predicted clean state before the reverse update. The authors spell out the consequence that most readers would otherwise miss: CG leaves the current action prediction unchanged, and the updated state influences the action only at the next prediction through state-to-action attention. CFG reinforces what was learned; CG steers where the prediction is going.
3. Method: One Policy, Two Steering Channels
3.1 Formalization: joint denoising conditioned on deployable observations only
The policy is defined as a single conditioned denoiser over a state stream and an action stream, each with its own per-position noise level:
$$(\hat{\mathbf{s}}^{0},\hat{\mathbf{a}}^{0})=F_{\theta}(\mathbf{s}^{\mathbf{k}_{s}},\mathbf{a}^{\mathbf{k}_{a}},\mathbf{k}_{s},\mathbf{k}_{a};\mathbf{o}_{t-\ell+1:t},\mathbf{z}) \tag{1}$$
Here $\mathbf{k}_{s},\mathbf{k}_{a}\in\{0,\ldots,T\}^{h}$ are the per-position noise levels, $T$ is the maximum diffusion level, $h=20$ is the horizon, $\mathbf{o}_{t-\ell+1:t}$ is the proprioceptive history window, and $\mathbf{z}$ is optional task context. The asymmetry between what conditions the model and what the model predicts is the whole point: full-body states supply training targets, while deployment requires only proprioceptive observations plus optional task context.
The evaluated tensors are concrete. The observation is 96-dimensional: 29 joint positions, 29 joint velocities, 3 projected-gravity components, 3 IMU channels, 29 previous actions, and 3 base linear velocity terms. The physical state is 192-dimensional, and in the released repository its layout is fixed by forward kinematics: bodies times 3 positions at [0:90], bodies times 3 linear velocities at [90:180], then root position, rotation, linear velocity and angular velocity at [180:192]. That 192-dimensional physical state is mapped to a 384-dimensional model representation by a fixed emphasis projection $E$ applied before the backbone and its pseudoinverse $E^{\dagger}$ applied after prediction. The evaluated transform concatenates a weighted symmetric random projection with an identity component; it is specified before training and is not learned by gradient descent, so it changes representation dimension without adding parameters or a trainable failure mode. Actions are 29 joint-controller commands converted to PD targets by the runtime, and the text-conditioned setting uses a 512-dimensional MotionCLIP semantic embedding. Normalization statistics are fitted on the training split and frozen at test time.
One sentence in the paper deserves to be pulled out because it is the design commitment everything else depends on: predicted states remain internal guidance variables, and only the selected action is issued to the joint controller. There is no second module, so there is no interface at which a generator can out-run a tracker.
3.2 The interleaved state-action transformer and its asymmetric attention
Figure 2 (paper Figure 3): conditioning, denoising and sampling from top to bottom. The encoder maps observation-history tokens and task tokens into a condition memory C; separate projections embed noisy state and action tokens with their noise levels and horizon positions; heads produce clean predictions for both streams.
Both streams are flattened into one token sequence with strictly alternating state and action slots:
$$[s^{\mathrm{tok}}_{1},a^{\mathrm{tok}}_{1},\ldots,s^{\mathrm{tok}}_{h},a^{\mathrm{tok}}_{h}] \tag{2}$$
The backbone is deliberately small: 2 layers, 4 heads, 256 dimensions. Its attention pattern is inherited from Diffuse-CLoC and is asymmetric in a way that encodes a causal claim about the robot rather than an engineering convenience. State queries attend to all state tokens but never to action tokens; action queries attend causally to both streams. So information flows one way, state into action, and the state stream is never polluted by the actions the model is simultaneously inventing. PredActor extends this backbone with cross-attention to the condition memory $\mathbf{C}=[c_{1},\ldots,c_{m}]$, which is how the encoded proprioceptive history and task context reach the denoising tokens, and separate heads emit the clean predictions $\hat{\mathbf{s}}^{0}$ and $\hat{\mathbf{a}}^{0}$.
In the released code the asymmetry is not a comment, it is mask arithmetic. PredActor/backbone/transformer_codiffuse.py lines 143-154 build odd and even token-index blocks and set the state-to-action direction to x_to_y_attn: no_attn, which is exactly the "state queries see states only" rule above. That is a useful thing to verify when reading a paper whose central mechanism is an attention pattern: the mask is checkable in a few lines.
3.3 The two steering channels: CFG and CG
Per diffusion level $t_{i}$, the backbone produces clean action and state predictions once:
$$(\hat{a}_{i}^{0},\hat{s}_{i}^{0})=F_{\theta}(a_{i},s_{i},t_{i},c_{\mathrm{enc}}) \tag{3}$$
When guidance is active, that same state prediction is adjusted by an analytic whole-body cost and then reused in the DDIM update:
$$\tilde{s}_{i}^{0}=\hat{s}_{i}^{0}-G(\hat{s}_{i}^{0};u_{v_{x}},u_{v_{y}},u_{v_{z}}),\qquad(a_{i-1},s_{i-1})=\mathcal{D}(a_{i},s_{i},\hat{a}_{i}^{0},\tilde{s}_{i}^{0};t_{i},t_{i-1}) \tag{4}$$
Two consequences follow, and both are about budget. Because $\hat{a}_{i}^{0}$ is not recomputed at level $t_{i}$, the evaluated path performs one backbone evaluation per diffusion step, or two in total for the two-step sampler. A variant that duplicated the backbone would run four forwards and would also change the guidance equation, so the paper excludes it from the exact acceleration ablation instead of quietly counting it as a fast configuration. Second, since the modified state enters $s_{i-1}$, the steering effect reaches the action only through the next backbone evaluation via state-to-action attention. Guidance is therefore one tick delayed by construction, which is a real property of the closed loop rather than an implementation artifact.
CFG needs no such gradient path: conditional, null, and partial-condition training gives the model a null prediction to mix against, which is what amplifies text-conditioned motion at inference.
The repository makes the CG side unusually inspectable. PredActor/modules/classifier_guidance.py lines 71-123 wrap the cost in torch.enable_grad() and call autograd.grad, with a guidance_scale playing the role of $\lambda$. config_files/g1_cond_diffuse.yaml sets guidance_mode: classifier under a key the authors label "paper-correct", routes vx_guide to the pelvis-torso cost and vy/wz/hz to the torso cost, and for the destination task uses proportional gains kp_x = 2.0, kp_y = 0.8, kp_w = 6.0 with an arrival_radius of 0.30 m. Those numbers are the difference between "we use classifier guidance" and a claim someone else can rerun.
3.4 Data and training: turning a labeled motion library into deployable observation windows
Figure 3 (paper Figure 4): (a) the motion library is built from generated motion and AMASS, BABEL supplies task labels, and MotionCLIP encodes them as conditioning latents; (b) an adaptive sampler feeds references into an asynchronous rollout pipeline where a frozen tracking teacher executes each reference under observation noise, action noise and external pushes.
Supervision comes from a frozen tracking teacher, and it exists only at training time. The teacher executes each reference under observation noise, action noise, and external pushes, seeding $\mathcal{D}_{0}$ with observation histories, future states, teacher actions, and task labels. Training then follows a DAgger-style iterative aggregation that alternates frozen-student rollout collection, isolated teacher trajectory shooting, and continued diffusion reconstruction. At round $j$ the student $\pi_{\theta_{j}}$ induces rollout origins $\mathcal{R}_{j}$, and from each origin $\xi$ an isolated teacher branch follows the aligned reference to produce a target state-action window $W_{E}(\xi)$ paired with the student's own history and task label:
$$\begin{aligned}\mathcal{D}_{j+1}&=\mathcal{D}_{j}\cup\{W_{E}(\xi):\xi\in\mathcal{R}_{j}\},\\ \theta_{j+1}&\approx\arg\min_{\theta}\mathcal{L}(\theta;\mathcal{D}_{j+1})\end{aligned} \tag{5}$$
The word "isolated" is doing real work. Teacher trajectory shooting happens in a separate simulation branch, so the teacher's actions enter the dataset as labels and never as interventions in the student's own rollout. That keeps the aggregated windows consistent with what the student will actually see at deployment, which is the usual failure mode of naive DAgger on a physical plant.
The objective is a three-term reconstruction loss over independently corrupted windows:
$$\begin{aligned}\mathcal{L}(\theta;\mathcal{D})=\mathbb{E}\bigg[&\sum_{i=1}^{h}w_{i}^{a}\lVert\hat{\mathbf{a}}_{i}^{0}-\mathbf{a}_{i}^{0}\rVert_{2}^{2}+\rho\sum_{i=2}^{h}\lVert\Delta\hat{\mathbf{a}}_{i}^{0}-\Delta\mathbf{a}_{i}^{0}\rVert_{2}^{2}\\ &+\gamma_{s}\sum_{i=1}^{h}w_{i}^{s}\lVert\hat{\mathbf{s}}_{i}^{0}-\mathbf{s}_{i}^{0}\rVert_{2}^{2}\bigg]\end{aligned} \tag{6}$$
The middle term is a consecutive-difference penalty on actions, which is what buys the smoothness the jerk column later measures; the third term supervises the internal future-state trajectory that guidance steers. Iterative aggregation changes which windows enter the replay buffer without replacing this reconstruction objective, so the state stream stays supervised by the same loss across all rounds.
Scale and hyperparameters are reported concretely: 152 ACCAD motions times 50 rollouts gives 7,600 episodes; 50 epochs at learning rate 1e-4 with cosine decay and 10k warmup steps, batch size 256, sampling rate 0.01, and three seeds (92025, 92026, 92027). Three seeds is what makes the checkpoint-mean-plus-SD reporting in Table 2 meaningful rather than decorative.
3.5 Rolling inference: schedule-matched reuse and delay compensation
Figure 4 (paper Figure 5): rolling denoising. Each horizon position can take its own DDIM jump, source and target levels are given by matrices S and S-prime, and only newly exposed tail positions receive fresh noise.
Restarting the full horizon every control tick is what makes naive diffusion control too slow for 50 Hz. PredActor instead carries a partially denoised buffer forward using the DDIM update
$$\hat{\epsilon}_{t}=\frac{u^{t}-\sqrt{\bar{\alpha}_{t}}\hat{u}^{0}}{\sqrt{1-\bar{\alpha}_{t}}},\qquad u^{t^{\prime}}=\sqrt{\bar{\alpha}_{t^{\prime}}}\hat{u}^{0}+\sqrt{1-\bar{\alpha}_{t^{\prime}}-\sigma_{t}^{2}}\hat{\epsilon}_{t}+\sigma_{t}\epsilon \tag{7}$$
$$\sigma_{t}=\eta\sqrt{\frac{1-\bar{\alpha}_{t^{\prime}}}{1-\bar{\alpha}_{t}}}\sqrt{1-\frac{\bar{\alpha}_{t}}{\bar{\alpha}_{t^{\prime}}}} \tag{8}$$
The evaluated setting takes $\eta=0$, hence $\sigma_{t}=0$ and the stochastic term disappears. Determinism is not a stylistic choice here: it is precisely what licenses carrying intermediate samples across control steps instead of resampling the whole horizon every tick. Source and target levels per horizon position are written as matrices $S,S^{\prime}\in\mathbb{Z}^{K\times h}$, broadcast along the feature dimension at each position, and the state and action streams may use different matrices. At runtime the policy reuses a stored intermediate sample whose source level exactly matches the level the next step needs, adds fresh noise only to the newly exposed tail positions, and denoises the shifted horizon under the latest observation. The paper names this property schedule-matched reuse and states its negation just as clearly: copying a clean prediction into an arbitrary noise level is wrong, because it falsifies the sample's diffusion coordinate.
Both pieces are locatable in the release. In PredActor/diffusion_policy/utils/noise_scheduler.py, get_denoising_matrix() returns the $(K,H)$ noise-level matrix used by both the DDPM path and the rolling DDIM path, and get_rolling_traj() extracts the $(B,H-1,D)$ rolling buffer of partially denoised trajectories, raising an error rather than silently degrading when no matching snapshot exists ("rolling trajectory has no matching denoising snapshot for ..."). build_ns_ddim_steps() constructs the NS-DDIM step sequence used by the evaluated profile, carrying four per-position levels (s_t_cur, s_t_prev, a_t_cur, a_t_prev) per step.
Execution adds one problem that only exists on hardware: the delay $\delta$ between "the policy computed an action" and "the motors move". PredActor does not build a separate delay model. It defines a delay coordinate in the same horizon frame, $u=(n_{\mathrm{obs}}-1)+\delta/\Delta t$, and linearly interpolates the two adjacent clean action slots to obtain delay-compensated PD targets. Lines 43-92 of utils/action_selection.py implement exactly this as coordinate = past_step + delay with past_step = n_obs_steps - 1, plus two validations: the interpolation endpoints must fall inside executable slots, and a fractional delay requires the terminal action noise level to be 0 (terminal_action_noise_levels()). So the silent error of interpolating a slot that is not yet fully denoised is blocked at the code level, not merely discouraged in the paper.
3.6 Onboard execution: C++ export and safety boundaries
Deployment exports the diffusion actor and its control interface to C++. The export artifact contains the TorchScript backbone, the diffusion schedule, normalization statistics, the actor configuration, and the deployment configuration. The C++ runtime normalizes observations and conditions, runs the actor, denormalizes the action trajectory, and interpolates the delay-compensated PD targets using the coordinate above. External conditions (text commands, joystick values) are latched by an asynchronous UDP receiver under a mutex, which decouples condition transport from the control callback so a late command cannot stall the 50 Hz tick. A Passive-Ready-Running state machine gates policy actions and falls back to Passive on a fall or torque fault. The evaluated Orin profile is a batch-one FP32 conditioned policy with two-step deterministic DDIM, packed self- and cross-attention projections, cached command embeddings, a precomputed schedule, stacked-axis whole-body guidance, eager finite checks, and LibTorch inference mode.
flowchart TB
OBS["Proprioceptive history o_t-l+1:t
96 dim: 29 q + 29 dq + 3 gravity + 3 imu
+ 29 prev action + 3 base lin vel"] --> ENC["Condition encoder
to memory C of m tokens"]
TXT["Optional task context z
512 dim MotionCLIP text embedding"] --> ENC
JOY["Test-time steering g
joystick velocity or destination"] --> COST["Analytic whole-body cost J
per-DOF guidance, stacked axes"]
ENC --> XATTN["Interleaved transformer 2 layers 4 heads 256 dim
tokens s1 a1 ... s20 a20
state queries see states only
action queries see both causally"]
XATTN --> PRED["One forward per level
clean preds s0_hat and a0_hat (Eq 3)"]
PRED --> CFG["Classifier-free guidance
mix conditional and null preds
amplifies the text condition"]
CFG --> CG["Classifier guidance
s_tilde = s0_hat - G(s0_hat; g) (Eq 4)
action pred NOT recomputed"]
PRED --> CG
CG --> DDIM["DDIM jump to target levels (Eq 7)
eta = 0 so sigma_t = 0, deterministic"]
DDIM --> ROLL["Rolling buffer shift
schedule-matched reuse, S and S_prime in Z^(K x h)
fresh noise only on newly exposed tail"]
ROLL -->|next control tick| XATTN
ROLL --> SEL["Delay compensation
u = (n_obs - 1) + delta / dt
interpolate two adjacent clean action slots"]
SEL --> PD["29 PD targets to the joint controller at 50 Hz
no motion reference tracker, no privileged state"]
PD --> ROBOT["Unitree G1 on Jetson Orin NX
callback p50 16.790 ms, p95 19.383 ms"]
ROBOT -->|new proprioception| OBS
ROBOT --> SAFE["Passive-Ready-Running state machine
fall or torque fault falls back to Passive"]
TEACHER["Frozen tracking teacher
training-time data source only"] -.->|labels, absent at deployment| XATTN
Figure 5: PredActor's closed-loop data flow. Solid edges are the inference path traversed once every 20 ms at deployment; the dashed teacher exists only during training. There is a single steering injection point, CG, and it modifies states rather than the current action, so its effect reaches the joints one denoising step later through state-to-action attention.
4. Experiments: The Paper's Own Three Questions
The evaluation is organized as three questions rather than three leaderboards. Q1: can PredActor hold destination CG steering and CFG text control at the same time without losing disturbance recovery? Q2: can exact guided inference fit a 20 ms onboard budget? Q3: can one and the same policy serve the simulation interface and run a physical G1? The policy controls 29 joints at 50 Hz and predicts a 20-step state-action horizon throughout.
4.1 Q1: seven policies on one calibrated MuJoCo plant
Seven policies are compared under matched push, destination, and text protocols in a single calibrated MuJoCo plant. Six are learned policies, each trained on 7,600 teacher-rollout episodes with three seeds; ARDY+SONIC is a single-model reference. Metrics cover survival, arrival within 0.6 m, navigation error, MotionCLIP text retrieval, jerk RMS, and latency. The authors' own reproduction of Diffuse-CLoC is labeled CLoC+FK: forward kinematics estimates the privileged-state inputs that method expects, and the suffix names exactly that observation interface. PredActor, by contrast, conditions on proprioceptive history with no FK-based reconstruction.
| Method | Jerk RMS lower is better ($10^{3}\,\mathrm{s}^{-3}$) | Latency lower is better (ms) | Push survival | Nav survival | Arrival rate | Last-valid error (m) | Text survival | Text retrieval |
|---|---|---|---|---|---|---|---|---|
| ARDY + SONIC (single-model reference) | 19.221 | N/E | 0.534 | 1.000 [2/15] | 0.500 [2/15] | 0.656 | 1.000 [23/45] | 0.237 [23/45] |
| DiffuseLoco w TextCFG | 15.899 ±6.870 | 12.141 ±0.145 | 0.563 ±0.009 | N/A | N/A | N/A | 0.947 ±0.047 | 0.424 ±0.159 |
| CLoC+FK | 266.605 ±32.412 | 6.102 ±0.029 | 0.303 ±0.031 | 0.064 ±0.005 | 0.000 ±0.000 | 3.194 ±0.112 | N/A | N/A |
| CLoC+FK w TextCFG | 101.682 ±2.110 | 8.788 ±0.141 | 0.296 ±0.020 | 0.450 ±0.059 | 0.333 ±0.067 | 2.217 ±0.450 | 0.782 ±0.033 | 0.416 ±0.106 |
| PredActor w/o TextCFG | 11.849 ±0.899 | 1.310 ±0.006 | 0.579 ±0.012 | 0.845 ±0.120 | 0.822 ±0.139 | 0.983 ±0.301 | N/A | N/A |
| PredActor w/o state | 12.448 ±0.675 | 1.050 ±0.004 | 0.487 ±0.033 | N/A | N/A | N/A | 0.778 ±0.145 | 0.397 ±0.023 |
| PredActor | 16.810 ±1.409 | 1.321 ±0.011 | 0.511 ±0.030 | 0.983 ±0.030 | 0.978 ±0.038 | 0.605 ±0.029 | 0.888 ±0.068 | 0.539 ±0.075 |
Table 2 (paper Table 2): checkpoint means with sample standard deviations. N/E means not evaluated, N/A means the policy does not support that interface at all. PredActor reaches 44 of 45 destination targets and 0.539 text retrieval over 119/135 eligible windows.
Figure 6 (paper Figure 6): MuJoCo traces for destination steering, text-conditioned motion, semantic interpolation and the push protocol, including direction-matched pelvis tilt under a 300 N push.
The steering result is the cleanest part of Q1. PredActor reaches 44 of 45 targets with a last-valid error of $0.605\pm0.029$ m, and it retrieves 0.539 raw top-1 MotionCLIP prompts over 119 of 135 eligible windows against 0.424 over 126 of 135 for DiffuseLoco w TextCFG. Text survival favors the baseline, 0.947 versus 0.888 over 135 trials, and the paper reports that plainly instead of choosing the metric that flatters it. The eligibility gap is worth understanding rather than glossing: retrieval is only counted over complete windows beginning at least 3 s after command activation, so a trial can contribute to survival and still have no eligible retrieval window.
Both ablations behave the way the architecture predicts, which is the strongest internal-consistency evidence in the paper. Removing CFG text conditioning (PredActor w/o TextCFG) drops arrival from 0.978 to 0.822 and worsens last-valid error from 0.605 to 0.983 m, while leaving push survival slightly higher at 0.579. Removing the state stream entirely (PredActor w/o state) removes the CG interface, so navigation columns become N/A, and text retrieval falls to 0.397 with text survival at 0.778. In other words the state stream is doing the semantic work the paper claims for it, and it is not merely a latency tax.
The ARDY+SONIC row needs a caveat that the paper supplies elsewhere: 22 of 45 generated references are rejected by the tracker's joint-range converter before rollout, and those are retained as non-evaluable rather than scored as failures or zeros. Its 1.000 text survival over 23 of 45 trials is therefore not comparable to a policy evaluated over all 135 trials, and the paper says so.
The matched state-stream ablation is reported separately with paired confidence intervals, under the primary no-CG contract and with whole-body guidance disabled outside the destination task. Parameter counts differ by 0.0161%, so this is a like-for-like retraining rather than a capacity argument:
| Metric (primary no-CG contract) | Matched PredActor (S+A) | PredActor w/o state (A) | Paired effect [95% CI] |
|---|---|---|---|
| Push survival | 0.535 ±0.009 | 0.487 ±0.033 | +0.048 [-0.024, 0.120] |
| Post-push recovery | 0.043 ±0.040 | 0.193 ±0.110 | -0.150 [-0.494, 0.194] |
| Text survival | 0.866 ±0.011 | 0.778 ±0.145 | +0.088 [-0.246, 0.421] |
| Text retrieval | 0.538 ±0.087 | 0.397 ±0.023 | +0.142 [-0.019, 0.302] |
| Jerk RMS | 15.684 ±2.368 | 12.448 ±0.675 | +3.236 [-2.200, 8.671] |
| Latency at batch 5 (ms) | 1.335 ±0.039 | 1.050 ±0.004 | +0.285 [0.199, 0.372] |
Table 3 (paper Table 8): matched-retraining state-stream ablation. Only the latency interval excludes zero, and it is unfavorable to the state stream. Under this contract the destination task is N/A and CG is off.
4.2 Q2: fitting guided diffusion inside 20 ms
Q2 is where the paper is most useful to anyone shipping iterative neural inference on embedded compute. It reports a cumulative acceleration path on one Jetson Orin NX, batch one, FP32, measuring the complete callback rather than a bare backbone forward:
| Composition / boundary | p50 (ms) | p95 (ms) | Saving vs previous (ms) | Calls over 20 ms |
|---|---|---|---|---|
| Structural baseline, sequential WBG | 21.461 | 22.036 | - | 600/600 |
| + stacked-axis WBG | 18.315 | 19.677 | 3.147 | 8/600 |
| + DDIM schedule precompute | 17.442 | 18.347 | 0.872 | 0/600 |
| + all-QKV packing | 16.331 | 17.372 | 1.111 | 0/600 |
| + condition cache | 15.930 | 16.728 | 0.401 | 0/600 |
| Same structure, runtime remeasured | 16.234 | 16.912 | - | 1/600 |
| + inference mode (final actor) | 13.175 | 13.400 | 3.058 | 0/600 |
| Final live callback-to-host | 16.790 | 19.383 | - | 7/600 |
Table 4 (paper Table 4): cumulative onboard acceleration path. The final row is the live callback including host-side transport, which is 3.6 ms above the actor-only measurement two rows up.
Figure 7 (paper Figure 7): the Jetson Orin NX board and the Unitree G1 used for onboard evaluation, with the callback latency distribution against the 20 ms budget.
Three details make this table trustworthy rather than promotional. First, the optimizations are explicitly computation-preserving: output drift stays below $1.073\times10^{-6}$, so what is being reported is the same policy running faster, not a different policy that happens to be quick. Second, a separate exact factorial study crosses four attention-projection layouts with condition-embedding reuse, schedule precomputation, and sequential versus stacked guidance axes over 32 valid combinations, and the paired attributions it yields (stacked WBG 2.845 ms, schedule precompute 1.672 ms, all-QKV packing 0.775 ms, condition reuse 0.190 ms, inference mode 3.166 ms) are stated to be conditional on the other settings and must not be summed. Cross-attention-only packing is reported as not consistently faster, which is the kind of negative result most acceleration tables omit. Third, the percentile columns are not summable either, and the paper marks the E0/E1/E2 boundaries so a reader cannot accidentally add p50s across rows.
The step-count reduction is reported alongside a behavioral cost rather than in isolation. Going from 10 to 2 NS-DDIM steps takes the p50 from 92.011 ms to 21.846 ms, and rolling push survival across those settings is 58% at two steps, 61% at three, and 45% at ten. Two steps is not simply "the fast option"; on this protocol it is also the more robust one, and the paper presents both numbers so the reader can see that the choice is not a straight speed-for-capability trade.
4.3 Q3: six settings on a physical G1
The real-robot evaluation is deliberately narrow and honestly labeled. Five trials per setting, four motion transitions and two push settings, with success defined as completing the transition or remaining upright after the push:
| Setting | Successes | Note |
|---|---|---|
| Stand to walk | 5/5 | joystick steering from rest |
| Walk to jog | 5/5 | velocity command increase while walking |
| Stand to squat | 5/5 | semantic transition |
| Squat to stand | 4/5 | one failure |
| Push while standing | 4/5 | one failure |
| Push while walking | 4/5 | one failure |
Table 5 (paper Table 3): real-robot trial outcomes, five trials per setting. The paper states that these counts are descriptive and attaches no inference to them.
Timing during these trials uses received robot-state messages without issuing actuation commands for the benchmark portion, which keeps the latency measurement from being contaminated by the transition being tested. The combined claim the authors make is modest and specific: 50 Hz onboard control at the measured p95, plus repeatable physical transitions and recovery. Nothing in Table 5 supports a stronger claim, and none is made.
Figure 8 (paper Figure 1): the four capabilities demonstrated inside one closed-loop policy: (a) joystick steering, (b) disturbance-responsive behavior, (c) semantic interpolation, (d) text-commanded control.
5. Limitations
Stated by the authors (Appendix D.3). The guidance interface is calibrated only over the tested objective range, and stronger guidance can move denoising outside the policy's training distribution. Text control is limited to learned semantic categories, because neither the conditioning model nor the coverage of the training data and motion library yet supports reliable interpretation of complex or compositional commands. The training library does not cover sustained multi-contact behaviors such as crawling or climbing. And the timing conclusion is bound to one Orin artifact, batch one, FP32: a different export or device requires requalification rather than rescaling.
My judgment one: the state-stream ablation is mostly inconclusive. In Table 8 five of six paired confidence intervals cross zero. Push survival +0.048 [-0.024, 0.120], post-push recovery -0.150 [-0.494, 0.194], text survival +0.088 [-0.246, 0.421], retrieval +0.142 [-0.019, 0.302], jerk +3.236 [-2.200, 8.671]. The single interval that excludes zero is latency, +0.285 [0.199, 0.372] ms, and it counts against the state stream. So under the primary no-CG contract, the honest reading is that the state stream costs a measurable 0.285 ms and its behavioral advantages are directionally positive but not established. The case for the state stream rests on the CG interface existing at all (the w/o state policy cannot do destination steering, arrival N/A), not on these paired effect sizes.
My judgment two: absolute disturbance robustness is still low. A push survival of 0.511 means roughly half the push trials hit a terminal condition, and these are 100-500 N horizontal pushes over 0.5 s, exactly the magnitudes a humanoid meets in daily operation. Terminals are triggered by root height below 0.55 m, absolute pelvis tilt at least 0.3 rad, or base angular speed at least 1 rad/s, with survival scored as the observed fraction of a 650-step horizon. Calling this "similar" to DiffuseLoco's 0.563 is accurate, but a reader can easily mistake parity for solved. The 0.303 and 0.296 of the two CLoC+FK rows, together with 0.487 for PredActor w/o state, suggest robustness at this magnitude is open for the whole family.
My judgment three: the real-robot sample cannot attribute failures. Five trials per setting with no confidence intervals, and one failure each in squat-to-stand and both push settings. Nothing distinguishes a policy failure from floor conditions, calibration drift, or a conservative safety-stop threshold. The paper does not hide this ("counts are descriptive"), but it does cap the real-robot conclusion at "these transitions are repeatably achievable".
My judgment four: the CLoC+FK baseline may be weakened by its reproduction interface. It is the authors' own implementation of Diffuse-CLoC, with forward kinematics estimating the privileged-state inputs the original method expects. FK reconstruction necessarily introduces error, and the paper cannot separate how much of the 101.682 jerk RMS or the 0.000 arrival rate belongs to the original method versus this substitute observation interface. Naming the variant and its suffix explicitly is good practice; the residual uncertainty still belongs in any side-by-side reading.
My judgment five: reproducibility currently stops at the evaluation layer. Shipping a browser MuJoCo harness and the PDP051 checkpoint already goes further than most humanoid-control papers, and it lets an outsider watch the policy behave on the same simulated plant. But with data collection and annotation, BC training, DAgger aggregation, and onboard deployment all unchecked, a third party cannot reproduce the training recipe (7,600 episodes, three seeds, aggregation rounds) or any part of the hardware result, and the hardware result is the paper's headline claim.
6. Conclusion and Outlook
The contribution separates cleanly into representation, system, and deployment layers. At the representation layer, PredActor shows that an internal predicted-state trajectory can serve simultaneously as the differentiable attachment point for CG and as the semantic carrier for CFG, and that this trajectory never has to leave the policy or become some tracker's reference. The old hierarchical failure mode, a generator proposing motion its tracker cannot execute, does not get mitigated here; it stops existing, because there is no second module. At the system layer, schedule-matched rolling denoising plus a set of computation-preserving runtime optimizations pull the complete two-step guided joint-diffusion callback inside 20 ms with output drift under $10^{-6}$. At the deployment layer, 96 proprioceptive dimensions including base linear velocity and 29 joint commands run at 50 Hz on a Unitree G1's Jetson Orin NX, which the authors present as the first fully onboard deployment of a joint state-action diffusion policy.
The paper also draws its own boundary clearly: guidance strength cannot exceed the calibrated range, text covers only learned semantic categories, the motion library contains no sustained multi-contact behavior, and the timing claim is bound to a specific export and device. Looking forward, the highest-value next steps are not another point on a metric. They are making guidance constrained or adaptive so that steering strength cannot push denoising out of the training distribution, extending language-action coverage to compositional commands, and completing the training and deployment release so that "fully onboard at 50 Hz" can be verified by someone else.
For teams building humanoid control, the most transferable part of this paper may not be the joint diffusion backbone at all. It is the acceleration table in Section 4.2: two-step sampling, rolling reuse, and five computation-preserving runtime optimizations, each with its benefit quantified, and with an explicit statement of which rows must not be added together and which number holds only for one specific artifact. That methodology travels directly to any robot system running iterative neural inference on Orin-class compute.
Golden Line
A predicted future state does not need to be executed; it only needs to exist, exist as a differentiable interface where objectives that were unknown at training time can land. PredActor keeps that trajectory locked inside the policy, so "steerable" and "directly executed" stop being a choice: guidance edits the state, the action catches up on the next denoising step, and within 20 ms exactly one thing is ever sent to the joints.