PAPER DEEP DIVE
CHOREO: Every Humanoid Skill as a Trajectory — training-free composition of 2,950 heterogeneous skills, 93.8% on 8-action long-horizon sequences on Unitree G1
CHOREO (Ocean University of China & CASIA) unifies outputs of RL policies, mocap data and generative models into SkillMotion trajectory assets, then composes long-horizon skills training-free via LLM planning, boundary-mismatch scoring, quintic polynomial seams and a standing bridge: 96.7% success over 552 skill boundaries, 4.6% fall rate (best baseline 41.5%), and 93.8% on 8-action sequences vs 41.7% for the strongest baseline; plus compatibility with 6 trackers and 100% ingestion of 3,069 heterogeneous instances. All experiments in simulation; real-robot deployment is left open.
One-line summary and background: why "knowing more skills" does not mean "being more capable"
Over the past two years, the skill list of humanoid robots has grown at an unprecedented pace: reinforcement learning has produced agile locomotion, motion-capture imitation has taught whole-body control, teleoperation has accumulated ground-truth demonstrations, and generative models can synthesize brand-new motions from text. Yet the team from Ocean University of China and the Institute of Automation, Chinese Academy of Sciences, in this arXiv paper posted in September 2026, points out an awkward reality: these skills are isolated islands. They are trained by different paradigms, stored in different representations, and depend on different controller interfaces — an action stream from an RL policy, a joint trajectory inside a mocap dataset, and a skeleton sequence synthesized by a diffusion model have no plug-and-play way to connect with one another. The paper puts it bluntly: simply accumulating independently trained capabilities does not naturally yield a more capable robot.
The mainstream solution is "train one bigger unified policy": feed increasingly diverse data into a single controller to cover a wider skill distribution. This works, but the paper identifies its fundamental limitation — every new skill is a new training problem: collect compatible data, adjust the training distribution, retrain an ever-larger policy. Meanwhile the producers of humanoid skills have become highly fragmented — different research communities, different simulators, different generative models each emit their own outputs, so retraining-based unification scales worse and worse. The authors therefore pose a complementary question: can heterogeneous pretrained humanoid capabilities be accumulated, reused and composed continually without retraining the underlying models?
CHOREO's answer rests on a plain observation: no matter how a skill was learned, it ultimately manifests as an executable robot motion trajectory. An RL policy generates trajectories through interaction; mocap data stores trajectories explicitly; a generative model's output can be converted into a trajectory. Rather than unifying controllers' internal architectures, unify their behavioral output. This reframing decouples skill generation from skill execution: integrating a new capability is just adding an asset to a trajectory library, without touching any existing controller. Expanding robot behavior turns from "repeatedly retraining ever more complex policies" into "growing a reusable motion-skill library."
Its pitch in one sentence: training-free integration, heterogeneous compatibility, scalable and continual — the three labels the paper's self-overview assigns to CHOREO, each backed by one of the three experiment groups below.
The CHOREO framework: three layers and the SkillMotion unified representation
CHOREO's runtime chains three components. Layer one is offline registration: a Source Adapter "materializes → canonicalizes → annotates" three source families — motion datasets (decoded directly), RL policies (rolled out in their native simulators for finite-horizon trajectories), and generative models (generated in their native environments, then converted). It ingests only the body-state trajectories themselves, never source-model parameters or policy actions. All trajectories are retargeted to a unified embodiment representation: the paper uses the 23-DoF Unitree G1 at 30 Hz, with unified joint ordering, coordinate conventions and units.
Each skill is represented as a SkillMotion asset M = (τ, z, b_in, b_out, e): τ is the T-frame trajectory (joint positions and velocities, root pose, foot contacts); z is the semantic descriptor (labels, motion attributes, provenance); b_in/b_out are structured boundary descriptors — entry/exit windows plus motion-profile summaries, the key to later composition; e is the execution descriptor recording target embodiment, tracker configuration and validation evidence. Missing velocities are derived by finite differencing; contact states are kept if present, inferred kinematically if possible, and otherwise marked "unknown, zero confidence" — the paper stresses that unknown contact must not be treated as "known no-contact."
Registration has a validation gate: schema legality, motion consistency, and a physical trackability check with the paired tracker. Repairable candidates are periodically corrected and re-checked; those that pass enter the library with validation evidence, failures are quarantined and never exposed to the runtime. The final library holds 2,950 admitted SkillMotion assets from heterogeneous sources: 2,901 motion trajectories, 16 RL policies and 33 generative models. Nearly three thousand trajectories from three very different paradigms appear to the runtime in one single format — the precondition for the multi-tracker and source-ingestion experiments below.
motion datasets / RL policies / generative models"] --> B["Source Adapter
materialize → canonicalize → annotate"] B --> C["Validation gate
schema / consistency / trackability"] C -->|pass| D["SkillMotion library
τ trajectory + z semantics + b boundary + e execution"] C -->|fail| E["Quarantine"] F["Natural-language instruction"] --> G["LLM task planner
semantic order only, no trajectories"] G --> H["Skill retrieval & sequence resolution"] D --> H H --> I{"Boundary mismatch d_i < δ ?"} I -->|yes| J["Quintic polynomial seam
6 constraints: pos/vel/acc on both ends"] I -->|no| K["Insert stable-stand bridge τ_st
two seams"] J --> L["Tracker adapter
reference/observation construction"] K --> L L --> M["GMT @ 50 Hz
Unitree G1 (MuJoCo)
23-D joint targets"]
Core method: boundary-mismatch scoring, quintic seams, and the stand bridge
Why does "frozen" matter so much? The whole value proposition hinges on it: if composition still required fine-tuning trackers or rerunning RL training, "training-free integration" would be void. Under the frozen protocol every new skill, source and composition enters the system with zero gradients — as the library grew from 2,901 trajectories to 3,069 instances, not a single backpropagation pass occurred. "Continual learning" changes shape: robot capability growth shows up not as weight updates but as asset accumulation that can be rolled back, audited and loaded on demand.
The hard part of composition: each skill verified in isolation does not mean the concatenation executes continuously. Online, an LLM task planner parses the natural-language instruction into a canonical skill-label sequence (e.g. walk_forward → horse_squat → walk_forward) — note the planner fixes only semantic order, generating no trajectories or control commands; the physical splicing is handled by the runtime.
For each adjacent pair M_i and M_{i+1}, CHOREO first aligns the successor in planar position and heading, selects splice endpoints within the annotated exit/entry windows, and computes a boundary-mismatch score: over four feature families F = {q, q̇, h, v} (joint positions, joint velocities, root height, root linear velocity), each dimension-normalized and weighted-summed, plus a confidence-weighted foot-contact mismatch term. The division of labor among weights is deliberate: joint-position difference measures "is the pose right," velocity difference distinguishes "similar pose but different motion trend," and the height and contact terms characterize vertical attitude and support configuration. This score is a reference-level filter, not a feasibility guarantee — below threshold δ, direct splicing is allowed.
On direct splicing, CHOREO replaces a short boundary window with a quintic polynomial seam: six endpoint constraints (displacement, velocity, acceleration on both sides) uniquely determine the polynomial coefficients; outside the seam, actions are preserved verbatim. Quintic (rather than cubic Hermite) guarantees continuous acceleration — the key to a jitter-free low-level controller. When the mismatch score exceeds the threshold, a pre-validated stable standing reference τ_st is inserted as a bridge — the bridge is not local smoothing but a genuine change of intermediate reference: the same seam operator first joins the source skill into standing, then standing into the successor. Every element (original skills, seams, bridge) flows through one execution interface, with per-tracker adapters handling coordinate transforms, time resampling and reference-window construction.
On the execution side the paper fixes the GMT tracker: references stored at 30 Hz are consumed by GMT at a 50 Hz control rate, outputting 23-D joint targets. Evaluation follows a "frozen protocol" — no source generator or source policy is invoked during online composition and execution; only pre-registered assets are used.
Experiment 1: long-horizon composition — the longer the sequence, the bigger the gap
The long-horizon benchmark contains 130 prompt-conditioned sequences: 20 two-action, 26 three-action, 36 five-action and 48 eight-action, totaling 552 planned skill boundaries. The success criterion is brutally strict: a sequence succeeds only if every skill and every transition completes — longer sequences give every method more chances to fail, and a single boundary mistake terminates the whole rollout. Four baselines — direct switching, fixed Hermite interpolation, fixed standing bridge, and Motion Matching — share the same task sequences, assets, initial states and frozen GMT.
| Method | 2-action | 3-action | 5-action | 8-action | Boundary SR | Fall rate | Δq (rad) |
|---|---|---|---|---|---|---|---|
| Direct switching | 75.0% | 61.5% | 58.3% | 14.6% | 67.2% | 54.6% | 0.0449 |
| Fixed Hermite | 85.0% | 69.2% | 44.4% | 22.9% | 68.8% | 52.3% | 0.0178 |
| Fixed standing bridge | 70.0% | 73.1% | 63.9% | 41.7% | 72.6% | 41.5% | 0.0180 |
| Motion Matching | 80.0% | 73.1% | 61.1% | 31.2% | 75.2% | 44.6% | 0.0171 |
| CHOREO | 100% | 100% | 91.7% | 93.8% | 96.7% | 4.6% | 0.0170 |
The defining shape of the results is a gap that widens with length: CHOREO completes all 46 two-/three-action sequences and scores 93.8% (45/48) on eight actions, while the strongest baseline drops to 63.9% at five actions and 41.7% at eight — CHOREO's lead grows from 15 points at two actions to 52.1 points at eight. The boundary-level metrics tell a consistent story: across all 552 planned boundaries CHOREO's switching success rate is 96.7% (21.5 points above Motion Matching), the fall rate drops from the safest baseline's 41.5% to 4.6%, while the joint-position discontinuity Δq = 0.0170 rad is simultaneously the lowest of all methods — "smooth" and "stable" are not in tension here.
The baseline analysis is genuinely insightful. Fixed Hermite squeezes Δq from direct switching's 0.0449 down to 0.0178 rad, yet its eight-action success rate is still only 22.9% with a fall rate above 50% — a smooth seam by itself cannot support reliable composition, because errors accumulate across boundaries. The fixed standing bridge does much better on long sequences (41.7% at eight actions), but its 41.5% fall rate shows a single intermediate pose cannot cover the wildly varied boundary pairs in this benchmark. Motion Matching is nearly as smooth but trails clearly in switching and sequence success. The conclusion: CHOREO's advantage is not "sewing smoothly" but simultaneously maintaining boundary feasibility, low fall rates and motion continuity, preventing transition errors from accumulating across boundaries.
Experiment 2: heterogeneous source ingestion and multi-tracker compatibility
The source-ingestion benchmark tests whether a universal interface can swallow structurally different outputs: 2,969 library trajectories, 50 RL-policy rollouts and 50 diffusion-generated motions — all 3,069 instances pass registration checks (schema legality, finite states, load-store consistency); of these, 256 sampled library motions plus all RL and diffusion motions — 356 instances — additionally undergo closed-loop physical rollouts. Result: ingestion 100% (3,069/3,069), physical execution 99.2% (353/356) — all three failures occur in the sampled library subset; the RL and diffusion sources pass 50/50. The experiment cleanly separates two claims: ingestion proves heterogeneous outputs convert into the unified representation; rollout proves the converted assets are truly executable under one frozen controller. The 100% on RL and diffusion subsets shows the interface is not bound to any single upstream representation.
The multi-tracker benchmark feeds the same 30 admitted trajectories to six frozen trackers: GMT, HoloMotion, SONIC and H-ACT each hit 100% (30/30), TWIST2 86.7% (26/30), OpenTrack 46.7% (14/30). The assets are byte-identical; all variation comes from each tracker's adapter and its own robustness. The paper reads these numbers with restraint: the four 100% results are evidence of interface compatibility, not a system ranking; the 46.7%–100% spread is a reminder that a unified motion representation cannot erase differences in low-level tracking capability itself.
Ablation: which component matters most
The transition-component ablation isolates three components on the same 130 tasks, same seeds, same frozen GMT. Removing "state-entry compatibility scoring" hurts most: eight-action success falls from 95.8% to 75.0%, switching success from 97.1% to 90.6%, and the fall rate rises from 5.4% to 16.2% — as transition errors accumulate over long sequences, "matching the robot's current state to a suitable skill entry" becomes ever more critical. Removing "motion-continuity scoring" leaves sequence success unchanged but raises Δq from 0.0162 to 0.0177 rad and slightly raises falls — its contribution is smoother, safer seams. Disabling "bridge validation" changes neither success nor falls on this fixed benchmark, but the paper explicitly warns against concluding validation is redundant: its safety value shows up in rejected bridge candidates, which are under-represented in this task set. This restrained reading of a null ablation is a model for many papers.
Honest boundaries
The limitations section is candid. First, the long-horizon benchmark uses fixed tasks run once each; performance variance, horizons beyond eight actions, and generalization to unseen task distributions are not characterized. Second, only 356 of 3,069 ingested instances received closed-loop rollouts. Third, the tracker comparison uses just 30 trajectories and measures motion-execution compatibility, not full system capability. Fourth, all experiments are in simulation — perception noise, actuation latency, model mismatch, contact uncertainty and real-world disturbances are all uncovered. The framework is also bounded by the coverage quality of existing skills and the bridge library, and transition scores rely on hand-designed kinematic/contact features. Future work lists five items: larger skill sets, repeated trials and out-of-distribution long compositions, real-robot deployment, learned transition-feasibility models, state-conditioned bridge generation, and closed-loop recovery under real uncertainty.
Why it is worth reading: bringing the harness mindset into humanoid control
First, unifying behavioral output is far cheaper than unifying internal architectures. This is the same principle as "standardize interfaces, not implementations" in software engineering. CHOREO asks no skill producer to change anything — RL policies, generative models and mocap data stay as they are and contribute only trajectories. With generation and execution decoupled, the marginal cost of "adding a skill" drops from "one training run" to "one registration" — the real reason it can claim scalability.
Second, the enemy of long horizons is error accumulation, not single-point failure. Fixed Hermite is the clearest exhibit: it makes every seam smooth (low Δq) yet scores only 22.9% on eight actions — smooth is not the same as feasible, and boundary errors still accumulate segment by segment. CHOREO's fix is to treat "can these connect" as a retrieval problem with an explicit score (continuity + compatibility + contact confidence), and to rewrite the intermediate reference via a standing bridge when the score fails, instead of forcing a seam. This holds for any system that chains individually verified capabilities into long tasks — including LLM-agent tool orchestration.
Third, the rigor of the evaluation protocol is itself a contribution. Sequence-level success with a one-vote-veto rule, failures kept in the denominator, a frozen protocol forbidding test-time updates, independent record-integrity verification in the ablation (520 episodes, zero duplicates and zero resource failures), and the restrained reading of the null bridge-validation ablation — these details make the 95.4%/93.8% numbers believable. Remember the boundaries too: simulation, a single G1 embodiment, GMT-dominated execution — whether the "standing bridge" survives disturbances on real hardware is the question this paper leaves for next year.
CHOREO is a collaboration between Ocean University of China (Ziyi Sun, Jingwen Chen — equal contribution) and the Institute of Automation, Chinese Academy of Sciences (Long Cheng, Zhaoxiang Zhang), supported by the National Natural Science Foundation of China (62606500) and the Shandong Provincial Natural Science Foundation. arXiv:2609.22274 (cs.RO, submitted September 11, 2026).
SOURCE LINKS



