
The ultimate guide to multi-harness RL: train models inside Claude Code and Codex — Hugging Face's four-harness experiment with a 2.6B model
The same model behaves like two different models depending on the agent harness around it: GLM-5.2 scores 23% in one harness and 52% in another on SWE-bench Pro; in a controlled experiment swapping the harness moved scores 13 points while swapping the model moved only 2.5–5. A Hugging Face team (Adithya S Kolavi, Joel Niklaus, Lewis Tunstall, Leandro von Werra and colleagues, with Liquid AI) publishes an open fix: run agentic RL inside the real harnesses — Claude Code, Codex, OpenCode, Mini-SWE-Agent — with zero harness code changes. The stack: OpenEnv as the shared interface, a capture proxy that mints a session id per rollout and records exact token ids, per-token behavior logprobs and loss masks at the model endpoint (full-distribution sampling, top_p=1.0), Harbor for 40+ harness adapters and 26 sandbox backends, and TRL's Async GRPO consuming the masked training sequences. The numbers: LFM2.5-2.6B trained across four harnesses rises from 42.2% to 54.2% average pass@1 with gains under all four, and uses 31% fewer tool calls on tasks both it and the base model solved, while the OpenCode-only arm's gains stay mostly at home. The article also reports the SFT comparison (RL 54.6% vs SFT 47.5%), the full attribution of the earlier Qwen rise-and-decline (output-budget exhaustion, working without submitting, and a reward term that collapsed held-out accuracy from 0.740 to 0.178), and engineering boundaries like the proxy's concurrency ceiling.
One model, several personalities: the harness is the hidden variable
If you use AI to write code, you have almost certainly run the same model through more than one tool — and noticed it does not behave the same way. It plans differently, reaches for different tools, and finishes tasks in one harness that it gets stuck on in another. The cause is the program wrapped around the model: the agent harness. It runs the loop, decides which tools the model gets, writes the context the model reads, parses what the model sends back, and decides when to stop. Change the harness and you change what the model sees and what it is allowed to do.
The difference is measurable. Joel Niklaus (Hugging Face) measured the same model, GLM-5.2, at 23% in one harness and 52% in another on SWE-bench Pro — a gap of nearly thirty points from the wrapper alone. Rankings do not carry over either: Codex ranks second of ten harnesses for GLM-5.2 and ninth for Gemma 4 26B-A4B.
The mismatch bites hardest for people who run open-weight models themselves: a model never trained in your harness calls tools your harness doesn't have, or writes output it can't read. And training in one harness doesn't fix it — the model learns that harness's habits and still struggles elsewhere. A long-form article by a Hugging Face team (Adithya S Kolavi, Joel Niklaus, Sergio Paniego Blanco, Leonie Monigatti, Amine Dirhoussi, Ben Burtenshaw, Lewis Tunstall, Leandro von Werra, with Liquid AI) proposes the direct fix: train with agentic RL inside the harnesses people actually use — Claude Code, Codex, OpenCode — with the harnesses running exactly as they ship, zero code changes. Using their open framework, they trained a 2.6B model across four harnesses at once: average pass@1 rose from 42% to 54%, and tool calls on solved tasks fell 31%.
The evidence: benchmark scores now come with a harness attached
This is 2026 model-card reality, not theory. GLM-4.7 promises gains "in mainstream agent frameworks such as Claude Code, Kilo Code, Cline, and Roo Code"; Kimi K2 reports Terminal-Bench twice — 25.0 under Terminus, 30.0 under Moonshot's own framework; MiniMax M2 names a harness for almost every benchmark; DeepSeek-V3.2's thinking mode wouldn't run under Terminus at all.
A controlled experiment makes the point brutally: three models within three points of each other on a public leaderboard (GLM-5.1, GPT-5.4, Kimi K2.6), run on the same 100 SWE-bench Verified tasks under three progressive harness configurations (Minimal → Improved → Full, adding context compression, retries, self-checks, rollback). Each model wins under a different configuration. Swapping the harness moved GLM-5.1 by 13 points; swapping the model inside a fixed harness moved it by only 2.5–5. Outside the lab it's the same: Claude Opus 4.5 scores 45.9% on Scale's SEAL SWE-bench Pro leaderboard and 55.4% inside Claude Code — identical weights.
Training-induced "harness lock-in" is even worse. The Orchard paper measured transfer: OpenSWE-32B, trained in OpenHands, drops 7.5 points moving to Mini-SWE-Agent and 58.8 points under never-seen Kimi-CLI — 3.6% on SWE-bench Verified, zero on Terminal-Bench 2.0. Scale-SWE stops producing valid tool calls anywhere but home. The paper names two failure classes: degraded resolve rate and catastrophic format failure.
KwaiKAT's team states the root cause plainly: trained on a single fixed harness, "the model often learns not 'how to solve the task' but 'how to solve the task under that particular harness's interface conventions.'" Three overfitting modes follow: action format, context structure, and control flow (retries, stop conditions). Harnesses differ along exactly these lines — API dialects (OpenAI chat-completions / Responses, Anthropic Messages, Gemini), tool calling vs prose edit formats (Aider), context compaction. In practice the failure is often an error, not a lower score: the model calls a tool by its training harness's name and the new harness rejects the call before the tool ever runs.
Frontier labs now train across harnesses deliberately: Poolside's Laguna mixes in 1.3B tokens of trajectories from OpenHands, OpenCode, and Mini-SWE-Agent; Kimi K3 builds harnesses from composable modules; Qwen3-Coder-Next generates data in six harnesses; Liquid AI picks a harness at random per task for LFM2.5-2.6B. OpenForgeRL showed variety costs no peak performance — the three-harness model wins even on the single harness's home turf (48.5 vs 46.0) and nearly doubles under Codex. The article ships a "paper receipts" viewer showing the original passages with page numbers from eight 2026 reports.
tool/argument names"] --> X["Rejected at the door by new harnesses"] F2["Context overfitting
compression & history"] --> Y["Lower resolve under new context organization"] F3["Control-flow overfitting
retry/stop habits"] --> Z["Stalls or never submits elsewhere"] S["Fix: harness scaling — vary the harness during training"] -.-> F1 & F2 & F3
White-box vs black-box: who owns the rollout loop
The article nails down its vocabulary first: the policy is the trained model (no loop, no memory); the agent is the policy inside a harness; the sandbox is where actions execute; a benchmark is a task set plus a comparison protocol.
The pivotal distinction: in a white-box environment the trainer owns the loop — it samples actions, calls env.step(), reads observations; every token is already in its hands (TRL's GRPOTrainer assumes this). In a black-box environment the harness owns the loop — it starts inside its sandbox, runs its own tools, compacts its own context, stops when it decides. The trainer stands outside and sees only a sequence of calls at a model endpoint. Microsoft's Agent Lightning team calls these "traditional agentic RL" vs "harnessed agentic RL." (A footnote untangles KAT-Coder's different usage of the same terms — context handling and loop ownership are separate choices.)
Why rewards are not enough: the token contract
The technical core is a plain observation: on-policy policy gradients need, per token, which token was sampled and with what probability — and a harness returns text plus a score. Three traps follow:
- Re-tokenization drift: different token sequences can produce identical text; re-tokenizing may yield different token IDs than the model sampled. TRL: "in RL, you optimize on the exact tokens the model produced."
- Tampered responses: harnesses add role markers, change whitespace, even repair malformed JSON when building the next prompt — training on edited text updates the model for a response it never generated.
- Probabilities must be recorded at generation time: in async RL the weights may have moved since sampling; recomputing logprobs with current weights makes the importance ratio 1 even when policies differ.
The fix: the sampler must return token ids and per-token logprobs, the trainer must keep them, and you must never re-encode decoded text — append template suffixes by id concatenation. That forces interception at the model endpoint, not the text boundary. The article also honestly records dissent: OpenForgeRL rebuilds training samples from text at the same proxy architecture, never mentions token ids, and still reports the strongest multi-harness results — the token-faithful vs text-reconstruction question is open.
Inside the framework: OpenEnv + capture proxy + Harbor + TRL
Three components: OpenEnv (a Meta PyTorch × Hugging Face collaboration, Gymnasium-style reset/step/state over HTTP, steered by a twelve-organization committee, BSD-3-Clause) is the shared interface; Harbor (from the Terminal-Bench team) supplies tasks and sandboxes — v0.22.0 ships adapters for 40+ harnesses and 26 execution backends, with ~80 task datasets in its format; TRL trains.
The capture proxy is the linchpin. A harness finds it like any model provider: base URL + API key — and the key is a session id minted per rollout, so one port serves every concurrent rollout (unregistered keys get 401). Four API dialects are identified by path, then headers, then body shape, using converters vendored from NVIDIA's Polar gateway; calls to the engine never stream — the proxy stores the full completion and replays it as a stream if asked. It records prompt token ids, sampled token ids, and per-token logprobs, sampling from the full distribution (top_p=1.0): truncating biases sampling away from the policy's own distribution (a known route to entropy collapse) and mismatches vLLM's processed logprobs; moving top_p to 1.0 improved the importance ratio from 0.985–0.993 to 0.9984–0.9999.
The proxy trusts no engine: a first-contact probe grades it, and only a "tokens"-level engine may supply training data — hosted APIs (OpenAI, Anthropic, HF Inference Providers) land at evaluation-only. Rollouts are stored as a call graph: each model call is a node, parented by the earlier call whose prompt+completion is the longest exact token prefix of its own prompt — retries are siblings that never continued; subagents and compacted contexts start new roots. Every root-to-leaf path becomes one training sequence, context tokens masked 0, sampled tokens 1; turns with missing or misaligned logprobs stay as context and are never targets.
The engineering detail is generous: four commands (openenv harbor info/rollout/serve/push) cover self-check, bare rollouts, serving, and Space deployment; every captured rollout is cross-checked call-by-call against Harbor's own ATIF trajectory file — which caught a harness sending an empty tools array, getting a 400, and leaving a well-formed-looking graph; the proxy now drops empty tools arrays before forwarding. Rewards are never combined by the integration (only a single score or one named "reward") — an earlier "+0.2 for submitting anything" term taught the policy to submit immediately and held-out accuracy fell from 0.740 to 0.178 while training reward looked healthy. The single-process proxy starts starving its health check around 200 concurrent sessions and crashes at 320 — recorded as a hard boundary.
Training small models: the LFM2.5-2.6B four-harness experiment
The experimental design is clean enough to copy: 1,000 SmolDataEnvs training tasks (400 medium, 600 hard, built from real Kaggle notebooks); 250 held-out test tasks with triple no-overlap (notebook, question, instruction); two arms — OpenCode-only vs four harnesses (OpenCode, Claude Code, Codex, Mini-SWE-Agent), each GRPO group of 8 rollouts using one harness; TRL Async GRPO for 1,000 steps on two H100s (one trains, one serves with vLLM), one E2B sandbox per rollout; evaluation every 100 steps across 4 harnesses × 250 tasks = 1,000 test cells. The baselines reproduce the problem: the same weights solve 62% under Mini-SWE-Agent but 33% under Claude Code.
Reward design: correctness 0/1 plus a tool-efficiency bonus (≤0.1, correct answers only). The bonus comes from the Qwen runs — with correctness alone, nothing tells the model to stop exploring, and tool calls per rollout crept from 13 to 41. The bonus is small but GRPO learns from within-group differences: when all 8 rollouts are correct, correctness provides no contrast and the bonus is the only signal, favoring shorter solutions. Audited counts: the bonus supplied the only reward contrast in 17.5% of OpenCode-only groups and 22.6% of multi-harness groups.
Results: both arms improve — OpenCode-only 42.2%→52.3%, multi-harness →54.2%; the 1.9-point overall gap is within noise. The per-harness split is the real finding: the OpenCode-only model gains almost entirely at home (OpenCode 34→58%), while the multi-harness model improves under all four (Claude Code 42→49%, Codex 43→54%, Mini-SWE-Agent ties). Efficiency mirrors it: on tasks both solved, the multi-harness model uses 31% fewer tool calls than the base model (vs 11% for OpenCode-only), with the biggest saving under Codex (~half); the OpenCode-only model instead generates more tokens than the base under never-trained Claude Code (up to 62% more at step 500). The article even discloses the confound that the two runs saw unequal data exposure (Claude Code expands each rollout into ~8 training rows) and labels the comparison observational, not a method ranking.
The SFT comparison: imitation is cheap and weaker
Using Qwen3.8-27B as a teacher across all four harnesses they collected 3,189 successful rollouts (888 tasks), converted to 17,929 SFT examples (tokenized with LFM's own template), and trained two SFT models. RL wins clearly: multi-harness RL 54.6% vs OpenCode SFT 47.5% vs multi-harness SFT 43.1% (11.5 points below RL). The subtle result is multi-harness SFT's "spread then give back": Claude Code +9.2, Codex +6.8, OpenCode +4.4, but Mini-SWE-Agent falls from 62.1% to 45.2% — canceling everything (reason unknown, acknowledged). SFT cuts tool calls 24.2% without being rewarded for it, but unevenly across harnesses.
The Qwen lessons: three declines and three recipe fixes
The earlier Qwen3.5-2B runs contributed equally useful failures: all three runs rose then fell (multi-harness 14.6%→37.0%→~26%). The traces explain why: multi-harness ran out of output budget (cells cut off at the 4,096-token evaluation limit rose from 9 to 556 of 1,000, while training allowed 16,384 — the model learned a habit evaluation would cut short); OpenCode-only kept working without submitting (calls 17→21, submission rate 69%→41%). Plus 35–58% of optimizer steps had no reward contrast, Claude Code produced 77% of training rows, and harder data did not rescue the decline. The three LFM recipe fixes — add the tool-call bonus, medium/hard tasks only, unify the 4,096 output cap across training and evaluation — grew directly out of these holes.
Four things worth taking away
- The harness is an underrated independent variable. A controlled 13-point harness swing vs a 2.5–5-point model swing, plus model cards now labeling scores with harnesses, means every coding benchmark number you quote deserves the question "measured in which harness?"
- The token contract is a hard constraint of black-box RL. Never re-tokenize, never train on repaired text, record probabilities at generation time — the model endpoint is the only interception point satisfying all three; the capture-level grading (hosted APIs evaluate, never train) is useful engineering knowledge in itself.
- "Harnesses as they ship" is feasible and rewarding. Zero harness code changes, one proxy with base-URL + session-id wiring, and Claude Code / Codex become training environments; the cross-harness model gains everywhere and cuts tool calls 31%, while single-harness gains stay mostly at home.
- The completeness of the failures is the article's second contribution. Qwen's rise-and-fall with trace-level attribution, the reward term that backfired (0.740→0.178), the proxy's concurrency ceiling, the resume bug replaying old tasks — every negative result comes with a diagnosis. For anyone reproducing this path, these save more than the positives.
Everything is open: the SmolDataEnvs task suite, the SFT dataset (per-harness configs, pre-tokenized variants, training script), both LFM RL models and both SFT models, a training-comparison dashboard, Harbor environment servers on HF Spaces, and a tutorial with HF Jobs and Slurm instructions. Larger runs are teased as ongoing.
Sources
- Original article: The ultimate guide to multi-harness RL (Hugging Face Space, Oct 1, 2026, by Adithya S Kolavi, Joel Niklaus, Sergio Paniego Blanco, Leonie Monigatti, Amine Dirhoussi, Ben Burtenshaw, Lewis Tunstall, Leandro von Werra)
- Repository: adithya-s-k/FineEnvs (article source under content/articles/multi-harness-rl/)
- Companion tutorial: FineEnvs repo, 05-multi-harness-rl/ (multi-harness and native OpenCode training scripts + REPRODUCE.md)


