PAPER DEEP DIVE
Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning
Meta's agentic meta-reasoning harness separates control from work: a controller runs Assess, Propose, Evaluate and Dispatch over a compact state and persistent artifact memory. On ProgramBench it reaches 71.5% with GPT-5.5, versus 58.0% for Codex.
Paper: Paras Dahal, Anton Bakhtin, Taco Cohen, Zhengxing Chen, Carole-Jean Wu, Rob Fergus, Scott Yih, Gabriel Synnaeve, Ruslan Salakhutdinov, Sanjeev Arora, Jason Weston and Anirudh Goyal, all at Meta Superintelligence Labs. Released on September 29, 2026 at https://arxiv.org/abs/2609.38147. Code status: no public code has been released, but the appendix reproduces every controller-stage prompt, the matched baseline prompt, the worker prompts and all tool schemas, which is enough to rebuild the skeleton of the harness.
In one sentence: the paper introduces agentic meta-reasoning, an inference-time harness in which a controller decides the next computation through four separate agentic stages (Assess, Propose, Evaluate, Dispatch), carries only a compact state between cycles and keeps full work products in a retrievable artifact memory. With the same models and the same call budget it beats direct control in all 12 matched comparisons, and on ProgramBench it reaches 71.5% with GPT-5.5, against 58.0% for Codex on the same model.
Background and Motivation
Language-model agents now spend a large number of model calls on a single problem, and a sizeable share of those calls are not spent advancing the task at all. They are spent deciding what to do next. The paper's running example is a proof: the agent holds a candidate whose main argument looks sound but whose central lemma has not been verified. It can try to prove the lemma, ask a second worker to check the argument it already has, or abandon the route and start again. Every option costs compute, and compute spent on the wrong option may be wasted entirely. For an agent working autonomously toward a goal, this choice is part of solving the problem, and it is a different capability from producing the next line of the proof.
The authors frame this as metacognitive control: assessing one's own progress and using that assessment to decide what to do next. Inside a single model call the object of control is the chain of thought. Inside an agent it is the body of work the run has already produced, so the agent has to decide which results to trust, what to build on and when to stop. A wrong control decision costs more than the compute it consumes. It can keep a failed line of attack alive for the rest of the run, or throw away a correct answer the agent had already found.
Most existing agents interleave these decisions with object-level work and make each one in a single step conditioned on an ever-growing history. ReAct, Reflexion, HuggingGPT, MemGPT and the Recursive Language Model all sit in this family. As a task accumulates intermediate artifacts, the context those decisions are made from becomes noisier, and control decisions are precisely the ones that depend most on an accurate picture of overall progress.
The central claim is that control decisions deserve an agentic reasoning process of their own. Before committing, an agent can deliberate and investigate: read back an earlier result, or dispatch a worker to check one. That deliberation has a structure of its own. It consolidates what the run has established, explores what could be done next, and assesses what each option is worth given the remaining budget. The authors call this agentic meta-reasoning: the agent reasons about, and acts on, its own inference process.
Three adjacent lines of work frame the contribution. Fixed-structure test-time compute (chain of thought, self-consistency, self-refinement, tree and graph of thoughts) commits to the number of samples, branches or refinement rounds before the problem reveals what work it needs. Learned or optimized orchestration (GPTSwarm, Meta-Harness, Conductor, and recent work that synthesizes harnesses automatically) trains or searches the coordinating system itself; this paper does neither, and instead isolates the control mechanism as its own agentic process. Work on metacognitive control and agent memory is closest in spirit, but differs in the object being controlled: not one reasoning chain, one tool trigger or one stopping rule, but a growing external computation built on persistent artifacts.
Preliminaries: Artifacts, Memory and the Worker Interface
An artifact is any stored output from a worker or from the controller: an attempt, a critique, or a note the controller writes for itself. Each artifact carries a stable identifier of the form round_worker, so 2_1 is the second worker of the third round. For a task $x$, let $M_t$ be the set of artifacts available at control cycle $t$.
A worker receives three things: the task, an instruction $g$ written by the controller, and a set of artifacts $C\subseteq M_t$ that the controller judges necessary for the assignment. Writing $W$ for worker execution and $y$ for the artifact it returns:
$$y=W(x,g,C)$$
The paper stresses that this is an interface rather than a pure function. A coding worker may read files outside $C$, modify the environment and behave stochastically. Workers never see the controller's private state or deliberation, only their instruction and the artifacts chosen for them. The memory design in the appendix follows the same principle from the other side: every controller stage sees memory only as an index of identifiers and titles, and must call read_memory to retrieve a body, while workers receive artifact bodies stripped of identifiers, so they cannot cite one back by name.
Method
1. The controller's action space
The controller can inspect stored work, record notes, launch workers, or stop with an answer. The first two are memory operations: for a set of identifiers $I$, $\mathrm{Read}(I)$ retrieves the corresponding artifacts, and for new content $u$, $\mathrm{Write}(u)$ stores it under a fresh identifier. Borrowing Kirsh and Maglio's term, the paper calls these epistemic actions: the controller leaves notes for itself to organize what it knows before acting on the task, for instance a warning that an earlier proof relied on an invalid assumption.
The third action launches a batch of $k$ parallel workers, each with its own instruction $g_i$ and context $C_i$, applying the worker interface $k$ times to the same task:
$$\mathrm{RunWorkers}\left(\{(g_i,C_i)\}_{i=1}^{k}\right)=\left\{W(x,g_i,C_i)\right\}_{i=1}^{k}$$
Returned artifacts enter memory together with the identifiers of their inputs. The same interface covers a fresh attempt, a targeted repair or a synthesis of earlier results; there are no fixed worker roles, and the assignment plus context determine what kind of work is done. The fourth action is $\mathrm{Stop}(y^{\star})$, where $y^{\star}$ must already exist in memory. It need not be the latest output, but it cannot be new: to submit a new answer, the controller must first have a worker produce it. The tool schemas make this concrete. The meta-reasoning controller's finish tool takes a memory_id, whereas the baseline's finish tool takes the answer text directly.
2. The four-stage control cycle
Each cycle runs Assess, Propose, Evaluate and Dispatch in sequence. Each stage is a full agentic loop with its own system prompt, its own context and its own permitted memory operations, and a stage may take several model calls to resolve because it can investigate before answering. Only a compact state persists between cycles; any stage can pull earlier artifacts from memory when it needs them.
Assess: what have we learned? When workers return, the controller updates its account of the run. With $s_{t-1}$ the previous assessment and $\Delta M_t$ the artifacts produced by the last batch of workers:
$$s_t=\mathrm{Assess}(x,s_{t-1},\Delta M_t;M_t)$$
The semicolon denotes access to persistent memory rather than inclusion of its contents in the prompt. On the text benchmarks the Assess prompt asks for a verdict on each new candidate (likely correct, has gaps, or fundamentally flawed) and a stop recommendation that defaults to continue, because workers tend to rate themselves highly on candidates with subtle errors. On ProgramBench the stage is entirely different: it maintains a cited factual snapshot with five sections, namely implemented, broken, unverified, unexplored and dead ends, and every bullet must cite a memory identifier or a commit hash. In the prompt's own words, worker prose claims without a commit or trace are not facts.
Propose: what could we do next? From the current assessment the controller lays out candidate computations before committing to any of them:
$$\mathcal{A}_t=\mathrm{Propose}(x,s_t;M_t)$$
One design choice here is deliberate: Propose is not told the remaining budget. Separating the generation of alternatives from the judgment of their affordability lets expensive but potentially valuable options enter the candidate set instead of being self-filtered. The prompt says so directly: think broadly, do not limit yourself to a few safe options.
Evaluate: which option is worth its cost? This is the only stage that sees the budget. With $b_t^{\mathrm{eval}}$ the budget remaining when evaluation begins, after earlier calls have been charged, the selected proposal is:
$$\tilde{a}_t=\mathrm{Evaluate}(x,s_t,b_t^{\mathrm{eval}},\mathcal{A}_t;M_t)$$
The authors are explicit that this is a prompted, qualitative assessment of computational value, not an exact optimization and not a learned value-of-computation estimator. The prompt rates each action high, medium or low value, emits a YAML block naming the chosen actions, the memory identifiers to pass as context and the number of parallel workers, and reminds the model that each deliberation turn already costs four calls before any worker runs. The bar for stopping is high: a candidate rated likely correct with no open questions touching it, only low-value proposals left, and a willingness to stake the run on that candidate.
Dispatch: what should the worker receive? Dispatch turns the selected proposal into an executable action:
$$a_t=\mathrm{Dispatch}(x,s_t,\tilde{a}_t;M_t)$$
For a lemma check it writes the verification instruction and attaches the proof plus any relevant critiques; a fresh attempt may receive no artifacts at all. The paper treats the choice of what a worker sees as part of choosing the computation. The ProgramBench Dispatch prompt also encodes a hard-won lesson: workers never see the deliberation, only their steering string, so each instruction must be self-contained and actionable, at least 200 to 500 characters, naming files, flags and success criteria. Passing a bare action identifier is called the single biggest failure mode of the dispatch step.
flowchart TD
T[Task x and call budget B] --> AS[Assess rewrite compact state s_t]
AS --> PR[Propose list candidate computations A_t budget hidden]
PR --> EV[Evaluate rate options against remaining budget only stage that sees budget]
EV --> DI[Dispatch write instructions and pick artifact context]
DI -->|launch parallel workers| W[Workers single call or coding agent]
DI -->|stop on an existing artifact| OUT[Submit answer already in memory]
W -->|new artifacts delta M_t| MEM[Persistent artifact memory M_t index holds ids and titles only]
MEM -->|read_memory on demand| PR
MEM -->|read_memory on demand| EV
AS -->|write_memory notes| MEM
W -->|full new artifacts reach next Assess| AS
Figure 1: The agentic meta-reasoning control cycle. Each stage is its own agentic loop; only a compact state crosses cycles, while full work products stay in persistent memory. Drawn from Section 3 and Appendices B and E of the paper.
Figure 2: Current agents interleave control with object-level work; agentic meta-reasoning makes control an explicit reasoning process over compact state, persistent memory and workers. Source: paper Figure 1.
3. The artifact graph: turning a run into something you can diagnose
The context artifacts chosen for each worker leave a record of how later work builds on earlier work. A worker that receives a proof and returns a critique creates an edge from the proof to the critique; a repair that receives both has two incoming edges. With artifact graph $G=(V,E)$ and $C_j$ the earlier artifacts supplied as context for $y_j$:
$$E=\{(y_i,y_j)\in V\times V:\ y_i\in C_j\}$$
Because inputs exist before a worker starts, the edges form a directed acyclic graph. Roots are work with no artifact inputs, branches are multiple follow-ups, and multiple incoming edges mark synthesis from several sources. The authors note that the graph only records controller-generated dependencies; it does not capture every information channel. A coding worker, for example, can read code another worker left on the shared filesystem without that dependency being recorded.
4. The Direct Control baseline and budget accounting
To isolate the control design, the paper builds a Direct Control Agent that uses the same workers and the same interfaces for delegation, context selection, artifact writing and stopping, but has no explicit separation of control: it chooses each action in a single turn over the accumulated history. Both systems receive the same nominal call budget $B$. If cycle $t$ starts with allowance $b_t$ and uses $c_t$ controller calls and $w_t$ worker calls:
$$b_{t+1}=b_t-c_t-w_t,\qquad b_0=B$$
Worker cost includes every call made inside a coding agent and sums calls across parallel workers; controller cost includes every call inside the four stages. The budget is enforced between cycles, so workers dispatched while still under the cap run to completion and can carry the total slightly past the allowance. The authors are careful to add that equal call counts do not mean equal tokens, latency or FLOPs, and that spending more is not inherently better.
5. Diagnostics: coverage, monitoring and selection
A final score cannot separate two kinds of failure: a run where no correct answer ever appeared, and a run where a correct answer sat in memory while the controller submitted a flawed revision. The paper separates them using correctness labels on intermediate artifacts. Let $\mathcal{C}$ mean the run contains a correct solution and $\mathcal{S}$ mean the submission is correct. A correct submission must come from the run, so $\mathcal{S}\subseteq\mathcal{C}$ and:
$$\Pr(\mathcal{S})=\Pr(\mathcal{C})\,\Pr(\mathcal{S}\mid\mathcal{C})$$
The first factor is coverage and the second is selection. Coverage can be tracked along the run: with $V_{\mathrm{sol},\leq k}$ the solution artifacts available within the first $k$ calls, $\mathrm{Coverage}(k)=\Pr\left(\exists y\in V_{\mathrm{sol},\leq k}:\ell(y)=1\right)$, which is the success rate a perfect selector could reach from the answers already on hand. Monitoring is measured with a Type-2 AUC: draw a correct candidate $Y^{+}$ and an incorrect candidate $Y^{-}$ and ask whether a confidence signal $r$ orders them correctly, with half credit for ties:
$$\mathrm{AUC}_2(r)=\Pr\left(r(Y^{+})>r(Y^{-})\right)+\tfrac{1}{2}\Pr\left(r(Y^{+})=r(Y^{-})\right)$$
This measures discrimination only, not calibration. Finally, the paper asks whether selection beats a structural baseline. The convergence frontier $F(G)$ is the set of deepest terminal artifacts, $q_F(G)$ is the probability that a uniform choice from the frontier is correct, and frontier selection gain is:
$$\mathrm{FSG}=\mathbb{E}\left[\ell(y^{\star})-q_F(G)\mid\mathcal{C}\right]$$
A positive gain means the agent picks better than choosing at random from its own frontier. The practical value of this decomposition is that it points at different remedies: coverage-bound failures call for better exploration or object-level problem solving, while selection-bound failures call for better assessment, verification and commitment.
Experimental Setup
Two settings are evaluated. On the three reasoning benchmarks, a worker is a single tool-free model call that returns a text artifact. IMO ProofBench-Advanced has 30 hard proof problems graded on a 0, 1, 6, 7 scale and reported as a percentage of the maximum total. ARC-AGI-2 has 120 abstract visual reasoning tasks scored by exact grid match. LongCoT-mini is a 507-problem split spanning logic, computer science, chemistry, chess and mathematics, where individual steps are easy and the difficulty is tracking state across many of them without dropping a constraint. On ProgramBench, 200 long-horizon program reconstruction tasks, the agent is given documentation and an execute-only reference binary and must write a codebase from scratch whose executable matches the reference on hidden tests; the score is the mean per-problem test-pass rate.
The underlying models are Gemini 3.1 Pro, GPT-5.5 and Opus 4.8. Nominal budgets are 25, 50 and 100 calls per problem on the reasoning benchmarks and 400, 800 and 1200 on ProgramBench, with the main comparison at the largest allowance. External baselines are the Recursive Language Model on the reasoning benchmarks and three complete coding agents on ProgramBench: mini-SWE Agent, Claude Code and Codex. Every system is told how much budget it has used and how much remains, and every system may stop early; the external baselines have no native notion of a call budget, so the authors add the same accounting and messages to each.
Results
Main results: ahead in all 12 matched comparisons
| Model | System | IMO ProofBench-Adv | ARC-AGI-2 | LongCoT-mini | ProgramBench |
|---|---|---|---|---|---|
| Gemini 3.1 Pro | Recursive Language Model | 73.3 | 76.7 | 46.5 | – |
| mini-SWE Agent | – | – | – | 42.0 | |
| Direct Control | 82.7 | 77.5 | 53.5 | 46.9 | |
| Meta-Reasoning | 91.3 | 84.2 | 62.7 | 48.7 | |
| GPT-5.5 | Recursive Language Model | 86.2 | 73.3 | 63.3 | – |
| mini-SWE Agent | – | – | – | 57.6 | |
| Codex | – | – | – | 58.0 | |
| Direct Control | 93.3 | 75.8 | 64.7 | 63.7 | |
| Meta-Reasoning | 94.6 | 79.2 | 65.1 | 71.5 | |
| Opus 4.8 | Recursive Language Model | 68.1 | 66.7 | 64.3* | – |
| mini-SWE Agent | – | – | – | 64.7 | |
| Claude Code | – | – | – | 65.5 | |
| Direct Control | 78.0 | 77.5 | 65.3 | 65.3 | |
| Meta-Reasoning | 80.2 | 80.0 | 66.5 | 67.2 |
Table 1: Performance (%) at the main budgets, 100 calls on the reasoning benchmarks and 1200 on ProgramBench, controller calls included. The asterisk marks the full-Python RLM variant. Source: paper Table 1.
The strongest evidence is the matched comparison with Direct Control: same models, same workers, same interfaces, same budget, and only the control differs. Meta-reasoning has the higher point estimate in all 12 pairings. Subtracting cell by cell gives the size of each gain and how it spreads across models:
| Benchmark | Gemini 3.1 Pro | GPT-5.5 | Opus 4.8 | Mean of three |
|---|---|---|---|---|
| IMO ProofBench-Adv | +8.6 | +1.3 | +2.2 | +4.0 |
| ARC-AGI-2 | +6.7 | +3.4 | +2.5 | +4.2 |
| LongCoT-mini | +9.2 | +0.4 | +1.2 | +3.6 |
| ProgramBench | +1.8 | +7.8 | +1.9 | +3.8 |
Table 2: Gain of meta-reasoning over Direct Control in points, computed from Table 1. The reasoning-benchmark means match the figures stated in the paper.
Two patterns stand out. First, the gain depends heavily on the model. Gemini 3.1 Pro benefits most on the reasoning benchmarks (its 9.2-point gain on LongCoT-mini is the largest single cell in the paper), GPT-5.5 benefits mostly on ProgramBench, and Opus 4.8 gains less but is positive everywhere. Second, ARC-AGI-2 is the most consistent benchmark, with every model gaining between 2.5 and 6.7 points, while LongCoT-mini has the widest spread, from 0.4 to 9.2.
The production-agent comparison is the number most readers will look for. On ProgramBench, meta-reasoning with GPT-5.5 reaches 71.5% against 58.0% for Codex on the same model, a 13.5-point gap; with Opus 4.8 it reaches 67.2% against 65.5% for Claude Code. It beats mini-SWE Agent with all three models by 2.5 to 13.9 points. The paper itself frames these as end-to-end reference points; the matched Direct Control comparison is what actually tests the effect of changing how the computation is controlled.
Figure 3: Performance across four benchmarks and three models against external research harnesses and production coding agents. Source: paper Figure 2.
More budget keeps helping, while direct control plateaus
A bigger allowance only matters if the agent finds a useful way to spend it. In Figure 4, the left panels show calls actually used and the right panels show final scores. Meta-reasoning spends more as its budget grows: at 1200 calls on ProgramBench it uses roughly 89% to 101% of the allowance across models. Direct control often stops early; with GPT-5.5 on ProgramBench it uses only about 18%.
The extra spending also buys better answers. As the ProgramBench allowance grows from 400 to 1200 calls, meta-reasoning with GPT-5.5 rises from 64.1% to 71.5%, while direct control stays near 64%. Stopping early is not the whole story: with Opus 4.8, direct control roughly doubles the calls it uses over the sweep, from 376 to 768, yet its score moves only from 62.7% to 65.3% and peaks at the middle budget. Continuing to work is not enough; what the agent does next matters too.
Meta-reasoning does not win at every budget. With Opus 4.8 it trails direct control at 400 calls, 56.6% against 62.7%, then leads at 1200 calls, 67.2% against 65.3%. GPT-5.5 is also slightly behind at the smallest ProgramBench allowance before pulling ahead. The authors read this as a fixed cost of staged control that needs enough budget to pay back, consistent with the four calls every deliberation turn spends before any worker runs.
Figure 4: Calls consumed (left; the diagonal is full use of the allowance) and final scores (right) as the nominal allowance grows. Reasoning results aggregate the three reasoning benchmarks; ProgramBench is shown separately. Source: paper Figure 3.
What the extra budget actually computes
At the main budget, meta-reasoning produces more worker artifacts than direct control on all three reasoning benchmarks, and reuse between them grows faster still. On IMO ProofBench-Advanced with Gemini 3.1 Pro, worker artifacts roughly double while recorded dependencies rise by an order of magnitude. Both agents choose context artifacts through the same interface, so direct control could build the same graph; it largely does not.
Graph size also grows with budget. On ProgramBench with GPT-5.5, going from 400 to 1200 calls yields more than 2.5 times as many nodes and nearly four times as many edges, and direct control shows no comparable growth. The matched examples in Figure 6 make the difference visible: on one ARC-AGI-2 problem, direct control produces six independent attempts and four one-hop follow-ups, while meta-reasoning explores more independent attempts, connects later work to earlier results and converges on a frontier of three candidates, combining exploration and reuse in the same run.
Figure 5: Worker-artifact topology on ProgramBench across budgets: nodes, edges, depth and width. Source: paper Figure 4.
Figure 6: Example artifact graphs from meta-reasoning and direct control on matched problems; the ringed node is the submitted artifact. Source: paper Figure 5.
Finding a correct answer is only half the problem
On the reasoning benchmarks the authors grade intermediate solution artifacts and compute coverage, monitoring and selection. The table below collects the values annotated in the paper's Figure 6 together with the figures quoted in the text:
| Model | Benchmark | Coverage gain (points) | FSG: Direct Control | FSG: Meta-Reasoning |
|---|---|---|---|---|
| Gemini 3.1 Pro | IMO ProofBench-Adv | +20 | – | – |
| Gemini 3.1 Pro | ARC-AGI-2 | +7 | – | – |
| Gemini 3.1 Pro | LongCoT-mini | +10 | +16 | +14 |
| GPT-5.5 | ARC-AGI-2 | +4 | +3 | +11 |
| GPT-5.5 | LongCoT-mini | +4 | +4 | +8 |
| Opus 4.8 | ARC-AGI-2 | +1 | +4 | +7 |
| Opus 4.8 | LongCoT-mini | −1 | +3 | +3 |
Table 3: Coverage and frontier selection gain (FSG). Coverage gain is meta-reasoning minus direct control; FSG is the accuracy gain of the actual submission over a uniform pick from the frontier. Source: annotations in paper Figure 6 and Section 6.4.
On coverage, meta-reasoning puts at least one correct candidate into more runs in most settings: 20 points more with Gemini 3.1 Pro on IMO ProofBench-Advanced and 10 points more on LongCoT-mini, with Opus 4.8 on LongCoT-mini roughly tied. On monitoring, worker self-confidence can be a weak signal. For Gemini 3.1 Pro on IMO ProofBench-Advanced its Type-2 AUC is 0.55, close to chance, while the Assess-stage verdict reaches 0.88. The gap is smaller for GPT-5.5 and Opus 4.8, and nearly absent for GPT-5.5 on LongCoT-mini.
The selection result is more modest. In 83% of the ARC-AGI-2 and LongCoT-mini runs the submission comes from the convergence frontier, yet about three quarters of those runs have several candidates there, so the agent still has to choose. With GPT-5.5 on ARC-AGI-2, meta-reasoning gains 11 points over a random frontier pick against 3 for direct control, but with Gemini 3.1 Pro on LongCoT-mini direct control is slightly ahead. A controller that ranks candidates better does not automatically turn that ranking into a better final answer.
Figure 7: (a) correct-candidate coverage; (b) Type-2 AUC of worker confidence and Assess-stage verdicts; (c) frontier selection gain over a uniform choice from each run's deepest terminal artifacts. Source: paper Figure 6.
Keeping useful history without replaying all of it
Artifact memory holds worker outputs and the controller's own notes side by side, and the controller re-reads its notes far more intensively than worker outputs. Among artifacts retrieved at least once, notes average 4.4 to 9.8 reads on ARC-AGI-2 and LongCoT-mini, against 1.9 to 3.3 for worker artifacts. Most artifacts are never read again, but the controller returns repeatedly to a small set of its own notes and uses them as standing organizational references.
Memory operations concentrate in Assess and Propose, and writes outnumber reads for every model. Assess writes steadily but almost never reads, because new artifacts already arrive in its prompt and an explicit read would mean fetching older work. Propose and Evaluate account for nearly all reads. Usage also varies sharply by model: Gemini 3.1 Pro uses memory lightly, GPT-5.5 more often, and Opus 4.8 writes many notes it never revisits.
The main payoff of persistent storage is a small controller state. Direct control's history exceeds one million characters on the longest Opus 4.8 IMO ProofBench-Advanced runs, while the meta-reasoning state stays in the thousands to tens of thousands of characters and often shrinks as a reasoning run progresses. On ProgramBench the gap narrows to a factor of a few and the meta-reasoning state grows over the run, because much of what the controller reads there is repository content it has to keep track of. The authors suggest that direct control's history reaches lengths at which models are known to struggle to find relevant information, which may explain why it stops improving with budget, but they did not test this directly.
Figure 8: Memory use of the Meta-Reasoning Agent on ARC-AGI-2 and LongCoT-mini, grouped by stage, with reads split into worker outputs and controller notes. Source: paper Figure 7.
Figure 9: Size of the representation persisted between control decisions (log scale): the accumulated message history for direct control and the compact Assess state for meta-reasoning. Source: paper Figure 8.
Code and Reproducibility
No code has been released, and a GitHub search turns up no official or third-party implementation. Reproduction therefore rests on the appendix: the full system prompts for all four stages in both the text and coding settings, the Direct Control prompt, the worker prompts, the memory rendering format, and JSON schemas for run_workers, finish, read_memory and write_memory. Two details matter most for anyone rebuilding it. First, ProgramBench workers share one filesystem with no branch isolation, so the Evaluate prompt forces any code-writing action to use a single worker and allows fan-out only for read-only actions. Second, on ProgramBench both of the authors' agents get a read-only git tool (log, show, diff, blame, cat-file and similar, output capped at 8000 characters) so the controller can read the code a worker actually committed rather than trusting its prose summary.
Two public codebases are worth reading alongside the paper. The first is the official ProgramBench repository, facebookresearch/ProgramBench. The metric reported in the paper, the mean per-problem test-pass rate, corresponds to the average pass rate in its batch evaluation summary:
# src/programbench/eval/eval_batch.py
@property
def average_pass_rate(self) -> float:
if not self.summaries:
return 0.0
return sum(s.score for s in self.summaries) / len(self.summaries)
The second is the reference implementation of the Recursive Language Model baseline, alexzhang13/rlm. Its entry point falls back to a plain model call once the maximum recursion depth is reached, which is the concrete form of the paper's description of a context held in a variable the model edits with code rather than in a transcript:
# rlm/core/rlm.py, RLM.completion
# If we're at max depth, the RLM is an LM, so we fallback to the regular LM.
if self.depth >= self.max_depth:
return self._fallback_answer(prompt)
For comparability the authors changed this baseline in two ways. They replaced the full Python REPL with a workspace validated against an abstract-syntax-tree allowlist that rejects loops, arithmetic, function definitions and imports, so every unit of real work must pass through a sub-model call. They also added the same budget accounting as the other systems, injecting a message at every tenth of the allowance and requiring submission at 90%.
Limitations
First, staged control costs compute (stated by the authors). The low-budget crossovers in Figure 4 show the design is not uniformly better than direct control; with Opus 4.8 at 400 calls it trails by 6.1 points. A wrong controller assessment can also propagate a misleading state or discard good partial work that is never re-read, and the compact state is lossy, relying on later memory reads to recover information that only becomes important later.
Second, the gains cannot be attributed to individual components (stated by the authors). The matched comparison changes the compact state, the staged agentic process and the controller's memory access all at once. It establishes the value of the design as a whole, not the necessity of any stage. Stage-removal, state-only and memory-interface ablations would be needed, and the paper does not run them.
Third, resources are matched only at the level of call counts (stated by the authors). Calls are easy to interpret but vary widely in token length and wall-clock time, so the paper draws no conclusions about token cost, latency or performance at equal cost. For deployment that is exactly the missing number: the four-stage controller spends four calls per turn before any worker runs, and each of those calls reads the state and the memory index, so the real bill need not scale with call counts.
Fourth, breadth and uncertainty (stated by the authors). Three frontier models and four benchmarks, one of them a 30-problem proof set; positive point estimates do not establish significance for every difference, nor generalization to weaker models, other tasks or much longer horizons. The LongCoT-mini chess subset is a concrete counterexample: meta-reasoning underperforms direct control there, and preliminary trajectory sampling points to unproductive reconsideration, where extra checking destabilizes an answer that was already correct.
Fifth, the production-agent comparison deserves a careful reading (our assessment). Claude Code and Codex run headless with their native tools disabled, and every action goes through a single container_bash MCP tool, so file edits happen only through heredocs and shell commands. They also do not receive the read-only git tool that both of the authors' agents get. The setup keeps the action channel identical, but it also means neither production agent runs in its native configuration. The 71.5% versus 58.0% result is best read as a comparison under one restricted container interface, not as a ranking of the shipped products.
Conclusion and Outlook
The contribution comes in three layers. Conceptually, the paper separates deciding what to compute next from doing the computation, and argues that meta-level cognition is itself work that needs multiple steps, dedicated tools and sub-tasks. As a system, it offers a concrete inference-time harness: four separate stages, a compact state backed by persistent artifact memory, and a stopping rule that can only submit artifacts that already exist. Methodologically, it introduces the artifact graph and the coverage, monitoring, selection and frontier-selection-gain diagnostics, which split one final score into three separately improvable questions: what computation happened, whether it produced correct work, and whether that work reached the final answer.
For engineers building coding agents and long-horizon harnesses, several practices transfer directly. Hide the budget from the stage that generates options, so expensive but valuable options are not filtered out early. Write controller state as a cited factual snapshot that separates implemented, broken, unverified, unexplored and dead-end items. Treat worker self-confidence as a weak signal and have an independent stage issue verdicts. Make every worker instruction self-contained instead of passing an internal identifier. Read a worker's actual commits rather than trusting its summary.
Open questions worth watching include component ablations that show which stages matter in which regimes, cost reported in tokens and latency rather than calls, tests on weaker models and longer horizons, and stopping and commitment rules that address the over-checking failure seen on chess. The paper's closing line states its position plainly: as runs grow longer, deciding what work is worth doing becomes work in its own right.
An agent's ceiling depends not only on how well it does the work, but on whether it can decide, before acting, which work is worth doing.
SOURCE LINKS



