Claude Fable 5.1: the long-horizon agentic tier that tops the ladder on a different ruler
claude-fable-5-1
The fixed version Anthropic positions for demanding reasoning and long-horizon agentic work (claude-fable-5-1, retirement no earlier than 2027-09-01): default thinking effort high at $10/$50 per million tokens, 2.5x sibling Opus 5.5, and default-high is itself a default cost behaviour. Its standing only holds once you switch rulers: #1 on net_improvement in the Arena Agent board (13.71, confirmed_success 19.83, praise 31.83) ahead of GPT-6 Astra (11.54/17.70/32.79), while Opus 5.5 misses the top eight; #2 on Terminal-Bench 4.0 at 57.88% (n=330), 0.3 points behind Astra at 58.18% but with pass@5 of 0.7879 against 0.7121, lower single-shot and more robust over five attempts; #4 on the intelligence index at 53.35 (max), below Opus 5.5. The picture is consistent: not first on composite intelligence, first on pushing a real task forward. A separate praise column means the board carries a human or judge component and is not an objective benchmark. Choose by whether the workload is hard from reasoning or from long-horizon consistency. Closed, API only. Graded A (confirmed); no long-task comparison on our own harness.
- CONFIDENCE
- Confirmed
- Two or more independent sources, or reproduced by our harness
- KEY METRIC
- Arena Agent 净改进
- Confirmed · 2026-09
- MATURITY
- Product
- research → demo → product → production
Our takeWe grade it A (confirmed): Anthropic's official documentation supplies the specs, three third-party boards synced into this site supply the readings, and the two are independent. Grade A guarantees only that the numbers are real, not that they move in the same direction on your workload - especially since Arena Agent is a judged protocol, and the presence of a
praisecolumn already tells you subjective scoring is baked in.The thing most worth remembering is that it tops the list only after the ruler changes. Two models from the same vendor and the same generation: on the intelligence index Opus 5.5 leads Fable 5.1 by 4.27 points, while on the agent board Fable 5.1 leads so decisively that Opus 5.5 does not appear in the top eight. That is not a contradiction, it is two boards measuring different things - composite intelligence versus long-horizon task progress. Our judgement is that for anyone putting an agent into a real workflow, the latter is closer to production truth, which is why it sits second on our reasoning ladder rather than being ranked by absolute score.
The Terminal-Bench pairing is unusually informative: single-shot accuracy of 57.88% is 0.3 points below the leader, yet pass@5 is 0.7879 against 0.7121. That means its failures are retryable failures (close misses that pass on another path) rather than systematic drift. Inside an agent pipeline, retryable failures are far cheaper than non-retryable ones, because retries are automatic while drift needs a human to catch. Hard boundaries: it is expensive, and expensive at the default tier; it is not first on the intelligence index, so picking the wrong evaluation frame means buying the wrong model; closed source with data leaving your perimeter.
The problem it solves: treating "pushing a real task forward" as its own objective
Most model leaderboards measure answer quality: given a question, was the reply right, was it good. Claude Fable 5.1 targets something else. Anthropic's official positioning reads demanding reasoning and long-horizon agentic work. Translated into production terms: it is not optimised for how well this turn reads, but for how far a real task has moved once it is handed over.
That difference is worth real money. A coding agent that runs for two hours reads code, edits code, runs tests, reads the failure, edits again - and whether the final diff is mergeable has almost nothing to do with the prose quality of any individual turn. What decides it is long-horizon coherence: did the goal drift, was the assumption from step four remembered at step thirty, and after a failure did it fall into the same hole twice.
Official specs
| Dimension | Official reading | Difference from Opus 5.5 |
|---|---|---|
| API ID | claude-fable-5-1 | Same generation, different positioning, both pinnable versions |
| Official positioning | demanding reasoning and long-horizon agentic work | Opus 5.5 is long-running agentic coding and knowledge work |
| Default thinking effort | high | Opus 5.5 defaults to medium; this is the substantive behavioural difference |
| Price | $10 in / $50 out | Opus 5.5 is $4 / $20, so Fable costs 2.5x more |
| Retirement commitment | Not before 2027-09-01 | Three weeks earlier than Opus 5.5, still a 12-month-class window |
The "default high" row deserves its own paragraph. It is not marketing language, it is default cost behaviour: out of the box, every Fable 5.1 call burns more thinking tokens than a default Opus 5.5 call. Confirm your task actually needs that effort before choosing it, otherwise you are paying 2.5x unit price multiplied by a higher thinking spend.
Third-party water level: first on the agent board, not first on the intelligence index
Only by putting two boards side by side does the positioning become legible - this model reaches the top by changing the ruler:
- Arena Agent board (net improvement): #1 Claude Fable 5.1 (Max),
net_improvement13.71,confirmed_success19.83,praise31.83. #2 is GPT 6 Astra (Max) at 11.54 / 17.70 / 32.79. Ranks 3 and 4 are Claude Opus 5 at high and max effort (10.25 / 10.16), #5 is Claude Fable 5 (High) at 8.81. Opus 5.5 does not appear in this board's top eight. - Artificial Analysis intelligence index: Fable 5.1 scores 53.35 at max (rank 4) and 53.20 at xhigh (rank 5) - below Opus 5.5's 57.62 / 55.99 / 53.58 at ranks 1, 2 and 3.
- Terminal-Bench 4.0: Fable 5.1 ranks #2 at 57.88% accuracy (n=330), only 0.3 points behind leader GPT-6 Astra's 58.18%; but its
pass@5is 0.7879 against Astra's 0.7121. Slightly lower single-shot accuracy, higher pass rate when given five attempts.
These three readings compose into one coherent picture: it is not first on the "composite intelligence" ruler, it is first on the "long-horizon task progress" ruler, and on "complete real engineering work in a terminal" it is effectively tied with the leader while being more robust under retry. For anyone putting an agent into a real workflow, the latter two rulers are the more relevant ones.
The semantics of net_improvement and confirmed_success also need spelling out: the first measures net forward progress against a baseline (making things worse subtracts), the second is the share of confirmed successes. Fable 5.1's 13.71 and 19.83 both exceed Astra's 11.54 and 17.70, yet Astra's praise (32.79) is higher - "looked good" and "actually made progress" are separate columns on this board, and that separation is precisely why it is more informative than a QA leaderboard.
Which workloads should pick it
- Pick it: long autonomous coding runs, refactors that must stay coherent across dozens of files, agent pipelines expected to "run and deliver" rather than pause for human confirmation at every step.
- Do not pick it: high-concurrency short QA, cost-sensitive batch work, workflows where a human reviews each step anyway - Opus 5.5's default tier covers these at 2.5x lower cost.
- Measure before deciding: is your task hard because of reasoning depth or hard because of long-range coherence? The intelligence index is more relevant for the former (Opus 5.5 leads), the agent board for the latter (Fable 5.1 leads). Getting that diagnosis wrong costs you more money and no improvement.
Boundaries
- Expensive, and expensive by default: $10/$50 plus a default of high effort means cost per task can exceed three times that of the same generation's workhorse tier. Retry-heavy agent loops amplify that multiple.
- Not first on the intelligence index: if your evaluation frame is composite intelligence rather than task progress, this is the wrong pick - Opus 5.5 at max scores 4.27 points higher.
- Arena Agent is a judged protocol: the existence of a
praisecolumn means subjective scoring is part of it. However solid net improvement looks, carry an evaluator-bias discount; do not read it as an objective benchmark. - Closed source, API only: no self-hosting, no fine-tuning, data leaves your perimeter.
Our verification status
Facts come from two independent sources: Anthropic's official model documentation (platform.claude.com models overview, read directly for positioning, default effort, pricing and retirement date) and the raw rows of three boards synced into this site - Arena Agent, the Artificial Analysis intelligence index and Terminal-Bench 4.0 (fetched 2026-09-22, field names and readings matched verbatim). Vendor documentation and third-party boards are independent of each other, hence confidence A (confirmed).
The limit of grade A still needs stating: we have not reproduced the capability claims. We have not run Fable 5.1 against Opus 5.5 or GPT-6 Astra on long-horizon tasks in our own harness, have not verified that Arena Agent's net_improvement moves in the same direction on our workloads, and have not measured the real token spend of the default high tier. The section on which workloads should pick it is inference drawn from board semantics, not a measured conclusion of ours. Reaching measured grade would require: a net-progress comparison of the three models on one fixed set of real tasks, a cost-effectiveness curve across default and high effort, and a recomputation of retry robustness (pass@5) on our own task set.