Claude Opus 5.5: the long-running agentic tier at $4/$20 with 1M context and adaptive thinking
claude-opus-5-5
The fixed version Anthropic positions for long-running agentic coding and knowledge work (claude-opus-5-5, retirement no earlier than 2027-09-22): 1M-token context, 128K-token single output, adaptive thinking defaulting to medium, at $4/$20 per million tokens, the cheapest of the three top rows we track. On the Artificial Analysis intelligence index this site syncs (2026-09-22) the top three rows are all its thinking tiers (max 57.62, xhigh 55.99, high 53.58), while the default medium tier sits #8 at 51.24, with sibling Fable 5.1 (53.35) and GPT-6 Astra (52.67) in between. Two discounts: those rows carry the with fallback qualifier while Astra reads (max), so the protocols differ, and the 53.94% at #8 on Terminal-Bench belongs to the previous-generation Opus 5. Closed, API only; adaptive thinking is a black box, so budget on p90 rather than the mean; 1M context is not 1M of effective attention, so whole-repo input still needs retrieval. Use it for most workloads and move to Fable 5.1 only when the highest tier is not enough: 2.5x the price for 4.3 index points. Graded A (confirmed); not benchmarked by us.
- CONFIDENCE
- Confirmed
- Two or more independent sources, or reproduced by our harness
- KEY METRIC
- AA 智能指数
- Confirmed · 2026-09
- MATURITY
- Production
- research → demo → product → production
Our takeWe grade it A (confirmed) on the strength of two mutually independent sources: Anthropic's official model documentation for specs and pricing, and the Artificial Analysis intelligence index we sync for third-party readings. This is the only evidence shape in our confidence system that earns an A naturally - specs are checkable, readings are recomputable, and the two do not share an origin. Grade A is responsible only for the factual layer, not for how the model behaves on your workload; we have not measured that half.
Its real value is not the 57.62 at the top, it is that the default effort tier already scores 51.24 and lands inside the top band. The production significance of this is badly underestimated: most teams do not hand-pick a thinking effort per call, they run the default. When the default is already at the top of the ladder, the gap between "no tuning at all" and "tuned to maximum" is 6.4 points rather than the dozen-plus points that used to be normal. Combined with $4/$20 pricing, it is currently the only flagship where "high capability" and "low cost tier" hold at the same time, which is why it sits first in our reasoning domain.
The hard boundaries need stating just as clearly: closed source, API only, data leaves your perimeter - none of which a compliance-sensitive team can route around. Adaptive thinking makes per-call cost imprecise, so budget against p90. Mid-context recall at 1M tokens remains an industry-wide weakness, so whole-repository stuffing is not a substitute for retrieval. One more trap worth flagging: the board rows carry a
with fallbackqualifier while GPT-6 Astra's do not, so subtracting 52.67 from 57.62 and calling it a five-point lead is not rigorous. Leave a protocol discount in that comparison.
The problem it solves: an agentic workhorse has to run long and run cheap
Frontier model competition in 2026 is no longer about how clever a single answer is. What actually blocks production is something else: an agent has to stay on track across dozens of tool calls, read in an entire repository, and notice on its own that step eleven was wrong and go back to fix it. Anthropic's official positioning for Claude Opus 5.5 is exactly that sentence - long-running agentic coding and knowledge work. Not "the smartest model", but "the one that keeps going".
That positioning drives every spec trade-off. Context goes to 1M tokens (a whole mid-size repository, not a handful of files), single-response output goes to 128K tokens (a module written in one pass instead of five), the thinking budget is adaptive rather than a fixed dial (cheap calls stay cheap, hard ones automatically think longer), and price sits at $4 / $20 per million tokens - 60% below the same-generation high tier. For an agent that runs thousands of times a day in CI, cost is not a footnote, it is the precondition for shipping at all.
Official specs: which numbers you can plan against
| Dimension | Official reading | Production meaning |
|---|---|---|
| API ID | claude-opus-5-5 | A pinnable version string, not a -latest alias |
| Context window | 1M tokens | Repository-scale input; long-running state need not be pushed out to a vector store every turn |
| Max output | 128K tokens | A complete module or a long report in one response, removing one stitching layer |
| Thinking control | Adaptive thinking, default medium | No manual effort selection per task; dial it down explicitly when cost matters more |
| Price | $4 in / $20 out per million tokens | The cheapest of the three entries at the top of this ladder; the other two are $10 / $50 |
| Retirement commitment | Not before 2027-09-22 | At least a 12-month window, enough to pin production dependencies to it |
Third-party water level: one model holding the top three index slots
On the Artificial Analysis intelligence index we sync (fetched 2026-09-22), the top three rows are the same model at different thinking efforts:
- #1: Claude Opus 5.5 (max with fallback), index 57.62
- #2: Claude Opus 5.5 (xhigh with fallback), 55.99
- #3: Claude Opus 5.5 (high with fallback), 53.58
- #8: Claude Opus 5.5 (medium with fallback), 51.24 - that is the reading for the default effort
Sandwiched between them are its own sibling Claude Fable 5.1 (max 53.35, xhigh 53.20, ranks 4 and 5) and OpenAI's GPT-6 Astra (max 52.67, xhigh 52.39, ranks 6 and 7). Put differently: take the default medium tier (51.24) and it is still only 6.4 points below the leader, while sitting above everything from rank 8 down - at 40% of the price. That is the direct basis for placing it at the top of our reasoning ladder: not a single highest point score, but simultaneously holding a top position and the low price on the same third-party ruler.
Two caveats stated plainly. First, the with fallback qualifier in those row names means these are retry/fallback readings, not single-sample ones; GPT-6 Astra's rows on the same board read (max) without that qualifier, so strictly the two are not measured under an identical protocol and cross-model comparison deserves a discount. Second, the row at rank 8 on the Terminal-Bench 4.0 board we sync is Opus 5 (53.94%) - the previous generation, not 5.5. We do not use a prior generation's reading to endorse a new model.
Model-selection guidance: when to step up a tier
Anthropic's official selection order is blunt: use Opus 5.5 for most workloads, and only move to Fable 5.1 when Opus 5.5 at its highest effort is still not enough. That order is worth copying verbatim, because the price gap is 2.5x ($4/$20 against $10/$50) while the intelligence-index gap runs the other way - Fable 5.1's max reading (53.35) is below Opus 5.5's max reading (57.62).
So what is Fable 5.1's premium buying? A different ruler: on the Arena Agent board, which measures how far a model actually pushes a real task forward, the number one for net improvement is Fable 5.1 (13.71), and Opus 5.5 does not appear in that board's top eight. This is therefore not a question of which model is stronger, but of whether your workload is composite intelligence or long-horizon autonomous progress. Treating the two as adjacent rungs on one ladder gets you both an overspend and the wrong tier.
Boundaries
- Closed source, API only: no self-hosting, no weight access, data leaves your perimeter. For teams whose compliance forbids that, it is a hard boundary rather than a configuration issue.
- Adaptive thinking is a black box: a default of medium means you cannot predict thinking-token spend per call. Budget against p90, not the mean.
- 1M context is not 1M effective attention: mid-document recall in very long contexts is broadly worse than at the edges, so dumping a whole repository in still needs retrieval as a backstop.
- Long-horizon failure is cumulative: the official positioning is long-running, but error recovery after dozens of turns still depends on the harness providing checkpoints and rollback. Do not assume the model catches itself.
Our verification status
Facts on this page come from two independent sources: Anthropic's official model documentation (the models overview on platform.claude.com, read directly for specs, pricing, thinking control and retirement date) and the Artificial Analysis intelligence-index rows already synced into this site (fetched 2026-09-22, with all four effort readings and their row-name qualifiers matched verbatim). The two sources are independent of each other, hence confidence A (confirmed).
Grade A covers only the layer that says the specs and the board readings are real. We have not reproduced the capability claims: we have not run Opus 5.5 against Fable 5.1 or GPT-6 Astra on long-horizon tasks in our own agent harness, have not quantified mid-context recall at 1M tokens, and have not measured the actual token-spend distribution of adaptive thinking under real load. Those measurements are the premise behind every "production meaning" inference on this page. Read them as assumptions, not as measurements.