Skip to content
←Back to the radar

SOTA · language

LLM

The foundation-model layer: the weights themselves plus what is tightly coupled to them - training recipes (pretraining, post-training, RL/RLHF), architecture and long context, and the inference serving stack with its cost per token

RULERIntelligence: Artificial Analysis Intelligence Index, GPQA, MMLU-Pro, Arena Elo. Capacity: effective long-context recall, modalities. Supply: tok/s, $ per 1M tokens, open-weight licence

The ladder

14
ReasoningDevelopmentTopA

Claude Opus 5.5: the long-running agentic tier at $4/$20 with 1M context and adaptive thinking

The fixed version Anthropic positions for long-running agentic coding and knowledge work (claude-opus-5-5, retirement no earlier than 2027-09-22): 1M-token context, 128K-token single output, adaptive thinking defaulting to medium, at $4/$20 per million tokens, the cheapest of the three top rows we track. On the Artificial Analysis intelligence index this site syncs (2026-09-22) the top three rows are all its thinking tiers (max 57.62, xhigh 55.99, high 53.58), while the default medium tier sits #8 at 51.24, with sibling Fable 5.1 (53.35) and GPT-6 Astra (52.67) in between. Two discounts: those rows carry the with fallback qualifier while Astra reads (max), so the protocols differ, and the 53.94% at #8 on Terminal-Bench belongs to the previous-generation Opus 5. Closed, API only; adaptive thinking is a black box, so budget on p90 rather than the mean; 1M context is not 1M of effective attention, so whole-repo input still needs retrieval. Use it for most workloads and move to Fable 5.1 only when the highest tier is not enough: 2.5x the price for 4.3 index points. Graded A (confirmed); not benchmarked by us.

57.62AA 智能指数Confirmed · 2026-09
ProductionAnthropicSite
Claude Opus 5.5: the long-running agentic tier at $4/$20 with 1M context and adaptive thinking
LLMDevelopmentTopC

Jev: TypeSafe's first System One Model types decisions and prices them two orders of magnitude lower

TypeSafe's first System One Model: outputs are type-safe structured values under three question primitives (Noul/Choice/Score) rather than strings, trained with RLCD for calibrated decision probabilities. The official 13-question parallel experiment makes one synthesized call ~12.2x cheaper and ~10.0x faster than 13 separate calls; four-workflow mean accuracy 67.8% at /tmp/run_ingest2.sh.0004 and 0.4s per instance - opus 5 / sol tier accuracy at almost two orders of magnitude less cost; input $0.042/M tokens, output free. Ships an OpenJev local reproduction path and an eight-item caveats list.

67.8%四工作流平均准确率(官方评测)Vendor Claim · 2026-09
ProductTypeSafeSiteRepo
Jev: TypeSafe's first System One Model types decisions and prices them two orders of magnitude lower
LLMDevelopmentTopA

MiMo-V2.6: Live RL as an auditable ledger of 30 steps, 750k trajectories, and dollars

Xiaomi's open-source MiMo-V2.6 Pro/Flash: 30 steps of Live RL, roughly 750k trajectories, public cost of $850k/$2.62M, and 7k+ RL environments released in the same batch; tops the open-weights camp on the Artificial Analysis intelligence index at 46 with +17/+14 points out-of-sample on DeepSWE v1.1; four demo lines (Vibe World, CUA, research, content creation) all open. Self-improvement becomes an auditable ledger - steps, trajectories, and dollars on the table for third-party recomputation.

46AA 智能指数(开源榜首)Confirmed · 2026-09
ProductXiaomiSiteRepo
MiMo-V2.6: Live RL as an auditable ledger of 30 steps, 750k trajectories, and dollars
LLMDevelopmentTopC

GLM-5.3: the open-weights coding flagship bought with post-training alone

Z.ai post-trained the very same base as GLM-5.2 and nothing else, producing the 744B-A40B GLM-5.3: 50 percent over 5.2 on the in-house Z.ai Code Bench, open-weights first on Terminal Bench 3.0 at 28.3 and Agents' Last Exam at 28.5, while AutomationBench 48.2 and GDPVal-AA v2 1769 beat every closed-source comparison on the official chart; CyberGym vulnerability discovery is SOTA and the exploitation-chain gains emerged without dedicated training. The companion 320B-A18B Flash starts from a new base and is the first GLM to mix sparse with linear attention, attacking long-context serving cost. Weights ship in FP8 and BF16, deployable on SGLang, vLLM, KTransformers and the Ascend stack.

28.3Terminal Bench 3.0(开源权重第一, 官方图)Vendor Claim · 2026-08
ProductZ.ai (Zhipu AI)SiteRepo
GLM-5.3: the open-weights coding flagship bought with post-training alone
ReasoningDevelopmentTopA

Claude Fable 5.1: the long-horizon agentic tier that tops the ladder on a different ruler

The fixed version Anthropic positions for demanding reasoning and long-horizon agentic work (claude-fable-5-1, retirement no earlier than 2027-09-01): default thinking effort high at $10/$50 per million tokens, 2.5x sibling Opus 5.5, and default-high is itself a default cost behaviour. Its standing only holds once you switch rulers: #1 on net_improvement in the Arena Agent board (13.71, confirmed_success 19.83, praise 31.83) ahead of GPT-6 Astra (11.54/17.70/32.79), while Opus 5.5 misses the top eight; #2 on Terminal-Bench 4.0 at 57.88% (n=330), 0.3 points behind Astra at 58.18% but with pass@5 of 0.7879 against 0.7121, lower single-shot and more robust over five attempts; #4 on the intelligence index at 53.35 (max), below Opus 5.5. The picture is consistent: not first on composite intelligence, first on pushing a real task forward. A separate praise column means the board carries a human or judge component and is not an objective benchmark. Choose by whether the workload is hard from reasoning or from long-horizon consistency. Closed, API only. Graded A (confirmed); no long-task comparison on our own harness.

13.71Arena Agent 净改进Confirmed · 2026-09
ProductAnthropicSite
Claude Fable 5.1: the long-horizon agentic tier that tops the ladder on a different ruler
ReasoningDevelopmentTopA

GPT-6 Astra: first on terminal engineering tasks, third tier on composite intelligence

The OpenAI flagship tier, officially positioned for hardest end-to-end work (gpt-6-astra): 1.05M-token context, 128K-token single output, knowledge cutoff 2026-04-30, with Functions, Web search, File search and Computer use built in, so search, retrieval and interface control ship with the model instead of an outer agent framework. Reasoning effort runs low to max in five steps while price stays fixed at $10/$50, so the cost lever is token consumption; siblings Sol ($2/$10) and Luna ($0.1/$0.5) make a 100x spread, so load splitting can stay inside one vendor. Readings (2026-09-22): #1 on Terminal-Bench 4.0 at 58.18% (n_trials=330, pass@5 0.7121), the protocol closest to a coding agent's daily life; #2 on Arena Agent with net_improvement 11.54 and confirmed_success 17.70, both below leader Fable 5.1, but the field's highest praise at 32.79; #6 and #7 on the intelligence index (max 52.67). The cutoff is nearly five months older than this page, so version numbers and API changes must go through Web search; computer-use reliability appears on no board; closed, API only. Graded A (confirmed); not benchmarked by us.

58.18%Terminal-Bench 准确率Confirmed · 2026-09
ProductOpenAISite
GPT-6 Astra: first on terminal engineering tasks, third tier on composite intelligence
LLMDevelopmentC

Laya: a 421M non-autoregressive System 1 decision model at 33 ms per forward pass

Convai Innovations' open-source non-autoregressive System 1 decision model: 421M parameters, ModernBERT backbone with [MASK] option extraction, one 33 ms forward pass per calibrated decision, Apache-2.0; trained with RLCD strictly proper scoring rules for honest probabilities; the model card ships a real benchmark against Jev plus an honest limitations list. Positioned for edge and high-concurrency small decisions (game NPCs, dialogue policy, request routing), not long reasoning or open generation.

33 ms单次前向延迟(421M,官方)Vendor Claim · 2026-09
ResearchConvai InnovationsSiteRepo
Laya: a 421M non-autoregressive System 1 decision model at 33 ms per forward pass
LLM

llama.cpp: seventeen backends and 1.5-to-8-bit quantisation, taking LLM inference anywhere a compiler exists

An MIT-licensed, dependency-free plain C/C++ inference implementation built on ggml. Seventeen backends span CUDA, Metal, HIP, Vulkan, SYCL, WebGPU, CANN, Hexagon, MUSA, zDNN and ZenDNN; integer quantisation runs 1.5 to 8 bits; CPU+GPU hybrid inference makes models larger than VRAM runnable at all; and GGUF, its output format, is a de facto standard that even vLLM consumes. Two lines - llama cli or llama serve - start an OpenAI-compatible server.

130kGitHub stars
llama.cpp: seventeen backends and 1.5-to-8-bit quantisation, taking LLM inference anywhere a compiler exists
LLM

vLLM: the open serving stack that grew out of PagedAttention, with 200+ architectures and five parallelism axes

An inference and serving library out of UC Berkeley's Sky Computing Lab with 2000+ contributors: paged KV memory, continuous batching and chunked prefill, CUDA/HIP graphs with torch.compile, quantisation across FP8/MXFP4/NVFP4/GGUF, pluggable attention kernels (FlashInfer, TRTLLM-GEN, FlashMLA), four speculative decoding routes (n-gram, suffix, EAGLE, DFlash), tensor/pipeline/data/expert/context parallelism with P/D disaggregation, and both OpenAI and Anthropic protocols.

93kGitHub stars
vLLM: the open serving stack that grew out of PagedAttention, with 200+ architectures and five parallelism axes
LLM

unsloth: the 2x-faster, 70%-less-VRAM fine-tuning library, now a desktop runtime that wires local weights into Claude Code and Codex

Self-described as the first desktop app to run and train models (Windows/macOS/Linux; multi-GPU across NVIDIA, AMD, Intel, CPU and Vulkan). Fine-tuning is claimed at 2x faster with 70% less VRAM, the post-training menu covers LoRA, QLoRA, full fine-tuning, pretraining, GRPO, DPO and FP8, and export reaches GGUF, NVFP4 and FP8. Unsloth Start connects Claude Code or Codex to local weights in one command; the model list includes Qwen3.8, GLM-5.3-Flash, Kimi K3, DeepSeek-V4, MiniMax-H3 and Gemma 4, with private web search, deep research, auto-compaction and RAG built in.

77kGitHub stars
unsloth: the 2x-faster, 70%-less-VRAM fine-tuning library, now a desktop runtime that wires local weights into Claude Code and Codex
LLM

SGLang: the serving framework that made prefix reuse a radix tree, and a rollout backend for frontier RL post-training

A high-performance serving framework hosted by LMSYS: RadixAttention prefix caching, a zero-overhead CPU scheduler, prefill-decode disaggregation, DFlash and Spec V2 speculative decoding, compressed finite state machines for structured output, and large-scale expert parallelism (96 H100 GPUs; 3.8x prefill and 4.8x decode on GB200 NVL72 part II). Its day-0 ledger covers Kimi K3, DeepSeek-V4, GLM5.2 NVFP4, Nemotron 3 and MiniMax M2, and AReaL, Miles, slime, Tunix and verl all use it as an RL rollout backend.

37kGitHub stars
SGLang: the serving framework that made prefix reuse a radix tree, and a rollout backend for frontier RL post-training
LLM

TensorRT-LLM: NVIDIA's specialised-kernel and rack-scale engine, with agentic serving as a first-class topic

NVIDIA's inference optimisation library for LLMs and visual generation models, fully open source since 2025-03-22. Its 28 tech blogs form an auditable engineering ledger: Skip Softmax and sparse attention for long context, the three-part expert parallelism series with one-sided AlltoAll over NVLink, DWDP for NVL72, guided decoding cooperating with speculative decoding, inference-time compute, and evaluating agentic serving with trace replay and job-level metrics. Claims include Llama 4 Maverick above 1,000 TPS per user and 40,000+ tok/s on B200; Bing and NAVER Place run it in production.

15kGitHub stars
TensorRT-LLM: NVIDIA's specialised-kernel and rack-scale engine, with agentic serving as a first-class topic
LLM

Dynamo: the datacenter-scale orchestration layer above inference engines, with KV-aware routing, tiered KVBM and an SLA-driven planner

NVIDIA's open-source (Apache-2.0) orchestration layer, Rust for the performance path and Python for extensibility. It explicitly does not replace SGLang, TensorRT-LLM or vLLM - it wires them into a coordinated multi-node system: disaggregated prefill/decode, KV-aware routing (2x faster TTFT on Qwen3-Coder 480B), KVBM offloading KV cache to CPU/SSD/remote, ModelExpress streaming weights over NVLink for 7x faster cold starts, an SLA-driven Planner (80% fewer breaches at 5% lower TCO in Alibaba production), and Grove for topology-aware NVL72 scheduling.

8.2kGitHub stars
Dynamo: the datacenter-scale orchestration layer above inference engines, with KV-aware routing, tiered KVBM and an SLA-driven planner
LLM

NInfer: A From-Scratch C++/CUDA Inference Engine for a Single RTX 5090

A from-scratch C++/CUDA inference engine for five explicitly registered Qwen checkpoints on one NVIDIA GeForce RTX 5090. Startup-frozen residency picks MTP or DFlash speculative decoding, Vision, and one of five KV storage formats; a shared Device KV pool plus pinned Host State/KV checkpoints reuse exact prompt prefixes across 240k-token contexts. Measured aggregate decode reaches 1,146.9 tok/s at concurrency 8 and 15,544.3 tok/s on a 7,680-token prefill.

1.4kGitHub stars
NInfer: A From-Scratch C++/CUDA Inference Engine for a Single RTX 5090

Evidence

12
2026-09-28unsloth: the 2x-faster, 70%-less-VRAM fine-tuning library, now a desktop runtime that wires local weights into Claude Code and CodexSelf-described as the first desktop app to run and train models (Windows/macOS/Linux; multi-GPU across NVIDIA, AMD, Intel, CPU and Vulkan). Fine-tuning is claimed at 2x faster with 70% less VRAM, the post-training menu covers LoRA, QLoRA, full fine-tuning, pretraining, GRPO, DPO and FP8, and export reaches GGUF, NVFP4 and FP8. Unsloth Start connects Claude Code or Codex to local weights in one command; the model list includes Qwen3.8, GLM-5.3-Flash, Kimi K3, DeepSeek-V4, MiniMax-H3 and Gemma 4, with private web search, deep research, auto-compaction and RAG built in.2026-09-28Dynamo: the datacenter-scale orchestration layer above inference engines, with KV-aware routing, tiered KVBM and an SLA-driven plannerNVIDIA's open-source (Apache-2.0) orchestration layer, Rust for the performance path and Python for extensibility. It explicitly does not replace SGLang, TensorRT-LLM or vLLM - it wires them into a coordinated multi-node system: disaggregated prefill/decode, KV-aware routing (2x faster TTFT on Qwen3-Coder 480B), KVBM offloading KV cache to CPU/SSD/remote, ModelExpress streaming weights over NVLink for 7x faster cold starts, an SLA-driven Planner (80% fewer breaches at 5% lower TCO in Alibaba production), and Grove for topology-aware NVL72 scheduling.2026-09-28TensorRT-LLM: NVIDIA's specialised-kernel and rack-scale engine, with agentic serving as a first-class topicNVIDIA's inference optimisation library for LLMs and visual generation models, fully open source since 2025-03-22. Its 28 tech blogs form an auditable engineering ledger: Skip Softmax and sparse attention for long context, the three-part expert parallelism series with one-sided AlltoAll over NVLink, DWDP for NVL72, guided decoding cooperating with speculative decoding, inference-time compute, and evaluating agentic serving with trace replay and job-level metrics. Claims include Llama 4 Maverick above 1,000 TPS per user and 40,000+ tok/s on B200; Bing and NAVER Place run it in production.2026-09-28llama.cpp: seventeen backends and 1.5-to-8-bit quantisation, taking LLM inference anywhere a compiler existsAn MIT-licensed, dependency-free plain C/C++ inference implementation built on ggml. Seventeen backends span CUDA, Metal, HIP, Vulkan, SYCL, WebGPU, CANN, Hexagon, MUSA, zDNN and ZenDNN; integer quantisation runs 1.5 to 8 bits; CPU+GPU hybrid inference makes models larger than VRAM runnable at all; and GGUF, its output format, is a de facto standard that even vLLM consumes. Two lines - llama cli or llama serve - start an OpenAI-compatible server.2026-09-28SGLang: the serving framework that made prefix reuse a radix tree, and a rollout backend for frontier RL post-trainingA high-performance serving framework hosted by LMSYS: RadixAttention prefix caching, a zero-overhead CPU scheduler, prefill-decode disaggregation, DFlash and Spec V2 speculative decoding, compressed finite state machines for structured output, and large-scale expert parallelism (96 H100 GPUs; 3.8x prefill and 4.8x decode on GB200 NVL72 part II). Its day-0 ledger covers Kimi K3, DeepSeek-V4, GLM5.2 NVFP4, Nemotron 3 and MiniMax M2, and AReaL, Miles, slime, Tunix and verl all use it as an RL rollout backend.2026-09-28vLLM: the open serving stack that grew out of PagedAttention, with 200+ architectures and five parallelism axesAn inference and serving library out of UC Berkeley's Sky Computing Lab with 2000+ contributors: paged KV memory, continuous batching and chunked prefill, CUDA/HIP graphs with torch.compile, quantisation across FP8/MXFP4/NVFP4/GGUF, pluggable attention kernels (FlashInfer, TRTLLM-GEN, FlashMLA), four speculative decoding routes (n-gram, suffix, EAGLE, DFlash), tensor/pipeline/data/expert/context parallelism with P/D disaggregation, and both OpenAI and Anthropic protocols.2026-09-12Sam Altman: The Next ChatGPT Will Watch Your Screen and Remember EverythingSam Altman says that within about six months the next generation of ChatGPT will watch your screen, listen to your meetings and calls, and quietly assemble a persistent context of your work. The real watershed, he argues, is no longer compute but memory.2026-09-06NInfer: A From-Scratch C++/CUDA Inference Engine for a Single RTX 5090A from-scratch C++/CUDA inference engine for five explicitly registered Qwen checkpoints on one NVIDIA GeForce RTX 5090. Startup-frozen residency picks MTP or DFlash speculative decoding, Vision, and one of five KV storage formats; a shared Device KV pool plus pinned Host State/KV checkpoints reuse exact prompt prefixes across 240k-token contexts. Measured aggregate decode reaches 1,146.9 tok/s at concurrency 8 and 15,544.3 tok/s on a 7,680-token prefill.

Assets you can use now

No installable skill wired to this domain yet.

Boundaries & failure modes

The boundary note for this domain is not written yet. It should answer three things: where it is stuck, under which conditions it breaks, and the order of magnitude of cost and latency. Until it lands, a ladder position only says this asset leads on our ruler — not that it works on your task.