Skip to content

SOTA RADAR

SOTA

Thirteen capability domains in four families. Each cell answers three things: what ruler measures this domain, who sits on top of the ladder right now, and how far that claim can be trusted. We do not keep a full tool directory — the long tail stays behind the inclusion gate, and only the ladder shows.

Tracked assets
0
Domains with a ladder
0/13
Ladder order: our pick first, then maturity, then recency of the measurement.

Language & Reasoning

4 domains · 0 tracked
Reasoningtracked 0

General reasoning, scientific QA, long-chain inference and planning

Nothing tracked yet

GPQA, MMLU-Pro, BBH, ARC-AGI
Reasoning
Codetracked 0

Code generation, repair, refactoring, software engineering tasks

Nothing tracked yet

SWE-bench Verified, LiveCodeBench
Code
Mathtracked 0

Competition math, proof assistance, numerical reasoning

Nothing tracked yet

AIME, FrontierMath, MathArena
Math
Multimodaltracked 0

Image-text understanding, chart/doc QA, visual reasoning

Nothing tracked yet

MMMU, MathVista, ChartQA, DocVQA
Multimodal

Vision & Generation

4 domains · 0 tracked
Image Generationtracked 0

Text-to-image, image editing, consistency and controllability

Nothing tracked yet

GenAI-Bench, DrawBench, arena preference, controllability
Image Generation
Video Generationtracked 0

Text/image-to-video, physical and temporal coherence

Nothing tracked yet

VBench, EvalCrafter, physical & temporal coherence
Video Generation
3D & World Modelstracked 0

3D asset generation, scene reconstruction, long-horizon world prediction

Nothing tracked yet

Geometric consistency, scene reconstruction, long-horizon prediction
3D & World Models
Speech & Audiotracked 0

TTS naturalness, voice cloning, music generation, ASR

Nothing tracked yet

TTS naturalness, voice cloning, music generation, ASR WER
Speech & Audio

Agents & Harness

3 domains · 0 tracked
Agentstracked 0

Multi-step tool use, web/OS tasks, long-horizon autonomy

Nothing tracked yet

WebArena, OSWorld, GAIA, multi-step tool tasks
Agents
Harness & Evaluationtracked 0

Evaluation infra, reproducibility, failure taxonomy, cost & latency

Nothing tracked yet

Completion rate, failure taxonomy, reproducibility, cost & latency
Harness & Evaluation
Games & Simulationtracked 0

Long-horizon policy in resettable environments, open-world game agents

Nothing tracked yet

Atari 100k, Procgen, MineDojo, BALROG
Games & Simulation

Embodied & Safety

2 domains · 0 tracked
Embodied AItracked 0

Real-robot manipulation, navigation, VLA and whole-body control

Nothing tracked yet

CALVIN, LIBERO, Open X-Embodiment, real robots
Embodied AI
Safety & Alignmenttracked 0

Jailbreak resistance, hallucination rate, tool misuse, long-task boundaries

Nothing tracked yet

Jailbreak rate, hallucination rate, tool misuse, long-task boundaries
Safety & Alignment

How confidence is graded

Applied only to SOTA claims and benchmark evidence. The letter on the seal says how far a claim is from something we reproduced ourselves — C and D are deliberately drawn dashed and grey so they never read as hard as an A.

ConfirmedTwo or more independent sources, or reproduced by our harness
ContestedSources disagree (data leakage, protocol mismatch)
Vendor ClaimOfficial model card or keynote only, no independent re-test
Needs ReproductionSingle source, not yet verified by us