MiMo-V2.6: Live RL as an auditable ledger of 30 steps, 750k trajectories, and dollars
xiaomi-mimo-v2-6
Xiaomi's open-source MiMo-V2.6 Pro/Flash: 30 steps of Live RL, roughly 750k trajectories, public cost of $850k/$2.62M, and 7k+ RL environments released in the same batch; tops the open-weights camp on the Artificial Analysis intelligence index at 46 with +17/+14 points out-of-sample on DeepSWE v1.1; four demo lines (Vibe World, CUA, research, content creation) all open. Self-improvement becomes an auditable ledger - steps, trajectories, and dollars on the table for third-party recomputation.
- CONFIDENCE
- Confirmed
- Two or more independent sources, or reproduced by our harness
- KEY METRIC
- AA 智能指数(开源榜首)
- Confirmed · 2026-09
- MATURITY
- Product
- research → demo → product → production
Our takeOur read: Xiaomi is currently the only open-weights lab publishing Live RL cost and trajectory ledgers at this granularity; AA index 46 and the DeepSWE out-of-sample gains are evidence from an independent board and an out-of-sample split respectively, so confidence is A. The gap is equally explicit: both tiers remain one rung below the closed-source top, and RL stability beyond the 30-step window has no public curve. For selection, Flash is the high-value local agentic-coding tier; Pro chases the lower edge of closed-source flagship workloads.
MiMo-V2.6 is Xiaomi's open-source Pro / Flash pair with an unusually granular public ledger: 30 steps of Live RL, roughly 750k trajectories, training cost of $850k / $2.62M (Pro / Flash), and 7k+ RL environments released alongside. It tops the open-weights camp on the Artificial Analysis intelligence index at 46, gains +17 / +14 points out-of-sample on DeepSWE v1.1, and ships four demo lines: Vibe World, CUA, research, and content creation.
The route matters because "self-improvement" becomes an auditable ledger: steps, trajectories, and dollars are all on the table so third parties can recompute cost-effectiveness on the same basis; 30 iterations of Live RL mean reward signal comes from real serving-time interaction rather than offline preference sets - rare public evidence for post-RLHF scaling routes.
The open release is complete: weights, environments, and eval scripts in one batch; Flash targets local and edge deployment while Pro aims at closed-source-tier agentic coding and tool-use workloads.
Boundary: both tiers still sit one rung below the closed-source ladder top, and Live RL's long-horizon stability (reward hacking, distribution drift) is only charted inside the 30-step window; behavior beyond it is on you to monitor.