Xiaomi released the MiMo-V2.6 models on September 22, 2026, and put a price on the training run in the release note. Six days of reinforcement learning cost about $850,000 for the Flash model and about $2.62 million for Pro, by Xiaomi’s own account. Labs almost never publish that number.
The figure is narrower than it sounds. It covers one phase, the final reinforcement-learning run, and not the pre-training, the earlier fine-tuning or the research that came before it. Read correctly, it is still one of the few cost figures any lab has published this year, and the only 2026 disclosure we found that prices a post-training run, which is where capability moved this year.
Every cost, benchmark and speed figure in this post is Xiaomi’s own, taken from its release note, pricing page and model cards on September 22. None of it has been independently replicated. We say that once here and attribute throughout.
- 01Xiaomi priced its RL run: about $850,000 for Flash and $2.62 million for Pro.Under six days, 30 steps per model, roughly 750,000 trajectories in total. The note frames it as the public end of half a year of research and engineering.
- 02The number is post-training only. It excludes pre-training, staff and failed runs.Treat it as the cost of one stated phase, not the cost of making the model. No hardware or GPU-hour figures are published alongside it.
- 03Prices did not move. V2.6 costs exactly what V2.5 cost.Pro is $0.435 / $0.87 per million tokens, Flash $0.14 / $0.28. The V2.5 models are deprecated on October 21, so the upgrade is a model-name change.
- 04Weights, the technical report, 7,000+ RL environments and the training framework are all released, under MIT.That is more than most 'open' releases ship. The benchmark claims, including the DeepSWE gains, remain Xiaomi's own runs on Xiaomi's harnesses.
01 — The releaseWhat shipped
- ModelsTwo native multimodal models, plus Pro served in a faster mode. API names are all-lowercase.
- mimo-v2.6-pro · mimo-v2.6-flash · mimo-v2.6-pro-ultraspeed
- ParametersSparse mixture-of-experts. Total / activated, per the Hugging Face model cards.
- Pro 1.02T / 42B · Flash 309B / 15B
- Context and inputsText, image, video and audio in; text out. Both models.
- 1M tokens
- UltraSpeedThe same Pro checkpoint on faster serving. Xiaomi's claim; no latency method published.
- “Up to 20x” inference speed
- Open weightsPro-RL, Flash-RL and a 9B distilled model, each tagged MIT on Hugging Face.
- 3 checkpoints · MIT
Xiaomi describes MiMo-V2.6-Pro as the flagship checkpoint of the series and UltraSpeed as the same Pro model at up to 20 times the inference speed. The release note says Pro scores 46 on the Artificial Analysis Intelligence Index, which Xiaomi presents as the strongest open-source result, ahead of Kimi K3 and Qwen3.8 Max and behind Claude Fable 5.1 and GPT-6 Astra. That ranking is Xiaomi’s reading of a third-party index on launch day, and we have not checked the board.
The OpenRouter listing for the three routes appeared on September 21 at 20:07 UTC, before the release note. A listing time is not a release date. Our September releases tracker dates each release from the vendor’s own announcement and treats listing times as corroboration only.
02 — The numberThe disclosed cost, and what it covers
The release note gives the training figures in one paragraph. We have split them into a table so that each number sits next to its scope. The technical report rounds the same two figures to $2.6 million and $0.9 million, and adds where Pro’s money went: 43.8% on rollouts, 43.5% on training and 12.7% on the grader. Exact-dollar figures circulate in press coverage; neither Xiaomi document prints them, so we do not either.
| Figure | Flash | Pro |
|---|---|---|
| Stated RL training cost | ≈ $850,000 | ≈ $2.62M |
| Wall-clock time | under 6 days | under 6 days |
| RL steps | 30 | 30 |
| Trajectories (both models combined) | ≈ 750,000 | ≈ 750,000 |
| Samples per update · tokens per step | 1,568 · 3.5–3.7B | 1,568 · 3.5–3.7B |
| Training-task pass rate, relative gain | +25% | +12% |
| DeepSWE v1.1, before → after RL | 48.8 → 65.7 | 58.4 → 72.6 |
Five things sit outside the number, and the note is candid about the first of them.
- Everything before the run. Xiaomi writes that behind the six days of live RL lie “half a year of foundational research accumulation and engineering trial and error.” None of that is priced.
- Pre-training and earlier fine-tuning. The starting checkpoints already scored 48.8 and 58.4 on DeepSWE before this run began. The report states the pre-training token counts, 48 trillion for Flash and 30 trillion for Pro, but never prices them.
- Hardware. The report says the run used “thousands of GPUs,” and nothing more: no GPU model, no GPU-hours, no rental rate. The dollar figure cannot be reconstructed or compared with another lab’s hardware table.
- People and failed runs. Research salaries and discarded experiments are the usual missing lines in any training-cost disclosure, and they are missing here too.
- Verification. The figure is self-reported. There is no auditor, no independent replication, and no method for how compute was priced.
With those limits stated, the disclosure is still unusual. Most labs publish either nothing or a pre-training compute figure. Xiaomi priced the post-training phase, which is where this year’s agent gains have come from, and attached the capability change to it: 30 steps moved DeepSWE by about 17 points for Flash and 14 for Pro, on Xiaomi’s harness. Our companion post on what labs disclose about training costs puts this row next to the few others that exist.
This path is slower, and far less visible. The 6 days of Live RL training for MiMo-V2.6 mark a public trek we’ve taken along this road; behind these 6 days lie half a year of foundational research accumulation and engineering trial and error.Xiaomi, MiMo-V2.6 release note, September 22, 2026
03 — The invoicePrices, batch and the deprecation date
The release note says the V2.6 series “adopts the same API pricing as the V2.5 series,” and the pricing page confirms it the simplest way possible: the old and new models share a row. Overseas rates, per million tokens:
| Model | Real-time | Cache hit | Batch |
|---|---|---|---|
| mimo-v2.6-pro | $0.435 / $0.87 | $0.0036 | $0.2175 / $0.435 |
| mimo-v2.6-flash | $0.14 / $0.28 | $0.0028 | $0.07 / $0.14 |
| mimo-v2.6-pro-ultraspeed | $4.35 / $8.70 | $0.036 | Not supported |
| mimo-v2.5-pro (deprecated Oct 21) | $0.435 / $0.87 | $0.0036 | Not supported |
| mimo-v2.5 (deprecated Oct 21) | $0.14 / $0.28 | $0.0028 | Not supported |
Three details on that page matter more than the headline rates.
- UltraSpeed costs ten times Pro. $4.35 / $8.70 against $0.435 / $0.87, on both input and output. That is our arithmetic on Xiaomi’s rates. It is excluded from the Batch API, so there is no half-price route to the fast tier.
- The V2.5 models retire on October 21, 2026 at 10:00 Beijing time. Since the prices and names line up, the migration is a string change, but it is a hard date. Our deprecation calendar covers how to plan for a hard sunset like this one.
- Batch applies to V2.6 only. The page lists batch rows for V2.6 Pro and Flash and states that the V2.5 models do not support the Batch API, so a workload that can wait costs half on the new names and full price on the old ones.
The note’s own value claim, that at the same intelligence level Pro costs “1/20 to 1/60” of overseas models, depends on which models and which index you pick, and we do not repeat it as a finding. Our frontier model price index lists every vendor’s rates side by side, and the routes on OpenRouter matched Xiaomi’s overseas prices exactly when we checked on September 22.
04 — The inventoryWhat is actually open
An open-weights release lets you run the model. An open training stack lets you change it. Xiaomi’s Hugging Face collection and release note describe both, which is rare. Our guide to what an open-source model release actually contains has the general checklist.
- Weights. MiMo-V2.6-Pro-RL (1.02T total parameters), MiMo-V2.6-Flash-RL (309B) and MiMo-V2.6-Distill-Qwen-9B. All three carry the MIT licence tag on Hugging Face.
- Technical report. A PDF linked from each model card, with the pre-training token counts, the RL setup and the cost split quoted above.
- 7,000+ RL task environments in four groups: software engineering, vulnerability reproduction, knowledge-intensive work, and web design and development.
- An end-to-end RL training framework built on verl, uni-agent and mini-swe-agent, covering environment interaction, trajectory collection, reward evaluation and policy optimisation.
- Composable harnesses. Minimal “mini-harnesses” that separate system prompts, tools and context management, so a training run can mix several agent frameworks.
Xiaomi’s evidence that the environments work is the 9B distilled model. Starting RL from that checkpoint, it reports gains on all 11 benchmarks it tested, including SWE-bench Verified from 61.1 to 66.2, Terminal Bench 2.1 from 37.1 to 52.8 and its own MiMo Visual Coding from 64.0 to 72.4. Those are Xiaomi’s runs, but they are the kind a reader with GPUs can attempt to reproduce, which is the point of releasing the environments. The wider context for why post-training has become the contested layer is in our piece on RL as the new moat.
05 — The claimsThe benchmark claims, with the harness caveat
The Pro model card publishes a 17-row table against MiMo-V2.5-Pro, Claude Opus 5, GPT-5.6 Sol and Claude Fable 5. The rows below are the ones a buyer of agent capacity would look at first, including the two where the new model trails badly.
| Benchmark | V2.6 Pro | V2.6 Flash | Opus 5 · GPT-5.6 Sol |
|---|---|---|---|
| DeepSWE v1.1 | 71.9 | 67.9 | 74.0 · 73.0 |
| AutomationBench v1.0.6 | 53.1 | 52.3 | 50.3 · 45.8 |
| Terminal Bench 2.1 | 89.9 | 87.6 | 89.1 · 88.8 |
| Terminal Bench 4.0 | 34.9 | 28.8 | 49.0 · 39.9 |
| OSWorld-Verified | 82.0 | 80.8 | 83.4 · 83.0 |
| ExploitBench | 47.9 | 25.3 | 70.0 · 78.5 |
Four cautions apply, and they are the same cautions that apply to every vendor table.
- The rival scores are Xiaomi’s table, not the rivals’. Opus 5 at 74.0 on DeepSWE here is not the figure Anthropic or OpenAI publish for their own models on their own harnesses, and Xiaomi’s technical report says only that baselines were “evaluated at the highest supported setting (max),” so it should not be compared with them. The same goes for the xAI DeepSWE figures in our Grok 4.7 post: different harness, different number.
- Two DeepSWE numbers exist for the same model. The release note reports the post-RL scores as 72.6 (Pro) and 65.7 (Flash); the model card reports 71.9 and 67.9. The technical report places the first pair at the end of RL, on its cost curve, and the second in its final results table, after a later on-policy distillation stage (MOPD2); it does not reconcile them further. Both are Xiaomi’s, and the gap is a reminder that the checkpoint stage alone moves a number.
- The security rows cut both ways. Xiaomi built its own cyber benchmarks, where V2.6 scores highly, and also printed ExploitBench, where Pro scores 47.9 against 70.0 and 78.5 for the closed models. Printing the losing rows is to Xiaomi’s credit; our guide to reading disclosed losses in vendor tables explains why that matters.
- The release note’s summary is broader than the table. Xiaomi describes Pro as level with Claude Opus 5 and GPT-5.6 Sol on most agent benchmarks. On the rows above that holds for Terminal Bench 2.1, AutomationBench and OSWorld, and does not hold for Terminal Bench 4.0 or ExploitBench. Read the table, not the sentence.
06 — The decisionThe one decision this changes: second-sourcing an agent workload
An open model at $0.435 / $0.87 that a vendor claims is close to Opus 5 on agent tasks is not a reason to move production. It is a reason to run a second source in parallel and measure. The disclosed training cost adds one thing to that decision: Xiaomi has shown it can improve the model in six-day increments, at a price it is willing to publish, which is a signal about how often the checkpoint will move. The earlier MiMo-V2-Pro launch in March made similar claims; the difference this time is that the training stack is public.
07 — ConclusionA priced training run is rarer than a benchmark win
Shadow-run one agent workload on Flash via batch, and rename any V2.5 calls before October 21
The dollar figures are Xiaomi’s, the benchmarks are Xiaomi’s, and nothing has been replicated. What is checkable is the price list, the deprecation date and the released stack, and those three are enough to justify a measured test. If you want help building the eval set that decides it, our AI transformation team does that work.