AI DevelopmentNew Release9 min readPublished September 22, 2026

About $850,000 for Flash · $2.62 million for Pro · six days of RL · the number is real, and it is not the whole bill

Xiaomi Published What Its AI Training Run Actually Cost

Xiaomi says six days of RL on MiMo-V2.6 cost about $850,000 for Flash and $2.62 million for Pro. What the figure covers, what it leaves out, and the prices.

DA
Digital Applied Team
Research and practical guidance
ReleasedSeptember 22, 2026
Sources checkedSeptember 22, 2026

Xiaomi released the MiMo-V2.6 models on September 22, 2026, and put a price on the training run in the release note. Six days of reinforcement learning cost about $850,000 for the Flash model and about $2.62 million for Pro, by Xiaomi’s own account. Labs almost never publish that number.

The figure is narrower than it sounds. It covers one phase, the final reinforcement-learning run, and not the pre-training, the earlier fine-tuning or the research that came before it. Read correctly, it is still one of the few cost figures any lab has published this year, and the only 2026 disclosure we found that prices a post-training run, which is where capability moved this year.

Every cost, benchmark and speed figure in this post is Xiaomi’s own, taken from its release note, pricing page and model cards on September 22. None of it has been independently replicated. We say that once here and attribute throughout.

Key takeaways
  1. 01
    Xiaomi priced its RL run: about $850,000 for Flash and $2.62 million for Pro.Under six days, 30 steps per model, roughly 750,000 trajectories in total. The note frames it as the public end of half a year of research and engineering.
  2. 02
    The number is post-training only. It excludes pre-training, staff and failed runs.Treat it as the cost of one stated phase, not the cost of making the model. No hardware or GPU-hour figures are published alongside it.
  3. 03
    Prices did not move. V2.6 costs exactly what V2.5 cost.Pro is $0.435 / $0.87 per million tokens, Flash $0.14 / $0.28. The V2.5 models are deprecated on October 21, so the upgrade is a model-name change.
  4. 04
    Weights, the technical report, 7,000+ RL environments and the training framework are all released, under MIT.That is more than most 'open' releases ship. The benchmark claims, including the DeepSWE gains, remain Xiaomi's own runs on Xiaomi's harnesses.

01 — The releaseWhat shipped

ModelsTwo native multimodal models, plus Pro served in a faster mode. API names are all-lowercase.
mimo-v2.6-pro · mimo-v2.6-flash · mimo-v2.6-pro-ultraspeed
ParametersSparse mixture-of-experts. Total / activated, per the Hugging Face model cards.
Pro 1.02T / 42B · Flash 309B / 15B
Context and inputsText, image, video and audio in; text out. Both models.
1M tokens
UltraSpeedThe same Pro checkpoint on faster serving. Xiaomi's claim; no latency method published.
“Up to 20x” inference speed
Open weightsPro-RL, Flash-RL and a 9B distilled model, each tagged MIT on Hugging Face.
3 checkpoints · MIT

Xiaomi describes MiMo-V2.6-Pro as the flagship checkpoint of the series and UltraSpeed as the same Pro model at up to 20 times the inference speed. The release note says Pro scores 46 on the Artificial Analysis Intelligence Index, which Xiaomi presents as the strongest open-source result, ahead of Kimi K3 and Qwen3.8 Max and behind Claude Fable 5.1 and GPT-6 Astra. That ranking is Xiaomi’s reading of a third-party index on launch day, and we have not checked the board.

The OpenRouter listing for the three routes appeared on September 21 at 20:07 UTC, before the release note. A listing time is not a release date. Our September releases tracker dates each release from the vendor’s own announcement and treats listing times as corroboration only.

02 — The numberThe disclosed cost, and what it covers

The release note gives the training figures in one paragraph. We have split them into a table so that each number sits next to its scope. The technical report rounds the same two figures to $2.6 million and $0.9 million, and adds where Pro’s money went: 43.8% on rollouts, 43.5% on training and 12.7% on the grader. Exact-dollar figures circulate in press coverage; neither Xiaomi document prints them, so we do not either.

Source: Xiaomi MiMo-V2.6 release note, September 22, 2026. All figures are Xiaomi’s own. The trajectory count is a combined total for both models. The technical report gives 2.7–3.7B tokens per step and a Flash starting score of 48.7.
FigureFlashPro
Stated RL training cost≈ $850,000≈ $2.62M
Wall-clock timeunder 6 daysunder 6 days
RL steps3030
Trajectories (both models combined)≈ 750,000≈ 750,000
Samples per update · tokens per step1,568 · 3.5–3.7B1,568 · 3.5–3.7B
Training-task pass rate, relative gain+25%+12%
DeepSWE v1.1, before → after RL48.8 → 65.758.4 → 72.6

Five things sit outside the number, and the note is candid about the first of them.

  1. Everything before the run. Xiaomi writes that behind the six days of live RL lie “half a year of foundational research accumulation and engineering trial and error.” None of that is priced.
  2. Pre-training and earlier fine-tuning. The starting checkpoints already scored 48.8 and 58.4 on DeepSWE before this run began. The report states the pre-training token counts, 48 trillion for Flash and 30 trillion for Pro, but never prices them.
  3. Hardware. The report says the run used “thousands of GPUs,” and nothing more: no GPU model, no GPU-hours, no rental rate. The dollar figure cannot be reconstructed or compared with another lab’s hardware table.
  4. People and failed runs. Research salaries and discarded experiments are the usual missing lines in any training-cost disclosure, and they are missing here too.
  5. Verification. The figure is self-reported. There is no auditor, no independent replication, and no method for how compute was priced.

With those limits stated, the disclosure is still unusual. Most labs publish either nothing or a pre-training compute figure. Xiaomi priced the post-training phase, which is where this year’s agent gains have come from, and attached the capability change to it: 30 steps moved DeepSWE by about 17 points for Flash and 14 for Pro, on Xiaomi’s harness. Our companion post on what labs disclose about training costs puts this row next to the few others that exist.

This path is slower, and far less visible. The 6 days of Live RL training for MiMo-V2.6 mark a public trek we’ve taken along this road; behind these 6 days lie half a year of foundational research accumulation and engineering trial and error.Xiaomi, MiMo-V2.6 release note, September 22, 2026

03 — The invoicePrices, batch and the deprecation date

The release note says the V2.6 series “adopts the same API pricing as the V2.5 series,” and the pricing page confirms it the simplest way possible: the old and new models share a row. Overseas rates, per million tokens:

Input / output per million tokens, overseas pricing. Source: Xiaomi MiMo pricing page, September 22, 2026. Cache writes are free for a limited time. Batch runs at half the real-time rate.
ModelReal-timeCache hitBatch
mimo-v2.6-pro$0.435 / $0.87$0.0036$0.2175 / $0.435
mimo-v2.6-flash$0.14 / $0.28$0.0028$0.07 / $0.14
mimo-v2.6-pro-ultraspeed$4.35 / $8.70$0.036Not supported
mimo-v2.5-pro (deprecated Oct 21)$0.435 / $0.87$0.0036Not supported
mimo-v2.5 (deprecated Oct 21)$0.14 / $0.28$0.0028Not supported

Three details on that page matter more than the headline rates.

  • UltraSpeed costs ten times Pro. $4.35 / $8.70 against $0.435 / $0.87, on both input and output. That is our arithmetic on Xiaomi’s rates. It is excluded from the Batch API, so there is no half-price route to the fast tier.
  • The V2.5 models retire on October 21, 2026 at 10:00 Beijing time. Since the prices and names line up, the migration is a string change, but it is a hard date. Our deprecation calendar covers how to plan for a hard sunset like this one.
  • Batch applies to V2.6 only. The page lists batch rows for V2.6 Pro and Flash and states that the V2.5 models do not support the Batch API, so a workload that can wait costs half on the new names and full price on the old ones.

The note’s own value claim, that at the same intelligence level Pro costs “1/20 to 1/60” of overseas models, depends on which models and which index you pick, and we do not repeat it as a finding. Our frontier model price index lists every vendor’s rates side by side, and the routes on OpenRouter matched Xiaomi’s overseas prices exactly when we checked on September 22.

04 — The inventoryWhat is actually open

Weights, environments and harness are different things

An open-weights release lets you run the model. An open training stack lets you change it. Xiaomi’s Hugging Face collection and release note describe both, which is rare. Our guide to what an open-source model release actually contains has the general checklist.

  • Weights. MiMo-V2.6-Pro-RL (1.02T total parameters), MiMo-V2.6-Flash-RL (309B) and MiMo-V2.6-Distill-Qwen-9B. All three carry the MIT licence tag on Hugging Face.
  • Technical report. A PDF linked from each model card, with the pre-training token counts, the RL setup and the cost split quoted above.
  • 7,000+ RL task environments in four groups: software engineering, vulnerability reproduction, knowledge-intensive work, and web design and development.
  • An end-to-end RL training framework built on verl, uni-agent and mini-swe-agent, covering environment interaction, trajectory collection, reward evaluation and policy optimisation.
  • Composable harnesses. Minimal “mini-harnesses” that separate system prompts, tools and context management, so a training run can mix several agent frameworks.

Xiaomi’s evidence that the environments work is the 9B distilled model. Starting RL from that checkpoint, it reports gains on all 11 benchmarks it tested, including SWE-bench Verified from 61.1 to 66.2, Terminal Bench 2.1 from 37.1 to 52.8 and its own MiMo Visual Coding from 64.0 to 72.4. Those are Xiaomi’s runs, but they are the kind a reader with GPUs can attempt to reproduce, which is the point of releasing the environments. The wider context for why post-training has become the contested layer is in our piece on RL as the new moat.

05 — The claimsThe benchmark claims, with the harness caveat

The Pro model card publishes a 17-row table against MiMo-V2.5-Pro, Claude Opus 5, GPT-5.6 Sol and Claude Fable 5. The rows below are the ones a buyer of agent capacity would look at first, including the two where the new model trails badly.

Source: MiMo-V2.6-Pro-RL model card, read September 22, 2026. All scores as printed by Xiaomi; the card does not say how the rival figures were obtained; the technical report says baselines were run at their highest reasoning setting. Rivals column: Claude Opus 5 · GPT-5.6 Sol.
BenchmarkV2.6 ProV2.6 FlashOpus 5 · GPT-5.6 Sol
DeepSWE v1.171.967.974.0 · 73.0
AutomationBench v1.0.653.152.350.3 · 45.8
Terminal Bench 2.189.987.689.1 · 88.8
Terminal Bench 4.034.928.849.0 · 39.9
OSWorld-Verified82.080.883.4 · 83.0
ExploitBench47.925.370.0 · 78.5

Four cautions apply, and they are the same cautions that apply to every vendor table.

  1. The rival scores are Xiaomi’s table, not the rivals’. Opus 5 at 74.0 on DeepSWE here is not the figure Anthropic or OpenAI publish for their own models on their own harnesses, and Xiaomi’s technical report says only that baselines were “evaluated at the highest supported setting (max),” so it should not be compared with them. The same goes for the xAI DeepSWE figures in our Grok 4.7 post: different harness, different number.
  2. Two DeepSWE numbers exist for the same model. The release note reports the post-RL scores as 72.6 (Pro) and 65.7 (Flash); the model card reports 71.9 and 67.9. The technical report places the first pair at the end of RL, on its cost curve, and the second in its final results table, after a later on-policy distillation stage (MOPD2); it does not reconcile them further. Both are Xiaomi’s, and the gap is a reminder that the checkpoint stage alone moves a number.
  3. The security rows cut both ways. Xiaomi built its own cyber benchmarks, where V2.6 scores highly, and also printed ExploitBench, where Pro scores 47.9 against 70.0 and 78.5 for the closed models. Printing the losing rows is to Xiaomi’s credit; our guide to reading disclosed losses in vendor tables explains why that matters.
  4. The release note’s summary is broader than the table. Xiaomi describes Pro as level with Claude Opus 5 and GPT-5.6 Sol on most agent benchmarks. On the rows above that holds for Terminal Bench 2.1, AutomationBench and OSWorld, and does not hold for Terminal Bench 4.0 or ExploitBench. Read the table, not the sentence.

06 — The decisionThe one decision this changes: second-sourcing an agent workload

An open model at $0.435 / $0.87 that a vendor claims is close to Opus 5 on agent tasks is not a reason to move production. It is a reason to run a second source in parallel and measure. The disclosed training cost adds one thing to that decision: Xiaomi has shown it can improve the model in six-day increments, at a price it is willing to publish, which is a signal about how often the checkpoint will move. The earlier MiMo-V2-Pro launch in March made similar claims; the difference this time is that the training stack is public.

You run a high-volume agent on a closed model and the bill is the problem
Replay a week of real traces through mimo-v2.6-flash on the Batch API at $0.07 / $0.14. Score completion, not benchmark points. If it holds on your tasks, the saving is large; if not, you have learned that in a week.
Flash, batch, shadow run
You need a second source for a long-horizon coding agent
Test mimo-v2.6-pro on the same harness you use in production, not on Xiaomi's table. DeepSWE-style tasks are where Xiaomi's own gains concentrate; Terminal Bench 4.0 is where the card shows the widest gap.
Pro, your harness
You need the fast tier
UltraSpeed is 10× the price of Pro with no batch route and no published latency method. Measure time to first token yourself before paying the multiple, and compare a smaller model at standard speed first.
Measure before paying 10×
You post-train your own models
The environments, framework and harnesses are the asset. The 9B distillation numbers are the reproducible starting point, and reproducing them is the first thing to attempt before trusting the larger claims.
Reproduce the 9B result

07 — ConclusionA priced training run is rarer than a benchmark win

What to do this week

Shadow-run one agent workload on Flash via batch, and rename any V2.5 calls before October 21

The dollar figures are Xiaomi’s, the benchmarks are Xiaomi’s, and nothing has been replicated. What is checkable is the price list, the deprecation date and the released stack, and those three are enough to justify a measured test. If you want help building the eval set that decides it, our AI transformation team does that work.

Digital Applied

Test an open model before you trust its table.

We replay your real agent traces through candidate models, score completion on your harness, and tell you where a cheaper model holds up.

Model evaluationSecond-source routingCost per task
Your next project

A second source you have measured

  • →Your traces, not the vendor's benchmarks
  • →Completion scored on your harness
  • →A routing rule you can defend
Questions and answers

The questions we get about MiMo-V2.6 and its training cost

Xiaomi says the final reinforcement-learning phase cost about $850,000 for MiMo-V2.6-Flash and about $2.62 million for MiMo-V2.6-Pro: under six days, 30 steps per model and roughly 750,000 trajectories combined. That excludes pre-training, earlier fine-tuning, staff and failed experiments, names no GPU model or pricing basis, and is self-reported.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

AI Development

What AI Labs Actually Disclose About Training Costs

A census of 21 training-cost figures labs published themselves, 2022 to 2026. Six give a dollar amount, none is audited, and each leaves something out.

September 22, 2026 · 8 minRead
AI Development

What "Open Source" Actually Includes for an AI Model

16 models in the open-model conversation, checked against the licence file, OSI's list and what else was released: weights, data, recipes, evaluation harness.

September 21, 2026 · 7 minRead
AI Development

Fireworks Ember-1: Kimi K3 Quality With Fewer Tokens?

Fireworks tuned Kimi K3 into Ember-1 and says it matches K3 with 35 to 50% shorter reasoning at the same price. The rows it loses, and the preview caveat.

September 23, 2026 · 5 minRead
AI Development

Gemini 3.8 Flash TTS: Voice Cloning and a Price That Doubles

Gemini 3.8 Flash TTS is generally available with voice cloning from a 30-second sample. The promotional price ends December 31 and doubles on January 1, 2027.

September 23, 2026 · 5 minRead
AI Development

Computer-Use Agents: Microsoft vs Anthropic vs Google

Microsoft GA, Anthropic public beta, and Google Gemini preview — OSWorld scores now 78% across frontier models above the ~72% human baseline. Routing guide.

May 22, 2026 · 16 minRead
AI Development

Agent Computer Use: Enterprise Automation Playbook

Enterprise playbook for deploying computer-use agents — a 40-point guardrails checklist spanning identity, audit, action boundaries, failures, and compliance.

May 22, 2026 · 17 minRead