AI DevelopmentReference8 min readPublished September 22, 2026

21 lab-published figures · 12 organisations · 6 in dollars · none audited, every one partial

What AI Labs Actually Disclose About Training Costs

A census of 21 training-cost figures labs published themselves, 2022 to 2026. Six give a dollar amount, none is audited, and each leaves something out.

DA
Digital Applied Team
Research and practical guidance
Rows21 lab + 2 academic
Data as ofSeptember 25, 2026

Every argument about what an AI model costs to train ends the same way: someone quotes a number, and nobody can say where it came from. So we collected only the figures labs have published themselves, in their own papers, model cards and release notes, and recorded for each what it covers and what the lab says it leaves out.

The result is 21 rows from 12 organisations between May 2022 and September 2026, plus two academic fine-tunes kept apart. Six rows give a dollar amount. The rest are GPU-hours, carbon, water or days-times-chips, and they cannot be added together. No dollar figure has been checked by anyone outside the lab. The four labs most people mean when they say “frontier” publish nothing for their frontier models.

The newest row, Xiaomi’s September 22 disclosure of its MiMo-V2.6 reinforcement-learning run, is also the only 2026 figure we found that prices post-training. Our post on that release covers it in full; here it takes its place in the table.

Key takeaways
  1. 01
    Six rows put a dollar figure on a training run. Five price GPU time at an assumed rental rate; Xiaomi does not say.DeepSeek-V3 ($5.576M), DeepSeek-R1 ($294K), MiniMax-M1 ($534,700), OLMo 3 ($2.75M), BLOOM ($2–5M) and Xiaomi ($850K and $2.62M). None counts staff, and only BLOOM's range includes preliminary experiments.
  2. 02
    Every figure excludes something, and the most candid labs say so themselves.Meta puts the true OPT cost at roughly twice the final run. BigScience says the final BLOOM run was 37% of emissions. Ai2 puts development at half of training again.
  3. 03
    Post-training is where the interesting numbers are, and there are only four of them.DeepSeek-R1 at $294K on a $5.576M base, MiniMax-M1 at $534,700, NVIDIA's 140K H100-hours, and Xiaomi's $2.62M. All centred on RL, all self-reported.
  4. 04
    OpenAI, Anthropic, Google DeepMind and xAI publish no figure for their frontier models.Only OpenAI's GPT-4 report states a reason. Third-party estimates for those models range from $30M to $191M depending on method, and should be cited by estimator.

01 — The findingWhat the labs have published, in four numbers

Rows
Lab-published figures
21

From 12 organisations, 2022 to 2026, each in the lab's own paper, card or release note. Two academic fine-tunes are listed separately.

Primary only
In dollars
Carry a dollar amount
6

Five are priced at a cloud-rental rate, usually $2 per H800 or H100 hour; Xiaomi states no basis. None is presented as a cost of ownership.

Rental-priced
Audited
Checked by an outside party
0

Peer review of a paper is not an audit of its cost line. Mistral's life-cycle analysis, reviewed by two external firms, is the nearest thing.

Self-reported
Silent
Frontier labs with no figure
4

OpenAI, Anthropic, Google DeepMind and xAI. Their current model cards and system cards give no compute, hardware-hour or cost line.

See section 06

02 — The datasetThe census: every self-published training-cost figure we found

One row per disclosure. The figure column keeps the lab’s own unit and, where a phrase is quoted, its own words. The last column says what the figure covers, because that is the part every secondary citation drops.

Sources: each lab’s paper, model card, release note or technical report, re-read at the primary URL on September 25, 2026. Rows marked “our arithmetic” multiply the lab’s stated chips by its stated days.
LabModel · publishedFigure as statedWhat it covers
XiaomiMiMo-V2.6 Flash / Pro · Sep 2026“around 850,000 and 2.62 million US dollars” (release note); “$0.9M” and “$2.6M” (technical report)RL post-training only: under 6 days, 30 steps each. Pre-training tokens stated (48T / 30T) but not priced.
DeepSeekDeepSeek-V3 · Dec 20242.788M H800 GPU-hours; “$5.576M” at an assumed $2 per GPU-hour rentalThe official run: pre-training ($5.328M), context extension ($0.238M), post-training ($0.01M).
DeepSeekDeepSeek-R1 · Sep 2025 (Nature)147K H800 GPU-hours; $294K at the same $2 rental assumptionReasoning RL stages plus SFT-data creation on top of the V3 base, which is not included.
DeepSeekDeepSeek-V2 · May 2024172.8K GPU-hours per trillion tokens (vs 300.6K for DeepSeek 67B)Pre-training rate only. No total, no dollar figure.
MiniMaxMiniMax-M1 · Jun 2025“a rental cost of just $534,700”; 512 H800s for three weeksThe RL run only, after 7.5T tokens of continued pre-training and SFT that are not costed.
MetaLlama 2 · Jul 20233.3M A100-80GB GPU-hours; 539 tCO2eq, offsetPre-training only. Fine-tuning and evaluation on third-party cloud, not quantified.
MetaLlama 3 · Apr 20247.7M H100-80GB GPU-hours; 2,290 tCO2eqPre-training only. No dollar figure.
MetaLlama 3.1 · Jul 202439.3M H100-80GB GPU-hours; 11,390 t location-based, 0 t market-based; 405B at 3.8 × 10²⁵ FLOPs“Training” per the card; the FLOPs figure is 405B pre-training. Up to 16K H100s.
MetaLlama 4 Scout / Maverick · Apr 20257.38M H100-80GB GPU-hours; 1,999 t location-basedPre-training only. The Behemoth teacher model is not in the table.
MetaOPT-175B · May 202275 t CO2eq for the final run; 992 A100sFinal run. Meta’s own footnote: with ablations, baselines and downtime, total is “roughly 2× higher.”
Ai2OLMo 2 · Dec 2024about 154 tCO2eq; about 1.1 million litres of waterFinal pre-training energy, measured at the node. Stated as a lower bound.
Ai2OLMo 3 Think 32B · Dec 202556 days on 1,024 dedicated H100s; “At a price of $2/H100 hour, this would cost $2.75M”Wall-clock for the whole pipeline: pre-, mid-, long-context and post-training. Research detours excluded.
Ai2 (Morrison et al.)OLMo development series · Mar 2025493 t CO2 and 2.769 million litres of waterDevelopment, hardware manufacturing and final runs together. Development alone was about half of training.
BigScience / Hugging FaceBLOOM-176B · Nov 20221,082,990 compute hours; 433 MWh; 25 t CO2eq; card: “Equivalent of $2-5M in cloud computing (including preliminary experiments)”The dollar range includes preliminary experiments. The final run was about 37% of overall emissions.
TIIFalcon-180B · Nov 202343,500 PF-days on up to 4,096 A100s; 3.5T tokensPre-training compute in PF-days. No hours, dollars or energy.
Google DeepMindGemma 2B / 7B · Mar 2024about 131 tCO2eq on TPUv5ePre-training, scaled to include data-centre overhead. No dollars or hours.
Google DeepMindGemma 2 · Jul 20241,247.61 tCO2eqPre-training, same method. Gemma 3 (2025) gives chip counts only.
Hugging FaceSmolLM3-3B · Jul 2025384 H100s for 24 days (about 221K GPU-hours, our arithmetic)Pre-training run. Ablations on the recipe not quantified.
MicrosoftPhi-4 · Dec 20241,920 H100-80G for 21 days on 9.8T tokens (about 968K GPU-hours, our arithmetic)Training run; the card does not split stages. No dollars or energy.
NVIDIALlama-3.1-Nemotron-Ultra-253B · May 2025“approximately 140k H100 hours” for the reasoning-RL stageReasoning RL only. The Llama 3.1 parent, pruning, distillation and SFT are outside the figure.
Mistral AIMistral Large 2 · Jul 202520.4 ktCO2e, 281,000 m³ of water and 660 kg Sb eq after 18 months of useA full life-cycle analysis: training plus 18 months of inference and hardware manufacturing. Not training-only.
Academic: Stanford et al.s1-32B · Jan 2025“Finetuning took 26 minutes on 16 NVIDIA H100 GPUs”SFT of Qwen2.5-32B on 1K examples. The widely quoted “$50” is not in the paper.
Academic: NovaSky, UC BerkeleySky-T1-32B-Preview · Jan 2025“trained for less than $450”; 19 hours on 8 H100s at Lambda rentalThe SFT run only. Data generation and the Qwen2.5 base are not costed.

Seven more model reports give hardware or relative compute but no absolute cost, hours or energy figure, and are not counted above: NVIDIA’s Nemotron-4 (768 DGX H100 nodes), Apple’s AFM-server (8,192 TPUv4 chips), Gemma 3 (chip counts per model), Kimi K2 (an H800 cluster), Kimi K3 (relative RL-compute plots only), Qwen3 (a relative one-tenth claim for distillation) and GLM-4.5 (nothing found).

03 — The unitsWhy the figures cannot be added up

The table has four families of unit, and a reader who wants one ranking has to convert between them with assumptions the labs did not make. We have not converted anything except chips times days, and we have not ranked the rows.

Unit 1
Rental dollars
GPU-hours × an assumed rate

DeepSeek and Ai2 assume $2 per H800 or H100 hour, MiniMax and NovaSky use rental prices, and BigScience gives a cloud-equivalent range. Xiaomi does not state its basis. A different rate moves every one of these numbers.

6 rows
Unit 2
GPU-hours or PF-days
No price attached

Meta's four Llama cards, DeepSeek-V2, NVIDIA and TII. Convertible to dollars only with a rental assumption the lab did not endorse, and only within one chip generation.

7 rows
Unit 3
Carbon, energy, water
tCO2eq, MWh, litres

Ai2, Google DeepMind, Mistral and Meta's OPT. Depends on grid mix, offsets and whether manufacturing is included. Meta reports 0 t market-based beside 11,390 t location-based.

6 rows
Unit 4
Chips times days
Wall-clock on a stated cluster

Microsoft, Hugging Face and the academic rows. Honest about time, silent about utilisation. Ai2 chose this unit for OLMo 3 deliberately, then priced it.

2 rows, plus both academic rows

04 — The exclusionsThe five things a disclosed cost almost always leaves out

The labs are more candid about this than their quoters. Read the exclusions in their own words first; the pattern follows.

Each statement is from the document named, read September 25, 2026. Quoted phrases are verbatim; the rest is paraphrased.
WhoWhat the lab says is outside the figure
DeepSeek, V3 paperThe costs “include only the official training of DeepSeek-V3, excluding the costs associated with prior research and ablation experiments on architectures, algorithms, or data.”
Meta, OPT paper“With ablations, baselines and downtime, our own estimates of total cost is roughly 2× higher.”
BigScience, BLOOM paperThe final training run was about 37% of overall emissions; intermediate runs and evaluation made up the other 63%.
Ai2, Morrison et al.Model development, which most developers do not disclose, amounted to about half of the final training’s impact.
Xiaomi, MiMo-V2.6 release note“Behind these 6 days lie half a year of foundational research accumulation and engineering trial and error.”
Ai2, OLMo 3 paperThe 56-day window “does not include any substantial modifications or research ideas that could expand the timeline substantially.” The recipe was developed at 7B or smaller.
Meta, Llama 2 and 3 cardsEmissions are reported for pre-training; fine-tuning, annotation and evaluation ran on third-party cloud compute and are not quantified.
Ai2, OLMo 2 paperThe estimate excludes embodied emissions, deployment and inference, so it “should be viewed as lower bounds.”
  1. The research before the run. Ablations, baselines and abandoned experiments. Meta’s OPT footnote doubles the total; BigScience’s share puts the final run at just over a third; Ai2 measured development at half of training again.
  2. The base model under a post-training figure. DeepSeek-R1’s $294K sits on DeepSeek-V3’s $5.576M. MiniMax-M1’s $534,700 follows 7.5 trillion tokens of continued pre-training. Xiaomi’s RL cost sits on 30 to 48 trillion tokens of pre-training it did not price.
  3. People, capital and data. Five of the six dollar rows are rental-priced GPU time, and Xiaomi does not say how it priced its run. Epoch AI’s estimate for GPT-4 puts the hardware at about $800 million to acquire against $40 million amortised for the run; staff at tens of millions. No lab row says it includes any of the three.
  4. Post-training, evaluation and fine-tuning. The Llama cards cover pre-training. Ai2 notes that repeated checkpoint evaluations and post-training consume a non-trivial share that pre-training hours alone do not capture.
  5. Inference and hardware life cycle. Only Mistral’s row includes use after training and the emissions of making the chips, and it is a different kind of number: 18 months of a deployed model, not a training run.
Given both the competitive landscape and the safety implications of large-scale models like GPT-4, this report contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar.OpenAI, GPT-4 Technical Report, March 2023

05 — The shiftPost-training is the number now, and there are four of them

Pre-training compute is the figure labs used to disclose, when they disclosed anything. The rows that matter for a 2026 buyer are the ones that price the reinforcement-learning phase, because that is where this year’s agent capability was added. There are four.

  • DeepSeek-R1: $294K, or 147K H800 GPU-hours, published with the Nature paper in September 2025. That is about 5% of the $5.576M DeepSeek stated for the V3 base, our arithmetic on two figures priced at the same $2 rate. Epoch AI had estimated the RL cost at around $1 million eight months earlier.
  • MiniMax-M1: $534,700 for three weeks on 512 H800s, June 2025. RL only.
  • NVIDIA Nemotron Ultra 253B: about 140K H100 hours for the reasoning-RL stage, May 2025. No dollar figure.
  • Xiaomi MiMo-V2.6: about $850K for Flash and $2.62M for Pro, six days, 30 steps, September 2026. Xiaomi ties the spend to a result: on its own harness, DeepSWE v1.1 moved from 48.8 to 65.7 for Flash and from 58.4 to 72.6 for Pro over those 30 steps.

Xiaomi’s report plots its benchmark score against cumulative dollars spent, which is why it is a more useful disclosure than the others even though it is the least specified on hardware. It is also the largest, and Xiaomi’s report says the run used thousands of GPUs. Kimi K3’s August 2026 report shows RL compute only as a relative plot, so the comparison a reader wants most, Xiaomi against Moonshot, is not possible from primaries. Our post on RL as the new moat covers why labs now guard this number.

06 — The silenceWho publishes nothing, and the estimates that fill the gap

Current model cards and system cards, read September 25, 2026. Only OpenAI’s 2023 report states a reason for withholding.
LabDocument checkedWhat it says
OpenAIGPT-4 Technical Report, Mar 2023An explicit decline: the report “contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar.”
OpenAIGPT-5 System Card, Aug 2025No training compute, hardware or cost figure. Mentions of compute refer to test-time compute.
AnthropicClaude Fable 5.1 and Mythos 5.1 System Card, Sep 2026Describes the data mix and post-training. No compute, hardware, energy or cost figure.
Google DeepMindGemini 3 Pro Model Card (updated May 2026)Hardware family only: trained on TPUs. No compute, energy or cost figure. Gemma is the exception, above.
xAIGrok 4 Model Card, Aug 2025No training compute, hardware or cost figure.

Where labs are silent, two estimators are quoted most. They disagree with each other by about 2× on GPT-4 and about 6× on Gemini Ultra, because they use different methods, and a citation should say which.

  • Epoch AI (Cottier et al., May 2024, revised February 2025) amortises hardware plus energy for the final run: GPT-4 about $40 million, Gemini Ultra about $30 million. The same paper puts GPT-4’s hardware at about $800 million to acquire and says frontier run costs have grown at 2.4× a year since 2016.
  • Stanford AI Index (2024 and 2025, using Epoch data at cloud-rental prices) gives GPT-4 about $78 to $79 million, Gemini Ultra $191 million and Llama 3.1-405B about $170 million. Meta’s own card for that model gives 39.3M GPU-hours and no dollar figure. The 2025 report adds that labs are disclosing less about their training processes, which makes the estimates harder.

Two numbers in wide circulation are not lab figures. The “$6 million DeepSeek model” rounds $5.576M and drops DeepSeek’s own exclusion of prior research; our post on DeepSeek’s first funding round covers the round it was reported to be raising. The “$50 reasoning model” for s1 does not appear in the s1 paper, which states 26 minutes on 16 H100s and nothing in dollars.

07 — The checklistHow to read the next one

Six questions before you quote a training cost

Which phase (pre-training, mid-training, RL, SFT)? Which unit, and at what rental rate if dollars? Which chips, how many, for how long? What does the lab itself say is excluded? Is there a base model under it? Has anyone outside the lab checked it? If the source cannot answer the first four, it is an estimate, and it should be cited by the estimator’s name.

You are quoting a figure in a deck or an article
Quote the lab's unit and the lab's scope in the same sentence: '$5.576M of rental-priced GPU time for the final run, excluding prior research, per DeepSeek.' A bare dollar number is the error every secondary makes.
Figure + scope + source
You want to compare two labs
Only compare rows in the same unit, on the same chip generation, covering the same phase. That leaves DeepSeek-V3 against OLMo 3 for full runs, and R1 against MiniMax-M1 for RL, both rental-priced; Xiaomi's RL figure has no stated pricing basis. Everything else needs an assumption you must state.
Same unit, same phase
You are budgeting your own fine-tune or RL run
Start from the academic rows and MiniMax, which are closest in scale, then apply the lab exclusions to yourself: double for experiments, add the evaluation compute, and price your own people. Our $200-a-month post covers the small end.
MiniMax and the academic rows
You need a figure for a frontier closed model
There is none from the lab. Cite Epoch AI or the AI Index by name, state the method (amortised or rental), and expect the two to differ by 2× or more.
Estimate, named

08 — MethodologyHow this census was built

Methodology

Primary-source census. Every row is a figure the organisation published itself; third-party estimates are separated and labelled.

What qualifies
A figure the lab or organisation published itself, in a paper, technical report, model card, official blog or release note, about the compute, energy or money cost of training a named model. Press reports, leaks and analyst estimates do not qualify for the main table.
Sources
arXiv abstracts and PDFs, the Nature supplementary PDF for DeepSeek-R1, GitHub and Hugging Face model cards, lab blogs and release notes, and Xiaomi’s technical report PDF. Quoted phrases are verbatim from the current version of each document.
As-of date
Every row was re-read at its primary URL on September 25, 2026. OLMo 2 and OLMo 3 have revised arXiv versions (October 2025 and April 2026); quotes were checked against the current version. The Sky-T1 page shows an update dated March 14, 2026.
Units and arithmetic
Each row keeps the lab’s own unit. The only conversions we made are chips multiplied by stated days, marked “our arithmetic,” and one ratio between two DeepSeek figures priced at the same stated rate. We did not rank rows across units.
Independent check
Means a named outside party reproduced, audited or reviewed the cost figure itself. Peer review of a paper does not count. No row meets the bar; Mistral’s externally reviewed life-cycle analysis is noted as the nearest.
Not found
No absolute cost, hours or energy figure for Kimi K2, Kimi K3, Qwen3, GLM-4.5, DeepSeek-V4, Nemotron-4 or Apple’s AFM. Those are listed as hardware-only or omitted, not estimated.
Refresh
Maintained by the Digital Applied Team and updated in place at this URL when a lab publishes a new figure. Rows added later will carry their own read date.

09 — ConclusionThe disclosed number is always the floor

What to do with this

Cite the lab’s figure with its scope attached, and treat any bare dollar number as an estimate until you find the paper

Twenty-one rows is the whole public record, and it is thinner than the argument it gets used in. What the rows do show is consistent: the final run is a fraction of the spend, post-training is now a cost line of its own, and the labs with the most compute say the least. For the price of running these models rather than training them, our frontier model price index is updated monthly, our post on what $200 a month of AI buys covers the small end, and our AI transformation team can help you size a fine-tune of your own.

Digital Applied

Size an AI investment on figures that trace to a source.

We build the cost models, eval sets and vendor comparisons that let you decide with numbers you can defend in the room.

Cost modellingVendor due diligenceModel evaluation
Your next project

A number with its scope attached

  • →Every figure traced to a primary
  • →Exclusions stated, not hidden
  • →A budget that survives the first question
Questions and answers

The questions we get about AI training costs

No frontier lab has published a figure for its current models. The most-cited numbers are estimates: Epoch AI puts GPT-4's final run at about $40 million on an amortised-hardware basis, and the Stanford AI Index puts it at $78 to $79 million at cloud-rental prices. Cite the estimator, not the lab.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

AI Development

Xiaomi Published What Its AI Training Run Actually Cost

Xiaomi says six days of RL on MiMo-V2.6 cost about $850,000 for Flash and $2.62 million for Pro. What the figure covers, what it leaves out, and the prices.

September 22, 2026 · 9 minRead
AI Development

What "Open Source" Actually Includes for an AI Model

16 models in the open-model conversation, checked against the licence file, OSI's list and what else was released: weights, data, recipes, evaluation harness.

September 21, 2026 · 7 minRead
AI Development

Fireworks Ember-1: Kimi K3 Quality With Fewer Tokens?

Fireworks tuned Kimi K3 into Ember-1 and says it matches K3 with 35 to 50% shorter reasoning at the same price. The rows it loses, and the preview caveat.

September 23, 2026 · 5 minRead
AI Development

Gemini 3.8 Flash TTS: Voice Cloning and a Price That Doubles

Gemini 3.8 Flash TTS is generally available with voice cloning from a 30-second sample. The promotional price ends December 31 and doubles on January 1, 2027.

September 23, 2026 · 5 minRead
AI Development

AI Agent Governance: Policy and Compliance 2026 Guide

AI agent governance framework for enterprises — access control, audit trails, data residency, and compliance with EU AI Act and SOC 2 requirements.

May 23, 2026 · 20 minRead
AI Development

Google AI Plans: Free vs Plus vs Pro vs Ultra 2026

Google's AI subscription tiers after I/O 2026 — AI Plus $7.99, AI Pro $19.99, AI Ultra $100 (new), AI Ultra $200 (was $250). Feature matrix and decision tree.

May 23, 2026 · 14 minRead