IMO 2026 will be remembered as the year AI hit a perfect score on the hardest math contest for high schoolers — but the six systems being credited with 42/42 did not earn that score the same way. Two were graded by IMO organisers through the competition’s official process. Four were tested independently by one venture capitalist, with Claude-based AI agents doing the grading. Untangling those two claims is the difference between reading this milestone correctly and repeating a category error most of the coverage already made.
The stakes go beyond one contest. A year ago, 35/42 — the exact gold-medal cutoff — was the AI ceiling at the IMO, and it was treated as a landmark. Twelve months later the flagship math eval is effectively saturated: multiple labs, multiple methods, multiple perfect scores. When that happens, the score itself stops carrying information. What starts carrying information is how the score was verified, what the run cost, and how much of the result belongs to the model versus the harness around it.
This post separates the officially graded results from the self-administered ones, lays all seven disclosed AI results side-by-side in one trust ladder, pulls the cost and repair-round data nobody else is quoting, and closes with a practical framework for evaluating models once their flagship benchmark stops discriminating.
- 01Only two perfect scores were officially IMO-graded.Huawei’s Celia and Xiaohongshu’s dots-note-3.0 went through the IMO’s formal channel — problems released after humans finished, fixed submission window, no human intervention, graded by IMO organisers. It’s the first flawless AI score under the competition’s own process.
- 02The famous ‘four AIs at 42/42’ were self-administered, not IMO-official.Menlo Ventures partner Deedy Das ran Claude Fable 5, GPT-5.6 Sol, Kimi K3, and Axiom Math’s AxiomProver against the same problems in his own harness. His repo is explicit: graders were Claude-based agents, and scores should be treated as strong but not authoritative.
- 03Cost and effort now separate the perfect scorers.The three verified 42/42 runs in the Das harness ranged from ~$20.54 to $51.05 and from 2.5 to 17.4 hours. Within one model family, effort setting alone swung the first-pass score by 11 points — same model, same prompt.
- 04The trajectory is silver → gold-at-cutoff → saturation.Google reached silver in 2024 over 2–3 days, two labs hit exactly 35/42 gold in 2025, and 2026 produced multiple claimed perfects within a fixed time limit. Google DeepMind’s own IMO-Bench suite shows the labs saw answer-only saturation coming.
- 05Post-saturation, judge models on evals you own — not headlines.Task-level evals on your real workload, harness quality, verification tier, and cost per correct answer are the metrics that still discriminate once the flagship benchmark is a wall of perfect scores.
01 — What HappenedTwo different events, one blurred headline.
The 67th International Mathematical Olympiad was held in Shanghai, with the two exam papers sat on July 15–16, 2026. Among 666 human contestants, exactly 7 achieved a perfect 42/42. Then two genuinely different AI stories happened in the same news cycle — and wire coverage promptly welded them together.
Event one — the official channel. Huawei’s system Celia and Xiaohongshu/RedNote’s dots-note-3.0 each scored 42/42 under the IMO’s formal evaluation process: the firms received the problems only after human contestants finished, submissions were due within a specified time limit, and the solutions were graded by IMO organisers themselves. Xiaohongshu’s statement put the protocol plainly: “During testing, any form of human intervention was strictly prohibited.” Per the AFP wire reporting on July 23, this is the first time any AI system has cleared a flawless score under the competition’s own official evaluation system.
Event two — the independent test. Separately, Menlo Ventures partner Deedy Das built his own minimal agent harness, fed it the same IMO 2026 problem set — outside the official channel — and reported that four more systems reached 42/42: Claude Fable 5, GPT-5.6 Sol, Kimi K3, and Axiom Math’s AxiomProver. The grading in his published repo was done by Claude-based AI agents, not human medalists and not IMO coordinators. His own README says to treat the scores as strong but not authoritative.
Both events are real, and both matter. But “graded by the IMO” and “graded by an AI agent in a VC’s side project” are different classes of claim, and any roster that lists all six systems in one breath — as most syndicated headlines did — is quietly upgrading four of them. It was also a striking week for AI mathematics generally: earlier that week, Claude Fable 5 helped a mathematician disprove the decades-old Jacobian conjecture for dimensions three and up — a capability story with its own verification caveats, which makes the grading-method discipline here doubly relevant.
Huawei Celia + dots-note-3.0
Problems released only after human contestants finished, fixed submission window, human intervention barred, solutions graded by the IMO itself. The first flawless AI scores under the competition’s official process.
Four more models, self-administered
Deedy Das ran Claude Fable 5, GPT-5.6 Sol, Kimi K3, and AxiomProver against the same problems in his own harness. Not run through the IMO; graders were Claude-based agents. Full audit trail published on GitHub.
"The frontier of AI has officially moved well past IMO math."— Deedy Das, Partner, Menlo Ventures, via AFP wire
02 — The Full RosterSeven results, three verification tiers — one table.
No published source lays out every disclosed IMO 2026 AI result side-by-side with the grading method as its own column — every existing writeup picks one storyline. So here is the full ladder: the two officially graded scores, the three verified 42/42 runs from the Das harness, and the two systems verified by other means — Axiom Math’s machine-checked Lean proofs and NVIDIA’s open-weight Nemotron 3 Ultra, whose solutions the IMO team graded at 30/42, above the 29-point gold threshold this year.
| System | Score (/42) | How it was graded | Cost + time disclosed | Weights |
|---|---|---|---|---|
| Tier 1 · Officially graded through the IMO’s channel | ||||
| Huawei “Celia” | 42/42 | Official IMO grading — problems released after humans finished, fixed window, no human intervention | Not disclosed | Closed |
| Xiaohongshu “dots-note-3.0” | 42/42 | Official IMO grading — same protocol | Not disclosed | Not stated |
| Tier 2 · Self-administered — Das harness, Claude-based agent graders | ||||
| Claude Fable 5 | 42/42 · first pass | Claude-based agent graders; no repair rounds needed | $51.05 · 2.5h total ($38.83 / 1.8h excluding infra retries) | Closed |
| GPT-5.6 Sol (xhigh) | 39 → 42 after repair | Claude-based agent graders; repair rounds used reviewer feedback | ~$20.54 · 3.8h total | Closed |
| Kimi K3 (Moonshot AI) | 36 → 42 after repair | Claude-based agent graders; multiple repair rounds | ~$31.40 · 17.4h total | Closed — weights promised by Jul 27, 2026 |
| Tier 3 · Self-run, verified by other methods | ||||
| Axiom Math “AxiomProver” | 42/42 | Machine-checked formal proofs in Lean 4 (Mathlib v4.31.0) — no evidence of official IMO sign-off | ~25h working time · ~8,000 lines of Lean | Not stated |
| NVIDIA Nemotron 3 Ultra | 30/42 | Run by NVIDIA under the contest time limit, no internet or external tools; solutions graded by the IMO team — above the 29-point gold threshold | Not disclosed | Open (released Jun 4, 2026) |
Read down the “how it was graded” column and the headline changes. Every 42/42 is real in the sense that somebody scored it — but the two tiers below the official one carry known limitations their own authors disclose. And one of the most interesting rows is not a perfect score at all: an open-weight model reaching gold-medal territory under IMO-team grading. AI researcher Ethan Mollick noted that this appears to be the first time an open model has reported gold-medal-level status at the IMO — a threshold he described as “a rather big threshold” when closed models first crossed it a year earlier.
03 — The ConflationHow the wire copy blurred official and self-graded.
The AFP wire story that most outlets syndicated on July 23 reported both events in the same piece — and its headline framing, that AI “caught up with humans” to score 100% under the competition’s official judging process, reads as if all six models were officially graded. Read closely, the wire’s own “independent verification” passage does separate Das’s four scores as self-reported via social media. But headlines travel and nuance doesn’t: many syndicated copies ran the roster flat, with no tier distinction at all.
Das himself is not the source of the confusion — his repo carries the caveat in plain text, and it is worth quoting exactly, because it is the single most load-bearing sentence in this entire story:
"Graders are Claude-based agents, not human medalists; treat scores as strong but not authoritative."— github.com/deedy/imo-2026 repository README
That sentence should be stapled to every secondary retelling of the “four AIs” claim, and mostly isn’t. Note also what a Claude-based grader implies for one specific row: Claude Fable 5’s 42/42 was scored, in part, by agents built on the same model family — a circularity the repo discloses but downstream coverage never mentions. None of this means the scores are wrong. It means they sit at a different rung of the trust ladder, and the market understood this before the media did: a Manifold prediction market on whether a lab would score a perfect IMO 2026 built its resolution criteria explicitly around verification quality — independent reports would need “more scrutiny,” and cherry-picking (running a model repeatedly and reporting only successes) wouldn’t count. Its implied odds climbed from 85% through 91% to 94% by July 19 and to 96% shortly after — informal test chatter pricing in the outcome well before any official confirmation landed.
Notably, neither Anthropic nor OpenAI had published any official announcement about their models’ IMO 2026 results as of this post’s publication — the entire “four AIs” storyline runs through one investor’s independent test plus wire pickup, not vendor newsrooms. When a capability claim this large has no primary vendor source, the verification tier is doing all the work.
04 — The TrajectorySilver to saturation in two years.
The speed of the ceiling collapse is the context that makes 2026 legible. In 2024, Google’s AI system reached silver-medal level, solving 4 of 6 problems — over two to three days, without a same-day time cap. In July 2025 at IMO 2025 in Queensland, Google DeepMind’s Gemini Deep Think and an OpenAI experimental reasoning model both scored exactly 35/42 — the precise gold-medal cutoff that year — each solving five of six problems and missing the hardest.
And here is the detail that makes the official-versus-self-graded split a recurring pattern rather than a 2026 one-off: only DeepMind’s 2025 run was graded by official IMO coordinators. OpenAI’s 2025 score was assessed independently by three former IMO medalists reaching unanimous consensus — credible, but not the official channel. Even then, IMO President Gregor Dolinar flagged the boundary of what official confirmation could cover: “Contest organizers could not verify how much computing power had been used by the AI models or whether there had been human involvement.” He also praised the graded solutions themselves: “Their solutions were astonishing in many respects. IMO graders found them to be clear, precise and most of them easy to follow.”
Silver medal — Google
Google’s system solved 4 of 6 problems working over 2–3 days, without the same-day time constraint humans face. Silver-medal level, and at the time a breakthrough.
Gold at the exact cutoff
Gemini Deep Think (officially IMO-graded) and an OpenAI experimental model (graded by three former IMO medalists) both landed on 35/42 — the precise gold threshold. Problem 6 defeated both.
Multiple claimed perfects
Two officially graded perfect scores, four self-administered claims, one Lean-formalized proof set, and an open-weight model at gold level. The flagship eval stopped discriminating at the top.
Two data points make a line; three make a trend. The trend here is not just “models got better” — it is that the labs themselves saw answer-only saturation coming. Google DeepMind built and published IMO-Bench, a four-part benchmark suite introduced at EMNLP 2025: 400 answer-verifiable problems, 60 proof-based problems, 1,000 grading examples, and 60 Lean-formalized problems — explicitly designed to grade proof quality, not just final answers. Researchers had likewise noted, before IMO 2026 concluded, that saturating easier proxy contests is not the same as a perfect IMO, and that the hardest olympiad problems are precisely where automated grading tends to break down. The institution that benefits most from “AI solved math” headlines was already building the harder eval — which tells you how much signal the labs themselves assign to a saturated one.
05 — The Independent TestInside the Das harness: repair rounds, effort tiers, and cost per run.
Whatever its verification tier, the Das repo is the richest public dataset to come out of IMO 2026 — 9 runs across 7 models, identical system prompts, a 150-minute cap per problem, and a deliberately minimal harness: a single-context agent loop with three tools (bash, write_file, read_file), no multi-agent orchestration, and no reviewer step on the initial attempt. That minimalism is what makes the cross-model comparison meaningful. Here is the field on first pass, before any repair rounds:
First-pass scores in the Das harness · IMO 2026, before repair rounds
Source: github.com/deedy/imo-2026 REPORT.md — Claude-agent graded, not IMO-officialThree findings in this data matter more than the perfect scores. First, the effort-tier swing. GPT-5.6 Sol at default effort scored 28/42 on first pass; the identical model at xhigh effort scored 39/42 — an 11-point swing from the effort setting alone, same model, same prompt. That is direct evidence that how much compute you spend per task is now a variable as large as which model you pick. Oddly, Sol’s max-effort run scored 30/42 — below xhigh — a reminder that effort tiers are not a monotonic dial either.
Second, the cost spread among the perfect runs. Claude Fable 5 solved everything in one attempt and was the fastest — $51.05 and 2.5 hours all-in ($38.83 and 1.8 hours excluding infrastructure retries). GPT-5.6 Sol was the cheapest at roughly $20.54 over 3.8 hours, needing repair rounds to close from 39 to 42. Kimi K3 got there too, but took 17.4 hours — about 4.6 times Sol’s wall-clock — and, in Das’s words, “a LOT of tokens,” at roughly $31.40. Three models, one score, three very different cost-and-latency profiles. That spread — roughly $20.54 to $51.05 per run — is invisible in every headline.
Third, the correlated failure. Meta Muse Spark 1.1 and DeepSeek V4 Pro independently produced the identical wrong answer to Problem 3 — c=(n+1)/(2n+1), false for all n≥2. Independent models converging on the same wrong answer is evidence the errors trace to shared training-data patterns rather than fully independent reasoning chains — and it is a stronger argument that evals can share blind spots than any generic saturation warning. This is the same lesson coding benchmarks taught earlier this year: in our SWE-Bench Verified analysis, the scaffolding around the model moved scores as much as the model itself. The Das data shows the same physics operating on math.
06 — VerificationThree ways to trust a proof.
IMO 2026 accidentally produced a clean natural experiment in verification methods. The same six problems were checked three different ways — and each way carries a different trust model.
Human institutional grading. The IMO organisers grading Celia and dots-note-3.0 is the gold standard for this contest — but even it has limits. As Dolinar’s 2025 comments made explicit, organisers cannot verify how much compute was burned or what happened before submission. Official grading verifies the output, not the process.
LLM-agent grading. The Das harness scales infinitely and publishes full audit trails, but the graders are themselves language-model agents — with all the shared-blind-spot risk the correlated P3 failure demonstrates. Strong evidence, not authoritative evidence, exactly as labeled.
Machine-checked formal proof. Axiom Math’s AxiomProver took a categorically different route: full formal proofs written in Lean 4 against Mathlib v4.31.0, mechanically verified by the proof assistant — roughly 8,000 lines of Lean across the six problems, about 25 hours of working time, heavily front-loaded on the two hardest problems. A Lean proof cannot be wrong in the way an LLM-graded solution can. But note the caveat that survives even here: mechanical checking validates the proof’s logical correctness, not that IMO organisers reviewed or accepted it — Axiom’s 42/42 shows no evidence of official IMO sign-off either. Formal verification and institutional verification are different axes, and this result maxes one while skipping the other.
Axiom Math itself is worth knowing: founded in 2022, it raised a $200M Series A led by Menlo Ventures — the same firm where Das is a partner, a connection worth noting when weighing the independent test’s framing — at a $1.6B valuation, is led by CEO Carina Hong with number theorist Ken Ono as founding mathematician, and had already posted a perfect 120/120 on Putnam 2025 against top humans’ 110. Meanwhile the open-weight story is converging on the same space: Kimi K3 — one of the four self-administered perfect scorers — shipped July 17 with open weights promised by July 27 but not yet released as of this writing, and NVIDIA’s already-open Nemotron 3 Ultra hit gold-level 30/42 under IMO-team grading.
07 — The FrameworkJudging models after the benchmark saturates.
Here is the practical payload. A saturated benchmark is not useless — it is a floor, not a differentiator. Once several models share the ceiling on a flagship eval, the buying decision moves to four questions the headline number cannot answer. This is the same eval-literacy discipline we laid out for coding models in our guide to reading SWE-Bench and Terminal-Bench, now with the IMO data to make each question concrete.
Task-level evals you own
The IMO measures olympiad math; your workload isn’t olympiad math. Build a 20–50 case eval from your real tickets, documents, or queries and run every candidate model against it. A saturated public benchmark tells you a model is in the frontier class — your eval tells you which one wins your work.
Harness quality
Das’s minimal three-tool loop, repair rounds, and reviewer feedback moved scores by multiple points independent of the model. Ask of any vendor number: what scaffolding produced it, how many attempts were allowed, and who graded the output.
Cost per correct answer
Three models tied at 42/42 — at ~$20.54, ~$31.40, and $51.05, across 2.5 to 17.4 hours. At equal accuracy, unit economics decide. And the 11-point effort-tier swing inside one model means cost tuning is now part of model selection, not an afterthought.
Verification tier
Officially graded, mechanically verified, or self-graded by AI agents? Record the tier next to every number your team quotes. IMO 2026 produced all three tiers in one week — and most coverage collapsed them into one claim.
Projecting forward: expect the 2027 version of this story to be about process, not scores. Proof-quality suites like IMO-Bench, formal-verification pipelines like Axiom’s, and disclosed cost-per-run audit trails like Das’s are the three formats that still produce signal once every frontier model clears the old bar. The models that look identical on a saturated leaderboard will separate sharply on your own tasks, your own latency budget, and your own cost ceiling — which is exactly the comparative evaluation work our AI transformation engagements start with: building the task-level eval that a public benchmark can no longer substitute for.
08 — ConclusionThe score stopped being the story.
When a benchmark saturates, the verification method becomes the benchmark.
IMO 2026 delivered a genuine milestone: the first flawless AI scores under the competition’s official grading — earned by Huawei’s Celia and Xiaohongshu’s dots-note-3.0, and by them alone. The four other perfect scores you read about — Claude Fable 5, GPT-5.6 Sol, Kimi K3, AxiomProver — came from one investor’s self-administered test, graded by Claude-based agents, published with an honest caveat that most retellings dropped. Both stories are impressive. They are not the same story.
The deeper shift is what saturation does to decision-making. A year after 35/42 was a landmark, a perfect score no longer separates frontier models — but grading method, harness design, effort tiers, and cost per correct answer separate them sharply. The richest findings of the week weren’t the 42s at all: an 11-point swing from an effort setting, two models sharing an identical wrong answer, and a 2.5×-cost spread among systems tied at the top.
So: four AIs scored a perfect 42/42 — so what? So this: stop quoting saturated benchmarks as buying signals, start recording who graded every number you cite, and move your evaluation budget to tests you own. The labs already have — that’s what IMO-Bench is. The press hasn’t. Your team can be ahead of one of them.