ARC Prize published independent verification results for Claude Opus 5 on July 24, 2026: 30.16% on ARC-AGI-3 at High reasoning effort — the highest score recorded on that benchmark to date. The number matters. But the process behind it matters more, because it is one of the very few launch-window benchmark results that a third party administered rather than the vendor itself.
The same 48 hours produced a second ARC-AGI-3 headline from a different source: Anthropic’s own launch page, which claims a result “three times as high as the next-best model” — with no methodology, no harness description, and no mention of ARC Prize anywhere on the page. Same benchmark, same week, two unlinked claims with very different evidentiary weight.
This guide uses that split-screen moment as a teaching case. We cover exactly what ARC Prize verified and how, what even ARC Prize does not disclose, the independent cross-checks that complicate the headline, and a practical scorecard for reading the next benchmark claim that lands in your feed.
- 01Third-party administration is the real story.ARC Prize ran the Opus 5 evaluation itself and published 30.16% on ARC-AGI-3, 97.5% on ARC-AGI-1, and 90.4% on ARC-AGI-2 at Max effort. Nearly every other launch-day figure that week was vendor-run on a vendor-chosen harness.
- 02The honest caveat is part of the result.ARC Prize states plainly that Max reasoning effort was not evaluated on ARC-AGI-3 due to a short testing window. A benchmark that publishes what it did not test is more credible than one that publishes only wins.
- 03The comparison figures are secondary-channel only.The widely shared numbers for GPT-5.6 Sol and Opus 4.8 on ARC-AGI-3 come from ARC Prize’s X account and secondary press summaries — not the primary results page. They deserve a hedge, not a citation.
- 04Even the gold standard has disclosure gaps.The ARC Prize results page publishes no attempt limit, no harness description, and no cost-per-task figure. Third-party verification is a spectrum, not a binary trust switch.
- 05Read every claim through five questions.Who administered it, on what harness, at which effort setting, with how many attempts, and is the score final or preliminary. Any claim that cannot answer all five is marketing until proven otherwise.
01 — The ResultWhat ARC Prize actually verified.
The ARC Prize results page for Claude Opus 5, published July 24, 2026, covers all three generations of the ARC-AGI benchmark family. On ARC-AGI-1’s 400-task Public Eval set, Opus 5 scored 97.5% at both High and Max reasoning effort — the two settings produced identical results on this benchmark. On ARC-AGI-2’s 120-task Public Eval set, it scored 90.4% at Max and 88.3% at High. And on the ARC-AGI-3 Public Demo set of 25 interactive environments, it scored 30.16% at High effort — rounded to 30.2% in ARC Prize’s own social posts, a secondary channel we could only corroborate through search snippets, and the highest score recorded on that benchmark to date.
Claude Opus 5 · ARC-AGI family, as verified by ARC Prize
Source: ARC Prize results page, July 24, 2026The run also cleared new ground. In ARC Prize’s words, Opus 5 “completed five additional Public Demo environments that no model had previously beaten, demonstrating strong logical reasoning.” Five environments out of a 25-environment set moving from never-solved to solved in a single evaluation is a meaningful capability signal — and it is a signal you can trust more than usual, precisely because the organization reporting it had no stake in the model’s launch week.
Note what the chart does not show: a Max-effort bar for ARC-AGI-3. That absence is deliberate, disclosed, and — as we will argue in Section 04 — one of the most instructive details on the whole page.
02 — MethodologyWhy this benchmark is built differently.
The ARC-AGI family’s stated design philosophy is “Easy for Humans, Hard for AI.” The methodology page targets fluid intelligence — reasoning over genuinely novel problems — rather than crystallized intelligence, the recall of knowledge absorbed during pretraining. Tasks are restricted to “core knowledge priors”: cognitive building blocks present at birth or acquired very early in human development. The point of that restriction is anti-contamination by construction — a massive pretraining corpus should confer no hidden advantage.
The three generations escalate that philosophy. ARC-AGI-3 is the sharpest break: instead of static input/output puzzles, agents are dropped into live interactive environments where they must, per ARC Prize, “explore novel environments, acquire goals on the fly, build adaptable world models, and learn continuously.” The benchmark explicitly measures “skill-acquisition efficiency over time” and “long-horizon planning with sparse feedback” — not just final-answer accuracy. ARC Prize, Inc. administers the benchmark with a public leaderboard, and participation also runs through an ARC Prize 2026 Track on Kaggle.
ARC-AGI-1
The original grid-puzzle format: novel abstract transformations, solved from a handful of examples. Frontier models now near-saturate it — Opus 5 posts 97.5% at both High and Max effort.
ARC-AGI-2
The same static format with a harder public eval set. The High-to-Max gap reappears here: 88.3% at High, 90.4% at Max — evidence that extra reasoning budget still buys accuracy on this set.
ARC-AGI-3
Live environments instead of puzzles. Agents acquire goals on the fly, build world models, and are scored on skill-acquisition efficiency and long-horizon planning under sparse feedback. The frontier is nowhere near saturation.
That design lineage is why a 30.16% reads as a record rather than a failure. Static benchmarks that frontier models saturate stop discriminating between them; an interactive benchmark where the best recorded score is under a third of the ceiling still has years of signal left. For a deeper treatment of contamination and saturation dynamics, see our guide on how to read a benchmark leaderboard.
03 — The Split ScreenOne benchmark, two headlines in 48 hours.
Here is the detail that makes this launch a teaching case. Anthropic’s own Opus 5 launch page, published the same day as ARC Prize’s results, cites an ARC-AGI-3 result of its own: “Opus 5’s score is three times as high as the next-best model.” That sentence carries no methodology, no harness description, and no third-party name — the page never references ARC Prize at all. The vendor-framed headline and the independently verified number arrived from two different, unlinked sources within the same 48 hours.
The rest of Anthropic’s launch-day benchmark story is fully vendor-run: results on Frontier-Bench v0.1 (Anthropic’s own eval), CursorBench 3.2, OSWorld 2.0, and Zapier AutomationBench were all produced by the vendor, on harnesses and effort settings the vendor selected. We covered those numbers — and the pricing story around them — in our Opus 5 launch analysis, so we will not re-derive them here. The point is the category difference: those are claims a buyer must take on trust; the ARC number is a claim a buyer can trace to an administrator with no launch to sell.
This is the subtle version of a common failure: even a credible administrator has informal channels, and numbers that appear only in social posts do not carry the same weight as numbers on the results page. A rigorous reader keeps three tiers in mind — primary page, administrator’s social channel, third-hand press — and lets a claim’s tier set its confidence, not its virality.
04 — Honest LimitsWhat even ARC Prize does not publish.
The most quietly instructive line on the results page is a limitation, stated without spin.
"Due to the short testing window, ARC-AGI-3 was evaluated only at High reasoning effort."— ARC Prize, Claude Opus 5 results page, July 24, 2026
Max reasoning effort — the setting that lifted ARC-AGI-2 from 88.3% to 90.4% — was simply not run on ARC-AGI-3. So the single most-quoted number of the launch window is, by its own administrator’s admission, potentially an undercount of the model’s ceiling. That kind of caveat is what honest benchmarking looks like: publish what was tested, name what was not, and resist extrapolating. Effort settings are not a footnote, either — High versus Max is exactly the cost-quality lever we quantified in our reasoning-effort benchmarks, and it can move scores by whole percentage points.
But the page’s transparency has edges, and naming them matters for calibration. The results page discloses no attempt limit per environment, no harness or scaffold description, no cost-per-task figure for the run, and no named auditor credential beyond ARC Prize itself. On cost, the page offers only a qualitative framing for Max-effort runs elsewhere: “This is competitive with previous frontier leaders, though at slightly higher cost.” No dollar figure for the ARC-AGI-3 run exists in any source we checked — so you will not find one in this post either.
The honest conclusion: third-party verification is a spectrum. An independent administrator with partial methodology disclosure sits well above a vendor’s own harness — and still short of fully open methodology. Both facts fit in one head.
05 — Counter-EvidenceThe cross-checks that complicate the story.
A single benchmark — even a well-administered one — is one measurement, not a verdict. The most useful counter-evidence this week comes from an independent eval that is methodologically adjacent to ARC-AGI-3: Guanghan Ning’s “Witness” benchmark, an interactive puzzle-game evaluation administered separately. As reported by The Decoder, Opus 5 scored 43.4 there — described as “statistically equivalent to competitors.”
The Decoder’s framing of that gap is worth quoting because it is the properly cautious read: the ARC-AGI-3 jump “may reflect targeted optimization rather than broad reasoning advances,” while “the evidence remains inconclusive.” Two interactive benchmarks, one dramatic outperformance, one statistical tie. That does not invalidate the ARC result — it bounds what the ARC result can be said to prove. A record on one benchmark is evidence of capability on that benchmark’s task distribution, not a general-intelligence coronation.
This is our core interpretive claim, and it cuts both ways: the same skepticism applies to any model that tops any single leaderboard this year, including ones we like. The practical discipline is to require two independent measurements before updating a routing decision — and to weight the measurement whose administrator, harness, and sample size you can actually name.
06 — Crowd EvalsThe Preliminary tag problem.
The third administration model — crowd preference — has its own failure mode, visible live this week. On the Arena.ai text leaderboard as of this week, several current frontier and near-frontier entries carry an explicit “Preliminary” tag with wide error bands tied directly to low vote counts: Meta’s muse-spark-1.1 at 1495±7 Elo on 7,927 votes, Moonshot’s kimi-k3 at 1486±10 on 3,619 votes, Google’s gemini-3.6-flash at 1485±9 on 4,747 votes, and Alibaba’s qwen3.7-max-preview at 1475±10 on 3,714 votes. Long-established entries carry no tag and vote counts an order of magnitude higher — grok-4.20-beta sits on 60,330 votes with a ±4 band.
Votes behind each Elo · Arena.ai text leaderboard, Jul 26 snapshot
Source: arena.ai/leaderboard/text, snapshot retrieved July 26, 2026 — Elo and vote counts move continuouslyThe pattern has a mechanism: early votes on a newly listed model skew toward fans and early adopters, and rankings visibly move once volume catches up. A dated example from the same window: Moonshot’s Kimi K3 ranked first on the Arena WebDev leaderboard with a “preliminary score of 1,679,” the source explicitly noting “Preliminary result based on 1,757 votes and likely to change” — a nine-day-old snapshot we cite purely as a methodology illustration, since the K3 release itself is covered separately. The lesson generalizes: a crowd eval’s sample size is part of the score, and any number quoted without its vote count is missing a significant digit.
07 — The ScorecardThe benchmark credibility scorecard.
Put every claim from this launch window side by side and the pattern teaches itself. The table below grades each published number on the five disclosure dimensions a buyer can actually check: who administered it, whether the harness is described on the public page, whether an attempt limit is disclosed, whether the sample size is shown, and whether the score is final or explicitly preliminary. Every cell reflects what the cited public page disclosed as of July 26, 2026 — except the Kimi K3 · Arena WebDev row, which reflects the dated July 17 snapshot referenced in Section 06.
| Published claim | Administered by | Harness disclosed | Attempts disclosed | Sample size shown | Status |
|---|---|---|---|---|---|
| Third-party administered | |||||
| ARC-AGI-3 · Opus 5, 30.16% | ARC Prize, Inc. | No | No | Yes — 25 environments | Final · High effort only |
| ARC-AGI-1 / -2 · Opus 5 | ARC Prize, Inc. | No | No | Yes — 400 / 120 tasks | Final · High and Max |
| Vendor-run | |||||
| Anthropic launch suite (Frontier-Bench v0.1, CursorBench 3.2, OSWorld 2.0, Zapier AutomationBench) | Anthropic (vendor) | Vendor-chosen, not independently described | No | Varies by eval | Final, as published |
| Anthropic’s own ARC-AGI-3 claim (“three times as high”) | Anthropic (vendor self-report) | No | No | No | No third party named on page |
| Crowd-preference platforms | |||||
| Arena.ai text Elo (Jul 26 snapshot) | Crowd platform | N/A — human preference | N/A | Yes — per-model vote counts | Several entries “Preliminary” |
| Kimi K3 · Arena WebDev, 1,679 | Crowd platform | N/A — human preference | N/A | Yes — 1,757 votes | Explicitly “Preliminary” |
Two readings jump out. First, no row scores clean across all five columns — not even ARC Prize. Second, the failure modes differ by administration model: third parties tend to under-disclose harness detail, vendors tend to under-disclose everything except the win, and crowd platforms disclose sample size honestly but cannot escape early-vote skew. Knowing which failure mode you are looking at is most of the skill.
08 — Buyer’s GuideHow to read the next benchmark claim.
Distill the launch window into a repeatable protocol and you get five questions. They would have correctly triaged every number discussed in this post.
Who administered the run?
Third party with no launch to sell > vendor self-report. ARC Prize running the eval itself is why 30.16% is citable; Anthropic’s unattributed ‘three times as high’ on the same benchmark, the same week, is not comparable evidence.
What harness, at what effort?
Scaffold and effort setting move scores by whole points — Opus 5 gained 2.1 points on ARC-AGI-2 just going High to Max. A score quoted without its setting is a range masquerading as a number.
How many attempts, on what sample?
ARC Prize shows sample sizes (400 / 120 tasks, 25 environments) but no attempt limit; Arena shows vote counts that are the score’s real error bar. If neither attempts nor sample are visible, confidence should drop sharply.
Is it final — and what was left untested?
Preliminary tags, short testing windows, and untested settings are part of the result. ARC Prize saying Max was not evaluated on ARC-AGI-3 is a feature; a page with only wins and no caveats is the red flag.
The fifth question is the cross-check: does a second, independent measurement agree? The Witness-benchmark tie next to the ARC record is this week’s live demonstration that one leaderboard is never enough. The same verification instinct applies beyond benchmarks — our buyer’s scorecard for agent claims runs the identical playbook against “agentic” marketing.
Looking forward, we expect the verification gap to become a competitive axis in its own right. As launch cadence accelerates, vendors that invite third-party administration early will increasingly earn a trust premium over vendors that publish only their own harnesses, because buyers are learning to price the difference. If your team is choosing models for production and wants evals run on your workloads rather than anyone’s marketing page, that comparative benchmarking is exactly where our AI transformation engagements start.
09 — ConclusionVerification is the scarce resource.
The score made headlines. The administration model is the news.
Claude Opus 5’s 30.16% on ARC-AGI-3 is a genuine record on a genuinely hard benchmark — five previously unbeaten environments fell in one run, verified by an administrator with no launch to sell. That combination, not the number alone, is what earned this result a different tier of credibility from everything else published in the same 48 hours.
The same week showed every other pattern in miniature: a vendor citing the same benchmark without naming its administrator, comparison figures circulating from social channels that never reached the primary page, an adjacent independent benchmark reading statistically flat, and crowd leaderboards honestly flagging their own scores as preliminary. None of this is scandal — it is the normal texture of a fast-moving eval ecosystem, and it rewards readers who ask structural questions.
So keep the five questions taped above the routing table: who administered it, on what harness, at which effort setting, with how many attempts, and is it final. Models will keep leapfrogging each other monthly. The discipline of reading their scorecards does not change — and it compounds.