AI DevelopmentMethodology20 min readPublished August 16, 2026

42 rows · 9 vendors · disclosure checked on every row · independent checks attempted on 15

Can a Buyer Reproduce a Vendor Benchmark Row?

We collected 42 benchmark rows that nine AI vendors published on their own launch pages and model cards, each dated on or before August 14, 2026, and scored each against a stated disclosure rubric: does the vendor’s own page name the benchmark version, the task subset, the evaluation harness, and the reasoning-effort setting? The full dataset, the rubric, and every source URL are in this post.

DA
Digital Applied Team
Senior strategists · Published Aug 16, 2026
PublishedAug 16, 2026
Read time20 min
Sources13 primary URLs
Rows collected
42
9 vendors · pages dated on or before Aug 14, 2026
Fully specified + independently confirmed
0
of 42, under the stated rubric
Rows naming an eval harness
10/41
of the 41 scored rows
Independent checks attempted
15/42
5 matched or near-matched

Vendor benchmark reproducibility has a simple test: could a buyer, starting from the vendor’s own launch page, work out exactly which test produced the number in the table? We collected 42 benchmark rows that nine vendors — DeepSeek, Google, xAI, Z.ai, Anthropic, OpenAI, Moonshot AI, NVIDIA, and Mistral AI — published on pages dated on or before August 14, 2026, and checked what each page actually discloses.

This is an audit of disclosure, not an accusation of dishonesty. A score that only the vendor has published is not evidence of a false score, and several vendors in this dataset disclose their methodology unusually well — those rows are shown too. What the dataset measures is narrower and checkable: whether the page states the benchmark’s version, the task subset, the evaluation harness, and the reasoning-effort setting. Those four fields were checked on all 42 rows. Independent locatability — whether a third-party source reports a matching score — was attempted on only 15 of the 42 rows, and the results are labeled accordingly.

The complete dataset is in this page, row by row, with the rubric stated in full so anyone can recompute the tallies. Rows that could not be resolved are reported as unresolved rather than dropped.

Key takeaways
  1. 01
    Zero of 42 rows were both fully disclosed and independently confirmed.Under the stated rubric — version, subset, harness, and effort all named on the vendor's own page, plus an independent source reporting a matching score — no row in the dataset clears both bars at once.
  2. 02
    The headline finding is about disclosure, not locatability.Disclosure was checked on all 42 rows. Independent locatability was attempted on only 15 of 42 within this pass's fetch budget, so the dataset supports a disclosure rate, not a reproducibility rate.
  3. 03
    The same benchmark name covers 26% and 88.4% for one model.xAI's own page reports Grok 4.6 at 26% on Terminal-Bench v3.0, while an independent evaluator separately measured the same model at 88.4% on the older v2.1 — two real numbers, two different tests, one benchmark name.
  4. 04
    DeepSWE v1.1 is the dataset's strongest positive result.Four vendors' self-reported DeepSWE v1.1 scores each fell within the independent leaderboard's stated uncertainty band. Even there, every vendor page omitted at least one disclosure field — most often the harness.
  5. 05
    Honest vendor-only is a real category, distinct from hidden.DeepSeek labels its internal test sets as internal; Z.ai labels its Code Bench private; Kimi K3's page is the most fully disclosed row in the set. Plain labeling tells a reader not to expect independent confirmation — which is itself disclosure.

01The FindingDisclosure, measured on every row.

Each of the 42 rows is a (vendor, model, benchmark, score) triple that a vendor itself published — in a launch blog post, a model card, or a linked developer guide. For every row we read the vendor’s own page for four fields: the benchmark version (where the benchmark has one), the task subset (where applicable), the evaluation harness, and the reasoning-effort setting. Those four fields decide whether a reader can even identify which test produced the number, before any question of re-running it.

The result: of the 41 rows we could score against the rubric — every row except Mistral’s unresolved one — only 10 name an evaluation harness, and those 10 come from just three vendors (DeepSeek, Moonshot AI, and NVIDIA). Sixteen of the 41 state a reasoning-effort setting. Twenty-nine of the 42 rows are missing at least one applicable disclosure field and are scored under-specified. And zero rows in the dataset are simultaneously fully disclosed and independently confirmed — though, as the methodology below explains, the independent-confirmation column was only attempted on 15 of the 42 rows, so that zero is a statement about this dataset under this rubric, not a claim that no such row could exist.

Dataset
Rows across 9 vendors
42

41 rows we could score against the rubric plus 1 row we could not resolve at all (Mistral's Shieldstral, chart-image only), kept in the table rather than dropped. Every row cites the primary page it came from.

On or before Aug 14, 2026
Harness named
Scored rows naming a harness
10/41

All 10 come from three vendors: DeepSeek (a page-level footnote naming the harness and its mode), NVIDIA (NeMo Gym / NeMo Evaluator SDK), and Moonshot AI's Kimi K3 page. Three further rows are harness-not-applicable.

3 of 9 vendors
Top tier
Fully specified and independently locatable
0

No row discloses version, subset, harness, and effort on the vendor's own page and also has an independently reported matching score. The nearest misses are the four DeepSWE rows — confirmed independently, but each missing a disclosure field.

Under the stated rubric

Two things this finding is not. It is not a ranking of vendors by trustworthiness — a 42-row sample selected by “most recent flagship launch” says what each specific page stated, not how each company behaves in general. And it is not evidence that any number is wrong: this research re-ran no benchmark. The pattern it does document is that, at the time of writing, the standard launch-page benchmark table leaves a buyer unable to name the exact test behind most of its rows — and that the gap is disclosure the vendor could close with a footnote.

02MethodThe rubric, stated in full.

The selection rule was fixed before any row was scored. For each of the nine vendors, we identified the single most recent flagship-model launch document dated on or before August 16, 2026, fetched it directly — never a third-party summary — and took rows from the vendor’s own comparison table in the order they appeared, skipping only competitor-model comparison columns and rows that were pricing, latency, or cost figures rather than capability scores. No row was excluded for looking strong, weak, or poorly documented.

Methodology
Collected: 42 (vendor, model, benchmark, score) rows — 41 we could score against the rubric, 1 we could not resolve at all — from 9 vendors’ own launch posts, model cards, and developer guides, each dated on or before August 14, 2026 and each fetched directly from the primary page. Checked on all 42 rows: whether the vendor’s own page states the benchmark version, task subset, evaluation harness, and reasoning-effort setting (a page-level footnote covering multiple rows counts as disclosure). Attempted on 15 of 42 rows only: locating the same (model, benchmark) pairing on an independent leaderboard, maintainer page, or paper — the locatability column is partial, and rows not attempted are labeled “not checked this pass,” never folded into “not locatable.” Excluded: safety and red-team capability evaluations, pricing, context-window, and latency figures. Known limits: this is one bounded research pass, not a standing audit; independent leaderboards re-rank continuously; two leaderboards returned only their top rows to our fetch method; and several vendor pages render tables as chart images that yield no machine-readable numbers. As-of: vendor rows are current as of each page’s own publish date, the most recent of which is August 14, 2026; independent checks reflect each leaderboard’s state at the time of writing.

Each row then receives one of four verdicts. The rubric’s most important design decision: independent reproducibility and complete disclosure are two different axes, and a row can score well on one and poorly on the other. A row that an independent leaderboard confirms can still be under-specified if the vendor’s own page never told a reader which version or harness produced the number.

Tier 1
Fully specified, locatable
0 of 42 rows

Version (if the benchmark has one), subset (if applicable), harness, and effort are all disclosed on the vendor's own page, AND an independent source reports a matching or closely matching score for the same model.

The empty tier
Tier 2
Specified but vendor-only
8 of 42 rows

Disclosure is complete, but no independent match was found — including benchmarks the vendor explicitly labels internal or private, where independent confirmation is not applicable by design.

Includes honest private benchmarks
Tier 3
Under-specified
29 of 42 rows

At least one applicable disclosure field is missing, ambiguous, or unreachable — regardless of whether an independent match exists. All four independently confirmed DeepSWE rows land here.

The largest bucket
Tier 4
Not locatable
5 of 42 rows

The page gives only a relative claim with no absolute number to audit, or the score could not be extracted from the fetched page at all (chart-image-only rendering).

4 relative-only + 1 access-blocked

The independent column distinguishes five outcomes: a match (within the independent source’s own stated uncertainty band, or within roughly 2 points where no band is given), a near-match (a larger but small gap, stated per row), inconclusive (an independent source exists and loaded, but our fetch did not surface the specific model’s row — a limit of our method, not proof the score is absent), absent (checked and not found), and not checked this pass (genuinely not attempted). Keeping those apart matters: collapsing “we did not check” into “nobody can find it” would be exactly the kind of imprecision this audit exists to flag. For the companion protocol on reading a single vendor table, see the four levers that decide what a vendor benchmark table means; this post applies that protocol as a dataset across nine vendors.

03The DataThe complete 42-row dataset.

Every row below cites the vendor page it came from; the group header for each vendor links the primary source. “Page footnote” means the field is disclosed in a single page-level footnote covering the vendor’s benchmark rows. “Via ‘High’ label” means the effort setting is implied by the vendor’s own column header naming the reasoning-effort SKU. Verdicts use the four-tier rubric above.

The complete 42-row vendor benchmark disclosure dataset. Vendor rows are current as of each page’s own publish date, the most recent of which is August 14, 2026; independent-leaderboard checks reflect each leaderboard’s state at the time of writing. Columns: benchmark, vendor-reported score, version, subset, harness, effort, independent check result, and verdict under the stated rubric.
BenchmarkScore (vendor page)VersionSubsetHarnessEffortIndependent checkVerdict
DeepSeek — V4-Pro-0813 · Hugging Face model card · harness (“DeepSeek Harness,” minimal mode) and effort (max; temperature 1.0, top_p 0.95) disclosed in one page-level footnote
Terminal Bench87.92.1N/APage footnotePage footnoteInconclusive — the v2.1 leaderboard loaded only its top 3 of 201 models for usVendor-only
DeepSWE62.7Not statedN/APage footnotePage footnoteMatch — 63%±6% on the independent leaderboardUnder-specified
Toolathlon-Verified74.1N/A“Verified”Page footnotePage footnoteNot checked this passVendor-only
DSBench-FullStack71.1N/ALabeled internal (†)Page footnotePage footnoteN/A — vendor-labeled internal test setVendor-only
HLE (with tools)60.0N/A“With tools” shownPage footnotePage footnoteNot checked this passVendor-only
AutomationBench (Public)31.8Not stated“(Public)”Page footnotePage footnoteNot checked this passVendor-only
Google — Gemini 3.7 Flash · launch post · version labels present; no harness or effort named anywhere on the page
FrontierCode43.6%1.1“Main”Not namedNot statedNot checked this passUnder-specified
DeepSWE65.3%v1.1N/ANot namedNot statedMatch — 65%±2% on the independent leaderboardUnder-specified
WebDev Arena1588 EloN/A (live arena)N/AN/A (arena is the harness)N/AInconclusive — the arena page failed to render a usable leaderboard for usVendor-only
GDP.pdf34.0%Not statedNot statedNot namedNot statedNot checked this passUnder-specified
AutomationBench30.4%Not statedNot stated — public subset or not is unsaidNot namedNot statedNot checked this passUnder-specified
xAI — Grok 4.6 · launch post · effort implied by the “Grok 4.6 High” column label; no harness named for its own scores
AA Intelligence Index61Index version not statedN/AN/A (composite of 9 evals)Via “High” labelInconclusive — only Claude Opus 5 rows were visible in what loaded for usUnder-specified
GDPval-AA1753v2N/ANot namedVia “High” labelNot checked this passUnder-specified
CursorBench69.9%v3.2N/ANot namedVia “High” labelN/A — vendor-run benchmark; we found no independent leaderboard for itUnder-specified
Terminal-Bench26%v3.0N/ANot namedVia “High” labelInconclusive for this row; an independent evaluator separately measured the same model at 88.4% on the older v2.1Under-specified
Harvey LAB (Vals)15.8%Not statedNot statedNot namedVia “High” labelNot checked this passUnder-specified
FrontierCode61.3%v1.1“Extended” — the full 150-task setNot namedVia “High” labelSubset meaning confirmed at the maintainer; the score itself not checkedUnder-specified
DeepSWE65.9%v1.1N/ANot namedVia “High” labelMatch — 67%±2% on the independent leaderboardUnder-specified
Z.ai — GLM-5.3 · developer guide · three tables of different disclosure character; no harness named in any of them
Terminal Bench28.33.0N/ANot namedNot statedInconclusive — the v3.0 track is confirmed to exist at the maintainer; no score list loaded for usUnder-specified
DeepSWE66.9v1.1N/ANot namedNot statedAbsent — the leaderboard snapshot predates GLM-5.3’s release by one day (a timing artifact, not a disclosure failure)Under-specified
Agents’ Last Exam28.5Not statedNot statedNot namedNot statedNot checked this passUnder-specified
Z.ai Code Bench34.5%N/A (private benchmark)N/ANot named“Max” shownN/A — vendor-labeled privateVendor-only
CyberGym84.5%Not statedNot statedNot namedNot statedNot checked this passUnder-specified
ExploitBench54.4%Not statedNot statedNot namedNot statedNot checked this pass — compare the GPT-5.6 Sol ExploitBench row belowUnder-specified
Anthropic — Claude Opus 5 · launch announcement · every headline evaluation is a relative claim; the page states no absolute benchmark scores
Frontier-BenchRelative claim onlyv0.1N/ANot namedNot statedN/A — no absolute number to checkNot locatable
CursorBenchRelative claim only3.2N/ANot named“Max effort” namedN/A — no absolute number to checkNot locatable
ARC-AGI-3Relative claim onlyN/AN/ANot namedNot statedN/A on this page; a maintainer results page exists and was not fetched this passNot locatable
Zapier AutomationBenchRelative claim onlyNot statedNot statedNot namedNot isolated for this claimN/A — a distinct benchmark from “AutomationBench” aboveNot locatable
OpenAI — GPT-5.6 Sol · GA announcement · per-cell version labels, no harness column, partial prose-only effort disclosure
Agents’ Last Exam52.7% in the table; 53.6 in proseNot statedNot statedNot namedPartial, prose onlyNot checked this passUnder-specified
GDPval-AA1,747.8 Elov2N/ANot namedNot statedNot checked this passUnder-specified
Terminal-Bench88.8%2.1N/ANot namedNot statedNear-match — 89.5% (xhigh) independently; the page does not label its own effortUnder-specified
AA Intelligence Index58.9v4.1N/AN/ANot tied to this figureInconclusive — only Claude Opus 5 rows were visible in what loaded for usUnder-specified
DeepSWE72.7%v1.1N/ANot namedNot statedMatch — 73%±3% on the independent leaderboardUnder-specified
FrontierMath89%v2Tier 1–3Not namedNot statedInconclusive — the hub page loaded without a score table for usUnder-specified
ARC-AGI-37.78%N/AN/ANot named (footnote unreachable)Not stated (footnote unreachable)Not checked this pass — a footnote marker sits on this row; its text was unreachable in our fetchesUnder-specified
ExploitBench73.5%Not statedNot statedNot namedNot statedNot checked this passUnder-specified
BrowseComp90.4%N/ANot statedNot namedNot statedNot checked; the identical figure is separately self-reported by Kimi K3 under a different disclosed conditionUnder-specified
Moonshot AI — Kimi K3 · technical blog · page-level effort disclosure plus per-benchmark footnotes naming which harness evaluated which model
BrowseComp90.4N/AStated: 1M-token context, no context managementDescribed on the pagePage footnote (max; temp 1.0, top-p 1.0)Not checked this passVendor-only — the most fully disclosed row in the dataset
NVIDIA — Nemotron 3.5 Lightning · Hugging Face model card · harness named explicitly (“NeMo Gym / NeMo Evaluator SDK”); thinking-toggle setting undisclosed
SWE-bench Verified51.56N/A“Verified” splitNamedNot stated — a thinking on/off toggle exists; which setting produced this number is unsaidNot checked this passUnder-specified
Terminal-Bench24.582.1N/ANamedNot statedInconclusive — placement not visible in the top rows that loaded for usUnder-specified
GDPval-AA832Not statedN/ANamedNot statedNot checked this passUnder-specified
Mistral AI — Shieldstral 1.0 · announcement and arXiv paper · comparison figures rendered as chart images on both pages
WildGuardTest (F1)Unresolved — chart-image onlyBoth pages present the comparison as chart images; no machine-readable number could be extractedNot locatable — our access limitation, not a vendor disclosure failure

Two dataset notes. The release dates used for DeepSeek V4-Pro-0813 (August 13) and GLM-5.3 (August 14) come from our earlier reporting, not from a date field on the primary pages — the Hugging Face card and the Z.ai developer guide show no publish date. And the rows marked inconclusive against Artificial Analysis pages are inconclusive because of our fetch method: the Terminal-Bench v2.1 page states it lists 30 of 201 models, but only its top three rows loaded for us, and the Intelligence Index page surfaced only Claude Opus 5 rows. Those scores may well be present on the live pages — we simply could not confirm them, and say so rather than counting them as absent.

Verdict distribution · 42 rows under the stated rubric

Source: this audit's 42-row dataset, tallied from the table above
Under-specifiedAt least one applicable disclosure field missing — includes all 4 independently confirmed DeepSWE rows
29 of 42
Specified but vendor-onlyComplete disclosure, no independent match found — includes vendor-labeled internal and private benchmarks
8 of 42
Not locatable4 relative-only claims plus 1 chart-image-only row
5 of 42
Fully specified and independently locatableComplete disclosure plus an independent matching score
0 of 42
Cite this
Digital Applied. “Can a Buyer Reproduce a Vendor Benchmark Row?” Digital Applied Blog, August 16, 2026. https://www.digitalapplied.com/blog/vendor-benchmark-reproducibility-audit-2026Vendor rows as published on pages dated on or before August 14, 2026; independent checks as of the time of writing. The rubric and full dataset are in this page so the tallies can be recomputed.

04Version TrapsSame name, different test.

Three patterns in the dataset show why a benchmark name without its qualifiers is close to meaningless — the first two confirmed at the benchmark maintainer’s own page, the third inferred by comparing two vendors’ footnotes.

Terminal-Bench v2.1 versus v3.0

xAI’s launch page reports Grok 4.6 at 26% on Terminal-Bench v3.0 — a genuinely distinct, newer track confirmed to exist at the maintainer’s leaderboard hub — while Artificial Analysis’s independent v2.1 evaluation separately measured the same model at 88.4% at “high” effort. Read side by side without the version label, those two numbers look like two different products. GLM-5.3 also reports on v3.0 (28.3), roughly in step with Grok’s v3.0 figure, while two of the three v2.1 self-reports in this dataset sit in the high 80s — DeepSeek at 87.9 and GPT-5.6 Sol at 88.8%. The third v2.1 self-report, NVIDIA’s Nemotron 3.5 Lightning at 24.58, sits down with the v3.0 figures instead, which is the limit of the version reading: the label tells a reader which test ran, not on its own how high the number will be. The maintainer’s hub currently lists five distinct tracks — 3.0, 2.1, 2.0, a legacy 1.0, and a Science track marked coming soon — so “Terminal-Bench” with no version number is presently ambiguous among at least four active or near-active tests.

FrontierCode’s nested subsets

The maintainer, Cognition, states the structure plainly: Diamond is the 50 hardest tasks, Main is the 100 hardest (including Diamond), and Extended is the full set of 150. Extended is therefore, by construction, the easiest-skewing subset — it includes the 50 tasks the harder subsets exclude. Grok 4.6’s 61.3% is on Extended; Gemini 3.7 Flash’s 43.6% is on Main. Both pages disclose their subset, which is to their credit — but the two numbers are not the same test, and a comparison table that lines them up as if they were would mislead without either vendor having misstated anything.

An “AutomationBench” name collision

DeepSeek and Google both report a score on “AutomationBench” — apparently the public, GitHub-hosted tool-automation benchmark that Kimi K3’s footnotes describe evaluating on a 600-task public subset. Anthropic’s page separately reports on “Zapier AutomationBench” — by its own description a business-workflow evaluation built with Zapier specifically. These read as two different evaluations sharing a near-identical name, not three vendors citing one benchmark, and the dataset keeps them separate. An all-pass-style metric on a workflow benchmark also compresses scores in a way that can look alarming without its methodology — another reason the name alone carries so little information.

None of this is new in kind — FrontierMath v2 exists because the maintainer corrected errors in a large share of the original problem set, and scores moved when it did. Versioning is healthy benchmark hygiene. The trap is only in citing a versioned benchmark without its version.

05Independent ChecksWhere the 15 attempted checks landed.

Independent locatability was attempted on 15 of the 42 rows. Seven reached a conclusion: four exact matches and one near-match (all on DeepSWE v1.1 and Terminal-Bench 2.1), one confirmed absence that is a timing artifact, and one access-blocked row. Eight were inconclusive — an independent source exists and loaded, but our fetch did not surface the specific model’s row. The remaining 27 rows were either explicitly not applicable (vendor-labeled private benchmarks, vendor-run benchmarks for which we found no independent leaderboard, or relative-only claims with no number to check) or genuinely not attempted within this pass’s budget, and are labeled as such.

The strongest positive result in the dataset is DeepSWE v1.1 — the one benchmark enough vendors reported in common to make a like-for-like check possible. The independent DeepSWE leaderboard publishes Pass@1 with an uncertainty band, average cost per task, and step count per model — a richer disclosure format than most of the vendor pages in this dataset — and four separate vendors’ self-reported scores each fell inside its band.

DeepSWE v1.1 cross-check: each vendor’s self-reported score against the independent leaderboard’s figure and uncertainty band, as of the time of writing.
ModelVendor’s own pageIndependent leaderboardReading
DeepSeek V4-Pro-081362.763%±6% [max]Match — inside the stated band
Gemini 3.7 Flash65.3%65%±2% [high]Match — inside the stated band
Grok 4.665.9%67%±2% [xhigh]Match — inside the stated band
GPT-5.6 Sol72.7%73%±3% [max]Match — inside the stated band
GLM-5.366.9Absent from the snapshotThe leaderboard snapshot predates the model’s release by one day — a timing artifact
Claude Opus 5Not reported on its launch page74%±4% [max] — ranked firstLeaderboard-only: the top model on the independent board never cited the benchmark itself

Even here, the disclosure gap persists: every vendor page in the match rows omitted at least one field — most often the harness — that the independent leaderboard had to supply. And the one near-match is instructive in the other direction: OpenAI’s page reports Terminal-Bench 2.1 at 88.8% without labeling which effort level produced it, while the independent evaluation lists 89.5% at “xhigh” — a 0.7-point gap that is impossible to interpret precisely because the vendor’s own effort setting is unstated. The FrontierMath row sits in the inconclusive bucket for the same methodological reason: the maintainer’s hub loaded for us without a score table, which is a statement about our fetch, not about the score.

A caveat two vendors both flagged
Two unrelated vendors independently flagged the same measurement caveat about a competitor model. DeepSeek’s comparison table headers a column “Fable-5 (w/ fallback),” and Kimi K3’s footnotes separately state, verbatim: “Claude Fable 5 hit fallbacks on 35% of the tasks in our evaluation, which may have negatively impacted its measured performance.” Kimi’s footnote does not state how many tasks that 35% is a share of, so as quoted the figure travels without its denominator. Two vendors evaluating the same third-party model on different benchmarks both disclosing the same class of caveat is the kind of methodological candor this audit is designed to surface — and it cuts in the disclosing vendors’ favor.

06By VendorThe vendor-by-vendor scorecard.

The per-vendor summary below is derived entirely from the dataset table — no new research, so every cell can be recomputed from the rows above. It is a summary of these specific rows, not a general trustworthiness ranking: each vendor is represented by one launch document, selected by recency, and a different document from the same vendor could score differently.

Disclosure scorecard by vendor, derived from the 42-row dataset: rows contributed, rows naming an evaluation harness, rows stating a reasoning-effort setting, and independent checks attempted with their outcomes.
Vendor (rows)Harness namedEffort statedIndependent checks attempted → outcome
DeepSeek (6)6 of 66 of 62 → 1 match, 1 inconclusive
Google (5)0 of 50 of 52 → 1 match, 1 inconclusive
xAI (7)0 of 77 of 7 — via the “High” column label3 → 1 match, 2 inconclusive
Z.ai (6)0 of 61 of 62 → 1 absent (timing artifact), 1 inconclusive
Anthropic (4 — all relative-only claims)0 of 41 of 40 attempted — no absolute number to check
OpenAI (9)0 of 90 of 9 — partial prose disclosure only4 → 1 match, 1 near-match, 2 inconclusive
Moonshot AI (1)1 of 11 of 10 attempted
NVIDIA (3)3 of 30 of 31 → 1 inconclusive
Mistral AI (1 — unresolved)1 → access-blocked (chart-image only)

Vendors that specify their rows fully are a real result, and worth naming. DeepSeek disclosed harness and effort for all six of its rows through a single page-level footnote — “the minimal mode of DeepSeek Harness,” at max reasoning effort with temperature 1.0 and top_p 0.95 — and marked its internal test sets with a dagger rather than presenting them like public benchmarks. Kimi K3’s single row is the most fully disclosed in the dataset: page-level effort settings, per-benchmark footnotes naming which of three harnesses evaluated which model, and explicit URLs to the independent sources used for competitor scores. NVIDIA named its harness on every row and its model card shows a comparison model outperforming its own model on 11 of the 14 listed benchmarks — a table that does not favor its own product on most rows shown.

Anthropic’s launch page is the outlier in the other direction, in a specific and unusual way: its four headline evaluations are all relative claims — “more than doubles,” “within 0.5% of,” “three times as high as,” “around 1.5x” — with no absolute number stated for any of them, which is why all four rows land in the not-locatable tier. The counterpoint from the independent side is striking: Claude Opus 5 ranks first on the independent DeepSWE leaderboard at 74%±4%, a benchmark its own launch page never mentions. A vendor under-claiming relative to independent measurement is exactly the kind of finding a disclosure-only reading would miss. And OpenAI’s page — one of the largest benchmark tables in the dataset, with per-cell version labels — states two different Agents’ Last Exam figures on the same page, 53.6 in prose and 52.7% in its table, with no effort label connecting them.

07LimitsWhat this table does not show.

A dataset like this invites over-reading, so the boundaries are worth stating as plainly as the findings.

  • Not evidence that any number is false. Every vendor-only or under-specified verdict is a disclosure finding, not an accuracy finding. This research re-ran no benchmark.
  • Not a reproducibility rate. Independent locatability was attempted on 15 of 42 rows. The other 27 are labeled not-applicable or not-checked — different claims from “not locatable,” kept apart throughout.
  • Not a vendor trustworthiness ranking. One launch document per vendor, selected by recency, is not a representative sample of any company’s publication history.
  • Not proof that a not-locatable score is unlocatable in general. The Claude Opus 5 ARC-AGI-3 row, for instance, points to a maintainer results page that exists and was not fetched in this pass.
  • Not a judgment about which harness or effort setting is correct. The audit asks whether a setting was disclosed, not whether the vendor chose the right one.

Three per-row caveats deserve their own line. Mistral’s WildGuardTest row stays unresolved because both the announcement and the linked arXiv abstract present the comparison as chart images — an access limitation of our text-based fetches, not a vendor disclosure failure, and the row is marked that way. OpenAI’s ARC-AGI-3 row carries a footnote marker whose text was not present in anything our fetches could load — so that row is scored as having an unreachable footnote, not as simply undisclosed. And the BrowseComp coincidence — Kimi K3 and GPT-5.6 Sol each self-report exactly 90.4 on the same benchmark under two different disclosed conditions — is reported as a coincidence a reader auditing these tables would want to find and judge for themselves. This research found no evidence either figure is wrong.

The rubric travels with the number
The zero-in-the-top-tier finding is only defensible because the rubric is published with it: version, subset, harness, and effort all disclosed on the vendor’s own page, and an independent source reporting a matching score. Quote the zero without the rubric and it becomes unfalsifiable — which would make it weaker, not stronger. If you cite this dataset, cite the rubric with it.

08For BuyersWhat to ask for before you buy.

For a team using benchmark tables in model procurement, the dataset reduces to four questions to put to any vendor number — the same four fields the audit scored. Which version of the benchmark, since the dataset shows one benchmark name covering scores from the mid 20s to the high 80s? Which subset, since nested subsets like FrontierCode’s make the full set the easiest one? Which harness, since only three of nine vendors named one anywhere? And which effort setting, since a 0.7-point gap between a vendor number and an independent one cannot be interpreted without it?

The realistic posture is not to distrust vendor tables but to treat them as claims pending qualification — and to prefer benchmarks with an independent leaderboard, where the check costs minutes. Our companion guide covers how to read a leaderboard without being fooled by contamination or cherry-picking, and the earlier CursorBench analysis explains why we found no independent leaderboard to check for that vendor-controlled benchmark — the reason those rows are marked not-applicable here rather than not-checked. This post is one of a pair of audits published together; the open-weight licence audit applies the same census method to model licences. For teams that want this class of verification run against their own shortlist before committing to a model, it is part of what our AI transformation engagements cover.

Looking forward, the incentive structure suggests disclosure will improve unevenly: benchmarks with strong independent leaderboards (DeepSWE v1.1 in this dataset) already show vendor self-reports converging on independently measured values, because divergence is checkable within minutes. Where no independent check exists, nothing external disciplines the footnote. If that holds, the gap this audit measures should close fastest exactly where independent infrastructure exists — which is an argument for funding leaderboards, not for distrusting vendors.

09ConclusionA disclosure gap a footnote could close.

The state of vendor benchmark disclosure, August 2026

The gap is not honesty. It is four missing fields.

Forty-two rows, nine vendors, one rubric: no row in this dataset is simultaneously fully disclosed on the vendor’s own page and independently confirmed. The finding that matters is the disclosure half — checked on every row — and it is fixable at the cost of a footnote: DeepSeek’s single page-level methodology note and Kimi K3’s per-benchmark harness footnotes show the format already exists in production.

The dataset’s positive results deserve equal billing. Four vendors’ DeepSWE self-reports landed inside the independent leaderboard’s uncertainty bands. Two vendors label their private benchmarks as private, plainly. One vendor’s model tops an independent leaderboard its own launch page never cited. Disclosure and accuracy are different axes, and this dataset documents plenty of the latter where it could be checked at all — which was 15 rows out of 42.

That last clause is the honest boundary of this audit, and the reason its headline is about what vendor pages state rather than what third parties can find. The table is in this page in full, with the rubric, so the next person to count can check ours.

Verify before you standardize

Benchmark tables are claims. Your workload is the test.

We run vendor claims — benchmarks, licences, pricing — through primary-source verification before clients commit to a model or platform, and build evaluation harnesses matched to the workloads that actually matter.

Free consultationExpert guidanceTailored solutions
What we work on

Model evaluation engagements

  • Primary-source verification of vendor benchmark claims
  • Task-level evals on your own repos and workloads
  • Model shortlisting with disclosure-quality scoring
  • Multi-vendor routing based on measured, not quoted, results
  • Procurement briefs your engineering team can defend
FAQ · Vendor benchmark audit

The questions worth asking about the numbers.

The audit collected 42 benchmark rows — (vendor, model, benchmark, score) triples — that nine AI vendors published on their own launch pages, model cards, and developer guides, each dated on or before August 14, 2026. For every row it checked four disclosure fields on the vendor's own page: the benchmark version (where one exists), the task subset (where applicable), the evaluation harness, and the reasoning-effort setting. Separately, for 15 of the 42 rows, it attempted to locate the same model-benchmark pairing on an independent leaderboard, maintainer page, or paper. Each row received one of four verdicts under a rubric stated in full in the post, so the tallies can be recomputed by anyone. The headline result: zero rows were both fully disclosed and independently confirmed; 29 of 42 were missing at least one applicable disclosure field.
Related dispatches

Continue exploring benchmark integrity.

AI Development

ARC Prize Verified Opus 5. That Is Rarer Than It Sounds.

ARC Prize independently administered Claude Opus 5's 30.16% ARC-AGI-3 result. Almost every other launch-day benchmark number is vendor-run on a vendor harness.

July 26, 2026 · 11 minRead
AI Development

FrontierMath v2: When AI Benchmarks Get Error-Corrected

Epoch AI found errors in 42% of FrontierMath problems and shipped v2 on June 12, 2026. Scores jumped, rankings held — here is what that means for model choice.

June 14, 2026 · 10 minRead
AI Development

Claude Opus 4.8, 48 Hours In: The Early Eval Roundup

Opus 4.8 tops the Artificial Analysis index, but GPT-5.5 still leads Terminal-Bench. An evidence-graded roundup of the first 48 hours of independent evals.

May 30, 2026 · 11 minRead
AI Development

Cost-Per-Successful-Task: A New AI Evaluation Metric

Why $/token is the wrong unit and $/successful-task is the right one. Formulas, worked examples across 6 task families, and a downloadable scoring template.

April 23, 2026 · 5 minRead
AI Development

Google Intelligent Eyewear: Gemini AI Glasses Fall 2026

Google announces Gemini-powered smart glasses with Samsung, Gentle Monster, and Warby Parker at I/O. Audio glasses ship fall 2026; display tier TBD.

May 20, 2026 · 18 minRead
AI Development

AI Search Agents Compared: Google, Perplexity, ChatGPT

Google's always-on information agents, Perplexity Pro, and ChatGPT Search compared. Which AI search agent delivers the best research results in 2026?

May 20, 2026 · 14 minRead