Vendor benchmark reproducibility has a simple test: could a buyer, starting from the vendor’s own launch page, work out exactly which test produced the number in the table? We collected 42 benchmark rows that nine vendors — DeepSeek, Google, xAI, Z.ai, Anthropic, OpenAI, Moonshot AI, NVIDIA, and Mistral AI — published on pages dated on or before August 14, 2026, and checked what each page actually discloses.
This is an audit of disclosure, not an accusation of dishonesty. A score that only the vendor has published is not evidence of a false score, and several vendors in this dataset disclose their methodology unusually well — those rows are shown too. What the dataset measures is narrower and checkable: whether the page states the benchmark’s version, the task subset, the evaluation harness, and the reasoning-effort setting. Those four fields were checked on all 42 rows. Independent locatability — whether a third-party source reports a matching score — was attempted on only 15 of the 42 rows, and the results are labeled accordingly.
The complete dataset is in this page, row by row, with the rubric stated in full so anyone can recompute the tallies. Rows that could not be resolved are reported as unresolved rather than dropped.
- 01Zero of 42 rows were both fully disclosed and independently confirmed.Under the stated rubric — version, subset, harness, and effort all named on the vendor's own page, plus an independent source reporting a matching score — no row in the dataset clears both bars at once.
- 02The headline finding is about disclosure, not locatability.Disclosure was checked on all 42 rows. Independent locatability was attempted on only 15 of 42 within this pass's fetch budget, so the dataset supports a disclosure rate, not a reproducibility rate.
- 03The same benchmark name covers 26% and 88.4% for one model.xAI's own page reports Grok 4.6 at 26% on Terminal-Bench v3.0, while an independent evaluator separately measured the same model at 88.4% on the older v2.1 — two real numbers, two different tests, one benchmark name.
- 04DeepSWE v1.1 is the dataset's strongest positive result.Four vendors' self-reported DeepSWE v1.1 scores each fell within the independent leaderboard's stated uncertainty band. Even there, every vendor page omitted at least one disclosure field — most often the harness.
- 05Honest vendor-only is a real category, distinct from hidden.DeepSeek labels its internal test sets as internal; Z.ai labels its Code Bench private; Kimi K3's page is the most fully disclosed row in the set. Plain labeling tells a reader not to expect independent confirmation — which is itself disclosure.
01 — The FindingDisclosure, measured on every row.
Each of the 42 rows is a (vendor, model, benchmark, score) triple that a vendor itself published — in a launch blog post, a model card, or a linked developer guide. For every row we read the vendor’s own page for four fields: the benchmark version (where the benchmark has one), the task subset (where applicable), the evaluation harness, and the reasoning-effort setting. Those four fields decide whether a reader can even identify which test produced the number, before any question of re-running it.
The result: of the 41 rows we could score against the rubric — every row except Mistral’s unresolved one — only 10 name an evaluation harness, and those 10 come from just three vendors (DeepSeek, Moonshot AI, and NVIDIA). Sixteen of the 41 state a reasoning-effort setting. Twenty-nine of the 42 rows are missing at least one applicable disclosure field and are scored under-specified. And zero rows in the dataset are simultaneously fully disclosed and independently confirmed — though, as the methodology below explains, the independent-confirmation column was only attempted on 15 of the 42 rows, so that zero is a statement about this dataset under this rubric, not a claim that no such row could exist.
Rows across 9 vendors
41 rows we could score against the rubric plus 1 row we could not resolve at all (Mistral's Shieldstral, chart-image only), kept in the table rather than dropped. Every row cites the primary page it came from.
Scored rows naming a harness
All 10 come from three vendors: DeepSeek (a page-level footnote naming the harness and its mode), NVIDIA (NeMo Gym / NeMo Evaluator SDK), and Moonshot AI's Kimi K3 page. Three further rows are harness-not-applicable.
Fully specified and independently locatable
No row discloses version, subset, harness, and effort on the vendor's own page and also has an independently reported matching score. The nearest misses are the four DeepSWE rows — confirmed independently, but each missing a disclosure field.
Two things this finding is not. It is not a ranking of vendors by trustworthiness — a 42-row sample selected by “most recent flagship launch” says what each specific page stated, not how each company behaves in general. And it is not evidence that any number is wrong: this research re-ran no benchmark. The pattern it does document is that, at the time of writing, the standard launch-page benchmark table leaves a buyer unable to name the exact test behind most of its rows — and that the gap is disclosure the vendor could close with a footnote.
02 — MethodThe rubric, stated in full.
The selection rule was fixed before any row was scored. For each of the nine vendors, we identified the single most recent flagship-model launch document dated on or before August 16, 2026, fetched it directly — never a third-party summary — and took rows from the vendor’s own comparison table in the order they appeared, skipping only competitor-model comparison columns and rows that were pricing, latency, or cost figures rather than capability scores. No row was excluded for looking strong, weak, or poorly documented.
Each row then receives one of four verdicts. The rubric’s most important design decision: independent reproducibility and complete disclosure are two different axes, and a row can score well on one and poorly on the other. A row that an independent leaderboard confirms can still be under-specified if the vendor’s own page never told a reader which version or harness produced the number.
Fully specified, locatable
Version (if the benchmark has one), subset (if applicable), harness, and effort are all disclosed on the vendor's own page, AND an independent source reports a matching or closely matching score for the same model.
Specified but vendor-only
Disclosure is complete, but no independent match was found — including benchmarks the vendor explicitly labels internal or private, where independent confirmation is not applicable by design.
Under-specified
At least one applicable disclosure field is missing, ambiguous, or unreachable — regardless of whether an independent match exists. All four independently confirmed DeepSWE rows land here.
Not locatable
The page gives only a relative claim with no absolute number to audit, or the score could not be extracted from the fetched page at all (chart-image-only rendering).
The independent column distinguishes five outcomes: a match (within the independent source’s own stated uncertainty band, or within roughly 2 points where no band is given), a near-match (a larger but small gap, stated per row), inconclusive (an independent source exists and loaded, but our fetch did not surface the specific model’s row — a limit of our method, not proof the score is absent), absent (checked and not found), and not checked this pass (genuinely not attempted). Keeping those apart matters: collapsing “we did not check” into “nobody can find it” would be exactly the kind of imprecision this audit exists to flag. For the companion protocol on reading a single vendor table, see the four levers that decide what a vendor benchmark table means; this post applies that protocol as a dataset across nine vendors.
03 — The DataThe complete 42-row dataset.
Every row below cites the vendor page it came from; the group header for each vendor links the primary source. “Page footnote” means the field is disclosed in a single page-level footnote covering the vendor’s benchmark rows. “Via ‘High’ label” means the effort setting is implied by the vendor’s own column header naming the reasoning-effort SKU. Verdicts use the four-tier rubric above.
| Benchmark | Score (vendor page) | Version | Subset | Harness | Effort | Independent check | Verdict |
|---|---|---|---|---|---|---|---|
| DeepSeek — V4-Pro-0813 · Hugging Face model card · harness (“DeepSeek Harness,” minimal mode) and effort (max; temperature 1.0, top_p 0.95) disclosed in one page-level footnote | |||||||
| Terminal Bench | 87.9 | 2.1 | N/A | Page footnote | Page footnote | Inconclusive — the v2.1 leaderboard loaded only its top 3 of 201 models for us | Vendor-only |
| DeepSWE | 62.7 | Not stated | N/A | Page footnote | Page footnote | Match — 63%±6% on the independent leaderboard | Under-specified |
| Toolathlon-Verified | 74.1 | N/A | “Verified” | Page footnote | Page footnote | Not checked this pass | Vendor-only |
| DSBench-FullStack | 71.1 | N/A | Labeled internal (†) | Page footnote | Page footnote | N/A — vendor-labeled internal test set | Vendor-only |
| HLE (with tools) | 60.0 | N/A | “With tools” shown | Page footnote | Page footnote | Not checked this pass | Vendor-only |
| AutomationBench (Public) | 31.8 | Not stated | “(Public)” | Page footnote | Page footnote | Not checked this pass | Vendor-only |
| Google — Gemini 3.7 Flash · launch post · version labels present; no harness or effort named anywhere on the page | |||||||
| FrontierCode | 43.6% | 1.1 | “Main” | Not named | Not stated | Not checked this pass | Under-specified |
| DeepSWE | 65.3% | v1.1 | N/A | Not named | Not stated | Match — 65%±2% on the independent leaderboard | Under-specified |
| WebDev Arena | 1588 Elo | N/A (live arena) | N/A | N/A (arena is the harness) | N/A | Inconclusive — the arena page failed to render a usable leaderboard for us | Vendor-only |
| GDP.pdf | 34.0% | Not stated | Not stated | Not named | Not stated | Not checked this pass | Under-specified |
| AutomationBench | 30.4% | Not stated | Not stated — public subset or not is unsaid | Not named | Not stated | Not checked this pass | Under-specified |
| xAI — Grok 4.6 · launch post · effort implied by the “Grok 4.6 High” column label; no harness named for its own scores | |||||||
| AA Intelligence Index | 61 | Index version not stated | N/A | N/A (composite of 9 evals) | Via “High” label | Inconclusive — only Claude Opus 5 rows were visible in what loaded for us | Under-specified |
| GDPval-AA | 1753 | v2 | N/A | Not named | Via “High” label | Not checked this pass | Under-specified |
| CursorBench | 69.9% | v3.2 | N/A | Not named | Via “High” label | N/A — vendor-run benchmark; we found no independent leaderboard for it | Under-specified |
| Terminal-Bench | 26% | v3.0 | N/A | Not named | Via “High” label | Inconclusive for this row; an independent evaluator separately measured the same model at 88.4% on the older v2.1 | Under-specified |
| Harvey LAB (Vals) | 15.8% | Not stated | Not stated | Not named | Via “High” label | Not checked this pass | Under-specified |
| FrontierCode | 61.3% | v1.1 | “Extended” — the full 150-task set | Not named | Via “High” label | Subset meaning confirmed at the maintainer; the score itself not checked | Under-specified |
| DeepSWE | 65.9% | v1.1 | N/A | Not named | Via “High” label | Match — 67%±2% on the independent leaderboard | Under-specified |
| Z.ai — GLM-5.3 · developer guide · three tables of different disclosure character; no harness named in any of them | |||||||
| Terminal Bench | 28.3 | 3.0 | N/A | Not named | Not stated | Inconclusive — the v3.0 track is confirmed to exist at the maintainer; no score list loaded for us | Under-specified |
| DeepSWE | 66.9 | v1.1 | N/A | Not named | Not stated | Absent — the leaderboard snapshot predates GLM-5.3’s release by one day (a timing artifact, not a disclosure failure) | Under-specified |
| Agents’ Last Exam | 28.5 | Not stated | Not stated | Not named | Not stated | Not checked this pass | Under-specified |
| Z.ai Code Bench | 34.5% | N/A (private benchmark) | N/A | Not named | “Max” shown | N/A — vendor-labeled private | Vendor-only |
| CyberGym | 84.5% | Not stated | Not stated | Not named | Not stated | Not checked this pass | Under-specified |
| ExploitBench | 54.4% | Not stated | Not stated | Not named | Not stated | Not checked this pass — compare the GPT-5.6 Sol ExploitBench row below | Under-specified |
| Anthropic — Claude Opus 5 · launch announcement · every headline evaluation is a relative claim; the page states no absolute benchmark scores | |||||||
| Frontier-Bench | Relative claim only | v0.1 | N/A | Not named | Not stated | N/A — no absolute number to check | Not locatable |
| CursorBench | Relative claim only | 3.2 | N/A | Not named | “Max effort” named | N/A — no absolute number to check | Not locatable |
| ARC-AGI-3 | Relative claim only | N/A | N/A | Not named | Not stated | N/A on this page; a maintainer results page exists and was not fetched this pass | Not locatable |
| Zapier AutomationBench | Relative claim only | Not stated | Not stated | Not named | Not isolated for this claim | N/A — a distinct benchmark from “AutomationBench” above | Not locatable |
| OpenAI — GPT-5.6 Sol · GA announcement · per-cell version labels, no harness column, partial prose-only effort disclosure | |||||||
| Agents’ Last Exam | 52.7% in the table; 53.6 in prose | Not stated | Not stated | Not named | Partial, prose only | Not checked this pass | Under-specified |
| GDPval-AA | 1,747.8 Elo | v2 | N/A | Not named | Not stated | Not checked this pass | Under-specified |
| Terminal-Bench | 88.8% | 2.1 | N/A | Not named | Not stated | Near-match — 89.5% (xhigh) independently; the page does not label its own effort | Under-specified |
| AA Intelligence Index | 58.9 | v4.1 | N/A | N/A | Not tied to this figure | Inconclusive — only Claude Opus 5 rows were visible in what loaded for us | Under-specified |
| DeepSWE | 72.7% | v1.1 | N/A | Not named | Not stated | Match — 73%±3% on the independent leaderboard | Under-specified |
| FrontierMath | 89% | v2 | Tier 1–3 | Not named | Not stated | Inconclusive — the hub page loaded without a score table for us | Under-specified |
| ARC-AGI-3 | 7.78% | N/A | N/A | Not named (footnote unreachable) | Not stated (footnote unreachable) | Not checked this pass — a footnote marker sits on this row; its text was unreachable in our fetches | Under-specified |
| ExploitBench | 73.5% | Not stated | Not stated | Not named | Not stated | Not checked this pass | Under-specified |
| BrowseComp | 90.4% | N/A | Not stated | Not named | Not stated | Not checked; the identical figure is separately self-reported by Kimi K3 under a different disclosed condition | Under-specified |
| Moonshot AI — Kimi K3 · technical blog · page-level effort disclosure plus per-benchmark footnotes naming which harness evaluated which model | |||||||
| BrowseComp | 90.4 | N/A | Stated: 1M-token context, no context management | Described on the page | Page footnote (max; temp 1.0, top-p 1.0) | Not checked this pass | Vendor-only — the most fully disclosed row in the dataset |
| NVIDIA — Nemotron 3.5 Lightning · Hugging Face model card · harness named explicitly (“NeMo Gym / NeMo Evaluator SDK”); thinking-toggle setting undisclosed | |||||||
| SWE-bench Verified | 51.56 | N/A | “Verified” split | Named | Not stated — a thinking on/off toggle exists; which setting produced this number is unsaid | Not checked this pass | Under-specified |
| Terminal-Bench | 24.58 | 2.1 | N/A | Named | Not stated | Inconclusive — placement not visible in the top rows that loaded for us | Under-specified |
| GDPval-AA | 832 | Not stated | N/A | Named | Not stated | Not checked this pass | Under-specified |
| Mistral AI — Shieldstral 1.0 · announcement and arXiv paper · comparison figures rendered as chart images on both pages | |||||||
| WildGuardTest (F1) | Unresolved — chart-image only | — | — | — | — | Both pages present the comparison as chart images; no machine-readable number could be extracted | Not locatable — our access limitation, not a vendor disclosure failure |
Two dataset notes. The release dates used for DeepSeek V4-Pro-0813 (August 13) and GLM-5.3 (August 14) come from our earlier reporting, not from a date field on the primary pages — the Hugging Face card and the Z.ai developer guide show no publish date. And the rows marked inconclusive against Artificial Analysis pages are inconclusive because of our fetch method: the Terminal-Bench v2.1 page states it lists 30 of 201 models, but only its top three rows loaded for us, and the Intelligence Index page surfaced only Claude Opus 5 rows. Those scores may well be present on the live pages — we simply could not confirm them, and say so rather than counting them as absent.
Verdict distribution · 42 rows under the stated rubric
Source: this audit's 42-row dataset, tallied from the table above04 — Version TrapsSame name, different test.
Three patterns in the dataset show why a benchmark name without its qualifiers is close to meaningless — the first two confirmed at the benchmark maintainer’s own page, the third inferred by comparing two vendors’ footnotes.
Terminal-Bench v2.1 versus v3.0
xAI’s launch page reports Grok 4.6 at 26% on Terminal-Bench v3.0 — a genuinely distinct, newer track confirmed to exist at the maintainer’s leaderboard hub — while Artificial Analysis’s independent v2.1 evaluation separately measured the same model at 88.4% at “high” effort. Read side by side without the version label, those two numbers look like two different products. GLM-5.3 also reports on v3.0 (28.3), roughly in step with Grok’s v3.0 figure, while two of the three v2.1 self-reports in this dataset sit in the high 80s — DeepSeek at 87.9 and GPT-5.6 Sol at 88.8%. The third v2.1 self-report, NVIDIA’s Nemotron 3.5 Lightning at 24.58, sits down with the v3.0 figures instead, which is the limit of the version reading: the label tells a reader which test ran, not on its own how high the number will be. The maintainer’s hub currently lists five distinct tracks — 3.0, 2.1, 2.0, a legacy 1.0, and a Science track marked coming soon — so “Terminal-Bench” with no version number is presently ambiguous among at least four active or near-active tests.
FrontierCode’s nested subsets
The maintainer, Cognition, states the structure plainly: Diamond is the 50 hardest tasks, Main is the 100 hardest (including Diamond), and Extended is the full set of 150. Extended is therefore, by construction, the easiest-skewing subset — it includes the 50 tasks the harder subsets exclude. Grok 4.6’s 61.3% is on Extended; Gemini 3.7 Flash’s 43.6% is on Main. Both pages disclose their subset, which is to their credit — but the two numbers are not the same test, and a comparison table that lines them up as if they were would mislead without either vendor having misstated anything.
An “AutomationBench” name collision
DeepSeek and Google both report a score on “AutomationBench” — apparently the public, GitHub-hosted tool-automation benchmark that Kimi K3’s footnotes describe evaluating on a 600-task public subset. Anthropic’s page separately reports on “Zapier AutomationBench” — by its own description a business-workflow evaluation built with Zapier specifically. These read as two different evaluations sharing a near-identical name, not three vendors citing one benchmark, and the dataset keeps them separate. An all-pass-style metric on a workflow benchmark also compresses scores in a way that can look alarming without its methodology — another reason the name alone carries so little information.
None of this is new in kind — FrontierMath v2 exists because the maintainer corrected errors in a large share of the original problem set, and scores moved when it did. Versioning is healthy benchmark hygiene. The trap is only in citing a versioned benchmark without its version.
05 — Independent ChecksWhere the 15 attempted checks landed.
Independent locatability was attempted on 15 of the 42 rows. Seven reached a conclusion: four exact matches and one near-match (all on DeepSWE v1.1 and Terminal-Bench 2.1), one confirmed absence that is a timing artifact, and one access-blocked row. Eight were inconclusive — an independent source exists and loaded, but our fetch did not surface the specific model’s row. The remaining 27 rows were either explicitly not applicable (vendor-labeled private benchmarks, vendor-run benchmarks for which we found no independent leaderboard, or relative-only claims with no number to check) or genuinely not attempted within this pass’s budget, and are labeled as such.
The strongest positive result in the dataset is DeepSWE v1.1 — the one benchmark enough vendors reported in common to make a like-for-like check possible. The independent DeepSWE leaderboard publishes Pass@1 with an uncertainty band, average cost per task, and step count per model — a richer disclosure format than most of the vendor pages in this dataset — and four separate vendors’ self-reported scores each fell inside its band.
| Model | Vendor’s own page | Independent leaderboard | Reading |
|---|---|---|---|
| DeepSeek V4-Pro-0813 | 62.7 | 63%±6% [max] | Match — inside the stated band |
| Gemini 3.7 Flash | 65.3% | 65%±2% [high] | Match — inside the stated band |
| Grok 4.6 | 65.9% | 67%±2% [xhigh] | Match — inside the stated band |
| GPT-5.6 Sol | 72.7% | 73%±3% [max] | Match — inside the stated band |
| GLM-5.3 | 66.9 | Absent from the snapshot | The leaderboard snapshot predates the model’s release by one day — a timing artifact |
| Claude Opus 5 | Not reported on its launch page | 74%±4% [max] — ranked first | Leaderboard-only: the top model on the independent board never cited the benchmark itself |
Even here, the disclosure gap persists: every vendor page in the match rows omitted at least one field — most often the harness — that the independent leaderboard had to supply. And the one near-match is instructive in the other direction: OpenAI’s page reports Terminal-Bench 2.1 at 88.8% without labeling which effort level produced it, while the independent evaluation lists 89.5% at “xhigh” — a 0.7-point gap that is impossible to interpret precisely because the vendor’s own effort setting is unstated. The FrontierMath row sits in the inconclusive bucket for the same methodological reason: the maintainer’s hub loaded for us without a score table, which is a statement about our fetch, not about the score.
06 — By VendorThe vendor-by-vendor scorecard.
The per-vendor summary below is derived entirely from the dataset table — no new research, so every cell can be recomputed from the rows above. It is a summary of these specific rows, not a general trustworthiness ranking: each vendor is represented by one launch document, selected by recency, and a different document from the same vendor could score differently.
| Vendor (rows) | Harness named | Effort stated | Independent checks attempted → outcome |
|---|---|---|---|
| DeepSeek (6) | 6 of 6 | 6 of 6 | 2 → 1 match, 1 inconclusive |
| Google (5) | 0 of 5 | 0 of 5 | 2 → 1 match, 1 inconclusive |
| xAI (7) | 0 of 7 | 7 of 7 — via the “High” column label | 3 → 1 match, 2 inconclusive |
| Z.ai (6) | 0 of 6 | 1 of 6 | 2 → 1 absent (timing artifact), 1 inconclusive |
| Anthropic (4 — all relative-only claims) | 0 of 4 | 1 of 4 | 0 attempted — no absolute number to check |
| OpenAI (9) | 0 of 9 | 0 of 9 — partial prose disclosure only | 4 → 1 match, 1 near-match, 2 inconclusive |
| Moonshot AI (1) | 1 of 1 | 1 of 1 | 0 attempted |
| NVIDIA (3) | 3 of 3 | 0 of 3 | 1 → 1 inconclusive |
| Mistral AI (1 — unresolved) | — | — | 1 → access-blocked (chart-image only) |
Vendors that specify their rows fully are a real result, and worth naming. DeepSeek disclosed harness and effort for all six of its rows through a single page-level footnote — “the minimal mode of DeepSeek Harness,” at max reasoning effort with temperature 1.0 and top_p 0.95 — and marked its internal test sets with a dagger rather than presenting them like public benchmarks. Kimi K3’s single row is the most fully disclosed in the dataset: page-level effort settings, per-benchmark footnotes naming which of three harnesses evaluated which model, and explicit URLs to the independent sources used for competitor scores. NVIDIA named its harness on every row and its model card shows a comparison model outperforming its own model on 11 of the 14 listed benchmarks — a table that does not favor its own product on most rows shown.
Anthropic’s launch page is the outlier in the other direction, in a specific and unusual way: its four headline evaluations are all relative claims — “more than doubles,” “within 0.5% of,” “three times as high as,” “around 1.5x” — with no absolute number stated for any of them, which is why all four rows land in the not-locatable tier. The counterpoint from the independent side is striking: Claude Opus 5 ranks first on the independent DeepSWE leaderboard at 74%±4%, a benchmark its own launch page never mentions. A vendor under-claiming relative to independent measurement is exactly the kind of finding a disclosure-only reading would miss. And OpenAI’s page — one of the largest benchmark tables in the dataset, with per-cell version labels — states two different Agents’ Last Exam figures on the same page, 53.6 in prose and 52.7% in its table, with no effort label connecting them.
07 — LimitsWhat this table does not show.
A dataset like this invites over-reading, so the boundaries are worth stating as plainly as the findings.
- Not evidence that any number is false. Every vendor-only or under-specified verdict is a disclosure finding, not an accuracy finding. This research re-ran no benchmark.
- Not a reproducibility rate. Independent locatability was attempted on 15 of 42 rows. The other 27 are labeled not-applicable or not-checked — different claims from “not locatable,” kept apart throughout.
- Not a vendor trustworthiness ranking. One launch document per vendor, selected by recency, is not a representative sample of any company’s publication history.
- Not proof that a not-locatable score is unlocatable in general. The Claude Opus 5 ARC-AGI-3 row, for instance, points to a maintainer results page that exists and was not fetched in this pass.
- Not a judgment about which harness or effort setting is correct. The audit asks whether a setting was disclosed, not whether the vendor chose the right one.
Three per-row caveats deserve their own line. Mistral’s WildGuardTest row stays unresolved because both the announcement and the linked arXiv abstract present the comparison as chart images — an access limitation of our text-based fetches, not a vendor disclosure failure, and the row is marked that way. OpenAI’s ARC-AGI-3 row carries a footnote marker whose text was not present in anything our fetches could load — so that row is scored as having an unreachable footnote, not as simply undisclosed. And the BrowseComp coincidence — Kimi K3 and GPT-5.6 Sol each self-report exactly 90.4 on the same benchmark under two different disclosed conditions — is reported as a coincidence a reader auditing these tables would want to find and judge for themselves. This research found no evidence either figure is wrong.
08 — For BuyersWhat to ask for before you buy.
For a team using benchmark tables in model procurement, the dataset reduces to four questions to put to any vendor number — the same four fields the audit scored. Which version of the benchmark, since the dataset shows one benchmark name covering scores from the mid 20s to the high 80s? Which subset, since nested subsets like FrontierCode’s make the full set the easiest one? Which harness, since only three of nine vendors named one anywhere? And which effort setting, since a 0.7-point gap between a vendor number and an independent one cannot be interpreted without it?
The realistic posture is not to distrust vendor tables but to treat them as claims pending qualification — and to prefer benchmarks with an independent leaderboard, where the check costs minutes. Our companion guide covers how to read a leaderboard without being fooled by contamination or cherry-picking, and the earlier CursorBench analysis explains why we found no independent leaderboard to check for that vendor-controlled benchmark — the reason those rows are marked not-applicable here rather than not-checked. This post is one of a pair of audits published together; the open-weight licence audit applies the same census method to model licences. For teams that want this class of verification run against their own shortlist before committing to a model, it is part of what our AI transformation engagements cover.
Looking forward, the incentive structure suggests disclosure will improve unevenly: benchmarks with strong independent leaderboards (DeepSWE v1.1 in this dataset) already show vendor self-reports converging on independently measured values, because divergence is checkable within minutes. Where no independent check exists, nothing external disciplines the footnote. If that holds, the gap this audit measures should close fastest exactly where independent infrastructure exists — which is an argument for funding leaderboards, not for distrusting vendors.
09 — ConclusionA disclosure gap a footnote could close.
The gap is not honesty. It is four missing fields.
Forty-two rows, nine vendors, one rubric: no row in this dataset is simultaneously fully disclosed on the vendor’s own page and independently confirmed. The finding that matters is the disclosure half — checked on every row — and it is fixable at the cost of a footnote: DeepSeek’s single page-level methodology note and Kimi K3’s per-benchmark harness footnotes show the format already exists in production.
The dataset’s positive results deserve equal billing. Four vendors’ DeepSWE self-reports landed inside the independent leaderboard’s uncertainty bands. Two vendors label their private benchmarks as private, plainly. One vendor’s model tops an independent leaderboard its own launch page never cited. Disclosure and accuracy are different axes, and this dataset documents plenty of the latter where it could be checked at all — which was 15 rows out of 42.
That last clause is the honest boundary of this audit, and the reason its headline is about what vendor pages state rather than what third parties can find. The table is in this page in full, with the rubric, so the next person to count can check ours.