BusinessFramework16 min readPublished August 12, 2026

One live launch table · four selection mechanisms · 4 questions before you trust a bolded row

How to Read a Vendor Benchmark Table That Shows a Loss

xAI shipped Grok 4.6 with a ten-row benchmark table in which every publicly checkable cell reproduces its cited source — and the table still steers the conclusion. The steering runs through four levers: which benchmark version, which task subset, what shape the scoring metric takes, and which rivals got a column at all. Here is the reading protocol, built from one live table.

DA
Digital Applied Team
Senior strategists · Published Aug 12, 2026
PublishedAugust 12, 2026
Read time16 min
Sources3 primary pages & boards
Same model, two versions
62.4pts
between the two Terminal-Bench readings
v3.0 26% vs v2.1 88.4% of tasks
Subset choice moves the gap
5.5pts
Fable 5 lead on FrontierCode v1.1 Main
vs ~2.3 on Extended
All-pass compression
2.5%
of Harvey LAB tasks for Sol — from 80.5% of criteria
Boards led by an absent model
7/8
checkable benchmarks, leader has no column

Reading a vendor benchmark table correctly is now a core buying skill, and the Grok 4.6 launch table xAI published on August 12, 2026 is a useful live teaching exhibit. Every cell we could check against a public leaderboard reproduces its cited source exactly — including rows where xAI's own model loses. The table is arithmetically honest. It still steers you.

The steering does not run through fabrication. It runs through four selection mechanisms that operate before a single number is typed: which version of a benchmark gets cited, which task subset gets reported, what shape the scoring metric takes, and which rivals get a column at all. Each mechanism is independently verifiable against a public board, and each one moves the conclusion a buyer walks away with.

This guide works through all four mechanisms using that one table, then compresses them into a four-question protocol you can run on the next vendor table you see — in roughly the time it takes to read the table itself. It pairs with our broader guide to reading AI leaderboards, which covers contamination and eval-gaming; this post is about what a completely honest table can still do.

Key takeaways
  1. 01
    The version decides the number.The same Grok 4.6 model reads 26% of tasks on the Terminal-Bench v3.0 figure xAI cites and 88.4% of tasks on Artificial Analysis' independent Terminal-Bench v2.1 board. Never accept a benchmark score without its version tag.
  2. 02
    The subset decides the gap.FrontierCode's subsets are nested, and xAI reported Extended — the full, easiest pool. On Cognition's harder public Main board, Fable 5's lead over Grok 4.6 widens from roughly 2.3 points to a confirmed 5.5 (53.5% vs 48.0% of tasks).
  3. 03
    The metric's shape decides the drama.Harvey's Legal Agent Benchmark grades all-pass: a task counts only if 100% of its rubric criteria pass. That is how GPT-5.6 Sol's 80.5% criteria pass rate becomes a 2.5%-of-tasks headline — real as measured, not a harness failure.
  4. 04
    The columns decide the winner.Claude Opus 5 leads seven of the eight benchmarks in xAI's table that have an independently checkable board — by 96 Elo, 138 Elo, and 17.5 points on the three we verified directly — and it appears in none of the table's four columns.
  5. 05
    Run the protocol, not the vibes.The table is not fabricated; it even prints seven rows its own model loses. Ask four questions instead of trusting bold: which version, which subset, does the metric zero partial credit, and who is missing a column.

01The ExhibitTen rows, four columns, one very careful footnote.

The exhibit: xAI's Grok 4.6 launch page, published August 12, 2026, carries an evaluation table with exactly four columns — Grok 4.6 (High), Grok 4.5 (High), GPT-5.6 Sol (Max), and Fable 5 (Max), as xAI labels them — across ten benchmark rows. The launch itself (pricing, availability, positioning) is covered in our companion breakdown of the Grok 4.6 launch; here we care only about how the table is built. One labeling note up front: the "(Max)" suffix on the Fable 5 column is xAI's own construction rather than an Anthropic effort tier — Artificial Analysis labels that model "Claude Fable 5 (with fallback)", a routed product.

Before any criticism, the credit. On the benchmarks in that table that have an independent public board — Artificial Analysis' GDPval-AA v2 and AA-Briefcase, plus the rival models' rows on Terminal-Bench v3.0 — every cell we checked reproduces its cited source exactly. Nothing we verified was invented, rounded up, or misquoted. The table also states its own sourcing rule in a footnote, and a caption on the same page describes where competitor figures come from.

The footnote, verbatim
"Best score per evaluation in bold. Third-party model scores are the best of self-reported or publicly available results."A separate chart caption on the same page adds: "Competitor figures are drawn from the respective developers' published system cards or benchmark leaderboards." — xAI, Grok 4.6 launch page, August 12, 2026.

That footnote is why this table is worth studying rather than dismissing. A fabricated table teaches you nothing except not to trust the vendor. An honest one teaches you where the real discretion lives: in the choices made before the numbers — version, subset, metric, columns. Those four choices are the rest of this post.

"A vendor table can be arithmetically honest and still steer the conclusion."— Digital Applied, editorial synthesis

02Mechanism 1 · VersionSame benchmark name, different test: the version lever.

xAI's table reports Grok 4.6 at 26% of tasks on "Terminal-Bench v3.0." That figure is vendor-only: at the time of writing, Grok 4.6 does not appear on the public Terminal-Bench v3.0 leaderboard at frontierbench.ai, which is run by Harbor and the Laude Institute. On that public board, Claude Opus 5 (max) leads at 43.5% of tasks (±1.7), with GPT-5.6 Sol (max) at 34.6% (±1.6) and Claude Fable 5 (max) at 34.1% (±1.7). Opus 5 — the board's leader — is not one of the table's four columns.

Now the version lever. On Terminal-Bench v2.1 — run independently by Artificial Analysis with the Terminus 2 harness in an e2b sandbox, pass@1 averaged over three repeats across 89 tasks — the same Grok 4.6 model scores 88.4% of tasks at high effort, within about a point of that board's leaders (GPT-5.6 Sol at xhigh, 89.5%, and Claude Opus 5 at max, 89.1%). Same model, same benchmark family name, a 62.4-point spread between the two readings — arithmetic we derived from the two published figures, not a number either board prints.

The two versions are genuinely different test sets, not a version bump on one continuous scale: tbench.ai, the benchmark's own home, still defaults to v2.1 at the time of writing, while v3.0 is the newer, separate task domain hosted at frontierbench.ai. Citing "Terminal-Bench" without a version tag is like citing "the SAT" without saying which century's scoring scale you mean.

Terminal-Bench scores for the same models on two different benchmark versions: the public v3.0 board at frontierbench.ai with its per-row agent harness disclosures, and Artificial Analysis' independent v2.1 board. Grok 4.6's v3.0 figure is vendor-stated only and absent from the public board at the time of writing.
Model (effort shown per board)v3.0 public board · % of tasksv3.0 harness (disclosed per row)v2.1, AA's board · % of tasks
Claude Opus 543.5% ± 1.7 (#1, max)mini-SWE-agent89.1% (max)
GPT-5.6 Sol34.6% ± 1.6 (#2, max)Codex89.5% (xhigh, #1)
Claude Fable 534.1% ± 1.7 (#3, max)Claude Codenot among the rows we read
Grok 4.515.7% ± 1.5 (#6, xhigh)Cursor CLInot among the rows we read
Grok 4.626% — xAI's table only; absent from the public board at the time of writingnot disclosed by xAI88.4% (high)

One more disclosure gap hides in that table's third column. Every row on the public v3.0 board states which agent harness produced the score — and they are all different: mini-SWE-agent for Opus 5, Codex for Sol, Claude Code for Fable 5, Cursor CLI for Grok 4.5. Harness choice materially shapes agentic benchmark results, which is exactly why the board discloses it per row. xAI discloses no harness at all for its own 26%-of-tasks figure. A score with no version tag and no harness disclosure is not wrong — it is simply unverifiable, and unverifiable is a property a buyer should price in.

03Mechanism 2 · SubsetNested subsets, and the gap that widens with difficulty.

FrontierCode is Cognition's coding benchmark, built by more than twenty named open-source maintainers, with each task graded by an ensemble of unit tests, rubrics, and verifiers. Its subsets are nested: Extended is the full task pool, Main is a harder subset drawn from inside it, and Diamond is the hardest tier — progressively narrower, progressively tougher.

How FrontierCode grades
"FrontierCode is the first benchmark to measure mergeability: would the maintainer actually merge this PR?" Cognition, FrontierCode methodology page. Runs that consult solution-bearing sources such as the original pull request are detected and scored zero.

xAI's table reports Grok 4.6 at 61.3% of tasks on "FrontierCode v1.1 (Extended)" — the full pool, which is also the easiest reading. On xAI's own Extended row, Fable 5 leads Grok 4.6 by roughly 2.3 points. But the subset Cognition publicly ranks is Main, the harder one — and there the same pairing reads Fable 5 (xhigh) at 53.5% of Main-subset tasks against Grok 4.6 (high) at 48.0%: a 5.5-point gap, more than double the Extended reading. Choosing the easier subset did not change who wins. It changed by how much — which, for a buyer comparing near-peers, is the entire question.

FrontierCode v1.1 Main · share of Main-subset tasks passed

Source: cognition.com/frontiercode, FrontierCode v1.1 Main public leaderboard (selected rows), read at the time of writing
Fable 5 (xhigh)#1 · $13.09 per rollout · 58.6k tokens
53.5%
Opus 5 (medium)#2 · $4.31 per rollout · 33.6k tokens
53.4%
Grok 4.6 (high)#3 · $2.88 per rollout · 36.8k tokens
48.0%
GPT-5.6 Sol (max)#4 · $6.29 per rollout · 33.2k tokens
47.5%
Opus 4.8 (max)#5 · $9.62 per rollout · 95.9k tokens
46.5%
Grok 4.5 (high)#9 · $1.30 per rollout · 15.3k tokens
42.4%

Two honest observations cut the other way. First, the score gaps are not an artifact of disqualified runs — the flag rate (the share of runs zeroed for consulting solution-bearing sources) sits between 0.0% and 0.7% of runs for the top rows shown. Second, the Main board publishes cost per rollout, and Grok 4.6 is markedly the cheapest per attempt among the top four — $2.88 per rollout against $13.09 for Fable 5 at xhigh and $6.29 for Sol at max. That is a genuine value story the table format alone never surfaces, and it is a better argument for Grok 4.6 than the subset choice was.

04Mechanism 3 · Metric ShapeAll-pass grading, where partial credit rounds to zero.

The strangest number in xAI's table is not one of its own. On Harvey's Legal Agent Benchmark (LAB), the table shows GPT-5.6 Sol at 2.5% — a score so low that most readers assume either a broken harness or a uselessly bad model. It is neither. It is what an 80.5% rubric-criteria pass rate looks like after the metric's shape gets done with it.

The background: Harvey introduced LAB in May 2026 as a set of partner-to-associate legal work assignments — vendor-stated at launch as "more than 1,200 agent tasks across 24 legal practice areas," graded against "over 75,000 expert-written rubric criteria." Harvey said at launch, "We're intentionally launching LAB without a leaderboard because we expect the dataset to evolve over time and we want to work with the community to ensure results are clear and intuitive in how they convey agent performance." The public leaderboard now live at Vals AI is that later step, and Vals documents the grading rule plainly: two LLM judges (GPT-5.5 at medium reasoning, Claude Sonnet 4.6 unmodified) each score every rubric criterion, and "a task passes only if 100% of its criteria pass." The final score averages the two judges' task pass rates.

That rule is the whole story. On Vals' board, Sol satisfies 80.5% of individual rubric criteria — roughly four-fifths of every fact, citation, conclusion, and formatting requirement across the set — but completes only 2.5% of tasks end to end with zero misses. Both numbers sit on the same leaderboard row, independently measured.

Vals' own explanation, verbatim
"The bottleneck is finishing, not the individual steps: top models satisfy roughly 85–95% of the individual criteria (Muse Spark 1.2 hits 94.52%), but a task scores only when every criterion passes — so one or two misses per task wipe out most of the credit."And on the ceiling: "Finishing one of these multi-step legal tasks end to end remains rare: the best agent, Muse Spark 1.2, manages it on 25.42% of tasks, and just two models clear 15% at all." — Vals AI, Harvey LAB leaderboard, updated August 10, 2026.
Harvey Legal Agent Benchmark results from Vals AI's public leaderboard, showing each model's share of individual rubric criteria passed against its share of tasks passed under all-pass grading, plus Grok 4.6's vendor-stated figure which is absent from the public board at the time of writing.
ModelShare of rubric criteria passedShare of tasks passed (all-pass)On Vals' public board?
Independently measured — Vals' public leaderboard
Muse Spark 1.294.52%25.42% (#1)Yes
Muse Spark 1.192.86%20.00% (#2)Yes
Grok 4.590.55%12.92% (#3)Yes
Claude Fable 590.48%11.25% (#4)Yes — 10.42% if 4 fallback tasks count as failures
GPT-5.6 Sol80.5%2.5%Yes
Vendor-stated — not on the public board
Grok 4.6not published15.8% — xAI's table onlyNo — absent at the time of writing

Three readings follow, and each matters to a buyer. First, the 2.5% is real as measured, not a harness failure. Vals went looking for a grading artifact that would explain the low scores, found one genuine but unrelated bug — DOCX tracked-changes were not being preserved when judges read submitted redline files, fixed upstream as pull #76 on the open harveyai/harvey-labs repo — and the all-pass compression persisted with the rubric unchanged. Two independent implementations show the same signature: Vals' own harness, and Artificial Analysis' separate "Harvey LAB-AA" run, which uses AA's Stirrup harness with Gemini 3.1 Pro as the grading judge on what AA describes as 120 private tasks spanning 24 legal practice areas. A score that reproduces across two independent pipelines is a measurement, not a glitch.

Second, it does not mean a five-fold capability gap. Sol passes 80.5% of individual criteria against Grok 4.5's 90.55% — a real difference, but nothing like the 2.5%-vs-12.92% task-level reading suggests. All-pass grading amplifies small per-step differences into dramatic headline ratios. And Sol is no legal incompetent: on Vals' separate Legal Research benchmark, the same model places third at the time of writing. Sample size amplifies the drama further: at the 120-task count AA describes for its implementation, one task flipping moves the all-pass score by roughly 0.83 percentage points — so a 2.5% score is about three completed tasks, and 15.8% is about nineteen.

Third, the bolded cell is not a crown. xAI's table shows Grok 4.6 at 15.8% of tasks as the best score among its four columns — a vendor-stated figure, since Grok 4.6 is absent from Vals' public board at the time of writing. But slot that 15.8 into the public board and it would rank third, behind both Muse Spark 1.2 (25.42% of tasks) and Muse Spark 1.1 (20.00%). Nothing in xAI's table is false. The frame just ends where the inconvenient models begin — the same caution we documented when another vendor's strongest rival row went unprinted.

05Mechanism 4 · ColumnsThe strongest rival never gets a column.

Version, subset, and metric shape each bend one row. Column selection bends the whole table at once, and it is the quietest mechanism of the four because nothing on the page is wrong — a model is simply not there. In xAI's table the missing model is Claude Opus 5. Across the three benchmarks where we directly confirmed a full independent board — GDPval-AA v2, AA-Briefcase, and Terminal-Bench v3.0 — Opus 5's published score beats the figure xAI's table cites for Grok 4.6 on all three, decisively.

GDPval-AA v2
Margin over the bolded 1753
96Elo

Claude Opus 5 (Adaptive Reasoning, Max Effort) leads AA's board at 1849 Elo (CI ±22) against Grok 4.6's 1753 (±21) — outside the overlapping confidence intervals. The xhigh setting also sits ahead at 1817.

artificialanalysis.ai/evaluations/gdpval-aa
AA-Briefcase
Margin over Grok 4.6's 1577
138Elo

Claude Opus 5 (max) leads at 1715 Elo against Grok 4.6's 1577 — well outside Grok 4.6's confidence interval. Grok 4.5 sits at 1313 on the same board, so the generational gain here is real and independently measured.

artificialanalysis.ai/evaluations/aa-briefcase
Terminal-Bench v3.0
Margin over the cited 26% of tasks
17.5pts

Claude Opus 5 (max) is the public board's #1 model at 43.5% of tasks — 17.5 points above the vendor-only figure xAI's table attributes to Grok 4.6, which is absent from that board at the time of writing.

frontierbench.ai

A scale note so the Elo figures land correctly: GDPval-AA v2's rating is a genuine Elo, computed from blind pairwise comparisons on real-world professional tasks across 44 occupations and 9 industries, anchored to a human-expert baseline of 1,000 — not a chatbot-arena popularity score. A 96-point gap on that scale is a real, repeated preference for one model's work product over the other's.

Beyond the three boards we verified directly, the published boards for the remaining checkable rows in xAI's table — Datacurve's DeepSWE v1.1 and the Mercor-and-Cognition APEX evaluations — show the same one-sided pattern, with Opus 5's published figures likewise above Grok 4.6's cited numbers. Net: of the eight benchmarks in xAI's table that have an independently checkable board, Claude Opus 5 leads seven — and appears in none of the four columns. Even on the composite Artificial Analysis Intelligence Index that xAI's own launch prose leans on — "Grok 4.6 achieves frontier intelligence across several agentic coding and knowledge work benchmarks. It matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, which is a composite score of nine benchmarks" (the nine-benchmark composition is xAI's characterization) — the tie it claims is not with the index leader: on AA's published index at the time of writing, both Claude Opus 5 and Claude Fable 5 sit above Grok 4.6 on that same index.

"Best score per evaluation in bold" is a true statement about a field of four. The reading protocol response is mechanical: for every benchmark a vendor cites, pull up that benchmark's own public board and check whether the top two or three models appear as columns. When the board's leader is missing from the table, the bold means "best among the models we chose to show you" — which is a different claim from the one your eye reads.

06The Fair ReadWhat the table gets right — and what sits inside noise.

A reading protocol that only finds sins is a hit piece, not a protocol. So score the table both ways. In xAI's favor: it prints seven rows in which its own model loses to a rival's printed score — losses disclosed on the vendor's own launch page, which is precisely the behavior buyers should reward. On CursorBench, xAI under-reported itself, quoting its own High-effort figure where a higher xhigh figure existed (we unpack that benchmark's vendor-controlled quirks in our CursorBench deep dive). And as the vendors' documentation reads to us, the table pairs Grok 4.6's own default effort tier against rival models' top documented settings — a framing that, if anything, undersells xAI's price-adjusted showing rather than inflating it. The effort ladders themselves differ per vendor and are worth understanding before comparing any two "max" labels; see our companion piece on the effort tiers behind these columns.

The cost story is also genuine. On FrontierCode Main, Grok 4.6 completed 48.0% of Main-subset tasks at $2.88 per rollout — against $13.09 for the board leader. The generation-over-generation gains are real on independently measured boards, too: Grok 4.5 to Grok 4.6 moves from 42.4% to 48.0% of Main-subset tasks on FrontierCode and from 1313 to 1577 Elo on AA-Briefcase. And the list pricing is aggressive for the tier.

The pricing line, verbatim
"Pricing starts at $2 per million input tokens and $6 per million output tokens. Additionally, there is a fast variant which is twice the price." xAI, Grok 4.6 launch page. The page names no separate model ID for the fast variant.

Now the other side of the fair read: the wins are the fragile part. Using Artificial Analysis' own published confidence intervals, none of the three benchmarks where xAI's table bolds Grok 4.6 as best — GDPval-AA v2, AA-Briefcase, Harvey LAB — shows a decisive, outside-the-noise lead. On GDPval-AA v2, Grok 4.6's 1753 (±21) overlaps both Fable 5's 1741 (±16) and Sol's 1728 (±16). On AA-Briefcase, Grok 4.6's 1577 (±11) overlaps Fable 5's 1574 (±11) almost entirely. And the Harvey LAB "win" rests on a vendor-only figure, absent from the public board, on a benchmark where — at the 120-task sample AA describes — the whole gap is a handful-of-tasks difference. The disclosed losses are solid; the bolded wins are statistical ties or unverifiable. That asymmetry, not any single number, is the table's real finding.

07Buyer ProtocolFour questions, two minutes each.

Everything above compresses into a protocol you can run on any vendor table before a purchase decision — no benchmark expertise required, just the benchmark's own public board and a browser tab. The worked column shows what each question surfaced on the Grok 4.6 table.

The four-question protocol for reading any vendor benchmark table, with the check to run in about two minutes for each question and the answer each question produced on xAI's Grok 4.6 launch table as a worked example.
The questionThe two-minute checkThis table, answered
1. Which version?Find the benchmark's own site and note its current default version, then search that version's public board for the model. A version tag missing from the vendor's citation is the tell.The cited "Terminal-Bench v3.0" figure (26% of tasks) is vendor-only, with no disclosed harness; the same model reads 88.4% of tasks on the independent v2.1 board.
2. Which subset?Look for subset names — Extended, Main, Diamond, "full," "verified," "hard" — and find which subset the benchmark's public leaderboard actually ranks.xAI cited Extended, the full and easiest pool; on the public Main board the rival's lead widens from roughly 2.3 points to 5.5 (53.5% vs 48.0% of tasks).
3. Does the metric zero partial credit?Read the grading rule on the leaderboard's methodology note. All-pass metrics compress high per-step rates into tiny headline scores; small task counts amplify every flip.Harvey LAB is all-pass: 80.5% of criteria becomes 2.5% of tasks. At the 120-task sample AA describes, one task moves a score by roughly 0.83 points.
4. Who's missing a column?Pull each cited benchmark's public board and check whether its top two or three models appear as columns in the vendor's table.Claude Opus 5 leads seven of the eight independently checkable benchmarks in the table — by 96 Elo, 138 Elo, and 17.5 points on the three we verified directly — and has no column.

The trend behind the protocol is worth naming. As agentic benchmarks replace single-answer quizzes, the industry is moving toward exactly the structures that make these four levers more powerful: versioned test sets that fork rather than increment, nested difficulty subsets, all-pass task grading, and per-harness scores that resist casual comparison. That is not vendors getting more devious — it is evaluation getting more realistic, with more degrees of freedom as a side effect. The same discipline applies well beyond coding models: we found the identical reading problems in voice AI vendor claims and in hallucination-rate benchmarks, where metric shape does most of the steering.

Looking forward, we expect the gap between vendor tables and public boards to keep widening through 2027 — launch-day tables will keep citing whichever version, subset, and column set reads best, while the independent boards (Artificial Analysis, frontierbench.ai, Vals, Cognition's public leaderboards) keep professionalizing with confidence intervals, per-row harness disclosure, and open grading code. The buyers who thrive are the ones who treat the vendor table as a list of claims to check rather than a verdict to accept. If your team is making a model decision on benchmark evidence and wants the comparative eval run on your own workload instead of a vendor's chosen subset, that is exactly what our AI transformation engagements start with.

08ConclusionRead the table's choices, not its arithmetic.

The reading protocol

Honest numbers, arranged: check the version, the subset, the metric, and the missing column.

The Grok 4.6 launch table is the rare exhibit that rewards close reading precisely because it is honest. Every publicly checkable cell reproduces its source; seven rows show the vendor's own model losing; and on one benchmark the vendor even quoted a lower figure for itself than it needed to. If this table misleads anyone, it will not be through a false number — it will be through a chosen one.

The four choices are now yours to check: a version whose independent board the model does not appear on, a subset easier than the one publicly ranked, a metric whose all-pass shape turns an 80.5% criteria rate into a 2.5%-of-tasks headline, and a column set that omits the model leading seven of the eight checkable boards. None of those checks needs more than the benchmark's own public leaderboard and about two minutes.

The general lesson outlives this table. Vendor benchmarks are becoming more honest in their cells and more editorial in their frames — and frames do not show up in a fact-check. Run the four questions on the next launch table you see, whatever the vendor and whichever direction the bold points. The numbers will usually be true. The arrangement is where the argument lives.

Benchmark claims, independently checked

Make the model decision on your workload, not the vendor's table.

We run comparative model evaluations on your actual workload — not a vendor's chosen subset — and turn the results into a defensible model-selection decision, delivered in days not quarters.

Free consultationExpert guidanceTailored solutions
What we work on

Model evaluation engagements

  • Comparative evals on your own tasks and repos
  • Vendor benchmark due diligence before contracts
  • Model routing by task class, cost per outcome
  • Eval harness design with all-pass and partial credit
  • Quarterly re-benchmarking as versions fork
FAQ · Reading vendor tables

The questions buyers ask about benchmark tables.

Because the two figures come from two genuinely different test sets that share a name. Terminal-Bench v2.1 — still the default on tbench.ai, the benchmark's own home — is run independently by Artificial Analysis using the Terminus 2 harness across 89 tasks, and Grok 4.6 scores 88.4% of tasks there at high effort. Terminal-Bench v3.0 is a newer, separate task domain hosted on a public board at frontierbench.ai, where the leading score at the time of writing is Claude Opus 5's 43.5% of tasks. xAI's 26%-of-tasks figure for Grok 4.6 cites v3.0, is vendor-stated only, discloses no agent harness, and does not appear on the public v3.0 board at the time of writing. The lesson: never accept a benchmark score without its version tag, and check whether the model actually appears on that version's independent leaderboard.
Related dispatches

Continue exploring vendor claims and evals.