Reading a vendor benchmark table correctly is now a core buying skill, and the Grok 4.6 launch table xAI published on August 12, 2026 is a useful live teaching exhibit. Every cell we could check against a public leaderboard reproduces its cited source exactly — including rows where xAI's own model loses. The table is arithmetically honest. It still steers you.
The steering does not run through fabrication. It runs through four selection mechanisms that operate before a single number is typed: which version of a benchmark gets cited, which task subset gets reported, what shape the scoring metric takes, and which rivals get a column at all. Each mechanism is independently verifiable against a public board, and each one moves the conclusion a buyer walks away with.
This guide works through all four mechanisms using that one table, then compresses them into a four-question protocol you can run on the next vendor table you see — in roughly the time it takes to read the table itself. It pairs with our broader guide to reading AI leaderboards, which covers contamination and eval-gaming; this post is about what a completely honest table can still do.
- 01The version decides the number.The same Grok 4.6 model reads 26% of tasks on the Terminal-Bench v3.0 figure xAI cites and 88.4% of tasks on Artificial Analysis' independent Terminal-Bench v2.1 board. Never accept a benchmark score without its version tag.
- 02The subset decides the gap.FrontierCode's subsets are nested, and xAI reported Extended — the full, easiest pool. On Cognition's harder public Main board, Fable 5's lead over Grok 4.6 widens from roughly 2.3 points to a confirmed 5.5 (53.5% vs 48.0% of tasks).
- 03The metric's shape decides the drama.Harvey's Legal Agent Benchmark grades all-pass: a task counts only if 100% of its rubric criteria pass. That is how GPT-5.6 Sol's 80.5% criteria pass rate becomes a 2.5%-of-tasks headline — real as measured, not a harness failure.
- 04The columns decide the winner.Claude Opus 5 leads seven of the eight benchmarks in xAI's table that have an independently checkable board — by 96 Elo, 138 Elo, and 17.5 points on the three we verified directly — and it appears in none of the table's four columns.
- 05Run the protocol, not the vibes.The table is not fabricated; it even prints seven rows its own model loses. Ask four questions instead of trusting bold: which version, which subset, does the metric zero partial credit, and who is missing a column.
01 — The ExhibitTen rows, four columns, one very careful footnote.
The exhibit: xAI's Grok 4.6 launch page, published August 12, 2026, carries an evaluation table with exactly four columns — Grok 4.6 (High), Grok 4.5 (High), GPT-5.6 Sol (Max), and Fable 5 (Max), as xAI labels them — across ten benchmark rows. The launch itself (pricing, availability, positioning) is covered in our companion breakdown of the Grok 4.6 launch; here we care only about how the table is built. One labeling note up front: the "(Max)" suffix on the Fable 5 column is xAI's own construction rather than an Anthropic effort tier — Artificial Analysis labels that model "Claude Fable 5 (with fallback)", a routed product.
Before any criticism, the credit. On the benchmarks in that table that have an independent public board — Artificial Analysis' GDPval-AA v2 and AA-Briefcase, plus the rival models' rows on Terminal-Bench v3.0 — every cell we checked reproduces its cited source exactly. Nothing we verified was invented, rounded up, or misquoted. The table also states its own sourcing rule in a footnote, and a caption on the same page describes where competitor figures come from.
That footnote is why this table is worth studying rather than dismissing. A fabricated table teaches you nothing except not to trust the vendor. An honest one teaches you where the real discretion lives: in the choices made before the numbers — version, subset, metric, columns. Those four choices are the rest of this post.
"A vendor table can be arithmetically honest and still steer the conclusion."— Digital Applied, editorial synthesis
02 — Mechanism 1 · VersionSame benchmark name, different test: the version lever.
xAI's table reports Grok 4.6 at 26% of tasks on "Terminal-Bench v3.0." That figure is vendor-only: at the time of writing, Grok 4.6 does not appear on the public Terminal-Bench v3.0 leaderboard at frontierbench.ai, which is run by Harbor and the Laude Institute. On that public board, Claude Opus 5 (max) leads at 43.5% of tasks (±1.7), with GPT-5.6 Sol (max) at 34.6% (±1.6) and Claude Fable 5 (max) at 34.1% (±1.7). Opus 5 — the board's leader — is not one of the table's four columns.
Now the version lever. On Terminal-Bench v2.1 — run independently by Artificial Analysis with the Terminus 2 harness in an e2b sandbox, pass@1 averaged over three repeats across 89 tasks — the same Grok 4.6 model scores 88.4% of tasks at high effort, within about a point of that board's leaders (GPT-5.6 Sol at xhigh, 89.5%, and Claude Opus 5 at max, 89.1%). Same model, same benchmark family name, a 62.4-point spread between the two readings — arithmetic we derived from the two published figures, not a number either board prints.
The two versions are genuinely different test sets, not a version bump on one continuous scale: tbench.ai, the benchmark's own home, still defaults to v2.1 at the time of writing, while v3.0 is the newer, separate task domain hosted at frontierbench.ai. Citing "Terminal-Bench" without a version tag is like citing "the SAT" without saying which century's scoring scale you mean.
| Model (effort shown per board) | v3.0 public board · % of tasks | v3.0 harness (disclosed per row) | v2.1, AA's board · % of tasks |
|---|---|---|---|
| Claude Opus 5 | 43.5% ± 1.7 (#1, max) | mini-SWE-agent | 89.1% (max) |
| GPT-5.6 Sol | 34.6% ± 1.6 (#2, max) | Codex | 89.5% (xhigh, #1) |
| Claude Fable 5 | 34.1% ± 1.7 (#3, max) | Claude Code | not among the rows we read |
| Grok 4.5 | 15.7% ± 1.5 (#6, xhigh) | Cursor CLI | not among the rows we read |
| Grok 4.6 | 26% — xAI's table only; absent from the public board at the time of writing | not disclosed by xAI | 88.4% (high) |
One more disclosure gap hides in that table's third column. Every row on the public v3.0 board states which agent harness produced the score — and they are all different: mini-SWE-agent for Opus 5, Codex for Sol, Claude Code for Fable 5, Cursor CLI for Grok 4.5. Harness choice materially shapes agentic benchmark results, which is exactly why the board discloses it per row. xAI discloses no harness at all for its own 26%-of-tasks figure. A score with no version tag and no harness disclosure is not wrong — it is simply unverifiable, and unverifiable is a property a buyer should price in.
03 — Mechanism 2 · SubsetNested subsets, and the gap that widens with difficulty.
FrontierCode is Cognition's coding benchmark, built by more than twenty named open-source maintainers, with each task graded by an ensemble of unit tests, rubrics, and verifiers. Its subsets are nested: Extended is the full task pool, Main is a harder subset drawn from inside it, and Diamond is the hardest tier — progressively narrower, progressively tougher.
xAI's table reports Grok 4.6 at 61.3% of tasks on "FrontierCode v1.1 (Extended)" — the full pool, which is also the easiest reading. On xAI's own Extended row, Fable 5 leads Grok 4.6 by roughly 2.3 points. But the subset Cognition publicly ranks is Main, the harder one — and there the same pairing reads Fable 5 (xhigh) at 53.5% of Main-subset tasks against Grok 4.6 (high) at 48.0%: a 5.5-point gap, more than double the Extended reading. Choosing the easier subset did not change who wins. It changed by how much — which, for a buyer comparing near-peers, is the entire question.
FrontierCode v1.1 Main · share of Main-subset tasks passed
Source: cognition.com/frontiercode, FrontierCode v1.1 Main public leaderboard (selected rows), read at the time of writingTwo honest observations cut the other way. First, the score gaps are not an artifact of disqualified runs — the flag rate (the share of runs zeroed for consulting solution-bearing sources) sits between 0.0% and 0.7% of runs for the top rows shown. Second, the Main board publishes cost per rollout, and Grok 4.6 is markedly the cheapest per attempt among the top four — $2.88 per rollout against $13.09 for Fable 5 at xhigh and $6.29 for Sol at max. That is a genuine value story the table format alone never surfaces, and it is a better argument for Grok 4.6 than the subset choice was.
04 — Mechanism 3 · Metric ShapeAll-pass grading, where partial credit rounds to zero.
The strangest number in xAI's table is not one of its own. On Harvey's Legal Agent Benchmark (LAB), the table shows GPT-5.6 Sol at 2.5% — a score so low that most readers assume either a broken harness or a uselessly bad model. It is neither. It is what an 80.5% rubric-criteria pass rate looks like after the metric's shape gets done with it.
The background: Harvey introduced LAB in May 2026 as a set of partner-to-associate legal work assignments — vendor-stated at launch as "more than 1,200 agent tasks across 24 legal practice areas," graded against "over 75,000 expert-written rubric criteria." Harvey said at launch, "We're intentionally launching LAB without a leaderboard because we expect the dataset to evolve over time and we want to work with the community to ensure results are clear and intuitive in how they convey agent performance." The public leaderboard now live at Vals AI is that later step, and Vals documents the grading rule plainly: two LLM judges (GPT-5.5 at medium reasoning, Claude Sonnet 4.6 unmodified) each score every rubric criterion, and "a task passes only if 100% of its criteria pass." The final score averages the two judges' task pass rates.
That rule is the whole story. On Vals' board, Sol satisfies 80.5% of individual rubric criteria — roughly four-fifths of every fact, citation, conclusion, and formatting requirement across the set — but completes only 2.5% of tasks end to end with zero misses. Both numbers sit on the same leaderboard row, independently measured.
| Model | Share of rubric criteria passed | Share of tasks passed (all-pass) | On Vals' public board? |
|---|---|---|---|
| Independently measured — Vals' public leaderboard | |||
| Muse Spark 1.2 | 94.52% | 25.42% (#1) | Yes |
| Muse Spark 1.1 | 92.86% | 20.00% (#2) | Yes |
| Grok 4.5 | 90.55% | 12.92% (#3) | Yes |
| Claude Fable 5 | 90.48% | 11.25% (#4) | Yes — 10.42% if 4 fallback tasks count as failures |
| GPT-5.6 Sol | 80.5% | 2.5% | Yes |
| Vendor-stated — not on the public board | |||
| Grok 4.6 | not published | 15.8% — xAI's table only | No — absent at the time of writing |
Three readings follow, and each matters to a buyer. First, the 2.5% is real as measured, not a harness failure. Vals went looking for a grading artifact that would explain the low scores, found one genuine but unrelated bug — DOCX tracked-changes were not being preserved when judges read submitted redline files, fixed upstream as pull #76 on the open harveyai/harvey-labs repo — and the all-pass compression persisted with the rubric unchanged. Two independent implementations show the same signature: Vals' own harness, and Artificial Analysis' separate "Harvey LAB-AA" run, which uses AA's Stirrup harness with Gemini 3.1 Pro as the grading judge on what AA describes as 120 private tasks spanning 24 legal practice areas. A score that reproduces across two independent pipelines is a measurement, not a glitch.
Second, it does not mean a five-fold capability gap. Sol passes 80.5% of individual criteria against Grok 4.5's 90.55% — a real difference, but nothing like the 2.5%-vs-12.92% task-level reading suggests. All-pass grading amplifies small per-step differences into dramatic headline ratios. And Sol is no legal incompetent: on Vals' separate Legal Research benchmark, the same model places third at the time of writing. Sample size amplifies the drama further: at the 120-task count AA describes for its implementation, one task flipping moves the all-pass score by roughly 0.83 percentage points — so a 2.5% score is about three completed tasks, and 15.8% is about nineteen.
Third, the bolded cell is not a crown. xAI's table shows Grok 4.6 at 15.8% of tasks as the best score among its four columns — a vendor-stated figure, since Grok 4.6 is absent from Vals' public board at the time of writing. But slot that 15.8 into the public board and it would rank third, behind both Muse Spark 1.2 (25.42% of tasks) and Muse Spark 1.1 (20.00%). Nothing in xAI's table is false. The frame just ends where the inconvenient models begin — the same caution we documented when another vendor's strongest rival row went unprinted.
05 — Mechanism 4 · ColumnsThe strongest rival never gets a column.
Version, subset, and metric shape each bend one row. Column selection bends the whole table at once, and it is the quietest mechanism of the four because nothing on the page is wrong — a model is simply not there. In xAI's table the missing model is Claude Opus 5. Across the three benchmarks where we directly confirmed a full independent board — GDPval-AA v2, AA-Briefcase, and Terminal-Bench v3.0 — Opus 5's published score beats the figure xAI's table cites for Grok 4.6 on all three, decisively.
Margin over the bolded 1753
Claude Opus 5 (Adaptive Reasoning, Max Effort) leads AA's board at 1849 Elo (CI ±22) against Grok 4.6's 1753 (±21) — outside the overlapping confidence intervals. The xhigh setting also sits ahead at 1817.
Margin over Grok 4.6's 1577
Claude Opus 5 (max) leads at 1715 Elo against Grok 4.6's 1577 — well outside Grok 4.6's confidence interval. Grok 4.5 sits at 1313 on the same board, so the generational gain here is real and independently measured.
Margin over the cited 26% of tasks
Claude Opus 5 (max) is the public board's #1 model at 43.5% of tasks — 17.5 points above the vendor-only figure xAI's table attributes to Grok 4.6, which is absent from that board at the time of writing.
A scale note so the Elo figures land correctly: GDPval-AA v2's rating is a genuine Elo, computed from blind pairwise comparisons on real-world professional tasks across 44 occupations and 9 industries, anchored to a human-expert baseline of 1,000 — not a chatbot-arena popularity score. A 96-point gap on that scale is a real, repeated preference for one model's work product over the other's.
Beyond the three boards we verified directly, the published boards for the remaining checkable rows in xAI's table — Datacurve's DeepSWE v1.1 and the Mercor-and-Cognition APEX evaluations — show the same one-sided pattern, with Opus 5's published figures likewise above Grok 4.6's cited numbers. Net: of the eight benchmarks in xAI's table that have an independently checkable board, Claude Opus 5 leads seven — and appears in none of the four columns. Even on the composite Artificial Analysis Intelligence Index that xAI's own launch prose leans on — "Grok 4.6 achieves frontier intelligence across several agentic coding and knowledge work benchmarks. It matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, which is a composite score of nine benchmarks" (the nine-benchmark composition is xAI's characterization) — the tie it claims is not with the index leader: on AA's published index at the time of writing, both Claude Opus 5 and Claude Fable 5 sit above Grok 4.6 on that same index.
"Best score per evaluation in bold" is a true statement about a field of four. The reading protocol response is mechanical: for every benchmark a vendor cites, pull up that benchmark's own public board and check whether the top two or three models appear as columns. When the board's leader is missing from the table, the bold means "best among the models we chose to show you" — which is a different claim from the one your eye reads.
06 — The Fair ReadWhat the table gets right — and what sits inside noise.
A reading protocol that only finds sins is a hit piece, not a protocol. So score the table both ways. In xAI's favor: it prints seven rows in which its own model loses to a rival's printed score — losses disclosed on the vendor's own launch page, which is precisely the behavior buyers should reward. On CursorBench, xAI under-reported itself, quoting its own High-effort figure where a higher xhigh figure existed (we unpack that benchmark's vendor-controlled quirks in our CursorBench deep dive). And as the vendors' documentation reads to us, the table pairs Grok 4.6's own default effort tier against rival models' top documented settings — a framing that, if anything, undersells xAI's price-adjusted showing rather than inflating it. The effort ladders themselves differ per vendor and are worth understanding before comparing any two "max" labels; see our companion piece on the effort tiers behind these columns.
The cost story is also genuine. On FrontierCode Main, Grok 4.6 completed 48.0% of Main-subset tasks at $2.88 per rollout — against $13.09 for the board leader. The generation-over-generation gains are real on independently measured boards, too: Grok 4.5 to Grok 4.6 moves from 42.4% to 48.0% of Main-subset tasks on FrontierCode and from 1313 to 1577 Elo on AA-Briefcase. And the list pricing is aggressive for the tier.
Now the other side of the fair read: the wins are the fragile part. Using Artificial Analysis' own published confidence intervals, none of the three benchmarks where xAI's table bolds Grok 4.6 as best — GDPval-AA v2, AA-Briefcase, Harvey LAB — shows a decisive, outside-the-noise lead. On GDPval-AA v2, Grok 4.6's 1753 (±21) overlaps both Fable 5's 1741 (±16) and Sol's 1728 (±16). On AA-Briefcase, Grok 4.6's 1577 (±11) overlaps Fable 5's 1574 (±11) almost entirely. And the Harvey LAB "win" rests on a vendor-only figure, absent from the public board, on a benchmark where — at the 120-task sample AA describes — the whole gap is a handful-of-tasks difference. The disclosed losses are solid; the bolded wins are statistical ties or unverifiable. That asymmetry, not any single number, is the table's real finding.
07 — Buyer ProtocolFour questions, two minutes each.
Everything above compresses into a protocol you can run on any vendor table before a purchase decision — no benchmark expertise required, just the benchmark's own public board and a browser tab. The worked column shows what each question surfaced on the Grok 4.6 table.
| The question | The two-minute check | This table, answered |
|---|---|---|
| 1. Which version? | Find the benchmark's own site and note its current default version, then search that version's public board for the model. A version tag missing from the vendor's citation is the tell. | The cited "Terminal-Bench v3.0" figure (26% of tasks) is vendor-only, with no disclosed harness; the same model reads 88.4% of tasks on the independent v2.1 board. |
| 2. Which subset? | Look for subset names — Extended, Main, Diamond, "full," "verified," "hard" — and find which subset the benchmark's public leaderboard actually ranks. | xAI cited Extended, the full and easiest pool; on the public Main board the rival's lead widens from roughly 2.3 points to 5.5 (53.5% vs 48.0% of tasks). |
| 3. Does the metric zero partial credit? | Read the grading rule on the leaderboard's methodology note. All-pass metrics compress high per-step rates into tiny headline scores; small task counts amplify every flip. | Harvey LAB is all-pass: 80.5% of criteria becomes 2.5% of tasks. At the 120-task sample AA describes, one task moves a score by roughly 0.83 points. |
| 4. Who's missing a column? | Pull each cited benchmark's public board and check whether its top two or three models appear as columns in the vendor's table. | Claude Opus 5 leads seven of the eight independently checkable benchmarks in the table — by 96 Elo, 138 Elo, and 17.5 points on the three we verified directly — and has no column. |
The trend behind the protocol is worth naming. As agentic benchmarks replace single-answer quizzes, the industry is moving toward exactly the structures that make these four levers more powerful: versioned test sets that fork rather than increment, nested difficulty subsets, all-pass task grading, and per-harness scores that resist casual comparison. That is not vendors getting more devious — it is evaluation getting more realistic, with more degrees of freedom as a side effect. The same discipline applies well beyond coding models: we found the identical reading problems in voice AI vendor claims and in hallucination-rate benchmarks, where metric shape does most of the steering.
Looking forward, we expect the gap between vendor tables and public boards to keep widening through 2027 — launch-day tables will keep citing whichever version, subset, and column set reads best, while the independent boards (Artificial Analysis, frontierbench.ai, Vals, Cognition's public leaderboards) keep professionalizing with confidence intervals, per-row harness disclosure, and open grading code. The buyers who thrive are the ones who treat the vendor table as a list of claims to check rather than a verdict to accept. If your team is making a model decision on benchmark evidence and wants the comparative eval run on your own workload instead of a vendor's chosen subset, that is exactly what our AI transformation engagements start with.
08 — ConclusionRead the table's choices, not its arithmetic.
Honest numbers, arranged: check the version, the subset, the metric, and the missing column.
The Grok 4.6 launch table is the rare exhibit that rewards close reading precisely because it is honest. Every publicly checkable cell reproduces its source; seven rows show the vendor's own model losing; and on one benchmark the vendor even quoted a lower figure for itself than it needed to. If this table misleads anyone, it will not be through a false number — it will be through a chosen one.
The four choices are now yours to check: a version whose independent board the model does not appear on, a subset easier than the one publicly ranked, a metric whose all-pass shape turns an 80.5% criteria rate into a 2.5%-of-tasks headline, and a column set that omits the model leading seven of the eight checkable boards. None of those checks needs more than the benchmark's own public leaderboard and about two minutes.
The general lesson outlives this table. Vendor benchmarks are becoming more honest in their cells and more editorial in their frames — and frames do not show up in a fact-check. Run the four questions on the next launch table you see, whatever the vendor and whichever direction the bold points. The numbers will usually be true. The arrangement is where the argument lives.