AI DevelopmentDecision Matrix14 min readPublished August 12, 2026

xAI’s default vs rivals’ ceilings · every checkable cell verified · the actual leader is missing from the table

Grok 4.6 vs Sol vs Opus 5 vs Fable 5: Read the Tiers

xAI’s Grok 4.6 launch table, published August 12, 2026, survives a cell-by-cell fact-check — all 31 publicly checkable numbers reproduce their cited sources. But it compares xAI’s own default effort rung against its rivals’ documented ceilings, and the model that leads seven of the eight independently checkable benchmarks, Claude Opus 5, has no column at all.

DA
Digital Applied Team
Senior strategists · Published August 12, 2026
PublishedAugust 12, 2026
Read time14 min
SourcesxAI docs · Artificial Analysis · Cursor · Vals AI
Grok 4.6 effort rungs
4
low · medium · high · xhigh — no max rung
default: high
AA Index gap · Grok 4.6 high vs Sol max
0.007pts
inside AA’s stated ±1% confidence interval
Checkable benchmarks led by Opus 5
7/8
absent from all four table columns
Checkable cells matching sources
31/31
zero fabrications found

xAI launched Grok 4.6 on August 12, 2026 with a ten-row benchmark table comparing it against GPT-5.6 Sol and Claude Fable 5 — and that table does something unusual for a vendor launch asset: every one of its 31 publicly checkable cells reproduces its cited source exactly. It is also a masterclass in why accurate numbers are not the same thing as a fair comparison.

Two levers shape what the table appears to say. The first is the effort tier: xAI benchmarked Grok 4.6 at its own default setting while quoting both rivals at their documented ceilings — an asymmetry that, counter to the cynical read, understates Grok 4.6 and even cost xAI a leaderboard win it could have claimed. The second is column selection: Claude Opus 5 leads seven of the eight benchmarks in this comparison that can be checked on independent boards, and it appears in none of the table’s four columns.

This analysis lays the three vendors’ effort ladders side by side, re-runs xAI’s rows against the public leaderboards they cite, separates the statistically real wins from the ones inside the error bars, and prices all four models by surface. If you route real budget on launch tables — as a follow-up to our previous-generation three-way comparison — this is the reading order that keeps you honest.

Key takeaways
  1. 01
    xAI benchmarked its default against rivals’ ceilings.Grok 4.6’s documented reasoning_effort ladder has four rungs — low, medium, high, xhigh — with high as the default and no max rung. OpenAI documents six rungs for GPT-5.6 Sol ending at max; Anthropic documents five for Opus 5 and Fable 5, also ending at max. The table quotes Grok 4.6 at high, and both rivals at max.
  2. 02
    The conservative rung cost xAI a board win.On Cursor’s public CursorBench v3.2 board, Grok 4.6 at Extra High scores 70.8% — rank 1, above Fable 5 Max’s 70.5% — at $2.81 per task versus $17.32. xAI printed its 69.9% High row instead. That is the opposite of score inflation.
  3. 03
    Zero fabrications — the lever is column selection.All 31 publicly checkable cells in the ten-row table reproduce their cited sources exactly. But Claude Opus 5 leads seven of the eight independently checkable benchmarks, tops Artificial Analysis’s index at 63.05, and has no column — so bolding the best score per row happens inside a field that excludes the leader.
  4. 04
    Two of Grok 4.6’s three row wins sit inside the error bars.Using Artificial Analysis’s own published 95% confidence intervals, the GDPval-AA v2 and AA-Briefcase leads are statistically insignificant. Opus 5’s leads on both do not overlap anyone’s interval — and Opus 5 is the model without a column.
  5. 05
    What the table genuinely supports is cost-efficiency.Grok 4.6 at its default effort statistically ties GPT-5.6 Sol at maximum effort on AA’s independent index — 60.92 vs 60.93 — at $0.84 per Index task versus $1.23. Read the rung, the benchmark version, the subset and the price surface before quoting any vendor table.

01What xAI PublishedTen rows, four columns, zero fabrications.

The launch table on x.ai/news/grok-4-6 lists ten benchmark rows across four model columns: Grok 4.6 (High), Grok 4.5 (High), GPT-5.6 Sol (Max) and Claude Fable 5 (Max), with the best score in each row set in bold. We checked every cell that has a public counterpart — 31 of them, spanning Terminal-Bench v3.0, DeepSWE v1.1, CursorBench v3.2, APEX-Agents, APEX-SWE, GDPval-AA v2 and AA-Briefcase — against each benchmark’s own public board: frontierbench.ai, deepswe.datacurve.ai, cursor.com/cursorbench, mercor.com/apex and artificialanalysis.ai/evaluations. Every checkable number reproduces its source exactly.

The table is not even arranged to hide defeats. Fable 5 wins five of the ten rows outright to Grok 4.6’s three, and xAI printed rows it loses badly — including Terminal-Bench v3.0 and DeepSWE v1.1. The launch mechanics, pricing and agentic results have their own write-up in our companion analysis of the Grok 4.6 launch; this post is about how the table frames those accurate numbers.

"Third-party model scores are the best of self-reported or publicly available results."— xAI methodology footnote, Grok 4.6 launch page

That footnote — together with its sibling line, “Competitor figures are drawn from the respective developers’ published system cards or benchmark leaderboards” — is the load-bearing caveat of the whole asset. It means the four columns were not run under one harness by one operator. Each cell is the best number available from somewhere, produced under whatever conditions that somewhere used. Accurate transcription, heterogeneous measurement.

The table
benchmark rows, four columns
10

Grok 4.6 (High), Grok 4.5 (High), GPT-5.6 Sol (Max), Claude Fable 5 (Max). The parenthetical tier labels are the story: one default, two ceilings.

x.ai/news/grok-4-6
Fact-check
checkable cells reproduce their sources
31/31

Every cell with a public counterpart matches it exactly across seven benchmark boards. The distortion in this table is selection — which models, which rungs, which subsets — not arithmetic.

verified against the named public boards
Row count
Fable 5 outscores Grok 4.6 on xAI’s own table
5–3

Of the ten rows xAI printed, Fable 5 takes five outright to Grok 4.6’s three. A vendor happy to print its own losses is not inflating scores; it is framing them.

xAI’s own launch table

02The Effort LaddersThree ladders, one asymmetric comparison.

Reasoning-effort tiers are the units of every 2026 benchmark number, and the three vendors build their ladders differently. Per xAI’s developer documentation, Grok 4.6’s reasoning_effort parameter has exactly four rungs — low, medium, high, xhigh — and the docs are explicit: “If not specified, reasoning_effort defaults to ‘high’. Reasoning cannot be disabled.” The top rung is documented as “xhigh — Maximum reasoning depth… available on grok-4.6.” There is no max rung, and Grok 4.5 does not support xhigh at all — requests sent to it with xhigh are silently treated as high.

OpenAI documents six effort rungs for GPT-5.6 Sol, from none up to a max ceiling, with medium as the default. Anthropic documents five rungs for Claude Opus 5 and Fable 5, from low up to max, with high as the default. Both rivals’ max ceilings are independently corroborated: Artificial Analysis’s board carries max-labelled runs for GPT-5.6 Sol and Claude Opus 5, and OpenAI’s live pricing page carries the tier. Lay the three ladders side by side — no vendor page does — and the launch table’s asymmetry becomes legible.

Documented reasoning-effort ladders for xAI Grok 4.6, OpenAI GPT-5.6 Sol, and Anthropic Claude Opus 5 and Fable 5, showing each ladder’s rungs, default, ceiling, and the rung xAI’s launch table used, compiled from each vendor’s developer documentation at the time of writing.
VendorModel(s)Documented rungsDefaultCeilingRung in xAI’s table
xAIGrok 4.6low · medium · high · xhighhighxhigh — no max rung existshigh — its default, 3rd of 4
OpenAIGPT-5.6 Solnone · low · medium · high · xhigh · maxmediummaxmax — its ceiling, 6th of 6
AnthropicClaude Opus 5 · Claude Fable 5low · medium · high · xhigh · maxhighmaxmax — Fable 5 only; Opus 5 has no column
Stated carefully
xAI compared its own default, third-of-four setting against GPT-5.6 Sol’s sixth-of-six and Fable 5’s fifth-of-five documented ceilings. That is an asymmetric comparison of the vendors’ own labels — not proof of intent either way. But the direction of the asymmetry matters: Grok 4.6’s showing at high understates what the model measures at its own xhigh ceiling. A vendor juicing its numbers would have done the opposite.

03The Measurable CostThe rung xAI chose cost it a win.

The asymmetry is not hypothetical — one public board measures it directly. Cursor’s CursorBench v3.2 leaderboard publishes scores by model and effort tier, so the same model appears at several rungs with a per-task cost attached. Grok 4.6 has no Max row there, which corroborates the four-rung ladder. What it does have is an Extra High row that xAI chose not to print for itself.

CursorBench v3.2 · score and cost per task, by model and effort tier

Source: cursor.com/cursorbench, retrieved at the time of writing · task suite is private
Grok 4.6 · Extra High$2.81 per task · the rung xAI didn’t print
70.8%
Rank 1
Fable 5 · Max$17.32 per task
70.5%
Opus 5 · Max$8.23 per task
70.0%
Grok 4.6 · High$2.34 per task · the row xAI published
69.9%
Opus 5 · Extra High$7.35 per task
69.3%
Fable 5 · Extra High$11.73 per task
68.4%
GPT-5.6 Sol · Max$5.69 per task
67.2%
Grok 4.6 · Medium$1.28 per task
67.1%
Grok 4.6 tiersRival models at their tiers

Read the top four bars. Grok 4.6 at Extra High scores 70.8% on Cursor’s private task suite — rank 1 on the board — above Fable 5 Max’s 70.5% at roughly a sixth of the per-task cost, $2.81 against $17.32. The figure xAI actually published is the 69.9% High row. By quoting its default instead of its ceiling, xAI gave up the rank-1 claim it could have made — the single clearest piece of evidence that the effort-tier asymmetry runs against xAI’s interest, not for it. How this board has evolved as a vendor-adjacent benchmark is a story we covered in our CursorBench v3.1 analysis; the version here is v3.2.

Caveat that stays attached
Cursor is Grok 4.6’s launch-day distribution partner — and Anysphere, the company behind Cursor, is the subject of a pending, press-reported SpaceX acquisition. CursorBench’s tasks are private, and Artificial Analysis has published no Grok 4.6 xhigh run of its own. So the xhigh headroom is measured by a corroborated primary board from a commercial partner — not by an independent audit. We print it with that label, and the conclusion it supports is directional, not decimal.

04Column SelectionThe missing column is the actual leader.

Artificial Analysis published its own independent Grok 4.6 score on launch day: 61 on the AA Intelligence Index, ranked #6 of the 183 models AA had evaluated (the median score is 34), labelled “Grok 4.6 (high).” That is genuine third-party corroboration, and it matches xAI’s rounded claim exactly. It also exposes what the four-column frame leaves out.

Artificial Analysis Intelligence Index scores and cost per Index task for Grok, GPT-5.6 Sol and Claude models by effort tier, with a column showing whether each row appears in xAI’s four-column launch table, based on artificialanalysis.ai as retrieved at the time of writing.
Model — AA’s own labelAA IndexCost per Index taskIn xAI’s four columns?
Claude Opus 5 (max)63.05 — board leader$2.34No
Claude Opus 5 (xhigh)62.52No
Claude Fable 5 (with fallback)62.07$3.14Yes — as “Fable 5 Max”
Claude Opus 5 (high)61.48No
GPT-5.6 Sol (max)60.93$1.23Yes
Grok 4.6 (high)60.92$0.84Yes — the vendor’s own row
GPT-5.6 Sol (xhigh)59.01No
GPT-5.6 Sol (high)57.33No
Grok 4.5 (high)55.76Yes

Claude Opus 5 at max tops Artificial Analysis’s entire board at 63.05 — above every model in every column of xAI’s table — and it is not in the table. The pattern repeats across the comparison’s independently checkable benchmarks: Opus 5 leads seven of the eight, including the AA Index (63.05), GDPval-AA v2 (an 1849 Elo rating against the benchmark’s human-expert anchor of 1000), AA-Briefcase (1715 Elo on the same style of scale), Terminal-Bench v3.0 (43.5% of its roughly 75 tasks), DeepSWE v1.1 (74% of its 113-task board), and the top position on both of Mercor’s public APEX leaderboards.

This is an omission, not a fabrication — and the distinction is the post’s whole point. Bolding the best score in each row is only meaningful relative to the field the table defines, and this field excludes the model that would take most of the bold. Opus 5’s own launch numbers, and the rarer story of one of them being independently administered, are covered in our Opus 5 launch analysis and our piece on ARC Prize’s verified Opus 5 result.

"Grok 4.6 buys Sol-class intelligence at roughly two-thirds the per-task price, still trails Anthropic’s best on capability — and the table is arranged so you don’t notice the second half."— Digital Applied editorial synthesis

05Statistical SignificanceTwo of three wins sit inside the error bars.

Grok 4.6’s three row wins on its own table — GDPval-AA v2, AA-Briefcase and Harvey LAB — are also the comparison’s most fragile numbers. Artificial Analysis publishes 95% confidence intervals for its evaluations, so significance is checkable rather than a matter of opinion.

On GDPval-AA v2 — a Bradley-Terry Elo rating anchored to human experts at 1000, not an arena-style crowd vote — Grok 4.6’s 1753 carries an interval of 1733–1774, which overlaps Fable 5’s 1724–1757 and Sol’s 1711–1744. The win is inside the noise. Opus 5’s 1849, interval 1826–1871, overlaps none of the three: its 96-point lead over Grok 4.6 is the row’s one statistically significant gap, and it belongs to the model with no column. On AA-Briefcase the same shape repeats: Grok 4.6’s 1576.59 (1565–1588) against Fable 5’s 1573.78 (1564–1584) is a 2.8-point gap well inside overlapping intervals, while Opus 5’s 1714.59 (1704–1727) is 138 points clear with no overlap.

The third win, Harvey LAB, is a legal-agent benchmark of 120 held-out tasks — so a single task is worth 0.83 percentage points. Grok 4.6’s claimed 15.8% is roughly 19 tasks. In a two-proportion test that clears significance against Sol’s 2.5% (p=0.0003) but not against Grok 4.5’s 12.92% (p=0.52) or Fable 5’s 11.25% (p=0.30) as scored on Vals AI’s public board. And Grok 4.6 itself is absent from that public board, where two versions of Meta’s Muse Spark score higher than its claimed figure — 25.42% and 20.00% of the same 120 tasks. On the board xAI’s row implies it tops, the claimed score would rank third.

Why Sol scores 2.5% — Vals AI’s own explanation
Harvey LAB grades all-pass: a task counts only when every rubric criterion passes. GPT-5.6 Sol (max) passes 80.5% of the individual rubric criteria across the 120 tasks but only 2.5% of tasks end-to-end — roughly 3 of 120. As Vals AI puts it: “the bottleneck is finishing, not the individual steps: top models satisfy roughly 85-95% of the individual criteria… but a task scores only when every criterion passes — so one or two misses per task wipe out most of the credit.” Sol is not weak at law in absolute terms — it ranks #3 at 48.1% on Vals’ separate Legal Research Bench. The all-pass metric compresses a ~1.1× criteria gap into a ~5–6× task-level gap.
AA Index
Grok 4.6 (high) vs Sol (max)
0.007pts

60.92 against 60.93 on Artificial Analysis’s independent index, whose stated 95% confidence interval is under ±1%. A statistical tie — reached with Grok 4.6 one rung below its own ceiling and Sol at its documented maximum.

artificialanalysis.ai · retrieved at the time of writing
GDPval-AA v2
Opus 5’s lead over Grok 4.6 — outside all intervals
+96Elo

Grok 4.6’s 12-Elo edge over Fable 5 on this row sits inside overlapping confidence intervals. Opus 5’s 1849 does not overlap any of the three printed models — the only significant gap on the row belongs to the absent model.

AA’s published 95% CIs
Harvey LAB
where the claimed 15.8% would rank
3rd

Grok 4.6 is not on Vals AI’s public Harvey LAB board. Muse Spark 1.2 (25.42%) and Muse Spark 1.1 (20.00%) both exceed the vendor-stated 15.8% of the benchmark’s 120 tasks — so the implied row win is a third-place score on the public board.

vals.ai/benchmarks/hlab · updated Aug 10, 2026

06Versions and SubsetsSame benchmark name, different test.

Effort rungs are one axis of selection. Benchmark versions and subsets are the other, and they are easier to miss because the names stay the same while the test changes underneath.

FrontierCode v1.1 defines three nested subsets: Extended is all 150 tasks, Main is the hardest 100, Diamond the hardest 50. xAI’s table reports the Extended subset — the easiest of the three — and Fable 5 led that row on xAI’s own table anyway. On Cognition’s public Main board, the best run per model scores: Fable 5 at 53.5%, Opus 5 at 53.4%, Grok 4.6 at 48.0% and Sol at 47.5% of the hardest 100 tasks. The ordering matches xAI’s row; the gap to Fable 5 stretches to 5.5 points on the harder subset. Scores are not comparable across subsets — which also means Anthropic’s self-published 53.5% (a Main figure) does not contradict xAI’s Extended row. They are different tests wearing one name.

Terminal-Bench is the sharpest version trap in the table. On v3.0 — a separate, roughly 75-task suite — the public board’s rival cells match xAI’s table exactly: Sol at 34.6%, Fable 5 at 34.1%. Grok 4.6’s 26% appears on no public v3.0 board — it is vendor-only. Each public row also runs under a different agent harness (Codex for Sol, Claude Code for Fable 5, Cursor CLI for Grok 4.5), and xAI discloses no harness for its own 26%. The board’s outright leader is, again, Opus 5 (max) at 43.5% — 17.5 points above Grok 4.6’s claimed score, and absent. Yet on Terminal-Bench v2.1, the version Artificial Analysis’s index actually uses, the story inverts into a near-tie: Sol xhigh 89.5%, Opus 5 max 89.1%, Grok 4.6 high 88.4%, Sol max 88.0%, Fable 5 84.6% of that suite’s tasks in AA’s runs. Any “crushed at terminal work” reading is v3.0-specific. State the version, every time.

Row-by-row comparison of what xAI’s Grok 4.6 launch table shows against what each benchmark’s own public or independent board shows for the same models, with the board to check, compiled at the time of writing.
BenchmarkxAI’s launch table showsThe public board showsWhere to check
AA Intelligence IndexGrok 4.6 at 61 matches Sol; Fable 5 leads the four columns at 62Opus 5 (max) tops the full 183-model board at 63.05 — no column in xAI’s tableartificialanalysis.ai/models/grok-4-6
CursorBench v3.2Grok 4.6 (High) 69.9%Grok 4.6 (Extra High) 70.8% — rank 1 at $2.81 per task, a tier xAI didn’t print for itselfcursor.com/cursorbench
FrontierCode v1.1Grok 4.6 61.3% on Extended — all 150 tasks, the easiest of three nested subsetsMain (hardest 100): Fable 5 53.5% · Opus 5 53.4% · Grok 4.6 48.0% — subsets are not comparable to each othercognition.com/frontiercode
Terminal-Bench v3.0Grok 4.6 26% — no harness disclosed for the cellRival cells match exactly, each under a different harness; Opus 5 (max) leads at 43.5% and is absentfrontierbench.ai
Harvey LABGrok 4.6 15.8% of 120 tasks — the row’s bolded winGrok 4.6 is not on the public board; two Muse Spark versions score higher, so 15.8% would rank #3vals.ai/benchmarks/hlab
One more label to read carefully
“Fable 5 Max” is the launch table’s column heading, not a product Anthropic exposes under that name in these runs. Artificial Analysis labels its entry “Claude Fable 5 (with fallback)” — a routed configuration that falls back to Opus 4.8 on failure. Vals AI documented the behavior on Harvey LAB: “Fable 5 fell back to Claude Opus 4.8 on 4 tasks; counting those as failures gives a no-fallback score of 10.42%” against its 11.25% headline. Fable 5’s own release benchmarks are covered in our Fable 5 release analysis.

07Price by SurfaceWhat the table genuinely supports: cost-efficiency.

Strip away the framing and a strong, defensible claim survives: at its default effort, Grok 4.6 statistically ties GPT-5.6 Sol at maximum effort on an independent index, at $0.84 per AA Index task against Sol max’s $1.23 — with Opus 5 max at $2.34 and Fable 5 at $3.14 on the same cost measure. At its own xhigh ceiling it takes rank 1 on CursorBench v3.2 for $2.81 a task against Fable 5 Max’s $17.32. The honest headline is not “new frontier leader” — it is Sol-class intelligence at roughly two-thirds of Sol’s per-task cost, while Anthropic’s best stays ahead on capability.

Per-task economics then depend entirely on which pricing surface your traffic hits — and none of these models is one price. Every figure below is an API list price per million tokens from the vendor’s own pricing page at the time of writing; how the two closest rivals price against each other is unpacked further in our Sol vs Fable 5 price-and-access comparison.

API list prices per million tokens by pricing surface for Grok 4.6, GPT-5.6 Sol, Claude Opus 5 and Claude Fable 5, from each vendor’s own pricing documentation at the time of writing.
SurfaceInput $/MtokCached inputOutput $/MtokNotes
xAI Grok 4.6 — API list, docs.x.ai · 500K context (xAI-documented)
Standard — prompt under 200K tokens$2.00$0.50$6.00Covers only the first 200K of the 500K window — 40% of it
Long context — prompt of 200K tokens or more$4.00$1.00$12.00Applies to every token in the request once crossed
Priority (“fast”)2× the applicable rateMultiplier per xAI; no separate fast model ID is documented
OpenAI GPT-5.6 Sol — API list, developers.openai.com · 1.05M context
Standard — short context$5.00$0.50$30.00Max output 128K tokens
Long context — prompt above 272K tokens$10.00$1.00$45.00Input doubles, output rises 1.5× vs standard
Batch / Flex$2.50$0.25$15.00Half the standard rate
Priority — short context$10.00$1.00$60.002× the standard rate
Anthropic — API list · 1M context, no long-context price tier
Claude Opus 5$5.00$0.50 cache read$25.00Batch 50% off standard · max output 128K tokens
Claude Fable 5$10.00$1.00 cache read$50.00Exactly 2× Opus 5 on every line, including cache
The repricing trap
Grok 4.6’s long-context rate is not marginal. Once a prompt reaches 200K tokens, the $4/$12 rate applies to the entire request — a 210K-token prompt bills every token at the long-context rate, not 200K at standard plus the overflow. Since the model’s documented window is 500K tokens, the headline $2/$6 price covers only the first 40% of the context you are buying. Note also that Grok 4.6’s cached-input price rose against Grok 4.5 — $0.30 to $0.50 standard, $0.60 to $1.00 long-context — while the context window and headline rates stayed identical: this generation is neither a context expansion nor a price cut.

One more surface rule: reseller listings do not reliably mirror vendor lists. On OpenRouter, Grok 4.6 and GPT-5.6 Sol match their vendors’ standard list prices, but some of Sol’s smaller siblings list at rates that match other OpenAI surfaces instead — so check each model’s listing against the vendor page rather than generalizing from one. Catalog hygiene of exactly this kind — prices, dates and surfaces — is the subject of our companion guide to reading model catalogs.

08Reading PlaybookHow to read the next vendor table.

The Grok 4.6 table is unusually instructive because it is honest at the cell level and shaped at every other level. That combination is the future of vendor benchmarking: as fact-checking gets easier, selection — of rungs, columns, versions and subsets — replaces fabrication as the lever. Four checks catch most of it.

Check 01
Read the rung

A score without its effort tier is a number without units. Grok 4.6 has four rungs with high as default and no max; Sol has six ending at max; Opus 5 and Fable 5 have five ending at max. Ask which rung produced each column — and whether defaults are being compared against ceilings.

Tier named? Now compare
Check 02
Read the version and subset

Terminal-Bench v2.1 and v3.0 produce opposite narratives for the same models. FrontierCode’s Extended, Main and Diamond subsets are nested and non-comparable. A benchmark name without its version and subset is a different test wearing the same jersey.

Version + subset, every time
Check 03
Read for the missing column

Open the boards the table itself cites and look for models that outscore every printed column. Here, that search takes minutes and finds Opus 5 leading seven of eight checkable benchmarks. The strongest competitor is often not misquoted — it is simply absent.

Find who is not in the table
Check 04
Read the price by surface

Headline rates cover part of the product. Grok 4.6’s $2/$6 applies below a 200K-token prompt — 40% of its 500K window — and crossing the threshold reprices the whole request. Label every figure by surface: standard, long-context, batch, priority, reseller.

Price = rate × surface

The deeper skill — reading what a vendor chose to disclose, and what the disclosure pattern itself tells you — is one we apply across launches in our companion piece on reading disclosed losses in vendor tables. And if your team is choosing models with real budget attached, benchmark tables should only ever be the shortlist stage: run your own evals on your own tasks before you commit. That comparative eval discipline is exactly where our AI transformation engagements usually start.

09ConclusionRead the tiers before you read the scores.

The bottom line

The rung is part of the number — and so is the column you don't see.

xAI fabricated nothing. All 31 publicly checkable cells in the Grok 4.6 launch table reproduce their sources, the table prints rows xAI loses, and the vendor quoted itself at its default rung while quoting rivals at their ceilings — a choice that cost it a rank-1 claim on CursorBench it could have made at xhigh. On the axis everyone assumes vendors cheat, this table runs the asymmetry against itself.

The shaping lives elsewhere. Claude Opus 5 leads seven of the eight independently checkable benchmarks in this comparison and appears in none of the four columns, two of Grok 4.6’s three row wins sit inside Artificial Analysis’s own confidence intervals, and the boldest-looking rows lean on the easiest subset or a version-specific board. What the table genuinely supports is still commercially important: Sol-class intelligence at roughly two-thirds of Sol’s per-task cost on an independent index — with Anthropic’s best still ahead on capability, at a higher price.

The forward lesson is bigger than one launch. Effort ladders now differ in length, default and ceiling across every major vendor, and benchmark families fork into versions and subsets faster than anyone’s intuition updates. That means the honest-cell, shaped-frame table is what most launches will look like from here. Treat every vendor table as a reading exercise — rung, version, subset, missing column, price surface — and then run your own tasks. The numbers can all be true and the conclusion still be yours to make.

Model evaluation without the vendor spin

Route budget on your own evals, not on someone else’s table.

Our team helps businesses pick, benchmark and route frontier models on their own tasks — reading past vendor framing to the effort tiers, surfaces and per-task costs that actually move the bill.

Free consultationExpert guidanceTailored solutions
What we work on

Frontier-model selection engagements

  • Comparative evals on your tasks — not vendor tables
  • Effort-tier and cost-per-task routing policies
  • Long-context billing audits across surfaces
  • Multi-vendor failover and fallback design
  • Benchmark-claim verification for procurement
FAQ · Effort tiers

The questions teams ask before routing budget.

The table quotes Grok 4.6 and Grok 4.5 at High, and GPT-5.6 Sol and Claude Fable 5 at Max. Per xAI’s own developer documentation, Grok 4.6’s reasoning_effort ladder has exactly four rungs — low, medium, high, xhigh — with high as the default and no max rung at all; reasoning cannot be disabled, and Grok 4.5 does not support xhigh. OpenAI documents six rungs for GPT-5.6 Sol ending at max, and Anthropic documents five rungs for Opus 5 and Fable 5, also ending at max. So xAI compared its own default, third-of-four setting against both rivals’ documented ceilings — an asymmetry that understates Grok 4.6 rather than inflating it, since the model scores higher at its own xhigh ceiling on the one public board that measures it.
Related dispatches

Continue exploring frontier comparisons.

AI Development

Grok 4.6 Lands: Same Price as 4.5, Frontier Parity

xAI's Grok 4.6 keeps Grok 4.5's headline rates, but the $2/$6 tier only holds below a 200K-token prompt: cross it and the whole request reprices.

August 12, 2026 · 13 minRead
AI Development

Effort Dials Arrive: Think Buttons and Sliders Explained

ChatGPT's effort slider moves a developer-API control into the chat window; a Think button is announced next week. What effort changes, how vendors label it.

August 7, 2026 · 17 minRead
AI Development

Open-Source Agent Memory: Mem0 vs Letta vs Zep Compared

Mem0, Letta and Zep bet differently on what remembering means. A licence, architecture and benchmark comparison, including why the published scores disagree.

August 4, 2026 · 17 minRead
AI Development

ARC Prize Verified Opus 5. That Is Rarer Than It Sounds.

ARC Prize independently administered Claude Opus 5's 30.16% ARC-AGI-3 result. Almost every other launch-day benchmark number is vendor-run on a vendor harness.

July 26, 2026 · 11 minRead
AI Development

Agent Computer Use: Enterprise Automation Playbook

Enterprise playbook for deploying computer-use agents — a 40-point guardrails checklist spanning identity, audit, action boundaries, failures, and compliance.

May 22, 2026 · 17 minRead
AI Development

State of AI Agents 2026: 200+ Data Points Compiled

The definitive State of AI Agents 2026 — 247 data points across adoption, ROI, autonomy, and governance, sourced from McKinsey, Stanford HAI, and Gartner.

May 22, 2026 · 16 minRead