AI DevelopmentDecision Matrix12 min readPublished August 14, 2026

19 benchmark rows · 5 models · a split decision, not a sweep

Gemini 3.7 Flash vs Sonnet 5 vs GPT-5.6 Terra: Real Wins

Google’s Gemini 3.7 Flash model card carries 19 benchmark rows across five models — not the five rows the launch coverage repeated. Read in full, it is a split decision: nine wins for 3.7 Flash, six for GPT-5.6 Terra, quiet wins for Claude Sonnet 5, and two rows where 3.6 Flash beats its own successor.

DA
Digital Applied Team
Senior strategists · Published Aug 14, 2026
PublishedAug 14, 2026
Read time12 min
SourcesModel card · AA · VentureBeat
Rows 3.7 Flash wins
9/19
of Google’s own table
AA Intelligence Index
56
same-vintage, v4.1.1
+4 vs 3.6 Flash
Rows Terra wins
6/19
incl. Terminal-bench, OSWorld-2.0
Intro price / 1M
$0.75
in · $3.75 out · to Dec 31

Gemini 3.7 Flash benchmarks look very different depending on which table you read. Google’s launch post led with five rows; the full model card publishes nineteen, compared across five models — Gemini 3.7 Flash, Gemini 3.6 Flash, Claude Sonnet 5, GPT-5.6 Terra, and Muse Spark 1.2. This post reads all nineteen.

The stakes are practical. Gemini 3.7 Flash shipped on August 13, 2026 at an introductory $0.75 per million input tokens and $3.75 per million output on Google’s API — a rate that undercuts every rival in its own comparison table, matched only by 3.6 Flash on the same introductory window. If the benchmark story holds, that price rewrites routing decisions. If it only holds on the five rows the vendor highlighted, it doesn’t. We covered the launch itself — pricing mechanics, positioning, and the asterisk on “half price” — in the full launch rundown; this post is the benchmark deep-dive.

What follows: the full 19-row table transcribed from the model card, a recomputed win tally, the rows where Sonnet 5 and even 3.6 Flash quietly beat the new model, the blank cell that most coverage repeated without investigating, the independent Artificial Analysis numbers that actually moved this time, and every price labelled by the surface it comes from. Every benchmark figure below is vendor-stated from Google’s card unless noted otherwise.

Key takeaways
  1. 01
    The real table is 19 rows and five models.Most launch coverage repeated the same five benchmarks Google led with. The model card publishes 19 rows — including rows Google loses, a fifth model column (Muse Spark 1.2), one blank cell, and one internal inconsistency.
  2. 02
    It’s a split decision, not a sweep.By our tally of Google’s own table: 3.7 Flash takes 9 rows, GPT-5.6 Terra 6, Sonnet 5 2 outright (3 against the headline trio), 3.6 Flash 1, and Muse Spark 1.2 tops GDPVal-AA outright.
  3. 03
    The independent signal finally moved.Artificial Analysis scores 3.7 Flash at 56 versus a recomputed 52 for 3.6 Flash at the same index vintage — a genuine +4 that breaks the Flash line’s “cheaper, not smarter” pattern from July.
  4. 04
    The blank Sonnet 5 cell is a version mismatch, not a dodge.Google tested OSWorld-2.0; Anthropic publishes a different variant, OSWorld-Verified, where Sonnet 5 reportedly scores 81.2%. The two numbers must never be compared as if they were the same benchmark.
  5. 05
    “$0.75 per million” is one of at least three live rates.Google’s intro price runs through December 31, 2026 and applies to 3.6 Flash too; the standard rate rises to $1.50/$7.50 from January 1, 2027 — the same rate 3.6 Flash’s standard pricing already carries. OpenRouter stacks its own time-limited 50%-off promo on top.

01The SetupFive rows in the press cut, nineteen in the card.

Google introduced Gemini 3.7 Flash as “our most intelligent workhorse model yet for coding and agents,” three weeks after 3.6 Flash’s launch carried nearly identical positioning. The announcement post and almost every piece of launch coverage we located anchor on the same five benchmarks: FrontierCode, Terminal-bench 2.1, GDM-MRCR, OSWorld-2.0, and Harvey LAB-AA.

The model card is a different document. It compares five models across nineteen rows — and the two extra dimensions matter. The fifth column, Muse Spark 1.2, wins one row outright and is silently dropped by most coverage. And eight of the ten rows 3.7 Flash doesn’t win sit in the fourteen the press cut left out.

The press cut
5 rows
FrontierCode · Terminal-bench 2.1 · GDM-MRCR · OSWorld-2.0 · Harvey LAB-AA

The five benchmarks Google led with in its announcement — and the same five that launch coverage repeated. Two of the five are rows Google loses to GPT-5.6 Terra; one is a Google-authored benchmark.

What most coverage shows
The model card
19 rows × 5 models
3.7 Flash · 3.6 Flash · Sonnet 5 · GPT-5.6 Terra · Muse Spark 1.2

The full published table: nineteen benchmark rows, a fifth model column, one blank cell on Sonnet 5’s computer-use row, and a 0.4pp inconsistency between the card and Google’s own announcement prose.

deepmind.google model card

None of this means the table is dishonest — publishing rows you lose is more disclosure than most vendors manage. It means the table rewards close reading. The rest of this post is that reading.

02The Full TableAll 19 rows, tallied honestly.

The table below transcribes every benchmark row from the Gemini 3.7 Flash model card, all five model columns included. Bold marks the best score in each row. Recomputing the tally across all nineteen rows: 3.7 Flash wins 9, GPT-5.6 Terra wins 6, Sonnet 5 wins 2 outright, 3.6 Flash wins 1 (against its own successor), and Muse Spark 1.2 wins 1. Against the two rivals Google names in its framing, Sonnet 5 takes three rows — GDPVal-AA, Agent’s Last Exam, and BioMysteryBench’s human-solvable subset — though on GDPVal-AA the whole table’s top score belongs to Muse Spark 1.2.

All 19 benchmark rows from the Gemini 3.7 Flash model card, comparing Gemini 3.7 Flash, Gemini 3.6 Flash, Claude Sonnet 5, GPT-5.6 Terra, and Muse Spark 1.2, grouped by category, with the best score in each row marked in bold.
Benchmark3.7 Flash3.6 FlashSonnet 5GPT-5.6 TerraMuse Spark 1.2
Coding & software agents
FrontierCode 1.1Production code quality43.6%34.4%42.7%41.3%
DeepSWE v1.1Long-horizon software engineering65.3%48.6% †53.8%69.6%54.9%
Terminal-bench 2.1Agentic terminal coding85.8%78.0%80.4%87.4%82.9%
Terminal-bench 3.0General agent capabilities14.9%5.4%14.6%20.8%
WebDev Arena (Elo)Arena.ai human preference, vendor-reported15881538154115231535
Enterprise & knowledge work
AutomationBenchWorkflow automation — Google’s private task set30.4%17.0%10.7%23.6%
GDPVal-AA v2 (Elo)Knowledge work15251422159815781628
Harvey LAB-AAComplex legal workflows90.7%85.1%90.1%85.2%
GDP.pdfExpert PDF document comprehension34.0%22.0%28.0%24.7%16.0%
Multimodal & long context
CharXiv Reasoning, no toolsChart reasoning84.5%85.2%77.0%85.9%
CharXiv Reasoning, with toolsChart reasoning88.7%89.4%88.3%
LVBenchLong video understanding85.4%84.2%68.5%78.9%
GDM-MRCR v2, 128k *Long-context retrieval — Google-authored97.0%91.8%81.5%93.5%
Computer use
OSWorld-2.0Agentic computer use47.9%33.8%Not published ‡50.2%
Agent’s Last ExamMultimodal desktop / OS tasks, pass rate26.3%24.2%33.3%28.0%
Science & expert reasoning
HLE-VerifiedMultidisciplinary expert reasoning53.6%51.2%31.0%51.1%
BioMysteryBench, human-solvableBiology mysteries — solvable subset87.1%80.6%87.5%83.8%
BioMysteryBench, human-difficultBiology mysteries — difficult subset43.5%41.2%34.1%49.4%
LABBench2Real-world biology research tasks82.1%76.1%80.1%81.2%

* GDM-MRCR is a Google DeepMind-authored benchmark, self-reported on every Gemini release — read it as a vendor-designed eval, not an independent standard. † Google’s model card lists 48.6% for 3.6 Flash on DeepSWE v1.1; Google’s own announcement prose, published the same day, states 49.0% for the identical metric. We use the card as the table of record and flag the 0.4pp inconsistency rather than silently reconciling it. ‡ Not published in Google’s table. Anthropic reports Sonnet 5’s computer use on a different benchmark variant, OSWorld-Verified — see Section 05. “—” means the model card lists no score for that cell.

How to read this table
Every number is Google’s own measurement of rival models, and the card does not disclose which reasoning-effort configuration produced each rival’s score. Where effort is visible elsewhere, the ladders are not equivalent: 3.7 Flash’s ceiling is “high” on a three-rung ladder, while GPT-5.6 Terra’s published ceiling is “max” and Muse Spark 1.2’s is “xhigh” — deeper ladders whose labels don’t map one-to-one onto Google’s. Treat margins of two points or less as ties, and treat every row as a claim to verify on your own workload, not a verdict.

03The WinsWhere 3.7 Flash genuinely wins.

Nine wins out of nineteen is a real result, and some of the wins are meaningful. The headline coding claim holds — narrowly: 43.6% on FrontierCode 1.1 against Sonnet 5’s 42.7%, a 0.9-point edge on production code quality. The WebDev Arena Elo of 1588 leads the field, and it is Arena.ai’s human-preference data rather than a private Google metric — though it reaches the card vendor-reported. The clearest daylight shows up on document and multimodal work: 34.0% on GDP.pdf against Sonnet 5’s 28.0%, 85.4% on LVBench long-video understanding against Sonnet 5’s 68.5% — video remains a Gemini strength that the text-first rivals don’t contest — and 53.6% on HLE-Verified, where Sonnet 5 posts a notably weak 31.0%, 22.6 points behind the leader.

Narrowest win
Harvey LAB-AA
+0.6pp

90.7% vs Sonnet 5’s 90.1% on complex legal workflows — close enough to call a statistical tie, though Google bolds it as a win. Terra sits at 85.2%, 3.6 Flash at 85.1%.

Effectively a tie
Clearest win
AutomationBench
+6.8pp

30.4% vs Terra’s 23.6% and Sonnet 5’s 10.7% on enterprise workflow automation. The caveat: this is Google’s private, non-public task set — the card’s own notes say so.

Private task set
Long context
GDM-MRCR v2 · 128k
+3.5pp

97.0% vs Terra’s 93.5% and Sonnet 5’s 81.5%. Decisive — but GDM-MRCR is Google DeepMind’s own benchmark, self-reported on every Gemini release. A vendor-authored eval, however transparently named.

Google-authored

Notice the pattern in those three caps: the narrowest win is on an independent benchmark, and the two most decisive wins are on a private task set and a Google-authored eval. That doesn’t make the numbers false — it makes them exactly the kind of numbers you weight down when comparing across vendors, the same discount you’d apply to any lab grading its own homework.

04The LossesThe rows Google didn’t highlight.

GPT-5.6 Terra is the strongest counterweight in Google’s own data: six rows, concentrated exactly where Google’s framing claims 3.7 Flash leads — agentic coding and computer use. Terra takes Terminal-bench 2.1 (87.4% vs 85.8%), Terminal-bench 3.0 (20.8% vs 14.9%, a hard new benchmark where the whole field scores low), DeepSWE v1.1 (69.6% vs 65.3%), and OSWorld-2.0 (50.2% vs 47.9%). Sonnet 5 takes Agent’s Last Exam outright at 33.3% — seven points clear of 3.7 Flash on multimodal desktop tasks — and edges BioMysteryBench’s human-solvable subset. On GDPVal-AA knowledge work, Sonnet 5’s 1598 Elo beats 3.7 Flash’s 1525 by 73 points, and Muse Spark 1.2 tops the whole row at 1628.

Margins over 3.7 Flash on the rows it loses

Source: Gemini 3.7 Flash model card, winning margins recomputed from Google’s own figures (GDPVal-AA’s Elo-scale losses excluded from the pp chart)
Agent’s Last ExamSonnet 5 33.3% vs 3.7 Flash 26.3%
+7.0pp
Terminal-bench 3.0Terra 20.8% vs 3.7 Flash 14.9%
+5.9pp
BioMysteryBench, human-difficultTerra 49.4% vs 3.7 Flash 43.5%
+5.9pp
DeepSWE v1.1Terra 69.6% vs 3.7 Flash 65.3%
+4.3pp
OSWorld-2.0Terra 50.2% vs 3.7 Flash 47.9%
+2.3pp
Terminal-bench 2.1Terra 87.4% vs 3.7 Flash 85.8%
+1.6pp
CharXiv Reasoning, no toolsTerra 85.9% vs 3.7 Flash 84.5%
+1.4pp
CharXiv Reasoning, with tools3.6 Flash 89.4% vs 3.7 Flash 88.7%
+0.7pp
BioMysteryBench, human-solvableSonnet 5 87.5% vs 3.7 Flash 87.1%
+0.4pp

The oddest entries are the two CharXiv rows, where Gemini 3.6 Flash beats its own successor — 85.2% vs 84.5% without tools and 89.4% vs 88.7% with tools. That is genuinely what Google’s table says, and it deserves credit for printing a same-vendor regression rather than trimming the rows. It is also a useful calibration: when a three-week successor loses to its predecessor on chart reasoning, the version bump bought capability in some places by spending it in others — normal for fast-cycle releases, invisible in a five-row press cut.

05Table LiteracyThe blank cell and the name collision.

Two traps in this table will produce confidently wrong takes, and most launch coverage walked past both. The first is the blank Sonnet 5 cell on OSWorld-2.0. It is not a zero, and it is not evidence that Anthropic declined to disclose computer-use performance. Anthropic publishes Sonnet 5’s computer-use score on a different benchmark variant — OSWorld-Verified, where it reportedly scores 81.2% — and disclosed that it changed its OSWorld-Verified methodology between Sonnet 4.6 and Sonnet 5, restating Sonnet 4.6 to 78.5% in the process. Google tested OSWorld-2.0. Different variant, different methodology, different scale: the 81.2% and the 47.9%–50.2% range in Google’s table must never be cross-compared as the same number.

Same name, different benchmark
The second trap is a name collision. AutomationBench in Google’s table (3.7 Flash: 30.4%) is Google’s private enterprise task set. AutomationBench-AA (3.7 Flash: 62.7%) is Artificial Analysis’s own, separate benchmark. Same-sounding name, different tasks, different scoring — merging them, or quoting the 62.7% as an improvement over the 30.4%, is a fabricated comparison.

Both traps are instances of the general failure mode we catalogued in our guide to reading vendor benchmark tables: version mismatches and missing cells get read as verdicts when they are artifacts of who tested what, under which harness. Add this card’s own contribution to the genre — the DeepSWE figure for 3.6 Flash that reads 48.6% in the table and 49.0% in the announcement prose published the same day — and the lesson generalizes: even first-party numbers disagree with themselves at the margins, which is precisely why margins under a point should never drive a routing decision.

06Independent SignalThe independent read: 52 → 56, same vintage.

This is the part of the story that vendor tables can’t settle, and it is where 3.7 Flash earns its most defensible claim. Artificial Analysis scores Gemini 3.7 Flash (high) at 56 on its Intelligence Index — a 4-point improvement over Gemini 3.6 Flash’s 52, with both scores computed at the same index vintage (v4.1.1, a nine-evaluation composite). That matters because of what came before: our July read of 3.6 Flash found a flat Intelligence Index — “cheaper, not smarter.” That pattern breaks here. AA also puts 3.7 Flash on its Intelligence-vs-Time Pareto frontier at 1.7 minutes per task, which it describes as “40% faster than GPT-5.6 Terra (max),” with output throughput around 340 tokens per second.

AA Intelligence Index · same-vintage scores by effort configuration

Source: Artificial Analysis Intelligence Index v4.1.1
Gemini 3.6 Flash (high)Recomputed at v4.1.1 vintage
52
Gemini 3.7 Flash (low)Lowest of three effort rungs
51
Gemini 3.7 Flash (medium)Middle rung
53
Gemini 3.7 Flash (high)Ceiling of a three-rung ladder
56
GPT-5.6 Terra (max)Ceiling of a deeper effort ladder
57
Muse Spark 1.2 (xhigh)Ceiling of a deeper effort ladder
57
Vintage warning
Our July post reported 3.6 Flash at a flat 50 — an earlier AA index vintage. AA has since updated its methodology to v4.1.1 and recomputed 3.6 Flash’s own score to 52. Comparing July’s 50 to today’s 56 conflates two vintages; the valid same-vintage comparison is AA’s own 52 → 56. The +4 is real either way — but only one arithmetic is honest.

Two qualifiers keep the independent read honest. First, the effort-ladder asymmetry from Section 02 applies here too: 3.7 Flash’s 56 is the top of a three-rung ladder (51/53/56), while Terra’s 57 comes at “max” and Muse Spark’s 57 at “xhigh” — each model’s respective ceiling, on ladders of different depth. Second, the ceiling is close but not reached: 56 is one point behind both, and VentureBeat notes Artificial Analysis places Claude Opus 5 at 63 on the same index — the frontier tier remains a different conversation. On AA’s sub-indices, as surfaced through OpenRouter’s AA-sourced panel, 3.7 Flash posts a Coding Index of 76.1 and an Agentic Index of 45.1; Arena.ai’s early human-preference results provisionally rank it ninth overall and eighth for web development. Independent evals exist this time — plural, and broadly consistent with each other.

"That is the metric enterprise teams will ultimately need to test: not price per million tokens in isolation, but cost per successfully completed task."— VentureBeat, Gemini 3.7 Flash launch coverage, August 13, 2026

07PricingOne model, three live rates — label every price by surface.

“Gemini 3.7 Flash costs $0.75 per million tokens” is true on exactly one surface. Google’s official API lists the introductory rate through December 31, 2026, and posts a standard rate of $1.50/$7.50 from January 1, 2027 — the same rate 3.6 Flash’s standard pricing already carries, and 3.6 Flash shares the intro window too, so the cliff as posted today reads Flash-line-wide rather than specific to the new model. The widely repeated “half price” framing is only true against the Flash line’s pre-August-13 workhorse rate. Meanwhile OpenRouter layers its own time-limited 50%-off promotion on top of Google’s already-discounted intro price. Each row below names its surface.

Live pricing surfaces for Gemini 3.7 Flash compared with the rival models in Google’s benchmark table, with input and output prices per million tokens and the conditions attached to each surface.
Pricing surfaceInput $/1MOutput $/1MConditions
Gemini 3.7 Flash — three live rates for one model
Google API, introductory$0.75$3.75Through December 31, 2026. The same intro rate applies to 3.6 Flash. Context caching $0.075/1M in the same window.
Google API, standard$1.50$7.50The rate Google’s pricing page posts for the end of the intro window on January 1, 2027 — double the intro price, and the same $1.50/$7.50 that 3.6 Flash’s standard pricing already carries.
OpenRouter promo$0.375$1.875OpenRouter’s own “50% off for a limited time” banner, stacked on Google’s intro rate and expiring on OpenRouter’s schedule, not Google’s.
The rivals in Google’s table — official list rates
Claude Sonnet 5$2.00$10.00Now the permanent standard price — the increase to $3/$15 scheduled for September 1 was cancelled by Anthropic. Batch: $1.00/$5.00.
GPT-5.6 Terra, short context$2.00$12.00Standard list. Batch and Flex: $1.00/$6.00; Fast mode: $4.00/$24.00.
GPT-5.6 Terra, long context$4.00$18.00Per OpenAI’s own pricing docs, whole requests above a 272K-token threshold reprice — 2× on input, 1.5× on output — a structure Google’s and Anthropic’s flat tables don’t have.
Muse Spark 1.2$1.25$4.25As listed in Google’s own model-card pricing row — cheaper than Sonnet 5 and Terra while tying Terra on the AA Intelligence Index.

One footnote deserves its own hedge: the model card marks the Flash prices with an asterisk whose footnote text we could not retrieve from the page itself. The introductory-window reading is consistent with Google’s announcement and its official pricing page, so we treat it as the intro-price marker — but it is a presumption, not a verbatim confirmation. And on whether the January 1 doubling actually lands, there is a fresh precedent pointing the other way: Sonnet 5’s own $2/$10 rate was announced at launch as introductory pricing with a scheduled step-up to $3/$15 — and Anthropic cancelled that increase, making the intro rate permanent. Vendors have now demonstrated that a scheduled price cliff is a plan, not a promise. Budget for $1.50/$7.50 from January; don’t be surprised if the cliff moves.

08Decision MatrixRouting the workloads this table actually settles.

Read as a whole, the nineteen rows don’t crown a single model — they partition the workload space. Here is how we’d translate the full table, the independent AA read, and the surface-labelled prices into routing defaults, treating every margin of two points or less as a tie to be broken by price.

Bulk coding at price
High-volume code generation & web dev

3.7 Flash wins FrontierCode narrowly and WebDev Arena’s Elo outright while costing a fraction of either rival on the intro rate. At $0.75/$3.75 through December, the price-per-win math is hard to argue with for volume work.

Pick Gemini 3.7 Flash
Hard agentic work
Terminal agents & computer use

Google’s own table gives GPT-5.6 Terra Terminal-bench 2.1 and 3.0, DeepSWE, and OSWorld-2.0 — the exact categories in 3.7 Flash’s positioning. Mind OpenAI’s published 272K long-context repricing on big agent transcripts.

Pick GPT-5.6 Terra
Knowledge work
Judgment-heavy documents & desktop tasks

Sonnet 5 beats both named rivals on GDPVal-AA and wins Agent’s Last Exam outright, at a $2/$10 rate that is now permanent after the cancelled September increase. Its weak HLE-Verified row cuts the other way — test your own mix.

Pick Claude Sonnet 5
Don’t decide on tables
Anything mission-critical

Eight of nineteen rows are decided by two points or less, effort configs aren’t comparable across vendors, and one cell is a version mismatch. Vendor tables shortlist candidates; your own task-level evals pick the winner.

Run your own evals

The projection worth making: three-week release cycles mean this table is a snapshot, not a standings board. The durable takeaways are structural — Google is now competitive enough on coding benchmarks to force per-task price comparisons, OpenAI holds the hardest agentic rows, Anthropic holds judgment-heavy knowledge work, and every vendor’s pricing now carries dated conditions that move independently of capability. If your stack still routes every workload to one default model, that is the actual finding of this table — and it’s the kind of routing decision our AI transformation engagements start with, benchmarked on your workloads rather than the vendor’s. For the wider market context beyond these three vendors, see our broader model comparison.

09ConclusionA split decision — and an index that finally moved.

The honest ledger, August 2026

Nine wins out of nineteen is a strong workhorse, not a coronation.

Read in full, Google’s own table says something more interesting than the launch framing: Gemini 3.7 Flash wins nine of nineteen rows, loses the hardest agentic benchmarks to GPT-5.6 Terra, cedes knowledge-work rows to Claude Sonnet 5 and Muse Spark 1.2, and even loses two chart-reasoning rows to its own predecessor. VentureBeat’s independent read matches ours: “In other words, Google’s own results do not show 3.7 Flash universally displacing higher-priced competitors. They instead suggest a model that has become substantially more competitive in coding and agent workloads while occupying a lower price tier.”

The strongest claim in the release isn’t in Google’s table at all — it’s Artificial Analysis moving its Intelligence Index from 52 to 56 at matched vintage, breaking the “cheaper, not smarter” pattern we documented in July. A real capability gain, delivered three weeks after the last one, at the lowest intro price in its comparison set, matched only by 3.6 Flash: that combination is the story, and it survives every caveat this post has raised.

The caveats still matter. Prices carry surfaces and expiry dates; benchmarks carry versions, authors, and effort configs; blank cells carry explanations. The teams that win with these releases aren’t the ones that pick whichever model’s vendor published the most flattering table this week — they’re the ones with the eval harness and the routing layer to re-test the frontier every few weeks and move traffic on their own numbers.

Route models on your numbers, not vendor tables

Benchmark tables shortlist models — your workloads pick the winner.

Our team builds model-routing and evaluation layers that re-benchmark frontier releases against your actual workloads — so price cuts and benchmark jumps become routing decisions, not guesswork.

Free consultationExpert guidanceTailored solutions
What we work on

Model evaluation & routing engagements

  • Task-level eval harnesses on your own workloads
  • Multi-vendor routing — Gemini / Claude / GPT / open weights
  • Cost-per-completed-task tracking across pricing surfaces
  • Promo-cliff and price-change monitoring for AI budgets
  • Agentic workflow benchmarking before production rollout
FAQ · Gemini 3.7 Flash benchmarks

The questions teams are asking this week.

The model card publishes 19 benchmark rows across five models: Gemini 3.7 Flash, Gemini 3.6 Flash, Claude Sonnet 5, GPT-5.6 Terra, and Muse Spark 1.2. Recomputing the winners across all rows: 3.7 Flash takes 9 (including FrontierCode, WebDev Arena Elo, GDP.pdf, LVBench, GDM-MRCR, HLE-Verified, and LABBench2), GPT-5.6 Terra takes 6 (DeepSWE, both Terminal-bench versions, CharXiv without tools, OSWorld-2.0, and BioMysteryBench’s difficult subset), Sonnet 5 takes 2 outright (Agent’s Last Exam and BioMysteryBench’s human-solvable subset), 3.6 Flash takes 1 (CharXiv with tools — against its own successor), and Muse Spark 1.2 takes 1 (GDPVal-AA). Most launch coverage repeated only the five rows Google highlighted, which is why the full tally reads so differently from the headlines.
Related dispatches

Continue exploring frontier releases.