Gemini 3.6 Flash benchmarks tell two different stories depending on which lab you ask. Google’s launch materials show gains on nearly every published eval; Artificial Analysis, benchmarking the model independently ahead of release, measured its composite Intelligence Index at 50 — exactly the score of the Gemini 3.5 Flash it replaces. Both are true, and the tension between them is the most useful lens on this release.
What did change, unambiguously, is the economics. Artificial Analysis clocked average time per task at 1.3 minutes, down from 2.7 — under half — and average cost per task at $0.50, down from $0.59. Google cut the output price from $9.00 to $7.50 per million tokens, about 17% off, while holding input at $1.50. In a month when Anthropic, OpenAI, and Moonshot AI all shipped their own mid-to-upper-tier contenders, the per-task price war is the axis that actually moved.
This guide covers what shipped, the honest independent read, every comparable benchmark across GPT-5.6 Terra, Claude Sonnet 5, and Kimi K3 — including where comparison is genuinely impossible — and a recomputed per-task cost table you will not find in any vendor’s launch post. For the raw capability race at the top of the market, see our last frontier-model shootout; this post is about the workhorse tier where volume economics decide.
- 01No composite intelligence gain — by design.Artificial Analysis scored Gemini 3.6 Flash at 50 on its Intelligence Index — identical to Gemini 3.5 Flash. Individual evals moved (GDPval-AA v2 up 72 Elo), but a regression on Humanity's Last Exam offset the gain in the composite. This is an efficiency release.
- 02Per task, everything got faster and cheaper.AA measured 1.3 minutes per task (from 2.7) and $0.50 per task (from $0.59). Output pricing fell from $9.00 to $7.50 per million tokens — about 17% — while input held at $1.50. Output speed hit 304 tokens/second in AA's pre-launch testing.
- 03Rivals cost 1.33× to 2.00× more on identical token loads.Recomputed from published list prices: Claude Sonnet 5's intro price is exactly 1.33× Gemini's on both input and output; Kimi K3 is exactly 2.00×; GPT-5.6 Terra floats between 1.73× and 1.87× depending on the input/output mix.
- 04The apples-to-apples evidence is thinner than headlines imply.Only SWE-Bench Pro is published by three of the four vendors. Kimi K3 publishes no SWE-Bench Pro or OSWorld score at all, and OpenAI's OSWorld 2.0 is a different benchmark version from Google's OSWorld-Verified — most cross-model claims compare different tests.
- 05This is Google's strongest available model — with bigger silicon behind it.Gemini 3.5 Pro remains unreleased, “currently testing with partners” per DeepMind, and Google says its most ambitious pre-training run yet — Gemini 4 — is already underway. 3.6 Flash is the placeholder flagship, priced to win volume now.
01 — What ShippedOne announcement, three models, one price cut.
Google launched Gemini 3.6 Flash on July 21, 2026, in a single announcement alongside Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber. Per the DeepMind model card, the new Flash keeps the 1M-token input context and 64K-token maximum output of its predecessor, carries a March 2026 knowledge cutoff, and accepts text, image, audio, and video inputs with text-only output. It was listed on OpenRouter the same day as google/gemini-3.6-flash.
Pricing is where the release plants its flag: $1.50 per million input tokens — unchanged — and $7.50 per million output tokens, cut from $9.00. We break down the launch itself, positioning, and rollout details in the full Gemini 3.6 Flash launch breakdown; this post stays on the numbers.
Per 1M output tokens
Cut from $9.00 — roughly 17% off — while input holds at $1.50 per million. The cheaper output price compounds with lower measured output-token use per task.
Input ceiling, 64K out
Same context ceiling as Gemini 3.5 Flash. Long-context retrieval scores 91.8% on GDM-MRCR v2 at 128K, falling to 54.0% at the full 1M window.
Text, image, audio, video
Full multimodal input with text-only output, per the DeepMind model card. Same-day availability on the Gemini API and OpenRouter as google/gemini-3.6-flash.
02 — Independent VerdictThe composite score that didn’t move.
Artificial Analysis benchmarked Gemini 3.6 Flash ahead of release across intelligence, time per task, and cost per task — and its headline finding is the one Google’s launch post does not lead with: the model scores 50 on the AA Intelligence Index v4.1, matching Gemini 3.5 Flash exactly. That places it just below Muse Spark 1.1 at xhigh effort (51) and GPT-5.6 Luna at max effort (51) on AA’s ladder.
The parity is not “nothing changed.” AA’s own writeup records a genuine +72 Elo gain on GDPval-AA v2 inside the composite — offset by a 3-point regression on Humanity’s Last Exam. Net composite movement: zero. What did move, decisively, is throughput and cost.
Gemini 3.6 Flash vs 3.5 Flash · AA independent measurement
Source: Artificial Analysis, Jul 21, 2026 — independent pre-release benchmarkingThe task-time cut is the sharpest single number in the release: 1.3 minutes average per task against 2.7 for the predecessor, driven by higher token efficiency and faster output — AA measured 304 tokens per second in pre-launch testing. Cost per task fell from $0.59 to $0.50, roughly 15% less on AA’s published dollar figures. Google’s own framing leans on the same source, claiming 17% fewer output tokens consumed on the Artificial Analysis Index — a vendor-stated figure that at least has the virtue of citing an independent measurement.
Our read: this is the clearest example yet of a trend we have been tracking all year — the mid-tier model race has quietly stopped being about composite intelligence scores and started being about unit economics. When an independent lab says your new model is exactly as smart as your old one and the launch still lands as a win, the market has told you what it is actually buying: tasks per dollar, not points per benchmark.
03 — Vendor BenchmarksGoogle’s published deltas, eval by eval.
Google’s own launch numbers are real improvements on the evals it chose to publish. The largest jumps cluster around agentic and applied work: DeepSWE v1.1 rose from 37% to 49%, MLE-Bench from 49.7% to 63.9%, and GDPval-AA v2 from 1,349 to 1,421 Elo — a +72-point delta that Artificial Analysis independently measured at the same magnitude. OSWorld-Verified computer use edged up from 78.4% to 83.0%, SWE-Bench Pro landed at 58.7%, and Terminal-Bench 2.1 at 78.0%.
Up from 37% on 3.5 Flash
A 12-point generational jump on Google's agentic software-engineering eval — the largest single-benchmark gain in the launch table.
Up from 49.7% on 3.5 Flash
Machine-learning engineering tasks, up 14.2 points. Notably, MLE-Bench is a Google-published metric no rival in this comparison reports at all.
Up from 1,349 on 3.5 Flash
Economically valuable knowledge work. The +72-Elo delta is the one vendor claim in this launch independently confirmed by Artificial Analysis's own measurement.
The community-arena picture is more humbling — and needs careful reading, because two different leaderboards produced the same rank number. As reported via Arena.ai’s rankings, Gemini 3.6 Flash entered the LMArena Text Arena at #12 with a score around 1,485 — behind Gemini 3.1 Pro, Gemini 3 Pro, GPT-5.6 Sol, and several Claude models, with Claude Fable 5 leading the board at 1,507. Separately, it also ranked #12 in Arena.ai’s Frontend Code Arena at 1,537 points — a genuine climb from Gemini 3.5 Flash’s #21 in that same category. Same rank, different boards, different meanings: on open text preference it sits mid-pack; on frontend coding it jumped nine places generation-over-generation.
Long context deserves its own caveat: 91.8% on GDM-MRCR v2 at 128K tokens degrades to 54.0% at the full 1M window. The million-token context is real, but retrieval quality at the far end of it is roughly a coin flip — plan pipelines accordingly.
04 — Original AnalysisThe benchmark coverage matrix: where comparison is even possible.
Here is the problem with most “Gemini vs GPT vs Claude vs Kimi” coverage this week: the four vendors publish four different baskets of evals. We cross-referenced each vendor’s own launch materials — Google’s model card, OpenAI’s GPT-5.6 GA tables, Anthropic’s Sonnet 5 announcement, and Moonshot AI’s Kimi K3 post — and mapped exactly which scores exist. The overlap is thin enough that the matrix itself is the finding.
| Benchmark | Gemini 3.6 Flash | GPT-5.6 Terra | Claude Sonnet 5 | Kimi K3 |
|---|---|---|---|---|
| Agentic coding suites | ||||
| SWE-Bench Pro | 58.7% | 63.4% | 63.2% | not published |
| Terminal-Bench 2.1 | 78.0% | 87.4% — aggregator-reported only; not in OpenAI’s GA tables | 80.4% — aggregator-reported; benchmark version unconfirmed | 88.3% (max effort) |
| DeepSWE | 49% (v1.1) | 69.6% (v1.1) | not published | 67.5% |
| Computer use · two different benchmark versions — never compare across this divider | ||||
| OSWorld-Verified | 83.0% | not published | 81.2% — aggregator-reported | not published |
| OSWorld 2.0 | not published | 50.2% | not published | not published |
| Knowledge work & ML engineering | ||||
| GDPval-AA v2 (Elo) | 1,421 | 1,593 | not published | 1,668 (max effort) |
| MLE-Bench | 63.9% | not published | not published | not published |
Read the matrix column by column and the honest comparisons shrink fast. SWE-Bench Pro is the only suite with three directly comparable primary-source scores: Terra 63.4%, Sonnet 5 63.2%, Gemini 3.6 Flash 58.7% — a 4.7-point gap between Gemini and the leader, with the two rivals effectively tied. GDPval-AA v2 gives three comparable Elo scores, and there Kimi K3 leads outright at 1,668 versus Terra’s 1,593 and Gemini’s 1,421 — 247 Elo above Gemini. On Terminal-Bench 2.1, only Gemini’s 78.0% and K3’s 88.3% are vendor-published; the Terra and Sonnet figures circulate via third-party aggregators and are not confirmed in either vendor’s primary text. Everything else is a single-vendor metric.
05 — Proprietary AnalysisPer-task cost, recomputed from list prices.
Every outlet repeated Google’s per-token price cut; none computed what it means at realistic token loads against the three rivals. So we did. The table below applies each model’s published list price — Gemini 3.6 Flash at $1.50 / $7.50, GPT-5.6 Terra at $2.50 / $15.00, Claude Sonnet 5 at its $2.00 / $10.00 intro rate, Kimi K3 at its $3.00 / $15.00 cache-miss rate — to three task profiles. The formula per cell is simply input tokens × input price plus output tokens × output price, at per-million list rates. No benchmark weighting, no vendor efficiency claims folded in.
| Scenario | Gemini 3.6 Flash | GPT-5.6 Terra | Sonnet 5 (intro) | Kimi K3 (cache-miss) |
|---|---|---|---|---|
| Cost per task = input tokens × input price + output tokens × output price · list prices per 1M tokens | ||||
| Quick agentic turn · 2,000 in / 500 out | $0.00675 | $0.0125 (1.85×) | $0.0090 (1.33×) | $0.0135 (2.00×) |
| Coding agent turn · 10,000 in / 3,000 out | $0.0375 | $0.0700 (1.87×) | $0.0500 (1.33×) | $0.0750 (2.00×) |
| Long-context research pull · 100,000 in / 5,000 out | $0.1875 | $0.3250 (1.73×) | $0.2500 (1.33×) | $0.3750 (2.00×) |
Two clean patterns fall out of the arithmetic that no vendor states directly. First, Sonnet 5’s intro price is exactly 1.33× Gemini’s on both dimensions — $2.00 over $1.50 and $10.00 over $7.50 both equal 1.33 — so the multiple never moves regardless of the token mix. Second, Kimi K3 is exactly 2.00× on both dimensions ($3.00 / $1.50 and $15.00 / $7.50), equally scenario-invariant. GPT-5.6 Terra is the only rival whose multiple shifts with the workload — 1.73× to 1.87× across our scenarios — because its input ratio to Gemini (1.67×) differs from its output ratio (2.00×): input-heavy jobs narrow Terra’s gap, output-heavy jobs widen it.
Across the board, the three rivals run 33% to 100% more per task than Gemini 3.6 Flash on identical token loads. One expiry date matters: Sonnet 5’s intro pricing runs through August 31, 2026 — at its standard $3.00 / $15.00 rate, its multiple becomes exactly 2.00×, identical to Kimi K3’s. And one asterisk favors K3: its $0.30 cache-hit input rate is far below anyone here, so cache-friendly, repetitive-context agents can land well under its cache-miss column.
06 — The FieldThe three rivals, in their own numbers.
Each competitor arrives with a different pitch. OpenAI’s GPT-5.6 reached general availability on July 9 with the fullest published eval table of the four — Terra’s SWE-Bench Pro 63.4%, DeepSWE v1.1 69.6%, and an Artificial Analysis Coding Agent Index of 77.4 all come from OpenAI’s own GA materials. Anthropic shipped Sonnet 5 on June 30 with the claim that “Sonnet 5 narrows the gap: its performance is close to that of Opus 4.8, but at lower prices” — its 63.2% SWE-Bench Pro sits 6 points under Opus 4.8’s 69.2%. Moonshot AI’s Kimi K3, launched July 17, is a 2.8-trillion parameter Stable LatentMoE design activating 16 of 896 experts per forward pass, and it posts the strongest agentic numbers in this field — where it publishes them at all.
GPT-5.6 Terra
The fullest primary-source eval table: SWE-Bench Pro 63.4%, DeepSWE v1.1 69.6%, BrowseComp 87.5%, GDPval-AA v2 1,593 Elo, OSWorld 2.0 50.2%. Batch and flex tiers halve the list price.
Claude Sonnet 5
Positioned as the most agentic Sonnet yet — plans, browsers, terminals, autonomous runs. SWE-Bench Pro 63.2% vs Opus 4.8's 69.2%. Computer-use scores circulate only via aggregators, not Anthropic's own text.
Kimi K3
2.8T params, 16 of 896 experts active, flat pricing across the full 1M window. Terminal-Bench 2.1 88.3%, DeepSWE 67.5%, GDPval-AA v2 1,668 Elo — all at max effort; no low/medium tier at launch.
K3 deserves special mention for candor: its own launch post concedes real weaknesses — thinking-history sensitivity, excessive proactiveness, and a user-experience gap against the top closed models — even while posting the highest GDPval Elo in this field. If you are evaluating it hands-on, start with our hands-on Kimi K3 setup guide; for the broader open-weight picture, how open-weight models stack up against Claude Opus covers the tier above.
“Despite being a highly competitive model overall, K3 nonetheless exhibits a noticeable gap in user experience compared with Claude Fable 5 and GPT 5.6 Sol.”— Moonshot AI, Kimi K3 launch post, July 17, 2026
07 — Decision MatrixWhich model for which workload.
Fold the coverage matrix and the cost table together and the routing logic writes itself. The honest version is per-workload, not per-headline — and it changes on August 31 when Sonnet 5’s intro pricing lapses.
Bulk agentic turns, cost-dominated
Gemini 3.6 Flash is the cheapest per task in every scenario we computed — rivals run 1.33× to 2.00× its list price — and AA measured under half the task time of its predecessor at 304 tokens/second. When volume dominates, per-task price wins.
SWE-Bench-Pro-shaped engineering
On the one suite with three primary-source scores, Terra (63.4%) and Sonnet 5 (63.2%) sit 4.7 points above Gemini (58.7%). For the toughest agentic coding, pay the multiple — and note Terra's DeepSWE 69.6% is the strongest published score on that suite.
GDPval-shaped professional tasks
Kimi K3 leads the only three-way comparable Elo table at 1,668 vs Terra's 1,593 and Gemini's 1,421 — and its $0.30 cache-hit input rate rewards repetitive-context designs. Weigh its self-declared UX gaps against the score.
Full-window 1M-token pulls
Gemini's 54.0% MRCR at 1M says the far end of the window is unreliable; its 91.8% at 128K is strong. Chunk to 128K where you can, and benchmark rivals on your own corpus — no vendor here publishes a comparable 1M retrieval score.
The meta-lesson for teams building on this tier: route by task class, re-run the arithmetic whenever a price changes, and treat vendor benchmark tables as marketing collateral until an independent lab or your own eval harness confirms them. This is exactly the kind of model-routing and cost-governance work our AI transformation engagements operationalize — the per-task cost table above is the first artifact we build for any client running agents at volume.
08 — What’s NextA placeholder flagship, with Gemini 4 in the oven.
The strangest fact about this launch is that Gemini 3.6 Flash is Google’s strongest available model. Gemini 3.5 Pro remains unreleased — “currently testing with partners,” per DeepMind’s comments reported by 9to5Google — which means the workhorse tier is, for now, the flagship tier. Never benchmark 3.5 Pro as a current competitor; it is not shipping.
And Google is already talking past it. In the same announcement cycle, DeepMind confirmed that pre-training for the next major generation has begun — with no release date and, notably, no capability claims attached.
“We have already started our most ambitious pre-training run yet, for Gemini 4, and can't wait to share more.”— Google DeepMind, July 21, 2026
Projecting forward: expect the per-task framing to harden into the default way this tier is sold. Artificial Analysis’s Intelligence Index composite now spans nine evaluations, and every vendor in this comparison publishes a different subset of single benchmarks around it — a structural incentive to market whichever basket flatters the release. If the July pattern holds — four vendors repricing the same tier inside three weeks, with Sonnet 5’s intro price expiring August 31 and K3’s cache economics rewarding specific architectures — the durable advantage goes to teams whose routing layer can re-price weekly, not to teams that picked a single vendor in July and locked in. Efficiency releases like this one make that discipline pay compounding returns.
09 — ConclusionAn efficiency release, priced to win volume.
The mid-tier race is now about tasks per dollar, not points per benchmark.
Gemini 3.6 Flash is the clearest efficiency release of the year: an independent composite score that did not move a single point, wrapped around a task loop that got twice as fast and meaningfully cheaper. Google’s published benchmark gains are real on the evals it chose; the honest comparison set against Terra, Sonnet 5, and Kimi K3 is far thinner than the week’s headlines implied.
The recomputed arithmetic is the durable takeaway: Sonnet 5 at a fixed 1.33× Gemini’s list price until August 31 and 2.00× after, Kimi K3 at a fixed 2.00× with a cache-hit escape hatch, Terra floating between 1.73× and 1.87× depending on your token mix. Those multiples — not any single benchmark score — are what a production routing decision at volume actually turns on.
Run your own evals on your own prompts, price your real token mixes, and revisit on September 1 when the Sonnet intro pricing lapses. In a tier where the smartest independent measurement says capability is flat, the spreadsheet is the benchmark.