AI DevelopmentDecision Matrix11 min readPublished July 22, 2026

Same intelligence score · under half the task time · rivals at 1.33×–2.00× the list price

Gemini 3.6 Flash Benchmarks: The Per-Task Price Shake-Up

Gemini 3.6 Flash launched July 21, 2026 with a stack of vendor-published benchmark gains — and an independent composite score that did not move at all. The real story is per-task economics: roughly half the task time, a cheaper output price, and list-price arithmetic that puts Sonnet 5, GPT-5.6 Terra, and Kimi K3 at 1.33× to 2.00× its cost on identical token loads.

DA
Digital Applied Team
Senior strategists · Published Jul 22, 2026
PublishedJul 22, 2026
Read time11 min
Sources10 primary & independent
AA Intelligence Index
50
identical to Gemini 3.5 Flash
±0 vs 3.5 Flash
AA time per task
1.3min
down from 2.7 min
−52 vs 3.5 Flash
AA cost per task
$0.50
down from $0.59
−15 vs 3.5 Flash
Output price / 1M
$7.50
down from $9.00
−17 vs 3.5 Flash

Gemini 3.6 Flash benchmarks tell two different stories depending on which lab you ask. Google’s launch materials show gains on nearly every published eval; Artificial Analysis, benchmarking the model independently ahead of release, measured its composite Intelligence Index at 50 — exactly the score of the Gemini 3.5 Flash it replaces. Both are true, and the tension between them is the most useful lens on this release.

What did change, unambiguously, is the economics. Artificial Analysis clocked average time per task at 1.3 minutes, down from 2.7 — under half — and average cost per task at $0.50, down from $0.59. Google cut the output price from $9.00 to $7.50 per million tokens, about 17% off, while holding input at $1.50. In a month when Anthropic, OpenAI, and Moonshot AI all shipped their own mid-to-upper-tier contenders, the per-task price war is the axis that actually moved.

This guide covers what shipped, the honest independent read, every comparable benchmark across GPT-5.6 Terra, Claude Sonnet 5, and Kimi K3 — including where comparison is genuinely impossible — and a recomputed per-task cost table you will not find in any vendor’s launch post. For the raw capability race at the top of the market, see our last frontier-model shootout; this post is about the workhorse tier where volume economics decide.

Key takeaways
  1. 01
    No composite intelligence gain — by design.Artificial Analysis scored Gemini 3.6 Flash at 50 on its Intelligence Index — identical to Gemini 3.5 Flash. Individual evals moved (GDPval-AA v2 up 72 Elo), but a regression on Humanity's Last Exam offset the gain in the composite. This is an efficiency release.
  2. 02
    Per task, everything got faster and cheaper.AA measured 1.3 minutes per task (from 2.7) and $0.50 per task (from $0.59). Output pricing fell from $9.00 to $7.50 per million tokens — about 17% — while input held at $1.50. Output speed hit 304 tokens/second in AA's pre-launch testing.
  3. 03
    Rivals cost 1.33× to 2.00× more on identical token loads.Recomputed from published list prices: Claude Sonnet 5's intro price is exactly 1.33× Gemini's on both input and output; Kimi K3 is exactly 2.00×; GPT-5.6 Terra floats between 1.73× and 1.87× depending on the input/output mix.
  4. 04
    The apples-to-apples evidence is thinner than headlines imply.Only SWE-Bench Pro is published by three of the four vendors. Kimi K3 publishes no SWE-Bench Pro or OSWorld score at all, and OpenAI's OSWorld 2.0 is a different benchmark version from Google's OSWorld-Verified — most cross-model claims compare different tests.
  5. 05
    This is Google's strongest available model — with bigger silicon behind it.Gemini 3.5 Pro remains unreleased, “currently testing with partners” per DeepMind, and Google says its most ambitious pre-training run yet — Gemini 4 — is already underway. 3.6 Flash is the placeholder flagship, priced to win volume now.

01What ShippedOne announcement, three models, one price cut.

Google launched Gemini 3.6 Flash on July 21, 2026, in a single announcement alongside Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber. Per the DeepMind model card, the new Flash keeps the 1M-token input context and 64K-token maximum output of its predecessor, carries a March 2026 knowledge cutoff, and accepts text, image, audio, and video inputs with text-only output. It was listed on OpenRouter the same day as google/gemini-3.6-flash.

Pricing is where the release plants its flag: $1.50 per million input tokens — unchanged — and $7.50 per million output tokens, cut from $9.00. We break down the launch itself, positioning, and rollout details in the full Gemini 3.6 Flash launch breakdown; this post stays on the numbers.

Output price
Per 1M output tokens
$7.50

Cut from $9.00 — roughly 17% off — while input holds at $1.50 per million. The cheaper output price compounds with lower measured output-token use per task.

Input: $1.50 / 1M
Context window
Input ceiling, 64K out
1Mtokens

Same context ceiling as Gemini 3.5 Flash. Long-context retrieval scores 91.8% on GDM-MRCR v2 at 128K, falling to 54.0% at the full 1M window.

Cutoff: March 2026
Modalities
Text, image, audio, video
4in

Full multimodal input with text-only output, per the DeepMind model card. Same-day availability on the Gemini API and OpenRouter as google/gemini-3.6-flash.

Output: text only
Release snapshot
Gemini 3.6 Flash shipped July 21, 2026 at $1.50 / $7.50 per million tokens (input / output), 1M-token context, March 2026 cutoff. It launched into a crowded three-week window: Claude Sonnet 5 arrived June 30, GPT-5.6 reached GA July 9, and Kimi K3 landed July 17 — four vendors repricing the same workhorse tier in under a month.

02Independent VerdictThe composite score that didn’t move.

Artificial Analysis benchmarked Gemini 3.6 Flash ahead of release across intelligence, time per task, and cost per task — and its headline finding is the one Google’s launch post does not lead with: the model scores 50 on the AA Intelligence Index v4.1, matching Gemini 3.5 Flash exactly. That places it just below Muse Spark 1.1 at xhigh effort (51) and GPT-5.6 Luna at max effort (51) on AA’s ladder.

The parity is not “nothing changed.” AA’s own writeup records a genuine +72 Elo gain on GDPval-AA v2 inside the composite — offset by a 3-point regression on Humanity’s Last Exam. Net composite movement: zero. What did move, decisively, is throughput and cost.

Gemini 3.6 Flash vs 3.5 Flash · AA independent measurement

Source: Artificial Analysis, Jul 21, 2026 — independent pre-release benchmarking
Intelligence Index — 3.5 FlashAA composite v4.1, 9 evaluations
50
Intelligence Index — 3.6 Flashidentical composite score
50
Time per task — 3.5 FlashAA-measured average
2.7 min
Time per task — 3.6 Flashunder half the predecessor
1.3 min
Cost per task — 3.5 FlashAA-measured average
$0.59
Cost per task — 3.6 Flashlower output-token use + cheaper output price
$0.50

The task-time cut is the sharpest single number in the release: 1.3 minutes average per task against 2.7 for the predecessor, driven by higher token efficiency and faster output — AA measured 304 tokens per second in pre-launch testing. Cost per task fell from $0.59 to $0.50, roughly 15% less on AA’s published dollar figures. Google’s own framing leans on the same source, claiming 17% fewer output tokens consumed on the Artificial Analysis Index — a vendor-stated figure that at least has the virtue of citing an independent measurement.

Our read: this is the clearest example yet of a trend we have been tracking all year — the mid-tier model race has quietly stopped being about composite intelligence scores and started being about unit economics. When an independent lab says your new model is exactly as smart as your old one and the launch still lands as a win, the market has told you what it is actually buying: tasks per dollar, not points per benchmark.

03Vendor BenchmarksGoogle’s published deltas, eval by eval.

Google’s own launch numbers are real improvements on the evals it chose to publish. The largest jumps cluster around agentic and applied work: DeepSWE v1.1 rose from 37% to 49%, MLE-Bench from 49.7% to 63.9%, and GDPval-AA v2 from 1,349 to 1,421 Elo — a +72-point delta that Artificial Analysis independently measured at the same magnitude. OSWorld-Verified computer use edged up from 78.4% to 83.0%, SWE-Bench Pro landed at 58.7%, and Terminal-Bench 2.1 at 78.0%.

DeepSWE v1.1
Up from 37% on 3.5 Flash
49%

A 12-point generational jump on Google's agentic software-engineering eval — the largest single-benchmark gain in the launch table.

+12 pts vs 3.5 Flash
MLE-Bench
Up from 49.7% on 3.5 Flash
63.9%

Machine-learning engineering tasks, up 14.2 points. Notably, MLE-Bench is a Google-published metric no rival in this comparison reports at all.

+14.2 pts vs 3.5 Flash
GDPval-AA v2
Up from 1,349 on 3.5 Flash
1,421Elo

Economically valuable knowledge work. The +72-Elo delta is the one vendor claim in this launch independently confirmed by Artificial Analysis's own measurement.

+72 Elo, AA-confirmed

The community-arena picture is more humbling — and needs careful reading, because two different leaderboards produced the same rank number. As reported via Arena.ai’s rankings, Gemini 3.6 Flash entered the LMArena Text Arena at #12 with a score around 1,485 — behind Gemini 3.1 Pro, Gemini 3 Pro, GPT-5.6 Sol, and several Claude models, with Claude Fable 5 leading the board at 1,507. Separately, it also ranked #12 in Arena.ai’s Frontend Code Arena at 1,537 points — a genuine climb from Gemini 3.5 Flash’s #21 in that same category. Same rank, different boards, different meanings: on open text preference it sits mid-pack; on frontend coding it jumped nine places generation-over-generation.

Long context deserves its own caveat: 91.8% on GDM-MRCR v2 at 128K tokens degrades to 54.0% at the full 1M window. The million-token context is real, but retrieval quality at the far end of it is roughly a coin flip — plan pipelines accordingly.

04Original AnalysisThe benchmark coverage matrix: where comparison is even possible.

Here is the problem with most “Gemini vs GPT vs Claude vs Kimi” coverage this week: the four vendors publish four different baskets of evals. We cross-referenced each vendor’s own launch materials — Google’s model card, OpenAI’s GPT-5.6 GA tables, Anthropic’s Sonnet 5 announcement, and Moonshot AI’s Kimi K3 post — and mapped exactly which scores exist. The overlap is thin enough that the matrix itself is the finding.

Benchmark coverage matrix comparing Gemini 3.6 Flash, GPT-5.6 Terra, Claude Sonnet 5, and Kimi K3 across seven benchmark suites, marking which vendors publish each score and which figures are only aggregator-reported.
BenchmarkGemini 3.6 FlashGPT-5.6 TerraClaude Sonnet 5Kimi K3
Agentic coding suites
SWE-Bench Pro58.7%63.4%63.2%not published
Terminal-Bench 2.178.0%87.4% — aggregator-reported only; not in OpenAI’s GA tables80.4% — aggregator-reported; benchmark version unconfirmed88.3% (max effort)
DeepSWE49% (v1.1)69.6% (v1.1)not published67.5%
Computer use · two different benchmark versions — never compare across this divider
OSWorld-Verified83.0%not published81.2% — aggregator-reportednot published
OSWorld 2.0not published50.2%not publishednot published
Knowledge work & ML engineering
GDPval-AA v2 (Elo)1,4211,593not published1,668 (max effort)
MLE-Bench63.9%not publishednot publishednot published

Read the matrix column by column and the honest comparisons shrink fast. SWE-Bench Pro is the only suite with three directly comparable primary-source scores: Terra 63.4%, Sonnet 5 63.2%, Gemini 3.6 Flash 58.7% — a 4.7-point gap between Gemini and the leader, with the two rivals effectively tied. GDPval-AA v2 gives three comparable Elo scores, and there Kimi K3 leads outright at 1,668 versus Terra’s 1,593 and Gemini’s 1,421 — 247 Elo above Gemini. On Terminal-Bench 2.1, only Gemini’s 78.0% and K3’s 88.3% are vendor-published; the Terra and Sonnet figures circulate via third-party aggregators and are not confirmed in either vendor’s primary text. Everything else is a single-vendor metric.

The OSWorld trap
OpenAI’s OSWorld 2.0 and Google’s OSWorld-Verified share a name but are different benchmark versions — Terra’s 50.2% and Gemini’s 83.0% are not scores on the same test, and any coverage placing them in one column is comparing incompatible evals. Kimi K3 publishes no SWE-Bench Pro or OSWorld score at all: its coding claims run through Program Bench, SWE Marathon, and DeepSWE. Where a number does not exist, the only honest cell is “not published.”

05Proprietary AnalysisPer-task cost, recomputed from list prices.

Every outlet repeated Google’s per-token price cut; none computed what it means at realistic token loads against the three rivals. So we did. The table below applies each model’s published list price — Gemini 3.6 Flash at $1.50 / $7.50, GPT-5.6 Terra at $2.50 / $15.00, Claude Sonnet 5 at its $2.00 / $10.00 intro rate, Kimi K3 at its $3.00 / $15.00 cache-miss rate — to three task profiles. The formula per cell is simply input tokens × input price plus output tokens × output price, at per-million list rates. No benchmark weighting, no vendor efficiency claims folded in.

Per-task cost comparison recomputed from published list prices for Gemini 3.6 Flash, GPT-5.6 Terra, Claude Sonnet 5 intro pricing, and Kimi K3 cache-miss pricing across three token-load scenarios, with each rival’s cost multiple versus Gemini.
ScenarioGemini 3.6 FlashGPT-5.6 TerraSonnet 5 (intro)Kimi K3 (cache-miss)
Cost per task = input tokens × input price + output tokens × output price · list prices per 1M tokens
Quick agentic turn · 2,000 in / 500 out$0.00675$0.0125 (1.85×)$0.0090 (1.33×)$0.0135 (2.00×)
Coding agent turn · 10,000 in / 3,000 out$0.0375$0.0700 (1.87×)$0.0500 (1.33×)$0.0750 (2.00×)
Long-context research pull · 100,000 in / 5,000 out$0.1875$0.3250 (1.73×)$0.2500 (1.33×)$0.3750 (2.00×)

Two clean patterns fall out of the arithmetic that no vendor states directly. First, Sonnet 5’s intro price is exactly 1.33× Gemini’s on both dimensions — $2.00 over $1.50 and $10.00 over $7.50 both equal 1.33 — so the multiple never moves regardless of the token mix. Second, Kimi K3 is exactly 2.00× on both dimensions ($3.00 / $1.50 and $15.00 / $7.50), equally scenario-invariant. GPT-5.6 Terra is the only rival whose multiple shifts with the workload — 1.73× to 1.87× across our scenarios — because its input ratio to Gemini (1.67×) differs from its output ratio (2.00×): input-heavy jobs narrow Terra’s gap, output-heavy jobs widen it.

Across the board, the three rivals run 33% to 100% more per task than Gemini 3.6 Flash on identical token loads. One expiry date matters: Sonnet 5’s intro pricing runs through August 31, 2026 — at its standard $3.00 / $15.00 rate, its multiple becomes exactly 2.00×, identical to Kimi K3’s. And one asterisk favors K3: its $0.30 cache-hit input rate is far below anyone here, so cache-friendly, repetitive-context agents can land well under its cache-miss column.

What this table is — and isn’t
This is list-price arithmetic, not a benchmark-weighted comparison — it says nothing about which model completes a task in fewer tokens, or completes it at all. Google frames the release as cheaper per task than rival mid-tier models, but published no methodology or comparison table for that claim — treat it as marketing framing. Our numbers above are reproducible from the four vendors’ public price pages alone.

06The FieldThe three rivals, in their own numbers.

Each competitor arrives with a different pitch. OpenAI’s GPT-5.6 reached general availability on July 9 with the fullest published eval table of the four — Terra’s SWE-Bench Pro 63.4%, DeepSWE v1.1 69.6%, and an Artificial Analysis Coding Agent Index of 77.4 all come from OpenAI’s own GA materials. Anthropic shipped Sonnet 5 on June 30 with the claim that “Sonnet 5 narrows the gap: its performance is close to that of Opus 4.8, but at lower prices” — its 63.2% SWE-Bench Pro sits 6 points under Opus 4.8’s 69.2%. Moonshot AI’s Kimi K3, launched July 17, is a 2.8-trillion parameter Stable LatentMoE design activating 16 of 896 experts per forward pass, and it posts the strongest agentic numbers in this field — where it publishes them at all.

OpenAI · GA Jul 9
GPT-5.6 Terra
$2.50 / $15.00 per 1M · $0.25 cached input

The fullest primary-source eval table: SWE-Bench Pro 63.4%, DeepSWE v1.1 69.6%, BrowseComp 87.5%, GDPval-AA v2 1,593 Elo, OSWorld 2.0 50.2%. Batch and flex tiers halve the list price.

openai.com/index/gpt-5-6
Anthropic · Jun 30
Claude Sonnet 5
$2 / $10 intro → $3 / $15 after Aug 31

Positioned as the most agentic Sonnet yet — plans, browsers, terminals, autonomous runs. SWE-Bench Pro 63.2% vs Opus 4.8's 69.2%. Computer-use scores circulate only via aggregators, not Anthropic's own text.

anthropic.com/news/claude-sonnet-5
Moonshot AI · Jul 17
Kimi K3
$3.00 / $15.00 per 1M · $0.30 cache-hit input

2.8T params, 16 of 896 experts active, flat pricing across the full 1M window. Terminal-Bench 2.1 88.3%, DeepSWE 67.5%, GDPval-AA v2 1,668 Elo — all at max effort; no low/medium tier at launch.

kimi.com/blog/kimi-k3

K3 deserves special mention for candor: its own launch post concedes real weaknesses — thinking-history sensitivity, excessive proactiveness, and a user-experience gap against the top closed models — even while posting the highest GDPval Elo in this field. If you are evaluating it hands-on, start with our hands-on Kimi K3 setup guide; for the broader open-weight picture, how open-weight models stack up against Claude Opus covers the tier above.

“Despite being a highly competitive model overall, K3 nonetheless exhibits a noticeable gap in user experience compared with Claude Fable 5 and GPT 5.6 Sol.”— Moonshot AI, Kimi K3 launch post, July 17, 2026

07Decision MatrixWhich model for which workload.

Fold the coverage matrix and the cost table together and the routing logic writes itself. The honest version is per-workload, not per-headline — and it changes on August 31 when Sonnet 5’s intro pricing lapses.

High-volume agent fleets
Bulk agentic turns, cost-dominated

Gemini 3.6 Flash is the cheapest per task in every scenario we computed — rivals run 1.33× to 2.00× its list price — and AA measured under half the task time of its predecessor at 304 tokens/second. When volume dominates, per-task price wins.

Pick Gemini 3.6 Flash
Hardest coding work
SWE-Bench-Pro-shaped engineering

On the one suite with three primary-source scores, Terra (63.4%) and Sonnet 5 (63.2%) sit 4.7 points above Gemini (58.7%). For the toughest agentic coding, pay the multiple — and note Terra's DeepSWE 69.6% is the strongest published score on that suite.

Pick Terra or Sonnet 5
Knowledge-work agents
GDPval-shaped professional tasks

Kimi K3 leads the only three-way comparable Elo table at 1,668 vs Terra's 1,593 and Gemini's 1,421 — and its $0.30 cache-hit input rate rewards repetitive-context designs. Weigh its self-declared UX gaps against the score.

Pick Kimi K3 — eyes open
Long-context retrieval
Full-window 1M-token pulls

Gemini's 54.0% MRCR at 1M says the far end of the window is unreliable; its 91.8% at 128K is strong. Chunk to 128K where you can, and benchmark rivals on your own corpus — no vendor here publishes a comparable 1M retrieval score.

Test on your corpus

The meta-lesson for teams building on this tier: route by task class, re-run the arithmetic whenever a price changes, and treat vendor benchmark tables as marketing collateral until an independent lab or your own eval harness confirms them. This is exactly the kind of model-routing and cost-governance work our AI transformation engagements operationalize — the per-task cost table above is the first artifact we build for any client running agents at volume.

08What’s NextA placeholder flagship, with Gemini 4 in the oven.

The strangest fact about this launch is that Gemini 3.6 Flash is Google’s strongest available model. Gemini 3.5 Pro remains unreleased — “currently testing with partners,” per DeepMind’s comments reported by 9to5Google — which means the workhorse tier is, for now, the flagship tier. Never benchmark 3.5 Pro as a current competitor; it is not shipping.

And Google is already talking past it. In the same announcement cycle, DeepMind confirmed that pre-training for the next major generation has begun — with no release date and, notably, no capability claims attached.

“We have already started our most ambitious pre-training run yet, for Gemini 4, and can't wait to share more.”— Google DeepMind, July 21, 2026

Projecting forward: expect the per-task framing to harden into the default way this tier is sold. Artificial Analysis’s Intelligence Index composite now spans nine evaluations, and every vendor in this comparison publishes a different subset of single benchmarks around it — a structural incentive to market whichever basket flatters the release. If the July pattern holds — four vendors repricing the same tier inside three weeks, with Sonnet 5’s intro price expiring August 31 and K3’s cache economics rewarding specific architectures — the durable advantage goes to teams whose routing layer can re-price weekly, not to teams that picked a single vendor in July and locked in. Efficiency releases like this one make that discipline pay compounding returns.

09ConclusionAn efficiency release, priced to win volume.

The workhorse tier, July 2026

The mid-tier race is now about tasks per dollar, not points per benchmark.

Gemini 3.6 Flash is the clearest efficiency release of the year: an independent composite score that did not move a single point, wrapped around a task loop that got twice as fast and meaningfully cheaper. Google’s published benchmark gains are real on the evals it chose; the honest comparison set against Terra, Sonnet 5, and Kimi K3 is far thinner than the week’s headlines implied.

The recomputed arithmetic is the durable takeaway: Sonnet 5 at a fixed 1.33× Gemini’s list price until August 31 and 2.00× after, Kimi K3 at a fixed 2.00× with a cache-hit escape hatch, Terra floating between 1.73× and 1.87× depending on your token mix. Those multiples — not any single benchmark score — are what a production routing decision at volume actually turns on.

Run your own evals on your own prompts, price your real token mixes, and revisit on September 1 when the Sonnet intro pricing lapses. In a tier where the smartest independent measurement says capability is flat, the spreadsheet is the benchmark.

Route models by task, not by headline

When capability is flat, per-task cost is the whole game.

Our team builds model-routing and cost-governance layers for businesses running AI agents at volume — benchmarking Gemini, GPT, Claude, and open-weight models on your actual workloads, delivered in days not quarters.

Free consultationExpert guidanceTailored solutions
What we work on

Model economics engagements

  • Per-task cost modeling across your real token mixes
  • Eval harnesses on your prompts — not vendor benchmarks
  • Multi-vendor routing — Gemini / GPT / Claude / open weights
  • Cache-architecture design for agent fleets
  • Quarterly repricing reviews as vendor lists shift
FAQ · Gemini 3.6 Flash benchmarks

The questions we get every week.

Gemini 3.6 Flash is Google's new mid-tier workhorse model, launched July 21, 2026 in a single announcement alongside Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber. Per the DeepMind model card it keeps a 1M-token input context with 64K max output, carries a March 2026 knowledge cutoff, and accepts text, image, audio, and video inputs with text-only output. Pricing is $1.50 per million input tokens (unchanged from 3.5 Flash) and $7.50 per million output tokens, cut from $9.00. It was listed on OpenRouter the same day as google/gemini-3.6-flash, and because Gemini 3.5 Pro is still unreleased, it is currently Google's strongest available model.
Related dispatches

Continue exploring model economics.