AI DevelopmentNew Release11 min readPublished July 22, 2026

Three Flash models in one drop · $7.50/M output · flat intelligence, half the time per task

Gemini 3.6 Flash: Google Ships Its New Workhorse Model

Google shipped Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber in a single July 21 drop — and conspicuously not Gemini 3.5 Pro. The headline model cuts output pricing from $9.00 to $7.50 per million tokens and, per Google, uses roughly 17% fewer output tokens per task. On the independent Artificial Analysis Intelligence Index, though, it scores a flat 50 — cheaper, not smarter.

DA
Digital Applied Team
Senior strategists · Published July 22, 2026
PublishedJuly 22, 2026
Read time11 min
SourcesGoogle, AA + 6 outlets
Output price / M tokens
$7.50
was $9.00 on 3.5 Flash
−16.7% list
Time per task (AA)
1.3min
down from 2.7 min on 3.5 Flash
−52%
AA Intelligence Index
50
unchanged vs 3.5 Flash
±0 pts
DeepSWE agentic coding
49%
3.5 Flash scored 37%
+12 pts

Gemini 3.6 Flash arrived on July 21, 2026 as the centerpiece of a three-model Google drop — alongside Gemini 3.5 Flash-Lite and the restricted Gemini 3.5 Flash Cyber. Google calls 3.6 Flash its “workhorse model” for coding, knowledge work, and multimodal tasks, and the pitch is economic: a lower output price, fewer tokens burned per task, and half the wall-clock time on independent measurement.

What the announcement didn’t include matters just as much. Gemini 3.5 Pro — the flagship of the 3.5 generation, teased back in May for a “next month” rollout — is still in partner testing with no date. And on the independent Artificial Analysis Intelligence Index, 3.6 Flash scores exactly what its predecessor scored. This is an efficiency release, not a capability leap, and reading it correctly changes what you should do about it.

This analysis covers what shipped and at what price, the counter-signal most coverage buries, the combined sticker-price-times-token-efficiency math that determines what your bill actually does, the sibling models, the 3.5 Pro delay timeline, and a practical routing take for teams deciding what to build on this quarter.

Key takeaways
  1. 01
    Three models shipped in one drop — and no 3.5 Pro.Gemini 3.6 Flash (the new workhorse), 3.5 Flash-Lite (high-throughput tier), and 3.5 Flash Cyber (restricted security model) all landed July 21, 2026. The 3.5 generation's flagship Pro tier remains unreleased.
  2. 02
    The price cut is real and stacks with token efficiency.Output pricing fell from $9.00 to $7.50 per million tokens (input unchanged at $1.50), and Google cites Artificial Analysis measurements showing ~17% fewer output tokens per task — the two multipliers compound on your actual bill.
  3. 03
    Raw intelligence is flat — this is not a capability jump.On the Artificial Analysis Intelligence Index, 3.6 Flash scores 50, unchanged from 3.5 Flash. Google's own benchmark gains (DeepSWE, MLE Bench, OSWorld-Verified) are applied agentic-task improvements, not composite-intelligence gains.
  4. 04
    Per-task economics improved on independent measurement.Artificial Analysis measured average time per task falling from 2.7 to 1.3 minutes and average cost per task from $0.59 to $0.50 — corroboration that the efficiency story holds outside Google's own numbers.
  5. 05
    The roadmap signal: 3.5 Pro slips, Gemini 4 pre-training begins.Bloomberg reported 3.5 Pro struggled to meet internal performance goals; Google says only 'as soon as it's ready.' Meanwhile Google confirmed its most ambitious pre-training run yet is underway for Gemini 4 — no timeline, no specs.

01What ShippedThree Flash models, zero flagships.

Google’s announcement post positions Gemini 3.6 Flash as the “workhorse model” — the default for coding, knowledge work, and multimodal tasks. It shipped immediately on the Gemini API via Google AI Studio, in Android Studio, Google Antigravity, the Gemini Enterprise Agent Platform, and the consumer Gemini app. Specs from the official model docs: a 1,048,576-token input context, 65,536-token max output, a March 2026 knowledge cutoff, and text, image, video, audio, and PDF inputs with text-only output, under the model ID gemini-3.6-flash.

All three models are closed and proprietary — no open weights, no self-hosting; access is metered through Google’s API and console. The trio slots into a lineup that spans from the consumer app tiers (we mapped those in our Google AI plans breakdown) up to the enterprise agent platform. What the drop pointedly lacks is a flagship: VentureBeat called the missing Gemini 3.5 Pro a “conspicuous omission,” and TechCrunch framed the release as notable for what it didn’t ship.

Workhorse
Gemini 3.6 Flash
$1.50 in · $7.50 out / 1M tokens

The new default. 1M-token context, 64K output, March 2026 cutoff. Output price cut from $9.00, ~17% fewer output tokens per task (Google-cited AA measurement), and vendor-reported gains on agentic benchmarks.

gemini-3.6-flash · live on the Gemini API
High Throughput
Gemini 3.5 Flash-Lite
$0.30 in · $2.50 out / 1M tokens

The subagent and bulk-processing tier: 350 output tokens/second, configurable thinking levels (minimal, low, higher), and a real capability gain over 3.1 Flash-Lite. Rolling out inside Google Search.

Fastest of the 3.5 line
Restricted
Gemini 3.5 Flash Cyber
Governments + trusted partners only

Fine-tuned on top of 3.5 Flash to find, validate, and patch software vulnerabilities, deployed inside Google's CodeMender agent. Not publicly purchasable; no public price — Google cites dual-use risk.

CodeMender limited-access pilot
Timing note
CNBC observed that the launch landed one day before Alphabet’s Q2 2026 earnings, framing it as Google working to show progress across a product pipeline facing delays and mounting competition — a backdrop that included capacity-constrained Kimi K3 subscriptions at Moonshot AI and Alibaba teasing Qwen 3.8 Max.

02The Counter-SignalCheaper, not smarter.

Here is the number most coverage buries a few paragraphs down: on the independent Artificial Analysis Intelligence Index, Gemini 3.6 Flash scores 50 — exactly what Gemini 3.5 Flash scored. Artificial Analysis put it bluntly: “Gemini 3.6 Flash does not improve in intelligence over 3.5 Flash.” If you are picking a default model on raw capability, nothing changed on July 21.

What did change is everything around the intelligence score. Artificial Analysis independently measured average time per task falling from 2.7 minutes to 1.3 minutes — a better-than-half reduction — and average cost per task falling from $0.59 to $0.50 under its own task-cost methodology, which is distinct from the sticker per-token price. Output speed came in at 304 tokens per second in AA’s pre-launch testing.

AA Intelligence Index
Flat vs 3.5 Flash
50

No composite-intelligence gain over the predecessor. Artificial Analysis's measurement is the key counter-signal to the launch framing: this is an efficiency release, not a capability jump.

±0 pts
Time per task
Down from 2.7 minutes
1.3min

Independently measured by Artificial Analysis across its task suite — a 52% reduction in average wall-clock time per task, the headline of AA's own launch analysis.

−52% · AA-measured
Output speed
AA pre-launch testing
304tok/s

Fast enough to compound with the ~17% fewer output tokens Google cites: fewer tokens, generated faster, at a lower list price per token.

Independent measurement

One framing trap to avoid: the flat Intelligence Index does not contradict the benchmark gains Google published (covered in section 04). The AA Index measures composite intelligence across a standard evaluation suite; Google’s DeepSWE, MLE Bench, and OSWorld-Verified numbers measure applied, agentic task performance — how efficiently the model executes multi-step work. Both can be true at once, and both are: same brain, meaningfully better workflow economics. For a business audience choosing a default model, that distinction matters more than any single benchmark delta.

03Cost AnalysisThe real cost math: two multipliers, one bill.

Every outlet reported the $9.00-to-$7.50 output price cut. Almost none combined it with the token-efficiency figure into a single answer to the question that actually matters: what happens to your bill? The two changes compound. The output list price fell $1.50 per million tokens — a 16.7% cut. Separately, Google cites Artificial Analysis Index measurements showing 3.6 Flash uses roughly 17% fewer output tokens per task than 3.5 Flash (on some benchmarks, like DeepSWE, Google reports the reduction reaches up to 65%). Multiply the two and a task’s output spend lands around 69% of what it was — roughly 31% lower output cost per task, by our arithmetic on Google’s token figure, before any workload-specific variation.

The table below pairs the vendor sticker prices with Artificial Analysis’s independently measured per-task figures — the combination nobody else has put in one place.

Sticker price versus real per-task cost for Google’s new Flash models against their predecessors: list pricing per million tokens from Google’s July 21, 2026 announcement and VentureBeat’s pricing table, alongside Artificial Analysis’s independently measured Intelligence Index, time per task, and cost per task. Delta rows are our arithmetic from the cited figures.
ModelList price ($/M in · out)AA Intelligence IndexAA time per taskAA cost per task
Workhorse Flash tier
Gemini 3.5 Flash (predecessor)$1.50 · $9.00502.7 min$0.59
Gemini 3.6 Flash (new)$1.50 · $7.50501.3 min$0.50
Generation delta (our arithmetic)in ±$0 · out −$1.50 (−16.7%)±0 pts−1.4 min (−52%)−$0.09 (−15%)
Lite tier
Gemini 3.1 Flash-Lite (predecessor)$0.25 · $1.5025— not published in cited coverage
Gemini 3.5 Flash-Lite (new)$0.30 · $2.5036— not published in cited coverage
Generation delta (our arithmetic)in +$0.05 · out +$1.00+11 pts

Two honest footnotes on that table. First, the AA cost-per-task delta computes to about 15% from the published $0.59 and $0.50 figures — the per-task saving is real but smaller than the headline token math suggests, because AA’s methodology captures full task behavior, not just output-token counts. Second, the Lite-tier upgrade is a price increase: 3.5 Flash-Lite costs more per token than the 3.1 Flash-Lite it succeeds, which — per VentureBeat’s pricing table — remains Google’s cheapest model at $0.25/$1.50, roughly half the speed. You are paying up for the +11-point capability gain and the throughput.

One claim to keep correctly labeled: Google says — as reported by CNBC — that 3.6 Flash is cheaper per task than GPT-5.6 Terra Max, Kimi K3, and Qwen 3.7 Max. That is a per-task economics claim built on token-efficiency multipliers, not a per-token sticker comparison — several competing models carry lower blended per-token totals on VentureBeat’s pricing table. Treat it as Google-stated positioning, and check per-task cost on your own workloads before repeating it.

04BenchmarksWhere 3.6 Flash does gain: applied agentic work.

Google’s launch benchmarks tell the applied-performance side of the story. On DeepSWE — Datacurve’s agentic software-engineering benchmark — 3.6 Flash scores 49% against 3.5 Flash’s 37%, which Google attributes to fewer unwanted code edits and reduced execution loops. MLE Bench (machine-learning research tasks) jumps from 49.7% to 63.9%. OSWorld-Verified (computer use) moves from 78.4% to 83.0%, and computer use is now a built-in client-side tool via the Gemini API and Gemini Enterprise. All of these are vendor-reported figures from Google’s announcement.

Google-reported agentic benchmarks · 3.6 Flash vs 3.5 Flash

Source: Google launch announcement, July 21, 2026 — vendor-reported
DeepSWE · 3.5 FlashAgentic software engineering · Datacurve
37%
DeepSWE · 3.6 FlashFewer unwanted edits, reduced execution loops
49%
MLE Bench · 3.5 FlashML research tasks
49.7%
MLE Bench · 3.6 Flash+14.2 pts vs predecessor
63.9%
OSWorld-Verified · 3.5 FlashComputer use
78.4%
OSWorld-Verified · 3.6 FlashComputer use now a built-in Gemini API tool
83.0%

Beyond the percentage benchmarks, Google reports a GDPval-AA v2 knowledge-work score of 1421 versus 1349 for 3.5 Flash, and the DeepMind model card lists 58.7% on SWE-Bench Pro. Early customers Hebbia and Harvey are cited for gains on document parsing, chart and data analysis, and report drafting — the knowledge-work profile Google is aiming the workhorse label at. For the cross-vendor picture — how these numbers stack against GPT-5.6, Sonnet 5, and Kimi K3, and what the recomputed per-task prices mean for a default pick — see our companion deep-dive on 3.6 Flash’s benchmarks versus the frontier field.

05The SiblingsFlash-Lite grows up; Cyber stays behind glass.

Gemini 3.5 Flash-Lite is the release’s quiet overachiever — and the one model in the drop with a genuine capability gain. Its Artificial Analysis Intelligence Index score of 36 is up 11 points from 3.1 Flash-Lite’s 25, it runs at 350 output tokens per second (the fastest of the 3.5 line), and Google positions it explicitly for subagent and high-throughput workloads like agentic search and document processing, with configurable thinking levels — minimal, low, and higher. It even beats the standard Gemini 3 Flash on some evals: 54.2% versus 49.6% on SWE-Bench Pro and 74.0% versus 65.1% on OSWorld-Verified, per Google. It is also rolling out inside Google Search.

AA Intelligence Index
Up from 25 on 3.1 Flash-Lite
36

A real capability gain — the contrast to 3.6 Flash's flat 50. The Lite tier got smarter and pricier; the workhorse tier got cheaper and faster at the same intelligence.

+11 pts
Output speed
Fastest of the 3.5 line
350tok/s

Google-cited Artificial Analysis figure. The predecessor 3.1 Flash-Lite remains cheaper per token but roughly half the speed, per VentureBeat's pricing table.

Built for subagents
Terminal-Bench 2.1
vs 31% for 3.1 Flash-Lite
54%

A +23-point jump on terminal-driven agentic tasks — the profile that matters for high-volume automation pipelines where a cheap model is invoked thousands of times.

+23 pts · vendor-reported

The economics of running fleets of cheap models as subagents — when Flash-Lite-class pricing beats one big-model call — is a deep enough topic that we broke it out into its own analysis of 3.5 Flash-Lite’s subagent economics. And if you are weighing the outgoing generation, our earlier look at Gemini 3.1 Flash-Lite — still Google’s cheapest model — covers the tier 3.5 Flash-Lite now replaces.

Gemini 3.5 Flash Cyber is a different animal entirely: a security-specialist model fine-tuned on top of 3.5 Flash to find, validate, and patch software vulnerabilities, paired with Google’s CodeMender agent. Inside CodeMender, multiple Flash Cyber agents run in parallel — invoked up to five times — with findings merged into one report. The design premise: a cheap model called many times beats one expensive call against a huge search space. Google is keeping it restricted, “exclusively available to governments and trusted partners” via a limited-access pilot, citing the dual-use nature of offensive security capability. No public price exists; Google says only that it runs at a lower price per token than larger models.

Vendor-run evaluation — label accordingly
On Google’s internal “Big Sleep” evaluation of the V8 JavaScript engine, 3.5 Flash Cyber found 55 unique confirmed issues versus 47 for mainline 3.5 Flash and 36 for Claude Opus 4.6 at a fixed number of invocations. This is a Google-run comparison, not an independent third-party benchmark — read it as directional. CNBC frames Flash Cyber as Google’s clearest answer yet to Anthropic’s lead in cybersecurity.

06The Missing ModelThe 3.5 Pro delay, as a dated timeline.

The oddity press coverage flagged: Google shipped a higher point-release number — 3.6 — in the mid-tier while the 3.5 generation’s flagship still hasn’t shipped. Nobody has assembled the Pro slippage as a single dated sequence, so here it is. The short version: a “next month” tease in May, a Bloomberg delay report in mid-July, and a launch day two months after that tease on which three Flash models shipped and Pro did not.

Chronological timeline of public statements and reporting on Gemini 3.5 Pro’s status, from its last public Preview update in February 2026 through the July 21, 2026 launch at which three Flash models shipped without it. Sources: TechCrunch, Bloomberg as cited by TechCrunch, Google’s announcement post, and Logan Kilpatrick’s X post.
DateSourceWhat was saidWhat actually happened
February 2026TechCrunch reportingGemini 3.5 Pro, as a Preview, receives its last public updateNo public Pro update in the five months since
May 2026Google, alongside the 3.5 Flash launchPro “already being used internally,” expected to roll out “next month”The teased window passed with no release
July 16, 2026Bloomberg (via TechCrunch)Google “struggled to meet internal performance goals” on 3.5 Pro, delaying launchIndependently reported; not vendor-confirmed
July 21, 2026Google launch post + Logan Kilpatrick on XPro “currently testing with partners,” broadly available “as soon as it’s ready”; Gemini 4 pre-training confirmed underway3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber shipped — no Pro, no date
"Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it's ready."— Logan Kilpatrick, Product Lead, Google DeepMind, July 21, 2026

Buried in the same announcement was the bigger roadmap item: Google confirmed, in its own words, “We have started our most ambitious pre-training run yet, for Gemini 4.” That is the full extent of what is confirmed — pre-training underway, no benchmarks, no specs, no timeline. Gemini 4 is not released, and nothing about it should be treated as imminent.

Our read on the version-number oddity: shipping the efficient mid-tier while holding the flagship is less an embarrassment than a strategy several labs have converged on — ship what’s ready, monetize the workhorse tier where the volume is, and let the flagship land when it clears internal bars. OpenAI’s staggered GPT-5.6 rollout followed a similar shape. The practical consequence for buyers is that “wait for Pro” is now an unbounded bet: per TechCrunch, the May promise has already slipped past two months with no new date offered.

07ImplicationsWhat this means for your model routing.

The launch changes the decision tree in specific, bounded ways. The workhorse tier got cheaper per task without getting smarter; the Lite tier got smarter and pricier; the flagship stayed vaporware; the security model is invitation-only.

Already on 3.5 Flash
Default workhorse workloads

Migrating to 3.6 Flash is close to a pure win: same intelligence class, lower output list price, fewer tokens per task, and half the AA-measured time per task. Verify on your own prompts, then switch the default.

Move to 3.6 Flash
High-volume automation
Subagent fleets

3.5 Flash-Lite's +11-point intelligence gain, 350 tok/s speed, and configurable thinking levels make it the throughput pick — but it costs more per token than 3.1 Flash-Lite, so re-run the math on thin-margin pipelines.

Pick 3.5 Flash-Lite
Waiting on 3.5 Pro
Flagship-dependent roadmaps

Partner testing, no date, one slipped promise, and one Bloomberg-reported internal miss. Do not gate a Q3 roadmap on Pro's arrival — build on what shipped and treat Pro as upside when it lands.

Stop waiting
Security tooling
Vulnerability discovery & patching

Flash Cyber is not purchasable — governments and trusted partners only via the CodeMender pilot. Watch the space; the parallel cheap-agent design pattern it validates is usable today with generally available models.

Watch, don't wait

The forward projection worth making: the frontier’s competitive axis is visibly shifting from headline intelligence to per-task economics. When a vendor ships a flat-intelligence release and leads with token efficiency — and an independent index validates the time and cost gains — the message is that the volume market is won on unit economics, not benchmark crowns. Expect rivals to answer on price-per-task, and expect “efficiency releases” to become a recognized category between major generations. If your team wants help benchmarking these models on your actual workloads and wiring per-task cost tracking into the decision, that comparative evaluation is exactly where our AI transformation engagements start.

08ConclusionThe workhorse thesis, priced.

The shape of the Gemini line, July 2026

Google sold efficiency as the product — and the independent numbers mostly back it.

Gemini 3.6 Flash is exactly what Google says it is: a workhorse. Not smarter than its predecessor — the flat Artificial Analysis Index score settles that — but meaningfully cheaper per unit of work, on both the vendor’s math and independent per-task measurement. The $9.00-to-$7.50 output cut compounds with the token-efficiency gain into roughly 31% lower output spend per task by our arithmetic, and AA’s measured halving of time per task is arguably the bigger operational win.

The asterisks matter too. The per-task-cheapness claims versus rivals are Google-stated positioning, not independent findings. The Lite tier’s upgrade is a price increase with a real capability gain attached. Flash Cyber’s eye-catching vulnerability numbers come from a Google-run evaluation. And the 3.5 Pro flagship is now two months past its teased window with nothing firmer than “as soon as it’s ready” — while Gemini 4 pre-training officially begins in the background.

The practical move is unchanged from every release this year: run your own evals on your own prompts, price your workloads per task rather than per token, and switch defaults only where the numbers hold. On this one, for teams already in the Gemini ecosystem, they mostly will.

Route models on economics, not headlines

The cheapest model is the one that finishes the task in fewer tokens.

Our team helps businesses benchmark frontier models on real workloads, wire per-task cost tracking into routing decisions, and ship agentic systems on the model tier the math actually supports.

Free consultationExpert guidanceTailored solutions
What we work on

Model-economics engagements

  • Per-task cost benchmarking — Gemini / GPT / Claude
  • Workhorse-tier migration audits (3.5 → 3.6 Flash)
  • Subagent architectures on Lite-tier pricing
  • Multi-vendor routing by task class
  • Token-spend observability & cost governance
FAQ · Gemini 3.6 Flash

The questions we get every week.

Gemini 3.6 Flash is Google's new mid-tier model, released July 21, 2026 alongside Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber. Google positions it as its 'workhorse model' for coding, knowledge work, and multimodal tasks. It was available immediately via the Gemini API in Google AI Studio, Android Studio, Google Antigravity, the Gemini Enterprise Agent Platform, and the consumer Gemini app, under the model ID gemini-3.6-flash. Like the rest of the drop, it is closed and proprietary — there are no open weights and no self-hosting option; access is metered through Google's API and console.
Related dispatches

Continue exploring frontier releases.