Gemini 3.6 Flash arrived on July 21, 2026 as the centerpiece of a three-model Google drop — alongside Gemini 3.5 Flash-Lite and the restricted Gemini 3.5 Flash Cyber. Google calls 3.6 Flash its “workhorse model” for coding, knowledge work, and multimodal tasks, and the pitch is economic: a lower output price, fewer tokens burned per task, and half the wall-clock time on independent measurement.
What the announcement didn’t include matters just as much. Gemini 3.5 Pro — the flagship of the 3.5 generation, teased back in May for a “next month” rollout — is still in partner testing with no date. And on the independent Artificial Analysis Intelligence Index, 3.6 Flash scores exactly what its predecessor scored. This is an efficiency release, not a capability leap, and reading it correctly changes what you should do about it.
This analysis covers what shipped and at what price, the counter-signal most coverage buries, the combined sticker-price-times-token-efficiency math that determines what your bill actually does, the sibling models, the 3.5 Pro delay timeline, and a practical routing take for teams deciding what to build on this quarter.
- 01Three models shipped in one drop — and no 3.5 Pro.Gemini 3.6 Flash (the new workhorse), 3.5 Flash-Lite (high-throughput tier), and 3.5 Flash Cyber (restricted security model) all landed July 21, 2026. The 3.5 generation's flagship Pro tier remains unreleased.
- 02The price cut is real and stacks with token efficiency.Output pricing fell from $9.00 to $7.50 per million tokens (input unchanged at $1.50), and Google cites Artificial Analysis measurements showing ~17% fewer output tokens per task — the two multipliers compound on your actual bill.
- 03Raw intelligence is flat — this is not a capability jump.On the Artificial Analysis Intelligence Index, 3.6 Flash scores 50, unchanged from 3.5 Flash. Google's own benchmark gains (DeepSWE, MLE Bench, OSWorld-Verified) are applied agentic-task improvements, not composite-intelligence gains.
- 04Per-task economics improved on independent measurement.Artificial Analysis measured average time per task falling from 2.7 to 1.3 minutes and average cost per task from $0.59 to $0.50 — corroboration that the efficiency story holds outside Google's own numbers.
- 05The roadmap signal: 3.5 Pro slips, Gemini 4 pre-training begins.Bloomberg reported 3.5 Pro struggled to meet internal performance goals; Google says only 'as soon as it's ready.' Meanwhile Google confirmed its most ambitious pre-training run yet is underway for Gemini 4 — no timeline, no specs.
01 — What ShippedThree Flash models, zero flagships.
Google’s announcement post positions Gemini 3.6 Flash as the “workhorse model” — the default for coding, knowledge work, and multimodal tasks. It shipped immediately on the Gemini API via Google AI Studio, in Android Studio, Google Antigravity, the Gemini Enterprise Agent Platform, and the consumer Gemini app. Specs from the official model docs: a 1,048,576-token input context, 65,536-token max output, a March 2026 knowledge cutoff, and text, image, video, audio, and PDF inputs with text-only output, under the model ID gemini-3.6-flash.
All three models are closed and proprietary — no open weights, no self-hosting; access is metered through Google’s API and console. The trio slots into a lineup that spans from the consumer app tiers (we mapped those in our Google AI plans breakdown) up to the enterprise agent platform. What the drop pointedly lacks is a flagship: VentureBeat called the missing Gemini 3.5 Pro a “conspicuous omission,” and TechCrunch framed the release as notable for what it didn’t ship.
Gemini 3.6 Flash
The new default. 1M-token context, 64K output, March 2026 cutoff. Output price cut from $9.00, ~17% fewer output tokens per task (Google-cited AA measurement), and vendor-reported gains on agentic benchmarks.
Gemini 3.5 Flash-Lite
The subagent and bulk-processing tier: 350 output tokens/second, configurable thinking levels (minimal, low, higher), and a real capability gain over 3.1 Flash-Lite. Rolling out inside Google Search.
Gemini 3.5 Flash Cyber
Fine-tuned on top of 3.5 Flash to find, validate, and patch software vulnerabilities, deployed inside Google's CodeMender agent. Not publicly purchasable; no public price — Google cites dual-use risk.
02 — The Counter-SignalCheaper, not smarter.
Here is the number most coverage buries a few paragraphs down: on the independent Artificial Analysis Intelligence Index, Gemini 3.6 Flash scores 50 — exactly what Gemini 3.5 Flash scored. Artificial Analysis put it bluntly: “Gemini 3.6 Flash does not improve in intelligence over 3.5 Flash.” If you are picking a default model on raw capability, nothing changed on July 21.
What did change is everything around the intelligence score. Artificial Analysis independently measured average time per task falling from 2.7 minutes to 1.3 minutes — a better-than-half reduction — and average cost per task falling from $0.59 to $0.50 under its own task-cost methodology, which is distinct from the sticker per-token price. Output speed came in at 304 tokens per second in AA’s pre-launch testing.
Flat vs 3.5 Flash
No composite-intelligence gain over the predecessor. Artificial Analysis's measurement is the key counter-signal to the launch framing: this is an efficiency release, not a capability jump.
Down from 2.7 minutes
Independently measured by Artificial Analysis across its task suite — a 52% reduction in average wall-clock time per task, the headline of AA's own launch analysis.
AA pre-launch testing
Fast enough to compound with the ~17% fewer output tokens Google cites: fewer tokens, generated faster, at a lower list price per token.
One framing trap to avoid: the flat Intelligence Index does not contradict the benchmark gains Google published (covered in section 04). The AA Index measures composite intelligence across a standard evaluation suite; Google’s DeepSWE, MLE Bench, and OSWorld-Verified numbers measure applied, agentic task performance — how efficiently the model executes multi-step work. Both can be true at once, and both are: same brain, meaningfully better workflow economics. For a business audience choosing a default model, that distinction matters more than any single benchmark delta.
03 — Cost AnalysisThe real cost math: two multipliers, one bill.
Every outlet reported the $9.00-to-$7.50 output price cut. Almost none combined it with the token-efficiency figure into a single answer to the question that actually matters: what happens to your bill? The two changes compound. The output list price fell $1.50 per million tokens — a 16.7% cut. Separately, Google cites Artificial Analysis Index measurements showing 3.6 Flash uses roughly 17% fewer output tokens per task than 3.5 Flash (on some benchmarks, like DeepSWE, Google reports the reduction reaches up to 65%). Multiply the two and a task’s output spend lands around 69% of what it was — roughly 31% lower output cost per task, by our arithmetic on Google’s token figure, before any workload-specific variation.
The table below pairs the vendor sticker prices with Artificial Analysis’s independently measured per-task figures — the combination nobody else has put in one place.
| Model | List price ($/M in · out) | AA Intelligence Index | AA time per task | AA cost per task |
|---|---|---|---|---|
| Workhorse Flash tier | ||||
| Gemini 3.5 Flash (predecessor) | $1.50 · $9.00 | 50 | 2.7 min | $0.59 |
| Gemini 3.6 Flash (new) | $1.50 · $7.50 | 50 | 1.3 min | $0.50 |
| Generation delta (our arithmetic) | in ±$0 · out −$1.50 (−16.7%) | ±0 pts | −1.4 min (−52%) | −$0.09 (−15%) |
| Lite tier | ||||
| Gemini 3.1 Flash-Lite (predecessor) | $0.25 · $1.50 | 25 | — not published in cited coverage | — |
| Gemini 3.5 Flash-Lite (new) | $0.30 · $2.50 | 36 | — not published in cited coverage | — |
| Generation delta (our arithmetic) | in +$0.05 · out +$1.00 | +11 pts | — | — |
Two honest footnotes on that table. First, the AA cost-per-task delta computes to about 15% from the published $0.59 and $0.50 figures — the per-task saving is real but smaller than the headline token math suggests, because AA’s methodology captures full task behavior, not just output-token counts. Second, the Lite-tier upgrade is a price increase: 3.5 Flash-Lite costs more per token than the 3.1 Flash-Lite it succeeds, which — per VentureBeat’s pricing table — remains Google’s cheapest model at $0.25/$1.50, roughly half the speed. You are paying up for the +11-point capability gain and the throughput.
One claim to keep correctly labeled: Google says — as reported by CNBC — that 3.6 Flash is cheaper per task than GPT-5.6 Terra Max, Kimi K3, and Qwen 3.7 Max. That is a per-task economics claim built on token-efficiency multipliers, not a per-token sticker comparison — several competing models carry lower blended per-token totals on VentureBeat’s pricing table. Treat it as Google-stated positioning, and check per-task cost on your own workloads before repeating it.
04 — BenchmarksWhere 3.6 Flash does gain: applied agentic work.
Google’s launch benchmarks tell the applied-performance side of the story. On DeepSWE — Datacurve’s agentic software-engineering benchmark — 3.6 Flash scores 49% against 3.5 Flash’s 37%, which Google attributes to fewer unwanted code edits and reduced execution loops. MLE Bench (machine-learning research tasks) jumps from 49.7% to 63.9%. OSWorld-Verified (computer use) moves from 78.4% to 83.0%, and computer use is now a built-in client-side tool via the Gemini API and Gemini Enterprise. All of these are vendor-reported figures from Google’s announcement.
Google-reported agentic benchmarks · 3.6 Flash vs 3.5 Flash
Source: Google launch announcement, July 21, 2026 — vendor-reportedBeyond the percentage benchmarks, Google reports a GDPval-AA v2 knowledge-work score of 1421 versus 1349 for 3.5 Flash, and the DeepMind model card lists 58.7% on SWE-Bench Pro. Early customers Hebbia and Harvey are cited for gains on document parsing, chart and data analysis, and report drafting — the knowledge-work profile Google is aiming the workhorse label at. For the cross-vendor picture — how these numbers stack against GPT-5.6, Sonnet 5, and Kimi K3, and what the recomputed per-task prices mean for a default pick — see our companion deep-dive on 3.6 Flash’s benchmarks versus the frontier field.
05 — The SiblingsFlash-Lite grows up; Cyber stays behind glass.
Gemini 3.5 Flash-Lite is the release’s quiet overachiever — and the one model in the drop with a genuine capability gain. Its Artificial Analysis Intelligence Index score of 36 is up 11 points from 3.1 Flash-Lite’s 25, it runs at 350 output tokens per second (the fastest of the 3.5 line), and Google positions it explicitly for subagent and high-throughput workloads like agentic search and document processing, with configurable thinking levels — minimal, low, and higher. It even beats the standard Gemini 3 Flash on some evals: 54.2% versus 49.6% on SWE-Bench Pro and 74.0% versus 65.1% on OSWorld-Verified, per Google. It is also rolling out inside Google Search.
Up from 25 on 3.1 Flash-Lite
A real capability gain — the contrast to 3.6 Flash's flat 50. The Lite tier got smarter and pricier; the workhorse tier got cheaper and faster at the same intelligence.
Fastest of the 3.5 line
Google-cited Artificial Analysis figure. The predecessor 3.1 Flash-Lite remains cheaper per token but roughly half the speed, per VentureBeat's pricing table.
vs 31% for 3.1 Flash-Lite
A +23-point jump on terminal-driven agentic tasks — the profile that matters for high-volume automation pipelines where a cheap model is invoked thousands of times.
The economics of running fleets of cheap models as subagents — when Flash-Lite-class pricing beats one big-model call — is a deep enough topic that we broke it out into its own analysis of 3.5 Flash-Lite’s subagent economics. And if you are weighing the outgoing generation, our earlier look at Gemini 3.1 Flash-Lite — still Google’s cheapest model — covers the tier 3.5 Flash-Lite now replaces.
Gemini 3.5 Flash Cyber is a different animal entirely: a security-specialist model fine-tuned on top of 3.5 Flash to find, validate, and patch software vulnerabilities, paired with Google’s CodeMender agent. Inside CodeMender, multiple Flash Cyber agents run in parallel — invoked up to five times — with findings merged into one report. The design premise: a cheap model called many times beats one expensive call against a huge search space. Google is keeping it restricted, “exclusively available to governments and trusted partners” via a limited-access pilot, citing the dual-use nature of offensive security capability. No public price exists; Google says only that it runs at a lower price per token than larger models.
06 — The Missing ModelThe 3.5 Pro delay, as a dated timeline.
The oddity press coverage flagged: Google shipped a higher point-release number — 3.6 — in the mid-tier while the 3.5 generation’s flagship still hasn’t shipped. Nobody has assembled the Pro slippage as a single dated sequence, so here it is. The short version: a “next month” tease in May, a Bloomberg delay report in mid-July, and a launch day two months after that tease on which three Flash models shipped and Pro did not.
| Date | Source | What was said | What actually happened |
|---|---|---|---|
| February 2026 | TechCrunch reporting | Gemini 3.5 Pro, as a Preview, receives its last public update | No public Pro update in the five months since |
| May 2026 | Google, alongside the 3.5 Flash launch | Pro “already being used internally,” expected to roll out “next month” | The teased window passed with no release |
| July 16, 2026 | Bloomberg (via TechCrunch) | Google “struggled to meet internal performance goals” on 3.5 Pro, delaying launch | Independently reported; not vendor-confirmed |
| July 21, 2026 | Google launch post + Logan Kilpatrick on X | Pro “currently testing with partners,” broadly available “as soon as it’s ready”; Gemini 4 pre-training confirmed underway | 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber shipped — no Pro, no date |
"Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it's ready."— Logan Kilpatrick, Product Lead, Google DeepMind, July 21, 2026
Buried in the same announcement was the bigger roadmap item: Google confirmed, in its own words, “We have started our most ambitious pre-training run yet, for Gemini 4.” That is the full extent of what is confirmed — pre-training underway, no benchmarks, no specs, no timeline. Gemini 4 is not released, and nothing about it should be treated as imminent.
Our read on the version-number oddity: shipping the efficient mid-tier while holding the flagship is less an embarrassment than a strategy several labs have converged on — ship what’s ready, monetize the workhorse tier where the volume is, and let the flagship land when it clears internal bars. OpenAI’s staggered GPT-5.6 rollout followed a similar shape. The practical consequence for buyers is that “wait for Pro” is now an unbounded bet: per TechCrunch, the May promise has already slipped past two months with no new date offered.
07 — ImplicationsWhat this means for your model routing.
The launch changes the decision tree in specific, bounded ways. The workhorse tier got cheaper per task without getting smarter; the Lite tier got smarter and pricier; the flagship stayed vaporware; the security model is invitation-only.
Default workhorse workloads
Migrating to 3.6 Flash is close to a pure win: same intelligence class, lower output list price, fewer tokens per task, and half the AA-measured time per task. Verify on your own prompts, then switch the default.
Subagent fleets
3.5 Flash-Lite's +11-point intelligence gain, 350 tok/s speed, and configurable thinking levels make it the throughput pick — but it costs more per token than 3.1 Flash-Lite, so re-run the math on thin-margin pipelines.
Flagship-dependent roadmaps
Partner testing, no date, one slipped promise, and one Bloomberg-reported internal miss. Do not gate a Q3 roadmap on Pro's arrival — build on what shipped and treat Pro as upside when it lands.
Vulnerability discovery & patching
Flash Cyber is not purchasable — governments and trusted partners only via the CodeMender pilot. Watch the space; the parallel cheap-agent design pattern it validates is usable today with generally available models.
The forward projection worth making: the frontier’s competitive axis is visibly shifting from headline intelligence to per-task economics. When a vendor ships a flat-intelligence release and leads with token efficiency — and an independent index validates the time and cost gains — the message is that the volume market is won on unit economics, not benchmark crowns. Expect rivals to answer on price-per-task, and expect “efficiency releases” to become a recognized category between major generations. If your team wants help benchmarking these models on your actual workloads and wiring per-task cost tracking into the decision, that comparative evaluation is exactly where our AI transformation engagements start.
08 — ConclusionThe workhorse thesis, priced.
Google sold efficiency as the product — and the independent numbers mostly back it.
Gemini 3.6 Flash is exactly what Google says it is: a workhorse. Not smarter than its predecessor — the flat Artificial Analysis Index score settles that — but meaningfully cheaper per unit of work, on both the vendor’s math and independent per-task measurement. The $9.00-to-$7.50 output cut compounds with the token-efficiency gain into roughly 31% lower output spend per task by our arithmetic, and AA’s measured halving of time per task is arguably the bigger operational win.
The asterisks matter too. The per-task-cheapness claims versus rivals are Google-stated positioning, not independent findings. The Lite tier’s upgrade is a price increase with a real capability gain attached. Flash Cyber’s eye-catching vulnerability numbers come from a Google-run evaluation. And the 3.5 Pro flagship is now two months past its teased window with nothing firmer than “as soon as it’s ready” — while Gemini 4 pre-training officially begins in the background.
The practical move is unchanged from every release this year: run your own evals on your own prompts, price your workloads per task rather than per token, and switch defaults only where the numbers hold. On this one, for teams already in the Gemini ecosystem, they mostly will.