Gemini 3.7 Flash launched on August 13, 2026 with a pitch built for headlines: a smarter workhorse model at half the price. The capability half of that pitch survives independent scrutiny. The pricing half needs an asterisk that almost no launch coverage printed — Google cut Gemini 3.6 Flash to the exact same introductory rate on the exact same day.
That means the honest framing is narrower than the marketing one. Gemini 3.7 Flash costs half of what the Flash workhorse tier cost before August 13 — $0.75 per million input tokens and $3.75 per million output tokens, against the $1.50/$7.50 that 3.6 Flash launched at in July. But a developer choosing between 3.6 and 3.7 Flash today pays identical per-token rates for either model until both step back to $1.50/$7.50 on January 1, 2027. “Half the price of 3.6 Flash” is never true of 3.6 Flash’s current price.
This analysis covers what actually shipped 23 days after the last Flash release, the full pricing ladder no other outlet has traced, the honest benchmark picture from Google’s own model card — including the four rows it loses to GPT-5.6 Terra — and the independent Artificial Analysis verdict that makes this launch different from the last one.
- 01Same price as 3.6 Flash today — that's the sharp fact.Both 3.6 and 3.7 Flash bill $0.75/$3.75 per million tokens through December 31, 2026. The half-price claim is only true against the workhorse tier's pre-August-13 rate of $1.50/$7.50 — which is also what both models revert to on January 1, 2027.
- 02Smarter this time, and independently verified.Artificial Analysis's August 13 article reports the Intelligence Index moving from 52 to 56 and places 3.7 Flash on its Intelligence vs. Time per Task Pareto frontier — a different verdict from the flat, price-led story that greeted 3.6 Flash in July.
- 03A 23-day release cadence, built on the same base.3.7 Flash shipped 23 days after 3.6 Flash (July 21 to August 13) and is explicitly based on 3.6 Flash per the model card — Google credits developer feedback and algorithmic innovations, not a new base model.
- 04It leads 7 of 13 rows on Google's own table — not the board.Google's comparison table shows 3.7 Flash losing Terminal-bench 2.1 (85.8 vs 87.4), Terminal-bench 3.0, OSWorld-2.0 (47.9 vs 50.2), and DeepSWE v1.1 to GPT-5.6 Terra, and GDPval-AA v2 to Muse Spark 1.2. This is a strong mid-tier release, not a frontier-leadership claim.
- 05From January 1, 2027, you pay the old rate for the smarter model.When the intro window closes, both Flash models list at $1.50/$7.50 — exactly what 3.6 Flash cost at its July launch. The durable gain is capability per dollar at an unchanged standard price, not a permanent price cut.
01 — What ShippedA 23-day turnaround on the workhorse tier.
Google published Gemini 3.7 Flash on August 13, 2026 — model ID gemini-3.7-flash, stable from day one with no preview suffix. It arrives exactly 23 days after Gemini 3.6 Flash, which shipped July 21 as part of a three-model drop we covered in our 3.6 Flash launch analysis. Press coverage rounds the gap to “three weeks” — VentureBeat’s phrasing, not Google’s — but the precise count is 23 days, and the cadence itself is the story: Google attributes the unusually short turnaround to developer feedback and algorithmic innovations rather than a bigger model.
The model card is explicit about lineage: Gemini 3.7 Flash is based on Gemini 3.6 Flash. This is an iteration on the existing workhorse, not a new base model — which makes the measured capability gains in Sections 03 and 04 more interesting, not less. It also lands in a crowded week; xAI shipped Grok 4.6 just one day earlier, on August 12.
One landscape note worth stating plainly, because the Flash cadence invites confusion: Gemini 3.5 Pro remains unreleased. VentureBeat’s August 13 coverage notes that Google’s latest released general-purpose Pro model remains Gemini 3.1 Pro, introduced in February, with 3.5 Pro still in partner testing as of July and no further timetable given at this announcement. The Flash line is where Google is iterating in public right now.
Jul 21 → Aug 13
The gap between 3.6 Flash and 3.7 Flash. Google credits developer feedback and algorithmic innovations; the model card confirms 3.7 Flash is built directly on 3.6 Flash.
1,048,576 in / 65,536 out
Text, image, video, audio, and PDF inputs; text output only. Function calling, structured output, code execution, context caching, batch API, and search grounding all supported at launch.
Per the model card
The card states a March 2026 cutoff, with the caveat that some domains may reflect knowledge closer to January 2025, in line with the Gemini 3 model family.
02 — PricingThe half-price claim, with its asterisk restored.
Google’s Antigravity team put the pitch in one sentence: “3.7 Flash is available through the end of the year at an introductory price of $0.75/1M input tokens and $3.75/1M output tokens - half the original 3.6 Flash cost per million tokens.” Read carefully — original is doing all the work. That sentence is arithmetically true against 3.6 Flash’s July 21 launch rate of $1.50/$7.50. It is not true against what 3.6 Flash costs now.
At the time of writing, Google’s own pricing page carries a separate 3.6 Flash block showing the identical introductory rate — $0.75 in, $3.75 out, through December 31, 2026 — and the 3.7 Flash model card’s own comparison table lists 3.6 Flash’s price cells with the same asterisked $0.75/$3.75 figures. The same Antigravity post confirms it directly: introductory pricing also applies to 3.6 Flash through December 31, 2026. Every launch-day outlet we checked repeated “half the price of 3.6 Flash” without noting the equalization.
The full ladder below is the table we could not find anywhere else in launch coverage — the Flash tier’s complete price history from 3.5 Flash through the January 2027 step-up, every cell from Google’s live pricing page or our own dated coverage of the 3.6 launch.
| Rate step | Input $/1M | Output $/1M | In effect | What changed |
|---|---|---|---|---|
| Before the workhorse era | ||||
| Gemini 3.5 Flash | $1.50 | $9.00 | Flat list, no expiry | The pre-3.6 baseline |
| The 3.6 Flash era — July 21 to August 12, 2026 | ||||
| Gemini 3.6 Flash at launch | $1.50 | $7.50 | Jul 21 – Aug 12, 2026 | Output cut $1.50 (−16.7%) vs 3.5 Flash; input unchanged |
| From August 13, 2026 — both workhorses, identical rates | ||||
| Gemini 3.6 Flash, repriced | $0.75 | $3.75 | Aug 13 – Dec 31, 2026 | Same-day intro cut — the step no coverage printed |
| Gemini 3.7 Flash, intro | $0.75 | $3.75 | Aug 13 – Dec 31, 2026 | 50% below the pre-Aug-13 workhorse rate |
| From January 1, 2027 — the dated cliff | ||||
| Both 3.6 & 3.7 Flash, standard | $1.50 | $7.50 | From Jan 1, 2027 | 2× the intro rate — exactly 3.6 Flash’s July launch price |
The supporting rates move on the same schedule. Context caching for 3.7 Flash is $0.075 per million tokens through December 31 — 10% of the intro input price — stepping to $0.15 from January 1, with cache storage at $0.50 per million tokens per hour doubling to $1.00 on the same date. The batch API takes 50% off whatever the current rate is: $0.375/$1.875 through December, $0.75/$3.75 after.
One labeling trap for anyone comparing marketplaces: OpenRouter lists google/gemini-3.7-flash with a headline $0.375/$1.875 badge marked 50% off. That is a Vertex-specific provider discount stacked on top of Google’s intro rate — not Google’s official price. OpenRouter’s Google AI Studio provider row shows $0.75/$3.75 with no badge, matching the official introductory rate exactly. Quote the provider row, not the badge.
Against the mid-tier field, the intro rate is aggressive even after the correction. Working from the model card’s own comparison row and the competitors’ live pricing pages: Claude Sonnet 5 lists at $2.00/$10.00 per million and GPT-5.6 Terra at $2.00/$12.00 — so at intro rates, 3.7 Flash runs 62.5% below Sonnet 5 on both input and output, and 68.75% below Terra on output. Even at the January 2027 standard rate, it stays 25% below Sonnet 5 on both sides. Meta’s Muse Spark 1.2 — listed at $1.25/$4.25 in Google’s own comparison table — is the one neighbor that flips: cheaper than the Flash standard rate, pricier than the intro. If your architecture fans out to subagent fleets, the intro-versus-standard doubling is the number to model — the same math we walked through in our subagent cost breakdown for Flash-tier models.
03 — BenchmarksGoogle’s own table: 7 of 13, not a clean sweep.
The model card publishes a five-model comparison across 13 benchmarks — 3.7 Flash, 3.6 Flash, Claude Sonnet 5, GPT-5.6 Terra, and Muse Spark 1.2. Our own count of Google’s table: 3.7 Flash posts the best score on 7 of 13 rows. GPT-5.6 Terra takes four — DeepSWE v1.1, both Terminal-bench versions, and OSWorld-2.0 — Muse Spark 1.2 wins GDPval-AA v2 outright, and Terra and Muse tie for the top Artificial Analysis Intelligence Index score at 57. That is a strong showing for a model priced at a fraction of its comparators, and it is not a frontier-leadership claim; Google printed its own losses. For where the previous generation stood, see how 3.6 Flash stacked up against Sonnet 5 and its rivals.
| Benchmark | 3.7 Flash | 3.6 Flash | Sonnet 5 | GPT-5.6 Terra | Muse Spark 1.2 |
|---|---|---|---|---|---|
| Coding & software engineering | |||||
| FrontierCode 1.1 | 43.6% | 34.4% | 42.7% | 41.3% | — |
| DeepSWE v1.1 | 65.3% | 48.6% | 53.8% | 69.6% | 54.9% |
| Code Arena (Elo) | 1588 | 1538 | 1541 | 1523 | 1535 |
| Agentic & computer use | |||||
| Terminal-bench 2.1 | 85.8% | 78.0% | 80.4% | 87.4% | 82.9% |
| Terminal-bench 3.0 | 14.9% | 5.4% | 14.6% | 20.8% | — |
| OSWorld-2.0 | 47.9% | 33.8% | — | 50.2% | — |
| AutomationBench | 30.4% | 17.0% | 10.7% | 23.6% | — |
| Knowledge & document work | |||||
| GDPval-AA v2 (Elo) | 1525 | 1422 | 1598 | 1578 | 1628 |
| Harvey LAB-AA (legal) | 90.7% | 85.1% | 90.1% | 85.2% | — |
| GDP.pdf (PDF comprehension) | 34.0% | 22.0% | 28.0% | 24.7% | 16.0% |
| Long context, video & composite | |||||
| GDM-MRCR v2 (Google-authored) | 97.0% | 91.8% | 81.5% | 93.5% | — |
| LVBench (long video) | 85.4% | 84.2% | 68.5% | 78.9% | — |
| AA Intelligence Index | 56 | 52 | 55 | 57 | 57 |
Three footnotes belong next to that table. First, an em-dash means no score published — the model card itself shows Sonnet 5’s OSWorld-2.0 cell as a bare dash, which is an absent number, never a zero or a loss. Second, GDM-MRCR is a Google DeepMind–authored long-context eval, not a third-party benchmark — weigh that row accordingly. Third, Google is inconsistent with itself on one cell: the model card’s table reads 48.6% for 3.6 Flash on DeepSWE v1.1 while Google’s own launch blog rounds the same figure to 49.0% — a small discrepancy, but a reminder that even vendor self-citations deserve checking.
The pattern in the wins is coherent. The biggest generational jumps land on production-shaped work: DeepSWE up 16.7 points over 3.6 Flash, AutomationBench up 13.4, GDP.pdf up 12.0, FrontierCode up 9.2. The losses cluster in exactly one place — the hardest agentic computer-and-terminal benchmarks, where GPT-5.6 Terra keeps the crown at more than three times the intro output price. That trade is the entire commercial argument of this release.
04 — Independent SignalThe verdict that changed since July.
When 3.6 Flash launched in July, our honest read was “cheaper, not smarter” — the price moved, the independent intelligence signal barely did. This launch earns a different verdict, and not from Google. Artificial Analysis’s August 13 article reports Gemini 3.7 Flash (high) scoring 56 on its Intelligence Index against 52 for 3.6 Flash — a genuine four-point gain — and states that this places Gemini 3.7 Flash on its Intelligence vs. Time per Task Pareto frontier. AA attributes the gain specifically to agentic evaluations rather than to the price cut.
The same article carries the one performance figure that exists anywhere for this launch: output throughput of roughly 340 tokens per second, which AA describes as nearly three times the output speed of GPT-5.6 Terra. That is a throughput measurement, not a latency number — no primary source publishes any latency figure, and the model card itself only discloses “occasional slowness or timeout issues” at launch, unquantified.
AA Intelligence Index · 3.7 Flash vs its price neighborhood
Source: Artificial Analysis Intelligence Index, as reported in AA's Aug 13, 2026 article and the Gemini 3.7 Flash model cardGoogle’s product page adds a set of vendor-curated testimonials from named engineering leads at Browser Use, Nunu.ai, Hebbia, Box, and Databricks. These are marketing, selected by Google — but the specifics are unusually concrete, and the Browser Use figure is the kind of claim a customer can be held to:
"The Gemini 3.7 Flash agent was 35% cheaper than 3.6 Flash, with a +8% observed prompt-cache hit rate and fewer tool errors."— Gregor Zunic, Co-founder & CTO, Browser Use — Google-curated testimonial
Read the two signals together and the interpretation writes itself. The vendor benchmark table says the model got meaningfully better at production-shaped agentic work; the one independent scorer agrees, on its own composite and its own speed measurements, and attributes the gain to the same category of evaluation. That alignment between vendor claim and third-party measurement is what was missing from the 3.6 Flash launch — and it is the reason this release deserves evaluation cycles rather than a shrug.
05 — Thinking LevelsThree thinking levels, priced identically.
Gemini 3.7 Flash exposes low, medium, and high thinking levels through the API — and, unlike some sibling models, minimal is explicitly unsupported and returns an API error. The interesting operational detail is how little intelligence you give up walking down the ladder: Artificial Analysis scores the three levels at 51, 53, and 56 respectively, against a median of 34 across similarly-priced reasoning models. Even the low setting clears the price tier’s median comfortably.
Thinking: low
The bulk-work setting for classification, extraction, and routing. Still 17 points above the 34-point median AA reports for similarly-priced reasoning models.
Thinking: medium
The balanced default for multi-step work where full reasoning depth is not worth the extra output tokens on every call.
Thinking: high
The benchmark configuration behind the launch numbers and the Pareto-frontier placement. Note: minimal is not supported at all — the API returns an error.
Day-one availability is broad: Google AI Studio, the Gemini API, Android Studio, Google Antigravity across IDE, CLI, and SDK, the Gemini Enterprise Agent Platform and Enterprise app, and Gemini Spark for AI Pro and Ultra subscribers in 160+ countries. If you work in Google’s agentic IDE, our guide to using Gemini models inside Google Antigravity covers the workflow, and Gemini Spark — where 3.7 Flash is now live for consumers — is the most visible consumer surface. Computer use ships in Preview status; the Live API is not supported. Search grounding and Google Maps grounding both work at launch, alongside Flex and Priority inference tiers.
06 — PlaybookWhat to do with a repriced workhorse tier.
The decision is not “switch or don’t.” It is four separate calls, one per workload class — and one calendar entry.
Upgrade the model ID, not the budget line
Same price, same base lineage, higher scores on 13 of 13 rows against its predecessor. There is no pricing reason to stay on 3.6 Flash — the switch is a config change. Re-run your own evals first; vendor deltas are directional, not guarantees.
High-volume agentic pipelines
At $0.75/$3.75 intro with batch halving that again, this is the aggressive quote for fan-out work — but model your 2027 budget at $1.50/$7.50, because the doubling lands on January 1 and applies to both workhorses.
Terminal, computer use, long-horizon SWE
Google's own table gives GPT-5.6 Terra the lead on Terminal-bench 2.1 and 3.0, OSWorld-2.0, and DeepSWE v1.1. Where those benchmarks resemble your workload, the premium model still earns its premium.
Document & analyst workflows
Sonnet 5, Terra, and Muse Spark 1.2 all out-rank 3.7 Flash on GDPval-AA v2. For work that lives and dies on knowledge-work quality rather than throughput cost, evaluate before consolidating on Flash.
Looking forward, the strategic read is about cadence economics, not this single release. Google has now shipped two workhorse iterations in 23 days, priced the newer one to remove any reason to stay on the older, and used an intro window to pull adoption forward ahead of a dated price step. If that pattern holds, the rational posture for teams is to keep model choice a configuration value rather than an architectural commitment — the vendors are iterating faster than procurement cycles. Building that kind of swap-friendly routing layer is exactly the sort of work our AI transformation engagements handle, from eval harnesses to cost-per-outcome dashboards.
07 — ConclusionSmarter is verified; half price is a frame.
The capability claim held up. The pricing claim needed an asterisk.
Gemini 3.7 Flash is the rare fast-follow release where the independent evidence is stronger than the marketing. A four-point gain on the Artificial Analysis Intelligence Index, a Pareto-frontier placement, and the biggest jumps landing on production-shaped agentic work — 23 days after the model it is built on.
The pricing story is where precision matters. Half price is true against the workhorse tier’s pre-August-13 rate, and only that. Today, 3.6 and 3.7 Flash cost the same $0.75/$3.75; from January 1, 2027, both cost the same $1.50/$7.50 — the exact rate 3.6 Flash launched at. What Google actually sold is more capability at an unchanged standard price, wrapped in an intro discount with a calendar date on it.
The practical takeaway fits in two lines. If you run Flash-tier workloads, the upgrade is free and the evals are worth running this week. And whatever you adopt, write January 1, 2027 into the budget now — the models got smarter, but the discount is the part that expires.