AI DevelopmentNew Release10 min readPublished August 13, 2026

23 days after 3.6 Flash · $0.75/$3.75 intro for both workhorses · +4 on the AA Index

Gemini 3.7 Flash: The Workhorse Gets Smarter — at Half the Old Price

Google shipped Gemini 3.7 Flash on August 13, 2026 at $0.75/$3.75 per million tokens — half what the workhorse tier cost before that day. The asterisk: Google quietly cut 3.6 Flash to the identical intro rate at the same time, so the two Flash workhorses bill the same until December 31. The capability claim fares better — an independent +4 on the Artificial Analysis Intelligence Index.

DA
Digital Applied Team
Senior strategists · Published Aug 13, 2026
PublishedAug 13, 2026
Read time10 min
SourcesModel card + independent
Intro price per 1M tokens
$0.75/$3.75
through Dec 31, 2026
= 3.6 Flash today
AA Intelligence Index
56
vs 52 for 3.6 Flash
+4 independent
Release gap
23days
Jul 21 → Aug 13
Model-card wins
7/13
benchmarks led on Google's own table

Gemini 3.7 Flash launched on August 13, 2026 with a pitch built for headlines: a smarter workhorse model at half the price. The capability half of that pitch survives independent scrutiny. The pricing half needs an asterisk that almost no launch coverage printed — Google cut Gemini 3.6 Flash to the exact same introductory rate on the exact same day.

That means the honest framing is narrower than the marketing one. Gemini 3.7 Flash costs half of what the Flash workhorse tier cost before August 13 — $0.75 per million input tokens and $3.75 per million output tokens, against the $1.50/$7.50 that 3.6 Flash launched at in July. But a developer choosing between 3.6 and 3.7 Flash today pays identical per-token rates for either model until both step back to $1.50/$7.50 on January 1, 2027. “Half the price of 3.6 Flash” is never true of 3.6 Flash’s current price.

This analysis covers what actually shipped 23 days after the last Flash release, the full pricing ladder no other outlet has traced, the honest benchmark picture from Google’s own model card — including the four rows it loses to GPT-5.6 Terra — and the independent Artificial Analysis verdict that makes this launch different from the last one.

Key takeaways
  1. 01
    Same price as 3.6 Flash today — that's the sharp fact.Both 3.6 and 3.7 Flash bill $0.75/$3.75 per million tokens through December 31, 2026. The half-price claim is only true against the workhorse tier's pre-August-13 rate of $1.50/$7.50 — which is also what both models revert to on January 1, 2027.
  2. 02
    Smarter this time, and independently verified.Artificial Analysis's August 13 article reports the Intelligence Index moving from 52 to 56 and places 3.7 Flash on its Intelligence vs. Time per Task Pareto frontier — a different verdict from the flat, price-led story that greeted 3.6 Flash in July.
  3. 03
    A 23-day release cadence, built on the same base.3.7 Flash shipped 23 days after 3.6 Flash (July 21 to August 13) and is explicitly based on 3.6 Flash per the model card — Google credits developer feedback and algorithmic innovations, not a new base model.
  4. 04
    It leads 7 of 13 rows on Google's own table — not the board.Google's comparison table shows 3.7 Flash losing Terminal-bench 2.1 (85.8 vs 87.4), Terminal-bench 3.0, OSWorld-2.0 (47.9 vs 50.2), and DeepSWE v1.1 to GPT-5.6 Terra, and GDPval-AA v2 to Muse Spark 1.2. This is a strong mid-tier release, not a frontier-leadership claim.
  5. 05
    From January 1, 2027, you pay the old rate for the smarter model.When the intro window closes, both Flash models list at $1.50/$7.50 — exactly what 3.6 Flash cost at its July launch. The durable gain is capability per dollar at an unchanged standard price, not a permanent price cut.

01What ShippedA 23-day turnaround on the workhorse tier.

Google published Gemini 3.7 Flash on August 13, 2026 — model ID gemini-3.7-flash, stable from day one with no preview suffix. It arrives exactly 23 days after Gemini 3.6 Flash, which shipped July 21 as part of a three-model drop we covered in our 3.6 Flash launch analysis. Press coverage rounds the gap to “three weeks” — VentureBeat’s phrasing, not Google’s — but the precise count is 23 days, and the cadence itself is the story: Google attributes the unusually short turnaround to developer feedback and algorithmic innovations rather than a bigger model.

The model card is explicit about lineage: Gemini 3.7 Flash is based on Gemini 3.6 Flash. This is an iteration on the existing workhorse, not a new base model — which makes the measured capability gains in Sections 03 and 04 more interesting, not less. It also lands in a crowded week; xAI shipped Grok 4.6 just one day earlier, on August 12.

One landscape note worth stating plainly, because the Flash cadence invites confusion: Gemini 3.5 Pro remains unreleased. VentureBeat’s August 13 coverage notes that Google’s latest released general-purpose Pro model remains Gemini 3.1 Pro, introduced in February, with 3.5 Pro still in partner testing as of July and no further timetable given at this announcement. The Flash line is where Google is iterating in public right now.

Release cadence
Jul 21 → Aug 13
23days

The gap between 3.6 Flash and 3.7 Flash. Google credits developer feedback and algorithmic innovations; the model card confirms 3.7 Flash is built directly on 3.6 Flash.

Press rounds to “three weeks”
Context window
1,048,576 in / 65,536 out
1M

Text, image, video, audio, and PDF inputs; text output only. Function calling, structured output, code execution, context caching, batch API, and search grounding all supported at launch.

Computer use in Preview
Knowledge cutoff
Per the model card
Mar2026

The card states a March 2026 cutoff, with the caveat that some domains may reflect knowledge closer to January 2025, in line with the Gemini 3 model family.

Vendor-stated

02PricingThe half-price claim, with its asterisk restored.

Google’s Antigravity team put the pitch in one sentence: “3.7 Flash is available through the end of the year at an introductory price of $0.75/1M input tokens and $3.75/1M output tokens - half the original 3.6 Flash cost per million tokens.” Read carefully — original is doing all the work. That sentence is arithmetically true against 3.6 Flash’s July 21 launch rate of $1.50/$7.50. It is not true against what 3.6 Flash costs now.

At the time of writing, Google’s own pricing page carries a separate 3.6 Flash block showing the identical introductory rate — $0.75 in, $3.75 out, through December 31, 2026 — and the 3.7 Flash model card’s own comparison table lists 3.6 Flash’s price cells with the same asterisked $0.75/$3.75 figures. The same Antigravity post confirms it directly: introductory pricing also applies to 3.6 Flash through December 31, 2026. Every launch-day outlet we checked repeated “half the price of 3.6 Flash” without noting the equalization.

The corrected frame
Three statements are honest; the popular fourth is not. Gemini 3.7 Flash costs half what the workhorse tier cost before August 13. It costs the same as 3.6 Flash today. And from January 1, 2027 you pay exactly the old 3.6 rate — $1.50/$7.50 — for the smarter model. What it never costs is half of 3.6 Flash’s current price.

The full ladder below is the table we could not find anywhere else in launch coverage — the Flash tier’s complete price history from 3.5 Flash through the January 2027 step-up, every cell from Google’s live pricing page or our own dated coverage of the 3.6 launch.

The Gemini Flash pricing ladder from 3.5 Flash through the January 2027 standard rate, showing input and output price per million tokens, the validity window of each rate, and the change each step represents.
Rate stepInput $/1MOutput $/1MIn effectWhat changed
Before the workhorse era
Gemini 3.5 Flash$1.50$9.00Flat list, no expiryThe pre-3.6 baseline
The 3.6 Flash era — July 21 to August 12, 2026
Gemini 3.6 Flash at launch$1.50$7.50Jul 21 – Aug 12, 2026Output cut $1.50 (−16.7%) vs 3.5 Flash; input unchanged
From August 13, 2026 — both workhorses, identical rates
Gemini 3.6 Flash, repriced$0.75$3.75Aug 13 – Dec 31, 2026Same-day intro cut — the step no coverage printed
Gemini 3.7 Flash, intro$0.75$3.75Aug 13 – Dec 31, 202650% below the pre-Aug-13 workhorse rate
From January 1, 2027 — the dated cliff
Both 3.6 & 3.7 Flash, standard$1.50$7.50From Jan 1, 20272× the intro rate — exactly 3.6 Flash’s July launch price

The supporting rates move on the same schedule. Context caching for 3.7 Flash is $0.075 per million tokens through December 31 — 10% of the intro input price — stepping to $0.15 from January 1, with cache storage at $0.50 per million tokens per hour doubling to $1.00 on the same date. The batch API takes 50% off whatever the current rate is: $0.375/$1.875 through December, $0.75/$3.75 after.

One labeling trap for anyone comparing marketplaces: OpenRouter lists google/gemini-3.7-flash with a headline $0.375/$1.875 badge marked 50% off. That is a Vertex-specific provider discount stacked on top of Google’s intro rate — not Google’s official price. OpenRouter’s Google AI Studio provider row shows $0.75/$3.75 with no badge, matching the official introductory rate exactly. Quote the provider row, not the badge.

Against the mid-tier field, the intro rate is aggressive even after the correction. Working from the model card’s own comparison row and the competitors’ live pricing pages: Claude Sonnet 5 lists at $2.00/$10.00 per million and GPT-5.6 Terra at $2.00/$12.00 — so at intro rates, 3.7 Flash runs 62.5% below Sonnet 5 on both input and output, and 68.75% below Terra on output. Even at the January 2027 standard rate, it stays 25% below Sonnet 5 on both sides. Meta’s Muse Spark 1.2 — listed at $1.25/$4.25 in Google’s own comparison table — is the one neighbor that flips: cheaper than the Flash standard rate, pricier than the intro. If your architecture fans out to subagent fleets, the intro-versus-standard doubling is the number to model — the same math we walked through in our subagent cost breakdown for Flash-tier models.

03BenchmarksGoogle’s own table: 7 of 13, not a clean sweep.

The model card publishes a five-model comparison across 13 benchmarks — 3.7 Flash, 3.6 Flash, Claude Sonnet 5, GPT-5.6 Terra, and Muse Spark 1.2. Our own count of Google’s table: 3.7 Flash posts the best score on 7 of 13 rows. GPT-5.6 Terra takes four — DeepSWE v1.1, both Terminal-bench versions, and OSWorld-2.0 — Muse Spark 1.2 wins GDPval-AA v2 outright, and Terra and Muse tie for the top Artificial Analysis Intelligence Index score at 57. That is a strong showing for a model priced at a fraction of its comparators, and it is not a frontier-leadership claim; Google printed its own losses. For where the previous generation stood, see how 3.6 Flash stacked up against Sonnet 5 and its rivals.

All 13 benchmarks from the Gemini 3.7 Flash model card comparison table, showing scores for Gemini 3.7 Flash, Gemini 3.6 Flash, Claude Sonnet 5, GPT-5.6 Terra, and Muse Spark 1.2, grouped by benchmark family, with the best score per row in bold.
Benchmark3.7 Flash3.6 FlashSonnet 5GPT-5.6 TerraMuse Spark 1.2
Coding & software engineering
FrontierCode 1.143.6%34.4%42.7%41.3%
DeepSWE v1.165.3%48.6%53.8%69.6%54.9%
Code Arena (Elo)15881538154115231535
Agentic & computer use
Terminal-bench 2.185.8%78.0%80.4%87.4%82.9%
Terminal-bench 3.014.9%5.4%14.6%20.8%
OSWorld-2.047.9%33.8%50.2%
AutomationBench30.4%17.0%10.7%23.6%
Knowledge & document work
GDPval-AA v2 (Elo)15251422159815781628
Harvey LAB-AA (legal)90.7%85.1%90.1%85.2%
GDP.pdf (PDF comprehension)34.0%22.0%28.0%24.7%16.0%
Long context, video & composite
GDM-MRCR v2 (Google-authored)97.0%91.8%81.5%93.5%
LVBench (long video)85.4%84.2%68.5%78.9%
AA Intelligence Index5652555757

Three footnotes belong next to that table. First, an em-dash means no score published — the model card itself shows Sonnet 5’s OSWorld-2.0 cell as a bare dash, which is an absent number, never a zero or a loss. Second, GDM-MRCR is a Google DeepMind–authored long-context eval, not a third-party benchmark — weigh that row accordingly. Third, Google is inconsistent with itself on one cell: the model card’s table reads 48.6% for 3.6 Flash on DeepSWE v1.1 while Google’s own launch blog rounds the same figure to 49.0% — a small discrepancy, but a reminder that even vendor self-citations deserve checking.

The pattern in the wins is coherent. The biggest generational jumps land on production-shaped work: DeepSWE up 16.7 points over 3.6 Flash, AutomationBench up 13.4, GDP.pdf up 12.0, FrontierCode up 9.2. The losses cluster in exactly one place — the hardest agentic computer-and-terminal benchmarks, where GPT-5.6 Terra keeps the crown at more than three times the intro output price. That trade is the entire commercial argument of this release.

04Independent SignalThe verdict that changed since July.

When 3.6 Flash launched in July, our honest read was “cheaper, not smarter” — the price moved, the independent intelligence signal barely did. This launch earns a different verdict, and not from Google. Artificial Analysis’s August 13 article reports Gemini 3.7 Flash (high) scoring 56 on its Intelligence Index against 52 for 3.6 Flash — a genuine four-point gain — and states that this places Gemini 3.7 Flash on its Intelligence vs. Time per Task Pareto frontier. AA attributes the gain specifically to agentic evaluations rather than to the price cut.

The same article carries the one performance figure that exists anywhere for this launch: output throughput of roughly 340 tokens per second, which AA describes as nearly three times the output speed of GPT-5.6 Terra. That is a throughput measurement, not a latency number — no primary source publishes any latency figure, and the model card itself only discloses “occasional slowness or timeout issues” at launch, unquantified.

AA Intelligence Index · 3.7 Flash vs its price neighborhood

Source: Artificial Analysis Intelligence Index, as reported in AA's Aug 13, 2026 article and the Gemini 3.7 Flash model card
Gemini 3.6 FlashThe predecessor baseline
52
Claude Sonnet 5$2.00 / $10.00 per 1M
55
Gemini 3.7 Flash (high)$0.75 / $3.75 per 1M intro
56
GPT-5.6 Terra$2.00 / $12.00 per 1M
57
Muse Spark 1.2$1.25 / $4.25 per 1M
57

Google’s product page adds a set of vendor-curated testimonials from named engineering leads at Browser Use, Nunu.ai, Hebbia, Box, and Databricks. These are marketing, selected by Google — but the specifics are unusually concrete, and the Browser Use figure is the kind of claim a customer can be held to:

"The Gemini 3.7 Flash agent was 35% cheaper than 3.6 Flash, with a +8% observed prompt-cache hit rate and fewer tool errors."— Gregor Zunic, Co-founder & CTO, Browser Use — Google-curated testimonial

Read the two signals together and the interpretation writes itself. The vendor benchmark table says the model got meaningfully better at production-shaped agentic work; the one independent scorer agrees, on its own composite and its own speed measurements, and attributes the gain to the same category of evaluation. That alignment between vendor claim and third-party measurement is what was missing from the 3.6 Flash launch — and it is the reason this release deserves evaluation cycles rather than a shrug.

05Thinking LevelsThree thinking levels, priced identically.

Gemini 3.7 Flash exposes low, medium, and high thinking levels through the API — and, unlike some sibling models, minimal is explicitly unsupported and returns an API error. The interesting operational detail is how little intelligence you give up walking down the ladder: Artificial Analysis scores the three levels at 51, 53, and 56 respectively, against a median of 34 across similarly-priced reasoning models. Even the low setting clears the price tier’s median comfortably.

Fast lane
Thinking: low
AA Index 51

The bulk-work setting for classification, extraction, and routing. Still 17 points above the 34-point median AA reports for similarly-priced reasoning models.

Cheapest effective tokens
Middle
Thinking: medium
AA Index 53

The balanced default for multi-step work where full reasoning depth is not worth the extra output tokens on every call.

Balanced default
Headline
Thinking: high
AA Index 56

The benchmark configuration behind the launch numbers and the Pareto-frontier placement. Note: minimal is not supported at all — the API returns an error.

Benchmark config

Day-one availability is broad: Google AI Studio, the Gemini API, Android Studio, Google Antigravity across IDE, CLI, and SDK, the Gemini Enterprise Agent Platform and Enterprise app, and Gemini Spark for AI Pro and Ultra subscribers in 160+ countries. If you work in Google’s agentic IDE, our guide to using Gemini models inside Google Antigravity covers the workflow, and Gemini Spark — where 3.7 Flash is now live for consumers — is the most visible consumer surface. Computer use ships in Preview status; the Live API is not supported. Search grounding and Google Maps grounding both work at launch, alongside Flex and Priority inference tiers.

06PlaybookWhat to do with a repriced workhorse tier.

The decision is not “switch or don’t.” It is four separate calls, one per workload class — and one calendar entry.

Already on 3.6 Flash
Upgrade the model ID, not the budget line

Same price, same base lineage, higher scores on 13 of 13 rows against its predecessor. There is no pricing reason to stay on 3.6 Flash — the switch is a config change. Re-run your own evals first; vendor deltas are directional, not guarantees.

Move to 3.7 Flash
Bulk & subagent fleets
High-volume agentic pipelines

At $0.75/$3.75 intro with batch halving that again, this is the aggressive quote for fan-out work — but model your 2027 budget at $1.50/$7.50, because the doubling lands on January 1 and applies to both workhorses.

Adopt at intro, budget at standard
Hardest agentic work
Terminal, computer use, long-horizon SWE

Google's own table gives GPT-5.6 Terra the lead on Terminal-bench 2.1 and 3.0, OSWorld-2.0, and DeepSWE v1.1. Where those benchmarks resemble your workload, the premium model still earns its premium.

Keep Terra in the mix
Knowledge-work Elo
Document & analyst workflows

Sonnet 5, Terra, and Muse Spark 1.2 all out-rank 3.7 Flash on GDPval-AA v2. For work that lives and dies on knowledge-work quality rather than throughput cost, evaluate before consolidating on Flash.

Test per workload

Looking forward, the strategic read is about cadence economics, not this single release. Google has now shipped two workhorse iterations in 23 days, priced the newer one to remove any reason to stay on the older, and used an intro window to pull adoption forward ahead of a dated price step. If that pattern holds, the rational posture for teams is to keep model choice a configuration value rather than an architectural commitment — the vendors are iterating faster than procurement cycles. Building that kind of swap-friendly routing layer is exactly the sort of work our AI transformation engagements handle, from eval harnesses to cost-per-outcome dashboards.

07ConclusionSmarter is verified; half price is a frame.

The shape of the workhorse tier, August 2026

The capability claim held up. The pricing claim needed an asterisk.

Gemini 3.7 Flash is the rare fast-follow release where the independent evidence is stronger than the marketing. A four-point gain on the Artificial Analysis Intelligence Index, a Pareto-frontier placement, and the biggest jumps landing on production-shaped agentic work — 23 days after the model it is built on.

The pricing story is where precision matters. Half price is true against the workhorse tier’s pre-August-13 rate, and only that. Today, 3.6 and 3.7 Flash cost the same $0.75/$3.75; from January 1, 2027, both cost the same $1.50/$7.50 — the exact rate 3.6 Flash launched at. What Google actually sold is more capability at an unchanged standard price, wrapped in an intro discount with a calendar date on it.

The practical takeaway fits in two lines. If you run Flash-tier workloads, the upgrade is free and the evals are worth running this week. And whatever you adopt, write January 1, 2027 into the budget now — the models got smarter, but the discount is the part that expires.

Put the right model on every workload

A workhorse model is only cheap if it lands on the right workloads.

Our team helps businesses evaluate, route, and operate frontier models — benchmarking Flash-tier workhorses against premium tiers on your actual workloads, with cost-per-outcome math built in from day one.

Free consultationExpert guidanceTailored solutions
What we work on

Model-routing engagements

  • Flash-tier vs premium-tier evals on your own prompts
  • Subagent fleet economics — intro vs standard rate modeling
  • Swap-friendly routing layers across Google, OpenAI, Anthropic
  • Agentic workflow buildouts with cost-per-outcome dashboards
  • January 2027 price-step budget planning
FAQ · Gemini 3.7 Flash

The questions we get every week.

Gemini 3.7 Flash is Google's workhorse-tier model, released August 13, 2026 — 23 days after Gemini 3.6 Flash, which it is explicitly built on per the model card. It ships as a stable model (ID gemini-3.7-flash, no preview suffix) with a 1,048,576-token input window and 65,536-token output limit, accepting text, image, video, audio, and PDF inputs with text output only. It launched day-one across Google AI Studio, the Gemini API, Android Studio, Google Antigravity, the Gemini Enterprise platform, and Gemini Spark for AI Pro and Ultra subscribers in 160+ countries. Google attributes the unusually fast turnaround to developer feedback and algorithmic innovations rather than a new base model.
Related dispatches

Continue exploring frontier releases.