On September 2, 2026 Google released Gemini 3.8 Flash, its third Flash model in six weeks by Google’s own count, at the same introductory price as the 3.7 Flash it replaces: $0.75 per million input tokens and $3.75 per million output tokens. The launch post carries a footnote our August coverage could only estimate: “Introductory price expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply.”
That is the doubling our August 14 post on budgeting for introductory pricing treated as the likely outcome. It is now Google’s stated plan, with a date. The second half of the story arrived the same day from Artificial Analysis, which measured the model at 59 on its Intelligence Index, three points above 3.7 Flash, and found that cost per task rose about 40% even though the per-token price did not move. This post puts the two facts together: the price you pay today, the price you pay in four months, and the token count that sits between them.
- 01Same intro price, and now a published end date.$0.75 input and $3.75 output per million tokens until December 31, 2026, then $1.50 and $7.50 from January 1, 2027, per Google’s launch post. Budget the second number.
- 02Per-token price flat, per-task cost up about 40%.Artificial Analysis measured $0.58 per Intelligence Index task against $0.40 for 3.7 Flash, driven by roughly 30% more output tokens per task and more turns on agentic tests. Google says the model “works harder” by design.
- 03Three points smarter on an independent index.59 on the Artificial Analysis Intelligence Index at high reasoning, level with GPT-5.6 Sol at xhigh and Grok 4.6 at medium. Google’s own numbers, including 54.9% on HLE-Verified, are vendor-run.
- 043.7 Flash stays, and lower effort is the cheap route.Google says 3.7 Flash “remains fully supported for efficiency-first workloads.” On 3.8, Artificial Analysis puts cost per task at $0.41 on medium and $0.24 on low reasoning.
01 — The releaseWhat shipped on September 2.
Gemini 3.8 comes in two variants that share one base model. Gemini 3.8 Flash is the general release, which Google describes as its “most intelligent workhorse model” with gains over 3.7 Flash in software engineering, agentic tasks and multi-step reasoning in specialised domains. Gemini 3.8 Flash Cyber is a cybersecurity variant with looser safety mitigations, available only through a new Fairwind Program for “trusted government authorities, as well as critical infrastructure operators and software maintainers.” This post is about the first one. The gated variant belongs with the other vetted-access security models, alongside the 3.5 Flash Cyber release it succeeds.
The general model is live where 3.7 Flash was: the Gemini API and AI Studio, Google Antigravity, Android Studio, Gemini Enterprise, and for Google AI Pro and Ultra subscribers the Gemini app, AI Mode in Search and Gemini in Sheets. The AI Mode change means the model answering Google searches has moved again, three weeks after our note on the 3.7 Flash swap. The context window is one million tokens, unchanged. The table below sets the two generations side by side with the source of each figure.
| Figure | Gemini 3.7 Flash | Gemini 3.8 Flash | Source |
|---|---|---|---|
| Release date | Aug 13, 2026 | Sep 2, 2026 | |
| Introductory price, input / output per M | $0.75 / $3.75 | $0.75 / $3.75 | |
| Price from January 1, 2027 | not stated at launch | $1.50 / $7.50 | Google, launch-post footnote |
| Intelligence Index, high reasoning | 56 | 59 | Artificial Analysis, Sep 2 |
| Cost per index task, high reasoning | $0.40 | $0.58 | Artificial Analysis, Sep 2 |
| Time per index task, high reasoning | 2.2 min | 2.5 min | Artificial Analysis, Sep 2 |
| Context window | 1M tokens | 1M tokens | Artificial Analysis; OpenRouter lists 1,048,576 |
| Status after the new release | “remains fully supported for efficiency-first workloads” | current Flash model |
02 — The number that mattersThe price line Google wrote down.
When 3.7 Flash launched on August 13 at $0.75 and $3.75, Google called the price introductory, and our launch coverage treated it as a half-price window on a model that would eventually cost more, without a standard rate to quote. The 3.8 Flash post settles it in a footnote: the introductory price expires on December 31, 2026, and from January 1, 2027 the model costs $1.50 per million input tokens and $7.50 per million output tokens. Both numbers are exactly double the current ones.
Two things follow for anyone running Flash-class models in production. First, the promotion now has an end date that a finance team can put in a spreadsheet, which is more than most vendors give you. Second, the per-token rate is only half the bill; the honest comparison is per task, and that is the subject of the next section. The row for 3.8 Flash on our frontier model API price index carries both dates.
The launch post does not state a new price for 3.7 Flash, which “remains fully supported.” Whether 3.7 keeps its $0.75 and $3.75 rate past December 31 is not stated anywhere we could find on September 2. Treat a 3.7 price as unchanged until Google publishes otherwise, and treat the January doubling as applying to 3.8 Flash specifically, because that is what the footnote says.
03 — The hidden line item“Works harder” means more tokens.
Google is unusually direct about the mechanism behind the gains: “3.8 Flash works harder. On complex tasks, it exhibits greater diligence, executing extra reasoning steps, and calling tools iteratively. At times, the model might use more tokens to maximize performance, especially at higher effort levels.” Artificial Analysis measured what that costs. At high reasoning, 3.8 Flash averaged about 48,000 output tokens per Intelligence Index task, a 30% increase on 3.7 Flash, and took more turns on the agentic evaluations. The result is a cost per task of $0.58 against $0.40 for 3.7 Flash, up roughly 40% on an unchanged per-token price, and a time per task of 2.5 minutes against 2.2.
Put the two sections together and the arithmetic is simple, so we will do it once and label it as ours. If the model’s token usage on a task stays where Artificial Analysis measured it, the same $0.58 task costs about $1.16 at the January rate, against $0.40 for the 3.7 Flash you may be running today. That is a near tripling of per-task cost across one model generation and one calendar boundary, none of which shows up in the headline “same price” line. Effort is the lever Google offers: Artificial Analysis puts 3.8 Flash at $0.41 per task on medium reasoning, where it scores 57, and $0.24 on low, where it scores 52 and matches 3.6 Flash at high in about a third of the time.
High
The headline configuration. Three points above 3.7 Flash, level with GPT-5.6 Sol at xhigh and Grok 4.6 at medium, and on the intelligence-versus-cost frontier per Artificial Analysis. Also the configuration where the extra tokens are most visible.
Medium
Matches GPT-5.6 Terra at max and Muse Spark 1.2 at xhigh on the index, at roughly the cost per task 3.7 Flash delivered on high. The natural setting for teams who want the new model at last month’s bill.
Low
Scores where 3.6 Flash scored on high, at 30% lower cost per task and about a third of the time per task. Artificial Analysis places it on the intelligence-versus-time frontier. Not a substitute for high on long agentic runs.
04 — EvidenceThe benchmarks, labelled.
Google’s launch post gives one number and several claims. The number is 54.9% on HLE-Verified, the verified version of Humanity’s Last Exam. The claims are that 3.8 Flash “outperforms most larger frontier models” on DeepSWE v1.1, a long-horizon software engineering test, and beats 3.7 Flash and other frontier models on Vals Finance Agent V2 and Harvey’s Legal Agent Benchmark. No comparison table accompanies the claims, so we print them as what they are: vendor-run and vendor-selected.
The independent picture comes from Artificial Analysis, which ran its full index on the day of release. The three-point gain over 3.7 Flash is “primarily driven by stronger performance on agentic evaluations”: tool use on τ³-Banking, where the model gains 12 points to 45%, coding on Terminal-Bench 2.1, and real-world tasks on GDPval-AA v2. Those are the same categories where the extra tokens are spent, which is consistent with Google’s “works harder” explanation rather than a jump in raw capability. Readers who want the 3.7 Flash baseline against Sonnet 5 and GPT-5.6 Terra can find it in our August benchmark comparison; the same caveats about benchmark versions apply.
05 — DecisionWhat to do before January.
The practical question is not whether 3.8 Flash is better. On the independent index it is, by three points. The question is which configuration you run and what you tell finance about Q1. The router below is how we would answer it for a team already on 3.7 Flash, using only the numbers above.
None of this needs a re-architecture. It needs a token counter on every task type, a model field that can change per route, and a calendar entry for December 31. Teams that built those three things for the 3.7 Flash window, as our ad-ops automation note recommended, get to make this decision with data. Teams that did not can start now; there are four months of intro pricing left, and our AI transformation practice builds exactly this kind of per-route cost visibility.
06 — ConclusionSame sticker, bigger basket.
The per-token price did not move. The tokens did, and in four months the price does too.
Gemini 3.8 Flash is a better model than 3.7 Flash by the one independent measure available on launch day, and it is sold at the same introductory rate. It also uses about 30% more output tokens per task, costs about 40% more per task as a result, and carries a published date on which the per-token price doubles.
The useful thing Google did was write the standard price down. The useful thing a buyer can do is measure tokens per task on their own workload this month, pick the reasoning level that fits, and forecast January at $1.50 and $7.50 rather than at the price on the sticker today.