AI DevelopmentNew Release8 min readPublished September 2, 2026

The intro price is unchanged. The date it stops being introductory is now in writing

Gemini 3.8 Flash Costs the Same Until It Doubles in January

Google released Gemini 3.8 Flash on September 2, three weeks after 3.7 Flash, at the same $0.75 and $3.75 per million tokens. The launch post also says when that price ends and what replaces it: $1.50 and $7.50 from January 1, 2027. An independent measurement found the new model already costs about 40% more per task than its predecessor at the same per-token rate.

DA
Digital Applied Team
Senior strategists · Published Sep 2, 2026
PublishedSep 2, 2026
Read time8 min
EventSep 2, 2026
Intro price, input / output per M
$0.75 / $3.75
identical to 3.7 Flash; expires December 31, 2026 (Google)
Standard price from January 1, 2027
$1.50 / $7.50
stated in Google’s launch post footnote
Artificial Analysis Intelligence Index
59
high reasoning; 3.7 Flash scored 56 (independent)
Cost per index task vs 3.7 Flash
+40%
$0.58 vs $0.40, same per-token price (Artificial Analysis)

On September 2, 2026 Google released Gemini 3.8 Flash, its third Flash model in six weeks by Google’s own count, at the same introductory price as the 3.7 Flash it replaces: $0.75 per million input tokens and $3.75 per million output tokens. The launch post carries a footnote our August coverage could only estimate: “Introductory price expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply.”

That is the doubling our August 14 post on budgeting for introductory pricing treated as the likely outcome. It is now Google’s stated plan, with a date. The second half of the story arrived the same day from Artificial Analysis, which measured the model at 59 on its Intelligence Index, three points above 3.7 Flash, and found that cost per task rose about 40% even though the per-token price did not move. This post puts the two facts together: the price you pay today, the price you pay in four months, and the token count that sits between them.

Key takeaways
  1. 01
    Same intro price, and now a published end date.$0.75 input and $3.75 output per million tokens until December 31, 2026, then $1.50 and $7.50 from January 1, 2027, per Google’s launch post. Budget the second number.
  2. 02
    Per-token price flat, per-task cost up about 40%.Artificial Analysis measured $0.58 per Intelligence Index task against $0.40 for 3.7 Flash, driven by roughly 30% more output tokens per task and more turns on agentic tests. Google says the model “works harder” by design.
  3. 03
    Three points smarter on an independent index.59 on the Artificial Analysis Intelligence Index at high reasoning, level with GPT-5.6 Sol at xhigh and Grok 4.6 at medium. Google’s own numbers, including 54.9% on HLE-Verified, are vendor-run.
  4. 04
    3.7 Flash stays, and lower effort is the cheap route.Google says 3.7 Flash “remains fully supported for efficiency-first workloads.” On 3.8, Artificial Analysis puts cost per task at $0.41 on medium and $0.24 on low reasoning.

01The releaseWhat shipped on September 2.

Gemini 3.8 comes in two variants that share one base model. Gemini 3.8 Flash is the general release, which Google describes as its “most intelligent workhorse model” with gains over 3.7 Flash in software engineering, agentic tasks and multi-step reasoning in specialised domains. Gemini 3.8 Flash Cyber is a cybersecurity variant with looser safety mitigations, available only through a new Fairwind Program for “trusted government authorities, as well as critical infrastructure operators and software maintainers.” This post is about the first one. The gated variant belongs with the other vetted-access security models, alongside the 3.5 Flash Cyber release it succeeds.

The general model is live where 3.7 Flash was: the Gemini API and AI Studio, Google Antigravity, Android Studio, Gemini Enterprise, and for Google AI Pro and Ultra subscribers the Gemini app, AI Mode in Search and Gemini in Sheets. The AI Mode change means the model answering Google searches has moved again, three weeks after our note on the 3.7 Flash swap. The context window is one million tokens, unchanged. The table below sets the two generations side by side with the source of each figure.

Gemini 3.7 Flash and 3.8 Flash as of September 2, 2026. Prices and dates from Google’s launch posts; index scores and cost per task from Artificial Analysis, measured at high reasoning.
FigureGemini 3.7 FlashGemini 3.8 FlashSource
Release dateAug 13, 2026Sep 2, 2026Google
Introductory price, input / output per M$0.75 / $3.75$0.75 / $3.75Google
Price from January 1, 2027not stated at launch$1.50 / $7.50Google, launch-post footnote
Intelligence Index, high reasoning5659Artificial Analysis, Sep 2
Cost per index task, high reasoning$0.40$0.58Artificial Analysis, Sep 2
Time per index task, high reasoning2.2 min2.5 minArtificial Analysis, Sep 2
Context window1M tokens1M tokensArtificial Analysis; OpenRouter lists 1,048,576
Status after the new release“remains fully supported for efficiency-first workloads”current Flash modelGoogle

02The number that mattersThe price line Google wrote down.

When 3.7 Flash launched on August 13 at $0.75 and $3.75, Google called the price introductory, and our launch coverage treated it as a half-price window on a model that would eventually cost more, without a standard rate to quote. The 3.8 Flash post settles it in a footnote: the introductory price expires on December 31, 2026, and from January 1, 2027 the model costs $1.50 per million input tokens and $7.50 per million output tokens. Both numbers are exactly double the current ones.

Two things follow for anyone running Flash-class models in production. First, the promotion now has an end date that a finance team can put in a spreadsheet, which is more than most vendors give you. Second, the per-token rate is only half the bill; the honest comparison is per task, and that is the subject of the next section. The row for 3.8 Flash on our frontier model API price index carries both dates.

What Google did not say

The launch post does not state a new price for 3.7 Flash, which “remains fully supported.” Whether 3.7 keeps its $0.75 and $3.75 rate past December 31 is not stated anywhere we could find on September 2. Treat a 3.7 price as unchanged until Google publishes otherwise, and treat the January doubling as applying to 3.8 Flash specifically, because that is what the footnote says.

03The hidden line item“Works harder” means more tokens.

Google is unusually direct about the mechanism behind the gains: “3.8 Flash works harder. On complex tasks, it exhibits greater diligence, executing extra reasoning steps, and calling tools iteratively. At times, the model might use more tokens to maximize performance, especially at higher effort levels.” Artificial Analysis measured what that costs. At high reasoning, 3.8 Flash averaged about 48,000 output tokens per Intelligence Index task, a 30% increase on 3.7 Flash, and took more turns on the agentic evaluations. The result is a cost per task of $0.58 against $0.40 for 3.7 Flash, up roughly 40% on an unchanged per-token price, and a time per task of 2.5 minutes against 2.2.

Put the two sections together and the arithmetic is simple, so we will do it once and label it as ours. If the model’s token usage on a task stays where Artificial Analysis measured it, the same $0.58 task costs about $1.16 at the January rate, against $0.40 for the 3.7 Flash you may be running today. That is a near tripling of per-task cost across one model generation and one calendar boundary, none of which shows up in the headline “same price” line. Effort is the lever Google offers: Artificial Analysis puts 3.8 Flash at $0.41 per task on medium reasoning, where it scores 57, and $0.24 on low, where it scores 52 and matches 3.6 Flash at high in about a third of the time.

Reasoning level
High
Index 59 · $0.58 per task · 2.5 min

The headline configuration. Three points above 3.7 Flash, level with GPT-5.6 Sol at xhigh and Grok 4.6 at medium, and on the intelligence-versus-cost frontier per Artificial Analysis. Also the configuration where the extra tokens are most visible.

Agentic work
Reasoning level
Medium
Index 57 · $0.41 per task

Matches GPT-5.6 Terra at max and Muse Spark 1.2 at xhigh on the index, at roughly the cost per task 3.7 Flash delivered on high. The natural setting for teams who want the new model at last month’s bill.

Most workloads
Reasoning level
Low
Index 52 · $0.24 per task · 0.8 min

Scores where 3.6 Flash scored on high, at 30% lower cost per task and about a third of the time per task. Artificial Analysis places it on the intelligence-versus-time frontier. Not a substitute for high on long agentic runs.

Latency-bound

04EvidenceThe benchmarks, labelled.

Google’s launch post gives one number and several claims. The number is 54.9% on HLE-Verified, the verified version of Humanity’s Last Exam. The claims are that 3.8 Flash “outperforms most larger frontier models” on DeepSWE v1.1, a long-horizon software engineering test, and beats 3.7 Flash and other frontier models on Vals Finance Agent V2 and Harvey’s Legal Agent Benchmark. No comparison table accompanies the claims, so we print them as what they are: vendor-run and vendor-selected.

The independent picture comes from Artificial Analysis, which ran its full index on the day of release. The three-point gain over 3.7 Flash is “primarily driven by stronger performance on agentic evaluations”: tool use on τ³-Banking, where the model gains 12 points to 45%, coding on Terminal-Bench 2.1, and real-world tasks on GDPval-AA v2. Those are the same categories where the extra tokens are spent, which is consistent with Google’s “works harder” explanation rather than a jump in raw capability. Readers who want the 3.7 Flash baseline against Sonnet 5 and GPT-5.6 Terra can find it in our August benchmark comparison; the same caveats about benchmark versions apply.

05DecisionWhat to do before January.

The practical question is not whether 3.8 Flash is better. On the independent index it is, by three points. The question is which configuration you run and what you tell finance about Q1. The router below is how we would answer it for a team already on 3.7 Flash, using only the numbers above.

Long agentic runs where 3.7 Flash was dropping tasks
Move to 3.8 Flash on high. Expect roughly 40% more per task now and budget double that from January 1. Measure output tokens per task on your own workload before and after; Artificial Analysis’s 30% is an average across its index.
3.8 Flash · high
Steady production traffic that was fine on 3.7 Flash
Stay on 3.7 Flash, which Google says remains fully supported, and watch for a 3.7 price statement. Trial 3.8 Flash on medium, which lands near 3.7’s cost per task with a one-point index gain.
3.7 Flash, or 3.8 · medium
Latency-bound work such as ad-ops, Sheets formulas, classification
3.8 Flash on low reasoning: 0.8 minutes per task on the index and $0.24 per task. Our earlier note on 3.7 Flash for ad-ops automation still applies; the January rise affects this tier least in absolute dollars.
3.8 Flash · low
Anyone forecasting 2027 spend on Flash-class models
Use $1.50 and $7.50 for 3.8 Flash from January 1, 2027, multiply by measured tokens per task rather than last quarter’s, and keep a column for a 3.7 Flash price change that Google has not announced.
Finance

None of this needs a re-architecture. It needs a token counter on every task type, a model field that can change per route, and a calendar entry for December 31. Teams that built those three things for the 3.7 Flash window, as our ad-ops automation note recommended, get to make this decision with data. Teams that did not can start now; there are four months of intro pricing left, and our AI transformation practice builds exactly this kind of per-route cost visibility.

06ConclusionSame sticker, bigger basket.

Gemini 3.8 Flash

The per-token price did not move. The tokens did, and in four months the price does too.

Gemini 3.8 Flash is a better model than 3.7 Flash by the one independent measure available on launch day, and it is sold at the same introductory rate. It also uses about 30% more output tokens per task, costs about 40% more per task as a result, and carries a published date on which the per-token price doubles.

The useful thing Google did was write the standard price down. The useful thing a buyer can do is measure tokens per task on their own workload this month, pick the reasoning level that fits, and forecast January at $1.50 and $7.50 rather than at the price on the sticker today.

Cost visibility per route

Know what a task costs before the price changes.

We put a token counter and a model field on every agent route we build, so that a price change or a hungrier model shows up in a forecast before it shows up on an invoice.

Free consultationExpert guidanceTailored solutions
What we work on

Model cost engagements

  • Per-route token and cost instrumentation
  • Reasoning-level sweeps on your own tasks
  • Model routing with a per-route model field
  • Intro-pricing calendars and Q1 forecasts
  • Migration between Flash-class generations
FAQ · Gemini 3.8 Flash

The questions we get about Gemini 3.8 Flash pricing.

$0.75 per million input tokens and $3.75 per million output tokens, the same introductory price as Gemini 3.7 Flash. Google’s launch post states the introductory price expires December 31, 2026 and that $1.50 input and $7.50 output per million apply from January 1, 2027. Cached input keeps the same 90% discount per Artificial Analysis.
Related dispatches

Continue exploring Gemini and model pricing.