BusinessCost Playbook12 min readPublished August 14, 2026

Two trap shapes · calendar and threshold · one budgeting rule

Budgeting AI Agents on Intro Pricing Built to Double

Gemini 3.6 Flash and 3.7 Flash share one posted price schedule: $0.75 per million input tokens and $3.75 output through December 31, 2026, then $1.50/$7.50 from January 1, 2027. That is a plan, not a certainty — Anthropic cancelled Sonnet 5’s own scheduled step-up on August 11. Here is how to budget an agent fleet when the headline rate has an expiry date.

DA
Digital Applied Team
Senior strategists · Published Aug 14, 2026
PublishedAug 14, 2026
Read time12 min
Sources5 vendor sources
Flash-line intro rate
$0.75/M
input · $3.75 output · to Dec 31
×2 posted Jan 1
Posted 2027 standard
$1.50/M
input · $7.50 output per 1M
Sonnet 5 rate
$2/$10
made permanent Aug 11
step-up cancelled
Terra repricing threshold
272K
whole request · 2× in, 1.5× out

Budgeting AI agents on introductory pricing is now a discipline of its own, because one of the most attractive workhorse rates on the market ships with a published expiry date. Gemini 3.7 Flash bills $0.75 per million input tokens and $3.75 per million output through December 31, 2026 — and Google’s own pricing page posts $1.50/$7.50 from January 1, 2027. Exactly double, on a date you can put in a calendar.

The stakes are simple: on the posted schedule, an agent fleet sized against the intro rate is a fleet that costs twice as much the week the window closes. And the calendar trap is only one of two mechanics currently separating the headline rate from the rate you actually pay — Grok 4.6 and GPT-5.6 Terra reprice an entire request the moment a prompt crosses a token threshold, no calendar required.

This playbook covers the full Gemini Flash schedule (including the detail most coverage missed: 3.6 Flash carries the identical step-up date), the one budgeting rule that makes intro pricing safe to use, the Anthropic reversal that proves step-ups are not inevitable, the threshold traps, and a trigger matrix you can hand to whoever owns the forecast.

Key takeaways
  1. 01
    The doubling is Flash-line-wide, not a 3.7 quirk.Google posts the same schedule for Gemini 3.6 Flash and 3.7 Flash: $0.75/$3.75 per 1M through December 31, 2026, then $1.50/$7.50 from January 1, 2027. Batch, Flex, and Priority tiers all step up in lockstep.
  2. 02
    Budget at the standard rate; treat the window as margin.Size the fleet against $1.50/$7.50. Every month the intro rate holds is found margin — and if the step-up executes, nothing in the forecast breaks.
  3. 03
    Step-ups sometimes get cancelled.Anthropic scrapped Sonnet 5's scheduled September 1 increase to $3/$15 on August 11, making $2/$10 the permanent standard price. Posted step-ups are plans under competitive pressure, not commitments.
  4. 04
    The second trap has no calendar: token thresholds.Grok 4.6 reprices the whole request at ≥200K prompt tokens ($2/$6 becomes $4/$12). GPT-5.6 Terra does it above 272K input tokens (2× input, 1.5× output). Different thresholds, different multipliers, same mechanic.
  5. 05
    Caching compresses the jump but doesn't cancel it.Gemini's cache-read price stays a constant 10% of the base input rate ($0.075/M now, $0.15/M from January), so heavy-caching workloads see a smaller relative increase on the cached share of tokens.

01The ScheduleOne schedule, two Flash models.

Start with what Google actually posted. The Gemini Developer API pricing page lists Gemini 3.7 Flash at $0.75 per million input tokens and $3.75 per million output “through December 31, 2026,” with a standard rate of $1.50/$7.50 “starting January 1, 2027.” The intro rate is exactly half the standard rate Google has posted for January — 3.7 Flash’s launch-day numbers cover how that price landed against the rest of the market on August 13.

The detail most launch coverage skipped: the same page posts the identical schedule for Gemini 3.6 Flash — same $0.75/$3.75 intro figures, same December 31 window, same $1.50/$7.50 from January 1. The January doubling is not a 3.7-Flash-specific event. It is a Google-Flash-line-wide repricing date, which means teams already running Gemini 3.6 Flash as their workhorse face the same step-up whether or not they ever migrate to 3.7.

Aug 13 release
Gemini 3.7 Flash
$0.75 / $3.75 → $1.50 / $7.50 per 1M

Google's newest Flash model. Intro pricing runs through December 31, 2026; the posted standard rate from January 1, 2027 is exactly double on both legs.

ai.google.dev/gemini-api/docs/pricing
The incumbent
Gemini 3.6 Flash
$0.75 / $3.75 → $1.50 / $7.50 per 1M

Identical schedule, identical dates, confirmed in its own section of the same pricing page. The doubling is a Flash-line event, not a new-model promo.

Same window · same step-up

The tier structure moves in lockstep, which matters for anyone modelling batch or priority traffic. Batch is a flat 50% of Standard at every point in time: $0.375/$1.875 through December 31, then $0.75/$3.75 from January on the posted schedule — meaning batch traffic in 2027 would cost the same as standard traffic does today. Flex pricing is numerically identical to Batch on the current page, and Priority runs a constant 1.8× Standard: $1.35/$6.75 now, $2.70/$13.50 from January. Because every tier is a fixed multiple of Standard, the doubling propagates through the whole surface — there is no tier that escapes it.

One neighbouring model worth noting: Gemini 3.5 Flash-Lite sits at $0.30 input / $2.50 output with no introductory language and no expiry date anywhere in its pricing section. If you need a Google-side rate you can forecast without a step-up scenario, that is currently the stable one.

02The Core RuleBudget at the standard rate. Bank the window.

The rule this whole playbook hangs on: size the fleet against the posted standard rate, and treat the intro window as margin. If the January step-up executes, your forecast already covers it. If it gets cancelled — a real possibility, as the next section shows — you keep the difference as found budget. The only losing move is the common one: sizing agent volume against $0.75/$3.75 and discovering in January that the same traffic costs twice as much.

To make that concrete, here is a blended-cost comparison. One explicit assumption first: the table assumes a 4:1 input:output token split — a round number chosen for illustration, not a measured average. No vendor publishes a “typical” agent workload ratio, and agent fleets vary enormously (tool-heavy loops skew input-heavy; generation-heavy ones do not). Re-run the arithmetic with your own ratio before using any of these blended figures.

Blended cost per one million total tokens across current model list prices, computed at an illustrative 4:1 input to output ratio, comparing Gemini 3.7 Flash introductory and posted standard rates against Claude, OpenAI, and xAI alternatives.
Model · surfaceInput $/1MOutput $/1MBlended $/1M total (4:1 assumed)
Gemini Flash line — intro vs posted standard
Gemini 3.7 Flash · Standard, through Dec 31$0.75$3.75$1.35
Gemini 3.7 Flash · Standard, posted from Jan 1, 2027$1.50$7.50$2.70
Gemini 3.7 Flash · Batch, through Dec 31$0.375$1.875$0.675
Gemini 3.7 Flash · Batch, posted from Jan 1, 2027$0.75$3.75$1.35
Alternatives at current list — no posted step-up
GPT-5.6 Luna · Standard, short context$0.20$1.20$0.40
Claude Haiku 4.5 · Standard$1.00$5.00$1.80
Grok 4.6 · Standard, prompts under 200K$2.00$6.00$2.80
Claude Sonnet 5 · Standard (permanent since Aug 11)$2.00$10.00$3.60
GPT-5.6 Terra · Standard, ≤272K input$2.00$12.00$4.00

Blended cost per 1M total tokens · illustrative 4:1 ratio

Vendor list prices, August 2026. Blended at an assumed 4:1 input:output split — an illustration, not a measured average.
GPT-5.6 Luna standard$0.20 / $1.20 per 1M
$0.40
3.7 Flash batch · intro$0.375 / $1.875 per 1M · to Dec 31
$0.675
3.7 Flash standard · intro$0.75 / $3.75 per 1M · to Dec 31
$1.35
Claude Haiku 4.5 standard$1.00 / $5.00 per 1M
$1.80
3.7 Flash standard · posted 2027$1.50 / $7.50 per 1M · from Jan 1
$2.70
Grok 4.6 standard · under 200K$2.00 / $6.00 per 1M
$2.80
Claude Sonnet 5 standard$2.00 / $10.00 per 1M · permanent
$3.60
GPT-5.6 Terra standard · ≤272K$2.00 / $12.00 per 1M
$4.00

Two honest readings of that table. First, the doubling is real money at fleet scale: at the illustrative 4:1 split, a fleet processing one billion total tokens a month — a hypothetical volume, chosen only to make the arithmetic visible — pays roughly $1,350 a month at the intro rate and $2,700 at the posted standard rate. That is about $16,200 a year of difference riding on one calendar date.

Second — and this is the caveat that keeps the framing honest — even the post-doubling rate is not expensive by current standards. At $1.50/$7.50, 3.7 Flash’s posted 2027 price still undercuts GPT-5.6 Terra’s current $2/$12 on both legs and Sonnet 5’s permanent $2/$10 on both legs. The trap is not that the standard rate is bad; it is that a forecast built on the intro rate silently assumes the discount lasts forever. Whether the model itself earns a place in your stack is a separate question — how it actually performs against Sonnet 5 and GPT-5.6 Terra is its own analysis.

03The Counter-PrecedentAnthropic just cancelled a scheduled doubling.

Here is why this post does not tell you the January doubling is certain: the most prominent scheduled step-up of the summer just failed to execute. Claude Sonnet 5 launched on June 30, 2026 at $2 per million input / $10 per million output, explicitly framed as introductory pricing through August 31 — with a posted increase to $3/$15 slated for September 1. On August 11, three weeks before it would have taken effect, Anthropic called it off. Its platform pricing documentation now records the reversal in plain terms.

Anthropic's own words · Aug 11, 2026
“The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur.” — a scheduled 50% step-up, cancelled three weeks before its effective date.

Read as a precedent, this cuts both ways. It proves posted step-ups are real enough to plan around — Anthropic published one, with dates and figures, and clearly intended it. And it proves they are soft enough to reverse — most plausibly under competitive pressure, in a month when Google priced a capable Flash model at $0.75 input. A budget that assumed Sonnet 5 would cost $3/$15 from September now has margin; a budget that assumed $2/$10 forever happened to be right, but only by luck.

One housekeeping note if you follow our pricing coverage: the August pricing tracker went to press on August 5 and frames Sonnet 5’s $2/$10 as a promo with a dated cliff — accurate when written, superseded by the August 11 reversal. The story moved; that is rather the point of budgeting this way. Treat every posted schedule, including Gemini’s, as the currently posted plan — an input to a forecast, not a fact about the future.

04The Second TrapNo calendar needed: the threshold trap.

Gemini’s trap is calendar-shaped: the rate changes on a date. The second trap is threshold-shaped: the rate changes when a single request gets big enough — and it reprices the whole request, not just the overflow. Two vendors run this mechanic today, with different numbers, and the difference matters.

Grok 4.6 bills $2/$6 per million with cached input at $0.50 for prompts under 200K tokens. At or above 200K, the rate becomes $4/$12 with cached input at $1.00 — and xAI’s pricing docs are explicit that “requests whose prompt reaches the listed token threshold are billed at the higher rate for all tokens in the request.” We covered Grok 4.6’s 200K-token whole-request repricing in depth at launch.

GPT-5.6 Terra runs the same structure with different parameters: $2/$12 per million for prompts up to 272K input tokens, and per OpenAI’s Terra model documentation, “Prompts with >272K input tokens are priced at 2x input and 1.5x output for the full request.” Keep the two distinct: Grok’s threshold is 200K with a 2× multiplier on both legs; Terra’s is 272K with 2× input but only 1.5× output. Same shape, different cliff heights.

Grok 4.6 cliff
Whole-request repricing
200K

Prompt reaches 200K tokens and the entire request bills at $4/$12 instead of $2/$6 — cached input doubles to $1.00/M too. A 2× multiplier on both legs.

docs.x.ai/docs/pricing
Terra cliff
2× input · 1.5× output
272K

Above 272K input tokens the full request reprices to $4/$18. Same whole-request mechanic as Grok, different threshold and a gentler output multiplier.

developers.openai.com
Luna label trap
Standard, not $0.10
$0.20/M

GPT-5.6 Luna's standard short-context rate is $0.20/$1.20. The $0.10/$0.60 figure floating around aggregators is the Batch rate — a surface label, not a discount you get on live traffic.

platform.openai.com pricing

The Luna card above is a third, quieter variant of the same lesson: the rate you see quoted is not always the rate for the surface you run. Aggregator listings have circulated Luna’s $0.10/$0.60 batch rate as if it were the standard price — OpenAI’s own table puts standard short-context at $0.20/$1.20. We hit the same surface-confusion pattern with DeepSeek V4 Flash list prices versus OpenRouter aliases in the cheap-tier repricing wave. Always price the surface you will actually call.

05Trigger MatrixThe hidden second price, by trigger.

Most pricing trackers list rates. What a budget owner actually needs is the trigger — the specific event that moves a workload from the headline rate to the second one. This matrix puts the two mechanics side by side, plus the two entries every forecast should carry as context: the step-up that got cancelled, and the rate that gets misquoted.

Repricing trigger matrix compiled from Google, Anthropic, xAI, and OpenAI pricing pages, showing each model’s headline rate, the mechanic that triggers its second price, the repriced rate, and the current status of each trigger.
ModelHeadline rate (in / out per 1M)TriggerRate once triggeredStatus
Calendar-triggered — a date you can plan around
Gemini 3.7 Flash$0.75 / $3.75Calendar: January 1, 2027$1.50 / $7.50Posted plan — budget for it; not yet certain
Gemini 3.6 Flash$0.75 / $3.75Calendar: January 1, 2027 (identical)$1.50 / $7.50Same posted schedule — the step-up is line-wide
Threshold-triggered — live today, per request
Grok 4.6$2.00 / $6.00Prompt reaches 200K tokens — whole request reprices$4.00 / $12.00Active now — cap prompts or eat a 2× line item
GPT-5.6 Terra$2.00 / $12.00Prompt exceeds 272K input tokens — full request at 2× in, 1.5× out$4.00 / $18.00Active now — different threshold and multipliers than Grok
No live trigger — but read the label
Claude Sonnet 5$2.00 / $10.00Was calendar: September 1, 2026 → $3/$15Cancelled August 11 — $2/$10 is now the permanent standard
GPT-5.6 Luna$0.20 / $1.20None — but $0.10/$0.60 is the Batch surface, not standardMislabelling risk, not a repricing risk

The pattern worth internalising: across four vendors, the headline rate and the effective rate have formally decoupled. Google decouples them with a date, xAI and OpenAI with a token count, and aggregators do it accidentally with surface labels. The common failure mode is the same in every case — a forecast keyed to the number on the marketing page rather than to the trigger conditions in the pricing documentation. Trackers list prices; budgets need triggers.

06The Caching NuanceCaching compresses the jump — it doesn’t cancel it.

Agent workloads are unusually cache-friendly — system prompts, tool schemas, and shared context repeat across every loop iteration — so Gemini’s caching schedule deserves its own budget line. Two separate prices exist here, in different units, and conflating them corrupts a forecast. The cache read price — what you pay when a cached token is served — is $0.075 per million tokens through December 31, doubling to $0.15 from January 1. The cache storage price — what you pay to keep tokens cached — is $0.50 per million tokens per hour through December 31, doubling to $1.00 per million per hour from January. One is per-token; the other is per-token-per-hour. They are different line items on the same page.

The genuinely useful nuance: the cache-read discount ratio is constant across the doubling. At $0.075 against a $0.75 input rate, cached reads cost 10% of base input; at $0.15 against $1.50, still exactly 10%. So while every price on the schedule doubles in absolute terms, a workload that serves a large share of its input from cache sees a smaller relative jump in blended cost than a cache-cold workload — the cached share was already riding at a tenth of the input rate and keeps that ratio. Caching does not cancel the doubling; it compresses it for the cached fraction of your tokens.

"The Gemini 3.7 Flash agent was 35% cheaper than 3.6 Flash, with a +8% observed prompt-cache hit rate and fewer tool errors."— Gregor Zunic, Co-founder & CTO, Browser Use — a vendor-curated testimonial on Google's Gemini Flash page

Treat that quote for what it is: a customer testimonial Google selected for its own product page, not an independent benchmark. But its shape is instructive — the cost win it describes comes partly from cache behaviour, and cache-hit rate is exactly the variable that will decide how hard the January step-up lands on any given fleet. Measuring your own hit rate now, while the intro window is open, is what turns the 2027 line of your forecast from a guess into arithmetic.

07The PlaybookFour moves before January.

The window runs about four and a half more months. Here is how we would spend them, by workload shape.

New agent fleets
Size against $1.50 / $7.50

Approve budgets at the posted standard rate and log the intro delta as explicit margin. If January executes, nothing breaks; if Google blinks the way Anthropic did, you bank the difference.

Budget at standard
Bulk & offline work
Route to the batch tier

Batch is a constant 50% of Standard — $0.375/$1.875 today, $0.75/$3.75 after the posted step-up. On the posted schedule, post-doubling batch equals today's standard rate, which makes it the softest landing of the Flash tiers.

Batch what can wait
Long-context agents
Guard the thresholds

If any route touches Grok 4.6 or GPT-5.6 Terra, enforce prompt caps below 200K and 272K respectively — one oversized request bills the entire prompt at the higher rate, not just the overflow.

Cap before the cliff
Every fleet
Calendar the re-check

Put a December pricing-page review in the calendar. Vendors both execute step-ups and cancel them; the only reliable source is the pricing page near the date, not coverage from launch week.

Verify in December

Model choice interacts with all of this. If the workload is genuinely subagent-shaped — high volume, narrow scope — the stable-priced tier below Flash may fit better than the discounted tier above it: the purpose-built subagent tier carries no expiry language at all, and bulk-workload cost math often favours models whose list price is boring. Boring is a feature in a forecast.

Looking forward, we expect more of this, not less. Dated intro windows let vendors buy market share without permanently repricing the line, and threshold repricing lets them defend margin on the long-context traffic that costs them the most to serve. Both mechanics reward the same operational habit: pricing review as a scheduled discipline rather than a launch-week event. If your team wants help building that discipline — fleet cost models, router policies, trigger-aware budget alerts — our AI transformation engagements start with exactly this kind of spend architecture.

08ConclusionPlan for the second price.

The shape of AI pricing, August 2026

The headline rate is an offer. The trigger is the contract.

Gemini’s Flash line is priced at $0.75/$3.75 today with $1.50/$7.50 posted for January 1, 2027 — and that schedule spans both 3.6 and 3.7 Flash, making it a line-wide repricing date rather than a single model’s promo. Budget at the standard rate, treat the intro window as margin, and the date loses its power over your forecast entirely.

Hold both precedents at once. Anthropic’s August 11 cancellation of Sonnet 5’s step-up proves posted increases sometimes never arrive; Grok 4.6 and GPT-5.6 Terra prove a request can be repriced today, mid-flight, with no calendar at all. The lesson underneath both is the same: the number on the launch page is the start of the pricing story, not the end of it.

The teams that get this right will not be the ones that guessed correctly about January. They will be the ones whose forecasts never depended on the guess — standard-rate budgets, batch routing for deferrable work, prompt caps under every threshold, and a December calendar entry to re-read the pricing page. None of that requires predicting what Google does next. That is the point.

Plan AI spend before the step-up

Budget at the standard rate and the intro window becomes pure margin.

We help teams build cost models, router policies, and budget guardrails for agent fleets — pricing-trigger audits, caching strategy, and vendor-mix planning, delivered in days not quarters.

Free consultationExpert guidanceTailored solutions
What we work on

Agent-fleet cost engagements

  • Fleet cost models keyed to standard rates, not promos
  • Pricing-trigger audits — calendar cliffs and token thresholds
  • Prompt-cap and caching strategy for long-context agents
  • Batch-tier routing for deferrable workloads
  • Vendor-mix planning across Google, Anthropic, OpenAI, xAI
FAQ · Intro-pricing budgeting

The questions we get every week.

Google's Gemini API pricing page lists 3.7 Flash at $0.75 per million input tokens and $3.75 per million output through December 31, 2026, with a standard rate of $1.50/$7.50 posted from January 1, 2027 — double on both legs. The tiers move in lockstep: Batch (and the numerically identical Flex tier) is a flat 50% of Standard at every point, so it goes from $0.375/$1.875 to $0.75/$3.75, and Priority runs a constant 1.8× Standard, going from $1.35/$6.75 to $2.70/$13.50. Context-cache read pricing doubles from $0.075 to $0.15 per million, and cache storage doubles from $0.50 to $1.00 per million tokens per hour. Because every surface is a fixed multiple of the standard rate, no tier escapes the step-up.
Related dispatches

Continue exploring agent economics.