DevelopmentPricing Tracker10 min readPublished August 14, 2026

Four frontier vendors, one pattern · 50 / 50 / 50 — and zero on Grok 4.6

LLM Batch APIs: The Half-Price Lane Nobody Budgets

Google, OpenAI and Anthropic each run an asynchronous batch lane at 50% of their synchronous rate — and on Gemini it stacks with the current introductory pricing. xAI is the outlier: a 20% discount confined to four older Grok models, and no batch discount at all on flagship Grok 4.6. Every number below is labeled by surface and traced to the page that publishes it.

DA
Digital Applied Team
Senior strategists · Published Aug 14, 2026
PublishedAugust 14, 2026
Read time10 min
Sources4 vendor pages + OpenRouter
Batch discount, big three
50%
Google · OpenAI · Anthropic
Grok 4.6 batch discount
0%
flagship not on the discount list
20% legacy only
Luna batch input
$0.10
per 1M · standard is $0.20
Gemini Flash batch input
$0.375
per 1M through Dec 31, 2026
−75% vs 2027 list

LLM batch API pricing is the most under-budgeted discount in the model market: Google, OpenAI and Anthropic all charge exactly half their synchronous rate for asynchronous work, yet most cost forecasts we review price every token at the sync list rate. That’s a 2× error on every workload that never needed a real-time answer in the first place.

The stakes compounded this quarter. The August wave of cheap-tier repricing pushed sync rates down across the board, and batch lanes halved them again — while one aggregator’s catalog quietly listed a batch rate as if it were a standard rate, a mislabeling we confirmed by comparing its own two Luna listings. Meanwhile one frontier vendor skipped the discount entirely on its flagship.

This tracker puts all four frontier labs’ batch programs on one grid with same-day numbers, each traced to the page that publishes it: who discounts what, by how much, what the discount actually costs you in mechanics, and which of the widely quoted prices is quietly wrong. Every price is labeled standard or batch, because mixing the two surfaces is exactly how the market got confused.

Key takeaways
  1. 01
    Three of the four frontier labs run a 50% batch lane.Google's Batch API, OpenAI's Batch API and Anthropic's Message Batches all charge half the synchronous rate — each vendor states the 50% figure on its own pricing or docs page.
  2. 02
    Grok 4.6 has no batch discount at all.xAI's batch discount is 20%, not 50% — and it covers exactly four older models (grok-4.3 and the grok-4.20 family). The flagship Grok 4.6 is excluded: “Models not listed above have no batch discount.”
  3. 03
    The GPT-5.6 Luna price on OpenRouter's non-batch listing is the batch rate.OpenAI's own model page lists Luna at $0.20/$1.20 per million standard. The $0.10/$0.60 figure is the Batch API rate — OpenRouter's non-batch listing surfaces the batch number.
  4. 04
    On Gemini, batch stacks with introductory pricing.Gemini 3.6 and 3.7 Flash batch is $0.375/$1.875 per million through December 31, 2026 — half the intro rate, and 75% below the $1.50/$7.50 standard sticker that takes effect January 1, 2027.
  5. 05
    The discount buys you a queue, not a slower model.Same models, same quality — you give up streaming, accept a completion window (24 hours at OpenAI; under an hour typical at Anthropic), and work within per-batch size caps. The engineering patterns live in our bulk-job guide.

01The PatternThe 50% standard — and the one vendor that skipped it.

Three vendors converged on the same number independently. Google’s Gemini API pricing page describes its batch lane as “Batch API (50% cost reduction)” across Gemini models. OpenAI’s Batch API guide calls it a “50% cost discount compared to synchronous APIs.” Anthropic’s Message Batches documentation charges all batch usage at 50% of standard API prices. Same models, same outputs — half the bill, in exchange for giving up the real-time response.

Then there’s xAI. Its pricing documentation does operate a Batch API — but the discount is 20%, not 50%, and it applies to exactly four named models, all older or smaller than the flagship. Grok 4.6 pays the full synchronous rate no matter how patient your workload is. Laid out as a share of the sync price, the asymmetry is hard to miss:

Batch price as a share of the synchronous rate

Source: Google, OpenAI, Anthropic and xAI pricing pages, August 2026
Synchronous list priceAny vendor · the rate most budgets assume
100%
Google Gemini Batch API50% cost reduction across Gemini models
50%
OpenAI Batch API50% off · chat, embeddings, images, video endpoints
50%
Anthropic Message BatchesAll usage at 50% of standard API prices
50%
xAI batch · grok-4.3 / 4.20 family20% discount, four older models only
80%
xAI Grok 4.6No batch discount on the flagship
100%

Why 50% and not some other number? The economics are legible enough: asynchronous jobs let vendors schedule inference into capacity troughs instead of provisioning for your peak, and a uniform, memorable discount is easier to sell than per-model haggling. Once one lab anchored the lane at half price, matching it became table stakes — which is what makes xAI’s 20%, legacy-models-only posture read less like a pricing decision and more like a capacity statement. That interpretation is ours, not the vendors’; what the pages themselves establish is only the pattern: 50, 50, 50, and an asterisk.

02Myth CorrectionThe Luna price you’ve seen is the batch rate.

GPT-5.6 Luna is the cheap workhorse of OpenAI’s current lineup, and the price attached to it on OpenRouter’s non-batch listing is $0.10 per million input tokens and $0.60 per million output. OpenAI’s own model page for GPT-5.6 Luna says otherwise: the standard synchronous rate is $0.20 input / $1.20 output per million. The $0.10/$0.60 figure is real — it’s the Batch API rate, exactly 50% of standard, listed on OpenRouter’s dedicated gpt-5.6-luna:batch SKU.

How did a batch rate become the “known” standard price? We compared the two OpenRouter listings side by side: the catalog page for the non-batch openai/gpt-5.6-luna listing surfaces $0.10/$0.60 — identical to its separate :batch SKU — while OpenAI’s own model card holds at $0.20/$1.20. Anyone copy-pasting a price from the aggregator instead of the vendor’s page inherits the batch number without the batch queue. To be clear, this is an aggregator-side listing quirk, not OpenAI publishing conflicting numbers — OpenAI’s surfaces are consistent.

Label every price by surface
A model no longer has a price — it has a standard price, a batch price, and sometimes an introductory price with an expiry date. GPT-5.6 Luna is $0.20 / $1.20 standard and $0.10 / $0.60 batch, per million input/output tokens. Quote either one without its label and someone downstream will build a forecast on the wrong lane. The same discipline applies to the August repricing wave we tracked in the cheap-tier repricing analysis.

The aggregator drift isn’t confined to Luna. OpenRouter also carries a dedicated gemini-3.7-flash:batch SKU at $0.1875 input / $0.9375 output per million — 50% off OpenRouter’s own standard listing for the model, and a different number again from Google’s console batch rate. The lesson generalizes: aggregators are useful for discovering that a batch SKU exists, and unreliable for the number attached to it. Price from the vendor page; route through the aggregator if you must.

03The ScorecardEvery batch rate, one grid.

No single published table puts all four vendors’ batch programs side by side — most coverage repeats one vendor’s announcement in isolation, which is precisely how the 50/50/50/0 pattern stays invisible. The grid below assembles same-day numbers, each from the vendor’s own pricing surface except the Luna batch row, which comes from OpenRouter’s dedicated :batch SKU — OpenAI publishes the 50% rule rather than a per-model batch rate. Standard rates marked with an asterisk are derived by doubling the vendor’s published batch rate under its stated 50% rule; everything else is printed as listed.

Batch discount scorecard comparing standard and batch per-million token rates, discount percentage, and completion window across OpenAI, Anthropic, Google and xAI models, assembled from each vendor’s own pricing page in August 2026, except the Luna batch row, which comes from OpenRouter’s :batch SKU.
ModelStandard $/M (in / out)Batch $/M (in / out)DiscountCompletion window
OpenAI — Batch API
GPT-5.6 Luna$0.20 / $1.20$0.10 / $0.6050%Fixed 24h, often quicker
Anthropic — Message Batches
Claude Sonnet 5$2.00 / $10.00*$1.00 / $5.0050%≤24h; most under 1h
Claude Opus 5 · Opus 4.8$5.00 / $25.00*$2.50 / $12.5050%≤24h; most under 1h
Claude Fable 5$10.00 / $50.00*$5.00 / $25.0050%≤24h; most under 1h
Google — Gemini Batch API
Gemini 3.6 / 3.7 Flash (intro, through Dec 31, 2026)$0.75 / $3.75$0.375 / $1.87550%Not published†
Gemini 2.5 Pro (≤200k-token prompts)$1.25 / $10.00*$0.625 / $5.0050%Not published†
Gemini 2.5 Flash$0.30 / $2.50*$0.15 / $1.2550%Not published†
xAI — Batch API
grok-4.3 · grok-4.20 family (3 variants)Varies by modelStandard less 20%20%Not published†
Grok 4.6$2.00 in (<200k)‡No batch discount0%

Reading the footnotes: * standard rate derived by doubling the vendor’s published batch rate under its stated 50% relationship. † the vendor’s pricing page states the discount without publishing a completion-window guarantee — don’t assume one. ‡ Grok 4.6’s listed input rate is $2.00 per million under 200k tokens ($0.50 cached), rising to $4.00 ($1.00 cached) above 200k; its standard output rate is not printed here because we did not verify it against xAI’s page. Gemini 2.5 Pro’s batch rate is tiered by context: $0.625/$5.00 per million up to 200k-token prompts, $1.25/$7.50 above. Note also that Anthropic’s $1/$5 Sonnet 5 batch row carries no expiry — Anthropic cancelled the previously planned September step-up on August 11, the reversal we documented in the intro-pricing budgeting guide.

04CompoundingBatch stacks on Gemini’s intro pricing.

The single most under-reported number in this landscape is what happens when Google’s two discounts compound. Gemini 3.6 and 3.7 Flash carry introductory standard pricing of $0.75 input / $3.75 output per million through December 31, 2026, doubling to $1.50 / $7.50 on January 1, 2027 — a Flash-line-wide step-up we covered in the Gemini 3.7 Flash launch analysis. The Batch API halves whichever rate is in force. So today, a batch token on 3.7 Flash costs $0.375 in / $1.875 out — half the intro rate, and exactly 25% of the standard sticker price that takes effect in January. The batch lane is, for the next four and a half months, a 75%-off lane against the 2027 baseline.

Batch input, today
Gemini 3.6 / 3.7 Flash per 1M
$0.375

Half of the $0.75 introductory standard rate. Output follows the same math: $1.875 batch against $3.75 standard. Both models bill identically.

through Dec 31, 2026
vs 2027 sticker
Effective discount vs post-intro list
75%

The Jan 1, 2027 standard rate is $1.50 / $7.50. Today's batch rate of $0.375 / $1.875 is 25% of that sticker — intro pricing and the batch discount compound.

$0.375 ÷ $1.50 = 25%
After the cliff
Batch rate doubles with the list
2×

From January 1, 2027 the batch rate steps to $0.75 / $3.75 — the same 50% relationship against the new $1.50 / $7.50 standard. Budget the doubling now, not in December.

Jan 1, 2027

The forecast implication cuts both ways. If you’re budgeting agent workloads on Flash — the scenario we worked through in our guide to budgeting around the introductory pricing — the batch lane is the hedge against the January doubling: a workload moved from sync-intro to batch-intro today absorbs the 2027 step-up and still lands below the old sync bill. But the compounding also means quotes built on today’s batch rate understate 2027 spend by 2× on their own terms. Any Flash number in a plan that survives past New Year needs two columns.

05The Fine PrintWhat the discount actually costs you.

Half price is not free money — it’s payment for flexibility you surrender. The two vendors that publish their batch mechanics in detail, OpenAI and Anthropic, structure the trade differently, and the differences matter more than the identical headline discount suggests. (Google and xAI state their discounts without publishing comparable completion-window mechanics on their pricing pages, so we make no turnaround claims for either.)

OpenAI
Batch API
50% off · fixed 24h window

Completion window is set to 24h only, with batches often finishing faster. Up to 50,000 requests or a 200 MB input file per batch, 2,000 batch creations per hour, and a separate rate-limit pool — batch usage does not consume your standard per-model limits. Eight endpoints qualify, including images and video; batch-generated videos stay downloadable for only 24 hours after completion.

developers.openai.com · Batch guide
Anthropic
Message Batches
50% off · most batches under 1h

Results arrive when all messages complete or at 24 hours, whichever comes first; batches that miss the ceiling expire. Up to 100,000 requests or 256 MB per batch, results downloadable for 29 days. Streaming is explicitly rejected inside batch requests — results come back as a file. Prompt caching stacks with the batch discount on a best-effort basis, with reported cache hit rates of 30–98% depending on traffic pattern.

platform.claude.com · batch processing

That last Anthropic detail deserves its own line item: batch and prompt caching are stacking discounts, not alternatives. A high-cache-hit batch workload on a Claude model pays the 50% batch rate on already-discounted cached reads — the compounding we priced out for the top-end model in our Fable 5 cache-and-batch cost engineering guide. The caveat is the phrase “best-effort”: concurrent batch processing makes cache hits probabilistic, which is why the vendor’s reported hit-rate range is as wide as 30–98%.

What none of this covers is the engineering discipline batch work demands — idempotency ledgers, schema validation, sample-based QA, retry design for jobs that expire at the 24-hour ceiling. We keep those patterns in a dedicated guide to running million-row LLM jobs without losing money; this post stays on the pricing surface.

06The OutlierGrok 4.6: the flagship with no cheap lane.

xAI’s pricing page introduces its batch lane the way every vendor does — “The Batch API lets you process large volumes of requests asynchronously at a discount to standard pricing” — and then quietly narrows it. The discount is 20%. It applies to exactly four models: grok-4.3, grok-4.20-0309-reasoning, grok-4.20-0309-non-reasoning, and grok-4.20-multi-agent-0309. Grok 4.6, the flagship, is not on the list, and the page closes the question in one sentence:

“Models not listed above have no batch discount.”— xAI pricing documentation, docs.x.ai

No press coverage we found frames this as what it is: the only frontier flagship excluded from its vendor’s batch discount. A bulk workload on Grok 4.6 pays the full synchronous rate — $2.00 per million input under 200k tokens, $4.00 above — whether it needs an answer in two seconds or two days. Against Sonnet 5’s $1.00 batch input or Luna’s $0.10, that’s not a price difference, it’s a category absence. Even the models xAI does discount get 20% where the rest of the market gives 50% of the bill back.

Projecting forward: this gap is unstable. Batch lanes exist because idle accelerator capacity is worth selling at any margin above marginal cost, and a vendor that cannot offer one on its flagship is signaling either constrained capacity or a deliberate bet that its demand is all latency-sensitive. Neither position tends to hold for long once bulk buyers start routing around it — classification, enrichment and evaluation workloads are the easiest traffic in the market to move. If a Grok 4.6 batch tier appears, expect it to match the 50% norm, because the 20% precedent has already proven uncompetitive at the margin. Until then, bulk work priced on Grok belongs in another vendor’s queue.

07Decision GuideWhich workloads belong in the batch lane.

The decision rule is simpler than most teams make it: if no human is waiting on the response, the workload is a batch candidate, and paying the synchronous rate for it is a standing 2× overspend. The matrix below sorts the common cases.

Bulk content & enrichment
Classification, tagging, embeddings, catalog work

The canonical batch workload — high volume, zero latency requirement. On OpenAI the embeddings endpoint batches directly; on Anthropic, caching stacks on repeated system prompts. Never pay sync rates here.

Always batch
Evaluation & QA runs
Model evals and regression suites

Nightly eval sweeps and prompt-regression suites tolerate a 24-hour window by definition. Half-price evals mean twice the coverage on the same budget — or the same coverage on a frontier model instead of a cheap one.

Batch by default
Agent & chat traffic
Interactive sessions, tool loops, streaming UX

A user is waiting and streaming is off the table in batch (Anthropic rejects stream: true outright). This traffic stays synchronous — the discount was never for it. Budget it at full rate and resist averaging the two lanes into one blended number.

Stay synchronous
Scheduled pipelines
Reports, digests, overnight generation

Anything cron-shaped fits the window; the only real work is idempotency and retry design for the expiry ceiling. If output must land by a fixed morning deadline, submit early enough that a full 24-hour window still makes the cutoff.

Batch with retry design

The organizational failure mode isn’t choosing the wrong lane — it’s never modeling the lanes separately at all. Most AI budgets we audit carry one blended per-token assumption, which means they overpay on bulk work and under-provision for interactive traffic simultaneously. Splitting the forecast into a sync line and a batch line is an afternoon of work that often surfaces material savings; it’s a standard step in our AI transformation engagements, and the batch lane is also what makes high-volume programs like our content engine economics work at scale.

08ConclusionBudget the lane, not the list price.

The batch landscape, August 2026

Half price is the market standard — and it's sitting unused in most budgets.

The pattern is now clean enough to state as a rule: Google, OpenAI and Anthropic all sell asynchronous inference at half the synchronous rate, on the same models, with the discount printed on their own pricing pages. The exceptions are exactly two — xAI discounts only four legacy models at 20%, and its flagship Grok 4.6 gets no batch discount at all.

The traps are equally specific. The GPT-5.6 Luna price on OpenRouter’s non-batch listing is the batch rate wearing a standard label — $0.10/$0.60 batch versus $0.20/$1.20 on OpenAI’s own model page. And Gemini Flash’s remarkable $0.375/$1.875 batch rate carries a date: it doubles with the rest of the Flash line on January 1, 2027. Price by surface, trace every number to the vendor’s page, and put expiry dates in the forecast, not the footnotes.

The forward bet is that the batch lane becomes more important, not less. As sync prices compress, vendors defend margin by selling their capacity troughs — and buyers who structure work asynchronously capture that spread. The teams that treat batch as an architectural default with a synchronous exception, rather than the reverse, will run the same workloads as their competitors at roughly half the marginal cost. That’s not a projection about models. It’s arithmetic that’s already on the pricing pages.

Cut your AI bill by routing the right lane

The same workload at half the marginal cost.

Our team builds and operates bulk LLM pipelines — batch-lane routing, cache stacking, idempotent job design and per-surface cost forecasting — delivered in days, not quarters.

Free consultationExpert guidanceTailored solutions
What we work on

AI cost engineering engagements

  • Batch-lane migration for bulk classification & enrichment
  • Two-column forecasts — sync vs batch, with expiry cliffs
  • Cache + batch stacking on Claude workloads
  • Vendor routing across OpenAI, Anthropic, Google & xAI
  • Idempotency & QA design for million-row jobs
FAQ · Batch API pricing

The questions we get every week.

For three of the four frontier vendors, exactly half. Google's Gemini pricing page describes its Batch API as a 50% cost reduction across Gemini models, OpenAI's Batch API guide states a 50% cost discount compared to synchronous APIs, and Anthropic's Message Batches documentation charges all batch usage at 50% of standard API prices. xAI is the outlier: its batch discount is 20% rather than 50%, and it applies only to four older models — grok-4.3 and three grok-4.20 variants. The flagship Grok 4.6 has no batch discount at all. The models and output quality are identical to the synchronous lane; what you give up is real-time delivery.
Related dispatches

Continue exploring AI economics.