LLM batch API pricing is the most under-budgeted discount in the model market: Google, OpenAI and Anthropic all charge exactly half their synchronous rate for asynchronous work, yet most cost forecasts we review price every token at the sync list rate. That’s a 2× error on every workload that never needed a real-time answer in the first place.
The stakes compounded this quarter. The August wave of cheap-tier repricing pushed sync rates down across the board, and batch lanes halved them again — while one aggregator’s catalog quietly listed a batch rate as if it were a standard rate, a mislabeling we confirmed by comparing its own two Luna listings. Meanwhile one frontier vendor skipped the discount entirely on its flagship.
This tracker puts all four frontier labs’ batch programs on one grid with same-day numbers, each traced to the page that publishes it: who discounts what, by how much, what the discount actually costs you in mechanics, and which of the widely quoted prices is quietly wrong. Every price is labeled standard or batch, because mixing the two surfaces is exactly how the market got confused.
- 01Three of the four frontier labs run a 50% batch lane.Google's Batch API, OpenAI's Batch API and Anthropic's Message Batches all charge half the synchronous rate — each vendor states the 50% figure on its own pricing or docs page.
- 02Grok 4.6 has no batch discount at all.xAI's batch discount is 20%, not 50% — and it covers exactly four older models (grok-4.3 and the grok-4.20 family). The flagship Grok 4.6 is excluded: “Models not listed above have no batch discount.”
- 03The GPT-5.6 Luna price on OpenRouter's non-batch listing is the batch rate.OpenAI's own model page lists Luna at $0.20/$1.20 per million standard. The $0.10/$0.60 figure is the Batch API rate — OpenRouter's non-batch listing surfaces the batch number.
- 04On Gemini, batch stacks with introductory pricing.Gemini 3.6 and 3.7 Flash batch is $0.375/$1.875 per million through December 31, 2026 — half the intro rate, and 75% below the $1.50/$7.50 standard sticker that takes effect January 1, 2027.
- 05The discount buys you a queue, not a slower model.Same models, same quality — you give up streaming, accept a completion window (24 hours at OpenAI; under an hour typical at Anthropic), and work within per-batch size caps. The engineering patterns live in our bulk-job guide.
01 — The PatternThe 50% standard — and the one vendor that skipped it.
Three vendors converged on the same number independently. Google’s Gemini API pricing page describes its batch lane as “Batch API (50% cost reduction)” across Gemini models. OpenAI’s Batch API guide calls it a “50% cost discount compared to synchronous APIs.” Anthropic’s Message Batches documentation charges all batch usage at 50% of standard API prices. Same models, same outputs — half the bill, in exchange for giving up the real-time response.
Then there’s xAI. Its pricing documentation does operate a Batch API — but the discount is 20%, not 50%, and it applies to exactly four named models, all older or smaller than the flagship. Grok 4.6 pays the full synchronous rate no matter how patient your workload is. Laid out as a share of the sync price, the asymmetry is hard to miss:
Batch price as a share of the synchronous rate
Source: Google, OpenAI, Anthropic and xAI pricing pages, August 2026Why 50% and not some other number? The economics are legible enough: asynchronous jobs let vendors schedule inference into capacity troughs instead of provisioning for your peak, and a uniform, memorable discount is easier to sell than per-model haggling. Once one lab anchored the lane at half price, matching it became table stakes — which is what makes xAI’s 20%, legacy-models-only posture read less like a pricing decision and more like a capacity statement. That interpretation is ours, not the vendors’; what the pages themselves establish is only the pattern: 50, 50, 50, and an asterisk.
02 — Myth CorrectionThe Luna price you’ve seen is the batch rate.
GPT-5.6 Luna is the cheap workhorse of OpenAI’s current lineup, and the price attached to it on OpenRouter’s non-batch listing is $0.10 per million input tokens and $0.60 per million output. OpenAI’s own model page for GPT-5.6 Luna says otherwise: the standard synchronous rate is $0.20 input / $1.20 output per million. The $0.10/$0.60 figure is real — it’s the Batch API rate, exactly 50% of standard, listed on OpenRouter’s dedicated gpt-5.6-luna:batch SKU.
How did a batch rate become the “known” standard price? We compared the two OpenRouter listings side by side: the catalog page for the non-batch openai/gpt-5.6-luna listing surfaces $0.10/$0.60 — identical to its separate :batch SKU — while OpenAI’s own model card holds at $0.20/$1.20. Anyone copy-pasting a price from the aggregator instead of the vendor’s page inherits the batch number without the batch queue. To be clear, this is an aggregator-side listing quirk, not OpenAI publishing conflicting numbers — OpenAI’s surfaces are consistent.
The aggregator drift isn’t confined to Luna. OpenRouter also carries a dedicated gemini-3.7-flash:batch SKU at $0.1875 input / $0.9375 output per million — 50% off OpenRouter’s own standard listing for the model, and a different number again from Google’s console batch rate. The lesson generalizes: aggregators are useful for discovering that a batch SKU exists, and unreliable for the number attached to it. Price from the vendor page; route through the aggregator if you must.
03 — The ScorecardEvery batch rate, one grid.
No single published table puts all four vendors’ batch programs side by side — most coverage repeats one vendor’s announcement in isolation, which is precisely how the 50/50/50/0 pattern stays invisible. The grid below assembles same-day numbers, each from the vendor’s own pricing surface except the Luna batch row, which comes from OpenRouter’s dedicated :batch SKU — OpenAI publishes the 50% rule rather than a per-model batch rate. Standard rates marked with an asterisk are derived by doubling the vendor’s published batch rate under its stated 50% rule; everything else is printed as listed.
| Model | Standard $/M (in / out) | Batch $/M (in / out) | Discount | Completion window |
|---|---|---|---|---|
| OpenAI — Batch API | ||||
| GPT-5.6 Luna | $0.20 / $1.20 | $0.10 / $0.60 | 50% | Fixed 24h, often quicker |
| Anthropic — Message Batches | ||||
| Claude Sonnet 5 | $2.00 / $10.00* | $1.00 / $5.00 | 50% | ≤24h; most under 1h |
| Claude Opus 5 · Opus 4.8 | $5.00 / $25.00* | $2.50 / $12.50 | 50% | ≤24h; most under 1h |
| Claude Fable 5 | $10.00 / $50.00* | $5.00 / $25.00 | 50% | ≤24h; most under 1h |
| Google — Gemini Batch API | ||||
| Gemini 3.6 / 3.7 Flash (intro, through Dec 31, 2026) | $0.75 / $3.75 | $0.375 / $1.875 | 50% | Not published† |
| Gemini 2.5 Pro (≤200k-token prompts) | $1.25 / $10.00* | $0.625 / $5.00 | 50% | Not published† |
| Gemini 2.5 Flash | $0.30 / $2.50* | $0.15 / $1.25 | 50% | Not published† |
| xAI — Batch API | ||||
| grok-4.3 · grok-4.20 family (3 variants) | Varies by model | Standard less 20% | 20% | Not published† |
| Grok 4.6 | $2.00 in (<200k)‡ | No batch discount | 0% | — |
Reading the footnotes: * standard rate derived by doubling the vendor’s published batch rate under its stated 50% relationship. † the vendor’s pricing page states the discount without publishing a completion-window guarantee — don’t assume one. ‡ Grok 4.6’s listed input rate is $2.00 per million under 200k tokens ($0.50 cached), rising to $4.00 ($1.00 cached) above 200k; its standard output rate is not printed here because we did not verify it against xAI’s page. Gemini 2.5 Pro’s batch rate is tiered by context: $0.625/$5.00 per million up to 200k-token prompts, $1.25/$7.50 above. Note also that Anthropic’s $1/$5 Sonnet 5 batch row carries no expiry — Anthropic cancelled the previously planned September step-up on August 11, the reversal we documented in the intro-pricing budgeting guide.
04 — CompoundingBatch stacks on Gemini’s intro pricing.
The single most under-reported number in this landscape is what happens when Google’s two discounts compound. Gemini 3.6 and 3.7 Flash carry introductory standard pricing of $0.75 input / $3.75 output per million through December 31, 2026, doubling to $1.50 / $7.50 on January 1, 2027 — a Flash-line-wide step-up we covered in the Gemini 3.7 Flash launch analysis. The Batch API halves whichever rate is in force. So today, a batch token on 3.7 Flash costs $0.375 in / $1.875 out — half the intro rate, and exactly 25% of the standard sticker price that takes effect in January. The batch lane is, for the next four and a half months, a 75%-off lane against the 2027 baseline.
Gemini 3.6 / 3.7 Flash per 1M
Half of the $0.75 introductory standard rate. Output follows the same math: $1.875 batch against $3.75 standard. Both models bill identically.
Effective discount vs post-intro list
The Jan 1, 2027 standard rate is $1.50 / $7.50. Today's batch rate of $0.375 / $1.875 is 25% of that sticker — intro pricing and the batch discount compound.
Batch rate doubles with the list
From January 1, 2027 the batch rate steps to $0.75 / $3.75 — the same 50% relationship against the new $1.50 / $7.50 standard. Budget the doubling now, not in December.
The forecast implication cuts both ways. If you’re budgeting agent workloads on Flash — the scenario we worked through in our guide to budgeting around the introductory pricing — the batch lane is the hedge against the January doubling: a workload moved from sync-intro to batch-intro today absorbs the 2027 step-up and still lands below the old sync bill. But the compounding also means quotes built on today’s batch rate understate 2027 spend by 2× on their own terms. Any Flash number in a plan that survives past New Year needs two columns.
05 — The Fine PrintWhat the discount actually costs you.
Half price is not free money — it’s payment for flexibility you surrender. The two vendors that publish their batch mechanics in detail, OpenAI and Anthropic, structure the trade differently, and the differences matter more than the identical headline discount suggests. (Google and xAI state their discounts without publishing comparable completion-window mechanics on their pricing pages, so we make no turnaround claims for either.)
Batch API
Completion window is set to 24h only, with batches often finishing faster. Up to 50,000 requests or a 200 MB input file per batch, 2,000 batch creations per hour, and a separate rate-limit pool — batch usage does not consume your standard per-model limits. Eight endpoints qualify, including images and video; batch-generated videos stay downloadable for only 24 hours after completion.
Message Batches
Results arrive when all messages complete or at 24 hours, whichever comes first; batches that miss the ceiling expire. Up to 100,000 requests or 256 MB per batch, results downloadable for 29 days. Streaming is explicitly rejected inside batch requests — results come back as a file. Prompt caching stacks with the batch discount on a best-effort basis, with reported cache hit rates of 30–98% depending on traffic pattern.
That last Anthropic detail deserves its own line item: batch and prompt caching are stacking discounts, not alternatives. A high-cache-hit batch workload on a Claude model pays the 50% batch rate on already-discounted cached reads — the compounding we priced out for the top-end model in our Fable 5 cache-and-batch cost engineering guide. The caveat is the phrase “best-effort”: concurrent batch processing makes cache hits probabilistic, which is why the vendor’s reported hit-rate range is as wide as 30–98%.
What none of this covers is the engineering discipline batch work demands — idempotency ledgers, schema validation, sample-based QA, retry design for jobs that expire at the 24-hour ceiling. We keep those patterns in a dedicated guide to running million-row LLM jobs without losing money; this post stays on the pricing surface.
06 — The OutlierGrok 4.6: the flagship with no cheap lane.
xAI’s pricing page introduces its batch lane the way every vendor does — “The Batch API lets you process large volumes of requests asynchronously at a discount to standard pricing” — and then quietly narrows it. The discount is 20%. It applies to exactly four models: grok-4.3, grok-4.20-0309-reasoning, grok-4.20-0309-non-reasoning, and grok-4.20-multi-agent-0309. Grok 4.6, the flagship, is not on the list, and the page closes the question in one sentence:
“Models not listed above have no batch discount.”— xAI pricing documentation, docs.x.ai
No press coverage we found frames this as what it is: the only frontier flagship excluded from its vendor’s batch discount. A bulk workload on Grok 4.6 pays the full synchronous rate — $2.00 per million input under 200k tokens, $4.00 above — whether it needs an answer in two seconds or two days. Against Sonnet 5’s $1.00 batch input or Luna’s $0.10, that’s not a price difference, it’s a category absence. Even the models xAI does discount get 20% where the rest of the market gives 50% of the bill back.
Projecting forward: this gap is unstable. Batch lanes exist because idle accelerator capacity is worth selling at any margin above marginal cost, and a vendor that cannot offer one on its flagship is signaling either constrained capacity or a deliberate bet that its demand is all latency-sensitive. Neither position tends to hold for long once bulk buyers start routing around it — classification, enrichment and evaluation workloads are the easiest traffic in the market to move. If a Grok 4.6 batch tier appears, expect it to match the 50% norm, because the 20% precedent has already proven uncompetitive at the margin. Until then, bulk work priced on Grok belongs in another vendor’s queue.
07 — Decision GuideWhich workloads belong in the batch lane.
The decision rule is simpler than most teams make it: if no human is waiting on the response, the workload is a batch candidate, and paying the synchronous rate for it is a standing 2× overspend. The matrix below sorts the common cases.
Classification, tagging, embeddings, catalog work
The canonical batch workload — high volume, zero latency requirement. On OpenAI the embeddings endpoint batches directly; on Anthropic, caching stacks on repeated system prompts. Never pay sync rates here.
Model evals and regression suites
Nightly eval sweeps and prompt-regression suites tolerate a 24-hour window by definition. Half-price evals mean twice the coverage on the same budget — or the same coverage on a frontier model instead of a cheap one.
Interactive sessions, tool loops, streaming UX
A user is waiting and streaming is off the table in batch (Anthropic rejects stream: true outright). This traffic stays synchronous — the discount was never for it. Budget it at full rate and resist averaging the two lanes into one blended number.
Reports, digests, overnight generation
Anything cron-shaped fits the window; the only real work is idempotency and retry design for the expiry ceiling. If output must land by a fixed morning deadline, submit early enough that a full 24-hour window still makes the cutoff.
The organizational failure mode isn’t choosing the wrong lane — it’s never modeling the lanes separately at all. Most AI budgets we audit carry one blended per-token assumption, which means they overpay on bulk work and under-provision for interactive traffic simultaneously. Splitting the forecast into a sync line and a batch line is an afternoon of work that often surfaces material savings; it’s a standard step in our AI transformation engagements, and the batch lane is also what makes high-volume programs like our content engine economics work at scale.
08 — ConclusionBudget the lane, not the list price.
Half price is the market standard — and it's sitting unused in most budgets.
The pattern is now clean enough to state as a rule: Google, OpenAI and Anthropic all sell asynchronous inference at half the synchronous rate, on the same models, with the discount printed on their own pricing pages. The exceptions are exactly two — xAI discounts only four legacy models at 20%, and its flagship Grok 4.6 gets no batch discount at all.
The traps are equally specific. The GPT-5.6 Luna price on OpenRouter’s non-batch listing is the batch rate wearing a standard label — $0.10/$0.60 batch versus $0.20/$1.20 on OpenAI’s own model page. And Gemini Flash’s remarkable $0.375/$1.875 batch rate carries a date: it doubles with the rest of the Flash line on January 1, 2027. Price by surface, trace every number to the vendor’s page, and put expiry dates in the forecast, not the footnotes.
The forward bet is that the batch lane becomes more important, not less. As sync prices compress, vendors defend margin by selling their capacity troughs — and buyers who structure work asynchronously capture that spread. The teams that treat batch as an architectural default with a synchronous exception, rather than the reverse, will run the same workloads as their competitors at roughly half the marginal cost. That’s not a projection about models. It’s arithmetic that’s already on the pricing pages.