How you call a model API moves what $200 buys by roughly an order of magnitude, with the model held constant. The conversion table below turns each vendor-published rate into input tokens at a fixed $200: Claude Opus 5 alone spans 40M tokens at the list rate to a derived 800M through the batch-plus-cache lane, a 20x range with the model held constant.
Two definitions carry the table. Cached input is not a discount you negotiate; it is a different price for a different operation — re-reading a prompt prefix the provider has already processed, billed at 0.1x the base input rate wherever the vendors below publish the multiplier. Batch is not a cheaper model; it is an asynchronous lane — the same weights answering within a service window instead of now, at 50% off both input and output where it exists.
- 01One model, one budget, a 20x spread.At $200, Claude Opus 5 spans 40M list tokens to a derived 800M with batch and caching combined — a 20x spread before you compare a single competing model.
- 02Cache is a 0.1x read multiplier; batch is 50% off.Anthropic, OpenAI and Alibaba each publish cache reads at one tenth of base input. Batch, where offered, halves both directions. Only some vendors print the combined price.
- 03DeepSeek prints an off-peak lane as a table cell.V4-Pro's cache-hit off-peak rate of $0.022 per million tokens is a printed table cell, not arithmetic. It turns $200 into roughly 9.1B input tokens — 60x its own peak cache-miss lane.
01 — The table$200, converted into tokens, lane by lane.
Every cell below is the same arithmetic: $200 divided by a vendor-published price per million input tokens. The rates themselves live in our maintained per-million-token price index; this page inverts them into the question a budget line actually asks — how many tokens does a fixed $200 buy, and which lever changes that most. Rows marked published cell are combined prices the vendor prints itself; the one row marked derived is our multiplication of two published multipliers, never a vendor figure.
| Model and lane | List input | Cached input | Batch input | Off-peak input |
|---|---|---|---|---|
| Anthropic — Claude API | ||||
| Claude Opus 5 | 40M | 400M | 80M | — |
| Claude Opus 5 · batch + cache derived | — | — | 800M* | — |
| Claude Sonnet 5 | 100M | 1,000M | 200M | — |
| OpenAI — short context, promotional rate | ||||
| GPT-5.6 Sol | 50M | 500M | 100M | — |
| GPT-5.6 Sol · batch cached input published cell | — | — | 1,000M | — |
| Google — Gemini API | ||||
| Gemini 3.7 Flash (promotional) | 267M | 2,667M | 533M | — |
| Gemini 3.1 Pro Preview (≤200K prompts) | 100M | 1,000M | 200M | — |
| DeepSeek — no batch lane; peak and off-peak published | ||||
| DeepSeek V4-Pro (peak) | 151.5M | 4,545M | none published | 303M |
| DeepSeek V4-Pro · cache-hit × off-peak published cell | — | 9,091M | — | — |
| Z.ai — cache lever only | ||||
| GLM-5.3 | 142.9M | 769.2M | none published | — |
| GLM-5.3-Flash (promotional) | 2,667M | 13,333M | none published | — |
| Moonshot — cache lever only | ||||
| Kimi K3 | 66.7M | 666.7M | none published | — |
| Alibaba — International endpoint | ||||
| Qwen3.8-Max† | 100M | 1,000M | 200M | — |
* Derived, not vendor-published: 0.5 (batch) × 0.1 (cache read) × $5.00 base input = $0.25 per million tokens. Anthropic’s FAQ confirms the two discounts combine, but the company never prints this cell the way OpenAI does one row up. † Alibaba documents cache and batch as two separate rules and nowhere confirms they stack, so no combined Qwen figure appears in this table — multiplying them would invent a number the vendor has not stated. The list lanes for GPT-5.6 Sol, Gemini 3.7 Flash and GLM-5.3-Flash are promotional, with vendor-announced end dates of November 21, 2026, December 31, 2026 and September 9, 2026 respectively.
02 — MethodWhat was computed, and what this page owns.
This page publishes one derivation and one borrowed axis, and it is deliberate about which numbers are whose. The rates belong to the vendors, the quality scores belong to Artificial Analysis, and the only thing that is ours is the inversion — tokens per fixed budget — plus the matrix of which levers each vendor publishes.
- What was computed
- For seven vendors’ current flagship models: $200 divided by each published input-token rate lane — list, cached input, batch, and time-of-day where published. Input side only; output rates are not tabulated here. Figures rounded to the nearest million tokens, or one decimal where finer.
- What this page owns
- The conversion and the quality-attached ranking, only. The per-million-token rates are owned by our maintained price index and are not restated as a rate card here. Batch pricing as a subject is owned by our batch API landscape; batch appears here strictly as one input to the conversion.
- Sources
- Each vendor’s own pricing documentation: Anthropic, OpenAI, Google, DeepSeek, Z.ai, Moonshot and Alibaba Cloud, plus the two vendors’ help-center articles on their $200 subscription tiers and Artificial Analysis’s leaderboard for the quality axis. No aggregator supplies any rate cell.
- Dates
- Collected 2026-09-01. Every pricing page cited is a living reference page, so figures are the current state as retrieved, not events on a date. Promotional end dates that extend past publication are vendor-announced schedules that were already public.
- Units
- Millions of input tokens (M) per $200. Output tokens are excluded from the table because cache reads apply to input only; batch, where offered, also halves output, which the interpretation notes but the cells do not encode.
- What is excluded
- The $200 subscription tiers, Claude Max and ChatGPT Pro, because neither vendor expresses them in a convertible unit (§06). Artificial Analysis’s own “Blended $/1M” cost column, an AA estimate that would corrupt the arithmetic if mixed with vendor list rates. Mainland-China endpoints, which this pass did not verify first-party.
- Known limitations
- The Anthropic batch-plus-cache cell is derived from two published multipliers, and is labelled so everywhere it appears. No stacked Qwen figure is computed. The Artificial Analysis page carries no visible as-of stamp, so its scores are reported as retrieved for this dataset, with no board date asserted.
03 — The leversWhich levers each vendor publishes — and which prices they print.
The conversion table only works because each vendor publishes a different set of levers, and the differences are structural, not cosmetic. Anthropic and Alibaba publish nearly identical percentages — cache reads at one tenth of base input, batch at half price — yet only Anthropic confirms in writing that the two combine. OpenAI goes further and prints the compounded cell in its own batch table. Google publishes both levers but states that batch context caching is billed at the same standard cache rate, so there is no extra stack to compute. DeepSeek has no batch lane at all and instead publishes the survey’s deepest lever: a full cache-by-time-of-day price matrix.
| Vendor · model | Cache read | Batch | Off-peak | Combined cell printed? |
|---|---|---|---|---|
| Anthropic · Opus 5, Sonnet 5 | Yes — 0.1x base input | Yes — 50% off both directions | not published | No. The FAQ confirms the discounts combine; no cell is printed. |
| OpenAI · GPT-5.6 Sol | Yes — 0.1x base input | Yes — 50% off both directions | not published | Yes. The batch table prints cached input at $0.20 per million tokens. |
| Google · Gemini 3.7 Flash, 3.1 Pro | Yes — per-model caching rates | Yes — 50% off | not published | Yes, but batch caching is billed at the standard cache rate — no extra stack. |
| DeepSeek · V4-Pro, V4-Flash | Yes — published cache-hit rows | none published | Yes — peak and off-peak rates printed | Yes. The full cache × time-of-day matrix is printed as table cells. |
| Z.ai · GLM-5.3 | Yes — $0.26 vs $1.40 (≈0.19x) | none published | not published | n/a — one lever only |
| Moonshot · Kimi K3 | Yes — $0.30 hit vs $3.00 miss (0.1x) | none published | not published | n/a — one lever only |
| Alibaba · Qwen3.8-Max | Yes — cache hits at 10% of input | Yes — 50% of real-time price | not published | No. Two separate rules; stacking is not confirmed anywhere on the page. |
Batch API and prompt caching discounts can be combined ... See prompt caching pricing for how the multipliers interact.Anthropic pricing documentation, FAQ
Prints the combined cell
OpenAI's batch table carries its own cached-input column, so the compounded price is a published number. DeepSeek prints a full cache-by-time-of-day matrix. In both cases the cheapest lane is a cell you can point at, not arithmetic you must defend.
Confirms stacking in prose
The FAQ states batch and caching combine, but no compounded cell is printed. The $0.25-per-million lane for Opus 5 is real per the vendor's own rules — and it is still your multiplication, which is why this page labels it derived every time it appears.
Levers that exclude each other
Anthropic's fast mode (research preview, $10/$50 for Opus 5) states plainly that it is not available with the Batch API. Levers are vendor-specific machinery, not a universal rulebook — some stack, some are mutually exclusive.
The mechanics behind the two big levers are covered in depth elsewhere: our prompt caching engineering guide explains what a cache read physically is, and the half-price batch lane maps who offers batch and on what terms. This page takes both as given and asks only what they do to a fixed budget.
04 — The findingThe same model spans 20x across its own lanes.
Hold the model constant and walk the lanes. Claude Opus 5 at $200 buys 40M input tokens at list, 80M through batch, 400M through cache reads, and a derived 800M with the two combined — the same weights answering the same prompts across a 20x range. Now compare across models at the list rate only: Opus 5’s 40M against Qwen3.8-Max’s 100M is a 2.5x gap, and against Gemini 3.7 Flash’s promotional 267M the gap is under 7x. Step further out and the between-model range is wider still: GLM-5.3-Flash’s promotional list rate returns 2,667M for the same $200. Both axes move the answer by more than most buyers assume, and only one of them appears on a pricing page.
DeepSeek makes the same point with printed cells rather than derivation. V4-Pro’s peak cache-miss rate is $1.32 per million tokens; its off-peak cache-hit rate is $0.022. Both are published table cells, and the ratio between them is 60x, entirely inside one model. At $200 that is the difference between 151.5M tokens and roughly 9.1B. The off-peak window is generous, too: DeepSeek’s definition leaves every hour outside two weekday UTC bands, and all of the weekend, at the lower rate.
Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday (all other hours are off-peak).DeepSeek API pricing documentation
One discipline keeps this table honest: a compounded price is only as real as its publication. Anthropic’s $0.25 batch-plus-cache lane for Opus 5 follows from two published multipliers and an FAQ that confirms they combine — so it appears here, marked derived. Where a vendor prints the cell, we cite the cell.
The practical reading order for a $200 budget follows directly. First ask which of your traffic is re-read prefix — system prompts, tool definitions, shared context — because the 0.1x cache-read multiplier is the single largest lever most vendors publish. Second, ask what fraction of your work can wait for an asynchronous window, because batch halves both directions where it exists. Only third ask which model, because by then the same model’s price has already moved by an order of magnitude under your first two answers.
05 — QualityThe quality axis, borrowed and attributed.
Tokens per dollar mean nothing without a quality axis, and the quality axis is not ours. The scores below are Artificial Analysis’s Intelligence Index, retrieved for this dataset from its public leaderboard, at the highest reasoning effort AA lists for each model. The ranking is AA’s judgment, not Digital Applied’s; we pair it with the cheapest lane each vendor publishes and draw one observation from the pairing.
| Model (AA’s label) | AA Intelligence Index | Cheapest published input lane | Tokens at $200 |
|---|---|---|---|
| Claude Opus 5 (max) | 63 | Cache read, $0.50 | 400M |
| GPT-5.6 Sol (max) | 61 | Batch cached input, $0.20 (printed) | 1,000M |
| Kimi K3 (max) | 60 | Cache-hit input, $0.30 | 666.7M |
| GLM-5.3 (max) | 60 | Cached input, $0.26 | 769.2M |
| Qwen3.8 Max | 58 | Cache-hit input, $0.20 | 1,000M |
| Gemini 3.7 Flash (high) | 56 | Context caching, $0.075 (promotional) | 2,667M |
| Claude Sonnet 5 (max) | 55 | Cache read, $0.20 | 1,000M |
| DeepSeek V4 Pro 0813 (max) | 53 | Cache-hit off-peak, $0.022 (printed) | 9,091M |
| Gemini 3.1 Pro Preview | 48 | Cached input, $0.20 (≤200K prompts) | 1,000M |
The observation the pairing earns: the quality gap between the top and bottom of AA’s table is fifteen index points, and the budget gap between lanes routinely dwarfs what that quality gap costs. Claude Opus 5’s derived batch-plus-cache lane ($0.25 per million) is cheaper than Gemini 3.1 Pro Preview’s list rate ($2.00 per million) — eight times cheaper, for a model AA scores fifteen points higher. On these numbers, learning to use a vendor’s published levers buys more capability per dollar than shopping down the quality table does.
06 — Negative resultThe two $200 plans that don’t convert.
There is a second way to spend $200 a month on frontier AI, and it deliberately resists this table. Claude Max at $200 is described in Anthropic’s help center only as “20 times more usage per session than the Pro plan” — no token count, no message count. ChatGPT Pro at $200 is described in OpenAI’s help center the same way: “Pro $100 unlocks 5x higher usage than Plus, while Pro $200 unlocks 20x usage than Plus.” OpenAI adds that “There is no setting to increase or bypass a model’s usage allowance.”
This is a finding, not a research gap. Both help-center articles were checked specifically for a token or message figure, and neither publishes one — the multiplier framing is the vendors’ own choice. The $200-a-month plans therefore cannot be merged into the $200-of-API table above, and any post that does merge them is inventing the conversion the vendors declined to publish. How those multipliers have shifted over the year is tracked in our coding-plan limit change ledger, and the broader included-versus-metered divide is the subject of our subscriptions-versus-API comparison — background here, because the conversion table has already made this page’s point without it.
One more absence worth stating precisely: no pricing page collected for this dataset publishes a prepaid volume-discount rate. Anthropic’s page says volume discounts “are negotiated on a case-by-case basis,” and its usage tiers govern rate limits, not price. If a fixed $200 is your whole lever budget, the published machinery — cache, batch, and DeepSeek’s clock — is the machinery there is. Mapping which of your workloads can actually ride those lanes is the kind of question our AI transformation engagements open with.
07 — ConclusionThe lane, not the logo.
How you call the API moves the answer by an order of magnitude. Which model you pick moves it less.
The conversion table says it in one row: Claude Opus 5 spans 40M to a derived 800M input tokens at the same $200, while the list-rate gap between Opus 5 and its nearest rivals is a small multiple. The same holds with printed cells at DeepSeek, where one model’s published lanes sit 60x apart. Budget conversations that start with “which model” are starting on the smaller variable.
The vendor differences that matter are about disclosure as much as price. OpenAI and DeepSeek print their compounded cheap lanes as table cells. Anthropic confirms its lanes combine but leaves the multiplication to you. A cheap lane you cannot trace to a vendor’s own page is not a lane; it is a guess.
And the two $200 subscriptions stay out of the table for the honest reason: neither vendor publishes the conversion. When a budget needs defending, use the arithmetic that traces — $200 divided by a published rate — and label every derived cell as derived, the way this page does.