AI DevelopmentCost Playbook11 min readPublished August 31, 2026

Same model, same $200 — Claude Opus 5’s published lanes span 10x in tokens, and 20x once the derived batch-plus-cache cell is included

What $200 a Month of AI Buys, and What Changes It Most

A conversion table that turns a fixed $200 into millions of input tokens at every rate lane seven vendors publish — list, cached input, batch, and DeepSeek’s off-peak — with a borrowed, attributed quality index alongside.

DA
Digital Applied Team
Senior strategists · Published Aug 31, 2026
PublishedAug 31, 2026
Read time11 min
Sources10
One model's spread
20x
Opus 5 at $200: 40M list tokens to 800M derived
Cache-read multiplier
0.1x
published by Anthropic, OpenAI and Alibaba alike
DeepSeek's printed cell
9.1B
V4-Pro tokens at $200, cache-hit off-peak
Prepaid volume discounts
0
published on the pricing pages collected here

How you call a model API moves what $200 buys by roughly an order of magnitude, with the model held constant. The conversion table below turns each vendor-published rate into input tokens at a fixed $200: Claude Opus 5 alone spans 40M tokens at the list rate to a derived 800M through the batch-plus-cache lane, a 20x range with the model held constant.

Two definitions carry the table. Cached input is not a discount you negotiate; it is a different price for a different operation — re-reading a prompt prefix the provider has already processed, billed at 0.1x the base input rate wherever the vendors below publish the multiplier. Batch is not a cheaper model; it is an asynchronous lane — the same weights answering within a service window instead of now, at 50% off both input and output where it exists.

Key takeaways
  1. 01
    One model, one budget, a 20x spread.At $200, Claude Opus 5 spans 40M list tokens to a derived 800M with batch and caching combined — a 20x spread before you compare a single competing model.
  2. 02
    Cache is a 0.1x read multiplier; batch is 50% off.Anthropic, OpenAI and Alibaba each publish cache reads at one tenth of base input. Batch, where offered, halves both directions. Only some vendors print the combined price.
  3. 03
    DeepSeek prints an off-peak lane as a table cell.V4-Pro's cache-hit off-peak rate of $0.022 per million tokens is a printed table cell, not arithmetic. It turns $200 into roughly 9.1B input tokens — 60x its own peak cache-miss lane.

01The table$200, converted into tokens, lane by lane.

Every cell below is the same arithmetic: $200 divided by a vendor-published price per million input tokens. The rates themselves live in our maintained per-million-token price index; this page inverts them into the question a budget line actually asks — how many tokens does a fixed $200 buy, and which lever changes that most. Rows marked published cell are combined prices the vendor prints itself; the one row marked derived is our multiplication of two published multipliers, never a vendor figure.

Input tokens a fixed $200 buys, in millions (M), by published rate lane. Formula for every cell: $200 ÷ the vendor’s published price per million input tokens; figures rounded. All prices are from the vendors’ own pricing documentation. The off-peak column is populated only where a vendor publishes a time-of-day rate — among the pages collected here, only DeepSeek does. Promotional lanes are labelled; end dates are vendor-announced.
Model and laneList inputCached inputBatch inputOff-peak input
Anthropic — Claude API
Claude Opus 540M400M80M
Claude Opus 5 · batch + cache derived800M*
Claude Sonnet 5100M1,000M200M
OpenAI — short context, promotional rate
GPT-5.6 Sol50M500M100M
GPT-5.6 Sol · batch cached input published cell1,000M
Google — Gemini API
Gemini 3.7 Flash (promotional)267M2,667M533M
Gemini 3.1 Pro Preview (≤200K prompts)100M1,000M200M
DeepSeek — no batch lane; peak and off-peak published
DeepSeek V4-Pro (peak)151.5M4,545Mnone published303M
DeepSeek V4-Pro · cache-hit × off-peak published cell9,091M
Z.ai — cache lever only
GLM-5.3142.9M769.2Mnone published
GLM-5.3-Flash (promotional)2,667M13,333Mnone published
Moonshot — cache lever only
Kimi K366.7M666.7Mnone published
Alibaba — International endpoint
Qwen3.8-Max†100M1,000M200M

* Derived, not vendor-published: 0.5 (batch) × 0.1 (cache read) × $5.00 base input = $0.25 per million tokens. Anthropic’s FAQ confirms the two discounts combine, but the company never prints this cell the way OpenAI does one row up. Alibaba documents cache and batch as two separate rules and nowhere confirms they stack, so no combined Qwen figure appears in this table — multiplying them would invent a number the vendor has not stated. The list lanes for GPT-5.6 Sol, Gemini 3.7 Flash and GLM-5.3-Flash are promotional, with vendor-announced end dates of November 21, 2026, December 31, 2026 and September 9, 2026 respectively.

02MethodWhat was computed, and what this page owns.

This page publishes one derivation and one borrowed axis, and it is deliberate about which numbers are whose. The rates belong to the vendors, the quality scores belong to Artificial Analysis, and the only thing that is ours is the inversion — tokens per fixed budget — plus the matrix of which levers each vendor publishes.

Methodology
What was computed
For seven vendors’ current flagship models: $200 divided by each published input-token rate lane — list, cached input, batch, and time-of-day where published. Input side only; output rates are not tabulated here. Figures rounded to the nearest million tokens, or one decimal where finer.
What this page owns
The conversion and the quality-attached ranking, only. The per-million-token rates are owned by our maintained price index and are not restated as a rate card here. Batch pricing as a subject is owned by our batch API landscape; batch appears here strictly as one input to the conversion.
Sources
Each vendor’s own pricing documentation: Anthropic, OpenAI, Google, DeepSeek, Z.ai, Moonshot and Alibaba Cloud, plus the two vendors’ help-center articles on their $200 subscription tiers and Artificial Analysis’s leaderboard for the quality axis. No aggregator supplies any rate cell.
Dates
Collected 2026-09-01. Every pricing page cited is a living reference page, so figures are the current state as retrieved, not events on a date. Promotional end dates that extend past publication are vendor-announced schedules that were already public.
Units
Millions of input tokens (M) per $200. Output tokens are excluded from the table because cache reads apply to input only; batch, where offered, also halves output, which the interpretation notes but the cells do not encode.
What is excluded
The $200 subscription tiers, Claude Max and ChatGPT Pro, because neither vendor expresses them in a convertible unit (§06). Artificial Analysis’s own “Blended $/1M” cost column, an AA estimate that would corrupt the arithmetic if mixed with vendor list rates. Mainland-China endpoints, which this pass did not verify first-party.
Known limitations
The Anthropic batch-plus-cache cell is derived from two published multipliers, and is labelled so everywhere it appears. No stacked Qwen figure is computed. The Artificial Analysis page carries no visible as-of stamp, so its scores are reported as retrieved for this dataset, with no board date asserted.

03The leversWhich levers each vendor publishes — and which prices they print.

The conversion table only works because each vendor publishes a different set of levers, and the differences are structural, not cosmetic. Anthropic and Alibaba publish nearly identical percentages — cache reads at one tenth of base input, batch at half price — yet only Anthropic confirms in writing that the two combine. OpenAI goes further and prints the compounded cell in its own batch table. Google publishes both levers but states that batch context caching is billed at the same standard cache rate, so there is no extra stack to compute. DeepSeek has no batch lane at all and instead publishes the survey’s deepest lever: a full cache-by-time-of-day price matrix.

Discount levers by vendor, from each vendor’s own pricing documentation: whether a cache-read rate, a batch lane, and a time-of-day rate are published, and whether the vendor prints the combined cache-plus-batch price or leaves the reader to multiply.
Vendor · modelCache readBatchOff-peakCombined cell printed?
Anthropic · Opus 5, Sonnet 5Yes — 0.1x base inputYes — 50% off both directionsnot publishedNo. The FAQ confirms the discounts combine; no cell is printed.
OpenAI · GPT-5.6 SolYes — 0.1x base inputYes — 50% off both directionsnot publishedYes. The batch table prints cached input at $0.20 per million tokens.
Google · Gemini 3.7 Flash, 3.1 ProYes — per-model caching ratesYes — 50% offnot publishedYes, but batch caching is billed at the standard cache rate — no extra stack.
DeepSeek · V4-Pro, V4-FlashYes — published cache-hit rowsnone publishedYes — peak and off-peak rates printedYes. The full cache × time-of-day matrix is printed as table cells.
Z.ai · GLM-5.3Yes — $0.26 vs $1.40 (≈0.19x)none publishednot publishedn/a — one lever only
Moonshot · Kimi K3Yes — $0.30 hit vs $3.00 miss (0.1x)none publishednot publishedn/a — one lever only
Alibaba · Qwen3.8-MaxYes — cache hits at 10% of inputYes — 50% of real-time pricenot publishedNo. Two separate rules; stacking is not confirmed anywhere on the page.
Batch API and prompt caching discounts can be combined ... See prompt caching pricing for how the multipliers interact.Anthropic pricing documentation, FAQ
Disclosure style one
Prints the combined cell
OpenAI · DeepSeek

OpenAI's batch table carries its own cached-input column, so the compounded price is a published number. DeepSeek prints a full cache-by-time-of-day matrix. In both cases the cheapest lane is a cell you can point at, not arithmetic you must defend.

Read it off the page
Disclosure style two
Confirms stacking in prose
Anthropic

The FAQ states batch and caching combine, but no compounded cell is printed. The $0.25-per-million lane for Opus 5 is real per the vendor's own rules — and it is still your multiplication, which is why this page labels it derived every time it appears.

You do the multiplying
The counterexample
Levers that exclude each other
Anthropic fast mode

Anthropic's fast mode (research preview, $10/$50 for Opus 5) states plainly that it is not available with the Batch API. Levers are vendor-specific machinery, not a universal rulebook — some stack, some are mutually exclusive.

Not everything stacks

The mechanics behind the two big levers are covered in depth elsewhere: our prompt caching engineering guide explains what a cache read physically is, and the half-price batch lane maps who offers batch and on what terms. This page takes both as given and asks only what they do to a fixed budget.

04The findingThe same model spans 20x across its own lanes.

Hold the model constant and walk the lanes. Claude Opus 5 at $200 buys 40M input tokens at list, 80M through batch, 400M through cache reads, and a derived 800M with the two combined — the same weights answering the same prompts across a 20x range. Now compare across models at the list rate only: Opus 5’s 40M against Qwen3.8-Max’s 100M is a 2.5x gap, and against Gemini 3.7 Flash’s promotional 267M the gap is under 7x. Step further out and the between-model range is wider still: GLM-5.3-Flash’s promotional list rate returns 2,667M for the same $200. Both axes move the answer by more than most buyers assume, and only one of them appears on a pricing page.

DeepSeek makes the same point with printed cells rather than derivation. V4-Pro’s peak cache-miss rate is $1.32 per million tokens; its off-peak cache-hit rate is $0.022. Both are published table cells, and the ratio between them is 60x, entirely inside one model. At $200 that is the difference between 151.5M tokens and roughly 9.1B. The off-peak window is generous, too: DeepSeek’s definition leaves every hour outside two weekday UTC bands, and all of the weekend, at the lower rate.

Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday (all other hours are off-peak).DeepSeek API pricing documentation
Derived numbers, labelled every time

One discipline keeps this table honest: a compounded price is only as real as its publication. Anthropic’s $0.25 batch-plus-cache lane for Opus 5 follows from two published multipliers and an FAQ that confirms they combine — so it appears here, marked derived. Where a vendor prints the cell, we cite the cell.

The practical reading order for a $200 budget follows directly. First ask which of your traffic is re-read prefix — system prompts, tool definitions, shared context — because the 0.1x cache-read multiplier is the single largest lever most vendors publish. Second, ask what fraction of your work can wait for an asynchronous window, because batch halves both directions where it exists. Only third ask which model, because by then the same model’s price has already moved by an order of magnitude under your first two answers.

05QualityThe quality axis, borrowed and attributed.

Tokens per dollar mean nothing without a quality axis, and the quality axis is not ours. The scores below are Artificial Analysis’s Intelligence Index, retrieved for this dataset from its public leaderboard, at the highest reasoning effort AA lists for each model. The ranking is AA’s judgment, not Digital Applied’s; we pair it with the cheapest lane each vendor publishes and draw one observation from the pairing.

Artificial Analysis Intelligence Index scores (AA’s labels and judgment, at the highest listed reasoning effort), paired with each model’s cheapest vendor-published input lane converted at $200. AA’s page carries no visible as-of stamp; scores are as retrieved for this dataset. AA’s own blended-cost column is deliberately not used anywhere on this page.
Model (AA’s label)AA Intelligence IndexCheapest published input laneTokens at $200
Claude Opus 5 (max)63Cache read, $0.50400M
GPT-5.6 Sol (max)61Batch cached input, $0.20 (printed)1,000M
Kimi K3 (max)60Cache-hit input, $0.30666.7M
GLM-5.3 (max)60Cached input, $0.26769.2M
Qwen3.8 Max58Cache-hit input, $0.201,000M
Gemini 3.7 Flash (high)56Context caching, $0.075 (promotional)2,667M
Claude Sonnet 5 (max)55Cache read, $0.201,000M
DeepSeek V4 Pro 0813 (max)53Cache-hit off-peak, $0.022 (printed)9,091M
Gemini 3.1 Pro Preview48Cached input, $0.20 (≤200K prompts)1,000M

The observation the pairing earns: the quality gap between the top and bottom of AA’s table is fifteen index points, and the budget gap between lanes routinely dwarfs what that quality gap costs. Claude Opus 5’s derived batch-plus-cache lane ($0.25 per million) is cheaper than Gemini 3.1 Pro Preview’s list rate ($2.00 per million) — eight times cheaper, for a model AA scores fifteen points higher. On these numbers, learning to use a vendor’s published levers buys more capability per dollar than shopping down the quality table does.

06Negative resultThe two $200 plans that don’t convert.

There is a second way to spend $200 a month on frontier AI, and it deliberately resists this table. Claude Max at $200 is described in Anthropic’s help center only as “20 times more usage per session than the Pro plan” — no token count, no message count. ChatGPT Pro at $200 is described in OpenAI’s help center the same way: “Pro $100 unlocks 5x higher usage than Plus, while Pro $200 unlocks 20x usage than Plus.” OpenAI adds that “There is no setting to increase or bypass a model’s usage allowance.”

This is a finding, not a research gap. Both help-center articles were checked specifically for a token or message figure, and neither publishes one — the multiplier framing is the vendors’ own choice. The $200-a-month plans therefore cannot be merged into the $200-of-API table above, and any post that does merge them is inventing the conversion the vendors declined to publish. How those multipliers have shifted over the year is tracked in our coding-plan limit change ledger, and the broader included-versus-metered divide is the subject of our subscriptions-versus-API comparison — background here, because the conversion table has already made this page’s point without it.

One more absence worth stating precisely: no pricing page collected for this dataset publishes a prepaid volume-discount rate. Anthropic’s page says volume discounts “are negotiated on a case-by-case basis,” and its usage tiers govern rate limits, not price. If a fixed $200 is your whole lever budget, the published machinery — cache, batch, and DeepSeek’s clock — is the machinery there is. Mapping which of your workloads can actually ride those lanes is the kind of question our AI transformation engagements open with.

07ConclusionThe lane, not the logo.

What $200 buys

How you call the API moves the answer by an order of magnitude. Which model you pick moves it less.

The conversion table says it in one row: Claude Opus 5 spans 40M to a derived 800M input tokens at the same $200, while the list-rate gap between Opus 5 and its nearest rivals is a small multiple. The same holds with printed cells at DeepSeek, where one model’s published lanes sit 60x apart. Budget conversations that start with “which model” are starting on the smaller variable.

The vendor differences that matter are about disclosure as much as price. OpenAI and DeepSeek print their compounded cheap lanes as table cells. Anthropic confirms its lanes combine but leaves the multiplication to you. A cheap lane you cannot trace to a vendor’s own page is not a lane; it is a guess.

And the two $200 subscriptions stay out of the table for the honest reason: neither vendor publishes the conversion. When a budget needs defending, use the arithmetic that traces — $200 divided by a published rate — and label every derived cell as derived, the way this page does.

Make a fixed AI budget go further

The rate is published. The lane is a choice.

Our team maps which of your AI workloads can actually ride the cheap lanes — cacheable prefixes, batchable jobs, off-peak windows — and what that does to a fixed monthly budget before anyone debates models.

Free consultationExpert guidanceTailored solutions
What we work on

AI cost engagements

  • Cache-hit modelling on real prompt traffic
  • Batch eligibility reviews for asynchronous workloads
  • Lane-by-lane budget conversion for committed spend
  • Promotional-expiry exposure in fixed budgets
  • Model selection after the levers are priced in
FAQ · AI token budgets

The questions we get about token budgets.

Not a published one, on any pricing page collected for this dataset. Anthropic states volume discounts are negotiated on a case-by-case basis, and its usage tiers govern rate limits rather than price. The published levers — cache reads, batch lanes, and DeepSeek's off-peak clock — are what move a fixed budget.
Related dispatches

Continue exploring AI model costs.