This is a maintained AI model API price index: what every current model costs per million tokens on input, cached input and output, plus batch, flex, off-peak, priority and fast-mode rates wherever the vendor publishes them, organised by the surface you buy on. Data as of August 30, 2026. Every figure traces to the vendor’s own pricing page. Promotional rates never share a column with list rates.
The reason to build it this way is not tidiness. In the last three weeks two separate frontier rates moved for reasons that were printed in a footnote and dropped by nearly everyone who quoted the number — including, in one case, us. A reader looking at two confident secondary sources had no way to tell which one to believe, because neither carried the footnote that settled it.
So the index leads with those two cases, then gives you the tables: the full rate card, the promotions with their end dates, the surface layer where the same model bills differently depending on where you bought it, the batch slugs that cost more than standard, the long-context cliffs that make any one-number-per-model table incomplete, and the one vendor charging in data rather than money.
- 01Claude Sonnet 5 is $2 / $10 and that is now the standard price.Anthropic announced the $2/$10 rate at launch as introductory pricing through August 31, 2026, then cancelled the scheduled September 1 increase to $3/$15 on August 10, 2026. Both the pricing docs and the Sonnet 5 launch post carry the reversal in writing. Any table still showing a September 1 cliff, or $3/$15 as a future Sonnet 5 rate, is wrong.
- 02GPT-5.6 Sol’s $4 / $20 is promotional, not a list price.OpenAI’s own model page states the promotional pricing is available at least through November 21, 2026. OpenAI publishes no figure for what follows. The vendor describes the cut as a 20% reduction in input and a 33% reduction in output, which back-solves to $5 / $30 — that is vendor arithmetic, not a published future price, and this index labels it as such.
- 03Promotional and list rates are in separate columns, always.Seven vendors here run a discount of some kind and only three publish a hard end date: Z.ai’s GLM-5.3-Flash promo ends at 24:00 on September 9, 2026 (UTC+8), Google’s Gemini 3.7 and 3.6 Flash promo ends December 31, 2026, and Tencent’s free Hy3 access is stated as running to September 30, year not printed on the page. Everything else is limited-time, temporarily, or — at MiniMax — permanently discounted.
- 04The same model bills differently depending on the surface.Anthropic bills through AWS Marketplace and Microsoft Foundry in Claude Consumption Units at $0.01 per CCU: same rate, different unit on the invoice. AWS Bedrock’s regional price list moves Gemma 3 27B input from $0.23 in US regions to $0.36 in London, a 56.5% premium AWS labels only as a region. Claude Fast mode is first-party only — not on Claude Platform on AWS or partner clouds.
- 05Rows we could not verify are marked and kept, never dropped.Three labels, three distinct styles: the vendor does not publish it, the page could not be located, or the page was located but could not be read. Cohere’s missing generative card is the first kind. Azure’s OpenAI per-token rates are the second. Amazon Nova’s region-selector tables are the third. They are not the same kind of gap.
01 — Why this existsTwo rates, two footnotes, and one of them was ours.
A price index earns the right to be cited when it can settle a disagreement that the reader cannot settle themselves. Two cases in August 2026 show exactly what that looks like, and they run in opposite directions from the same underlying mistake.
The first is Anthropic’s Claude Sonnet 5. When Sonnet 5 launched, $2 per million input tokens and $10 per million output tokens were published as introductory pricing through August 31, 2026, with a scheduled increase to $3 / $15 on September 1. Two secondary sources contradicted each other this month: one reported the increase had been cancelled, one reported it was landing. Only the vendor could settle it, and the vendor did. Anthropic’s pricing documentation now carries a callout stating that the $2 / $10 pricing “is now the standard price” and that “the previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur.” The Sonnet 5 launch post carries a dated edit note to the same effect, marked August 10, 2026.
Our own August 5 pricing tracker reported the expiry footnote accurately, quoting Anthropic’s own wording, and built a section on the September 1 step-up. Five days later the vendor withdrew the increase. The post was not careless; it was a correct snapshot of a rate whose footnote the vendor then deleted. That is the strongest argument for a maintained index we can make, and it is about our own work, which is why it belongs at the top of this page rather than buried in a correction note.
From Anthropic’s pricing documentation, fetched August 30, 2026: “The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur.”
The same vendor’s launch post carries a dated edit: “Edit August 10, 2026: Sonnet 5’s introductory pricing of $2 per million input tokens and $10 per million output tokens is now permanent.” Anthropic also annotated its own cost-performance charts as stale, because they were plotted at $3 / $15. A vendor telling you which of its own published numbers is out of date is rarer than it should be.
The second is OpenAI’s GPT-5.6 Sol, and it is the same error running backwards. Sol’s current $4 / $20 was widely reported — and recorded in our own internal anchors — as a list price. It is not. OpenAI’s model page states plainly that “GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.” A promotional decrease was mistaken for a permanent one, where Sonnet 5 was a scheduled increase mistaken for an inevitable one. Both mistakes come from taking a rate without its footnote.
The two vendors also differ sharply in what they disclose about the other side of the promotion. Google publishes both halves for its Gemini 3.7 and 3.6 Flash promotion — the promotional rate, the end date, and the exact figure that applies afterwards. OpenAI publishes a soft floor and no post-promotional rate at all. What it does publish is a characterisation: the cut is “a 20% reduction in input pricing and a 33% reduction in output pricing.” Run that backwards and you land on $5 / $30, which is GPT-5.5’s rate. This index will report that only as vendor arithmetic, and never in a price cell, because OpenAI has not published it as a price.
A scheduled increase, cancelled
Sonnet 5 launched with introductory pricing through August 31, 2026 and a scheduled step-up to $3 / $15. On August 10 Anthropic cancelled the increase and made $2 / $10 the standard price. Two secondary sources disagreed about this all month; the vendor's own docs and launch post settle it in two independent places.
A promotional decrease, mislabelled
GPT-5.6 Sol's rate is promotional and carries a vendor-published floor of at least through November 21, 2026. No post-promotional figure is published. The $5 / $30 that the vendor's own percentages back-solve to is arithmetic, not a price, and appears nowhere in this index's rate columns.
A rate without its footnote is not a rate
Every promotional row in this index carries its published end date, or an explicit note that the vendor published none. Merging a promotional rate into a list column produces a table that cannot be falsified, which for four of the seven discounting vendors here is exactly what would happen.
The habit of reading a rate together with its footnote is the same literacy our model catalogue literacy guide argues for across prices, dates and surfaces. This page is that argument turned into a dataset.
02 — MethodWhat was collected, what is missing, and who keeps it current.
The gaps do as much work here as the prices. A price index that silently drops the rows it could not verify looks more complete than it is, and quietly teaches the reader that everything absent from it does not exist. Every unverified row below is kept in and labelled, and the three labels mean three different things.
What was collected. For every current model with a publicly published API price, the input rate, the cached-input rate, the output rate, and every discount or premium tier the vendor publishes — batch, flex, off-peak, priority, fast mode, long-context bands and regional endpoints. Rows are organised by surface: the vendor’s own direct API, the cloud resellers that meter the same model, and the coding plans that include it without publishing a token rate.
Sources. Every price traces to the vendor’s own pricing page, model page or documentation. Where a figure could only be obtained from a second vendor’s page — Google’s Vertex tables carry cache-read rates that the Gemini Developer API page does not render — the row says which page it came from. No figure in the rate columns comes from an aggregator or a price-tracker site.
Dates. Pages were fetched on August 30, 2026. The as-of date is August 30, 2026. Nothing dated after that appears as an event. Scheduled future prices — Google’s January 1, 2027 reversion is the clearest — are recorded as plans, not outcomes. The Sonnet 5 case above is the reason that distinction is enforced rather than assumed.
Currency. Every row states its currency and its unit. Where a vendor publishes a native non-USD price, both are reported as published and neither is converted. DeepSeek’s CNY table is not a conversion of its USD table — the implied cross-rate is about 6.82 CNY per USD on the main rows and about 7.14 on the V4-Flash cache-hit row, because each locale is rounded to its own tidy numbers. Tencent publishes $0.834 / $2.501 / $0.042 on its English release and ¥6 / ¥18 / “as low as ¥0.3” on its Chinese one; both are vendor-published regional list prices and neither is our arithmetic.
The three gap labels. Rows that could not be verified are kept in the tables with one of three marks, rendered three different ways so the distinction survives a screenshot: not published means the vendor does not publish this figure; page not located means the page that would carry it was not found; and not machine-readable means the page was located but its tables did not render to any fetch method used. The first is a fact about the vendor. The second and third are facts about this collection pass.
What is excluded. Non-token meters, except where a vendor prices a frontier product only that way — Cohere’s instance-hour Model Vault rates and MiniMax’s per-second video rates are in because their absence would misrepresent the vendor. Fine tuning, image, audio, embedding and tool-call pricing are out of the main table; they belong to a different index. Coding-plan subscriptions appear only as a surface with no published token rate, because that is the accurate description of them.
Refresh owner and cadence. Maintained by the Digital Applied Team, reviewed monthly and out of cycle when a frontier vendor moves a headline rate. The slug carries no date deliberately: each refresh updates the tables in place and re-dates the page through its modified time, so the URL stays stable and citable while the as-of date moves.
Our archive already holds four dated pricing snapshots. This URL is the one that gets updated from here on. Those four stand as historical records of what the market published on their own dates, and should be cited that way — not as current rates:
LLM API pricing index and cost tracker (March 26, 2026) · LLM API pricing index, Q2 2026 (April 12, 2026) · AI model API pricing tracker, Q2 2026 (April 23, 2026) · AI API pricing, August 2026: cuts, promos and traps (August 5, 2026).
Each of those four carries an update block pointing here, so the cross-link runs both ways. The August 5 post also carries a dated correction to its Sonnet 5 section, for the reason set out above.
Every rate column in this index resolves to one of the pages below. They are listed rather than footnoted so a reader can re-derive any cell, and so a later refresh has an explicit list to re-fetch.
| Vendor | What it supplies | Primary page |
|---|---|---|
| Anthropic | Model, batch, fast-mode and cache rates; CCU billing on AWS and Microsoft Foundry; data-residency multiplier | platform.claude.com pricing |
| OpenAI | Standard, Batch, Flex and Fast mode tables; the Sol promotional footnote; the 272,000-token long-context rule | platform.openai.com pricing |
| Standard, batch, flex and priority tabs; the promotional rates and their December 31, 2026 end date | ai.google.dev pricing · Vertex AI pricing | |
| DeepSeek | Peak and off-peak tables in both USD and CNY; the peak-window definition | api-docs.deepseek.com pricing |
| Z.ai | GLM list and promotional rates with the September 9 end date; the coding-plan credit multipliers | docs.z.ai pricing |
| Moonshot AI | Per-model cache-hit, cache-miss and output rates; the batch tier and its exclusions | platform.kimi.ai pricing |
| MiniMax | M3 Standard and Priority tabs, the permanent-discount badge, and the per-second H3 video rates | platform.minimax.io pricing |
| Alibaba Cloud | Every Qwen SKU with its tiering, promotional labels, batch and cache rules | Model Studio billing |
| Tencent | The Hy4 preview rates in both native currencies, on the vendor’s two language releases | Hy4 preview release |
| xAI | The full Grok price table, the 200k long-context rule, and the per-model batch support statements | docs.x.ai models |
| Meta | Standard and Contributor tier rates, the tier definition, and the no-long-context-premium statement | Meta Model API pricing |
| Mistral | Per-model cards and rates; the Priority Tier multiplier and its SLA | docs.mistral.ai priority tier |
| Cohere | Model Vault per-instance rates, and the absence of a current frontier per-token card | cohere.com pricing |
| Amazon Web Services | The Bedrock regional price list, the Priority and Flex multipliers, and the per-model cards | Bedrock pricing |
| Microsoft | Deployment types, the Data Zone and Government premiums, and the long-context restatement | Foundry deployment types |
| OpenRouter | The passthrough statement, the platform and BYOK fees, and the live slug rates including the batch inversions | openrouter.ai pricing |
03 — The IndexEvery current model, by surface.
Read the columns carefully. List holds the standard rate. Promotional holds a discounted rate together with its published end date, or the vendor’s own words where no date exists. They are never merged, and a blank promotional cell means the vendor publishes no promotion on that row — not that we did not look.
Discount tiers collects batch, flex and off-peak. Premium tiers collects fast mode, priority, peak rates and the long-context bands. Both are per-vendor vocabularies rather than a shared one: “batch” means a 50% discount at Anthropic, OpenAI, Google, Alibaba, Mistral and AWS; 60% of standard at Moonshot; a 20% discount at xAI on one model and nothing at all on its flagship; and, on three OpenRouter slugs, a price higher than standard. Section 06 takes that apart.
| Model | Surface | Currency and unit | List input | List cached input | List output | Promotional rate and end date | Discount tiers | Premium tiers and bands |
|---|---|---|---|---|---|---|---|---|
| Anthropic — Claude API, first party | ||||||||
| Claude Fable 5 | Claude API direct | USD / 1M tokens | $10.00 | $1.00 | $50.00 | None published | Batch $5.00 / $25.00 | Cache write $12.50 for 5 minutes, $20.00 for 1 hour. No fast mode row. |
| Claude Mythos 5 (limited availability) | Claude API direct | USD / 1M tokens | $10.00 | $1.00 | $50.00 | None published | Batch $5.00 / $25.00 | On Bedrock, AWS states access “is gated and requires approval.” Capabilities are not documented on the pricing page and are not asserted here. |
| Claude Opus 5 | Claude API direct | USD / 1M tokens | $5.00 | $0.50 | $25.00 | None published | Batch $2.50 / $12.50 | Fast mode $10.00 / $50.00 — exactly 2× standard, Claude API first party only. Cache write $6.25 / $10.00. |
| Claude Opus 4.8 | Claude API direct | USD / 1M tokens | $5.00 | $0.50 | $25.00 | None published | Batch $2.50 / $12.50 | Fast mode $10.00 / $50.00, first party only. Applies across the full context window. |
| Claude Opus 4.7, 4.6 and 4.5 | Claude API direct | USD / 1M tokens | $5.00 | $0.50 | $25.00 | None published | Batch $2.50 / $12.50 | No fast mode. On 4.7 a fast request returns an error; on 4.6 it runs at standard speed and bills at standard rates. |
| Claude Opus 4.1 (retired, except on Bedrock and Google Cloud) | Claude API direct | USD / 1M tokens | $15.00 | $1.50 | $75.00 | None published | Batch $7.50 / $37.50 | Retired on the first-party API; still sold on two reseller surfaces at these rates. |
| Claude Opus 4 (retired, except on Google Cloud) | Claude API direct | USD / 1M tokens | $15.00 | $1.50 | $75.00 | None published | Batch $7.50 / $37.50 | One surviving surface. |
| Claude Sonnet 5 | Claude API direct | USD / 1M tokens | $2.00 | $0.20 | $10.00 | None. The launch-era introductory pricing became the standard price on August 10, 2026, and the scheduled increase was cancelled. No expiry. | Batch $1.00 / $5.00 | Cache write $2.50 / $4.00. No fast mode row. |
| Claude Sonnet 4.6 and 4.5 | Claude API direct | USD / 1M tokens | $3.00 | $0.30 | $15.00 | None published | Batch $1.50 / $7.50 | Both older Sonnets cost 50% more on input than Sonnet 5 and 50% more on output. |
| Claude Sonnet 4 (retired, except on Bedrock and Google Cloud) | Claude API direct | USD / 1M tokens | $3.00 | $0.30 | $15.00 | None published | Batch $1.50 / $7.50 | — |
| Claude Haiku 4.5 | Claude API direct | USD / 1M tokens | $1.00 | $0.10 | $5.00 | None published | Batch $0.50 / $2.50 | Cache write $1.25 / $2.00. |
| Claude Haiku 3.5 (retired, except on Bedrock and Google Cloud) | Claude API direct | USD / 1M tokens | $0.80 | $0.08 | $4.00 | None published | Batch $0.40 / $2.00 | The cheapest Claude row anywhere, and it is retired on the first-party API. |
| Every Claude model above | Claude Platform on AWS · Claude in Microsoft Foundry | Claude Consumption Units, $0.01 per CCU (100 CCU = $1.00) | Token usage is rated in USD at the standard per-model rates above, discounts applied, then converted to CCUs and reported hourly. The AWS or Azure bill shows one CCU line item. | None published | Batch discount applies at the token-rating step, before conversion. | Fast mode is not available here. Data residency: inference_geo: "us" applies a 1.1× multiplier to every token category. | ||
| OpenAI — direct API, Standard tier, prompts at or below 272,000 input tokens | ||||||||
gpt-5.6-sol | OpenAI API direct | USD / 1M tokens | not publishedOpenAI publishes no post-promotional figure for Sol. | not published | not published | $4.00 in · $0.40 cached · $20.00 out“available at least through November 21, 2026.” The vendor calls the cut “a 20% reduction in input pricing and a 33% reduction in output pricing,” which back-solves to $5 / $30 — vendor arithmetic, not a published price. | Batch $2.00 / $10.00 · Flex identical · cache write $2.50 | Fast mode $8.00 / $40.00. Above 272,000 input tokens: $8.00 in, $30.00 out, for the whole request. |
gpt-5.6-terra | OpenAI API direct | USD / 1M tokens | $2.00 | $0.20 | $12.00 | None — no promotional language on its model page | Batch and Flex $1.00 / $6.00 · cache write $1.25 | Fast mode $4.00 / $24.00. Above 272,000 in: $4.00 / $18.00. |
gpt-5.6-luna | OpenAI API direct | USD / 1M tokens | $0.20 | $0.02 | $1.20 | None — no promotional language on its model page | Batch and Flex $0.10 / $0.60 | Fast mode $0.40 / $2.40. Above 272,000 in: $0.40 / $1.80. |
gpt-5.6-cyber (Daybreak) | OpenAI API direct only — absent from Azure’s model list | USD / 1M tokens | $12.50 | $1.25 | $75.00 | None published | None. No batch row, no flex row, and v1/batch is listed as “Not supported.” | No fast mode row. Cache write $15.625. The highest headline price in the index with no published way to reduce it. |
gpt-5.5-cyber (Daybreak) | OpenAI API direct only | USD / 1M tokens | $12.50 | $1.25 | $75.00 | None published | None published | Requires separate approval and provisioning. |
gpt-5.4-cyber (Daybreak) | OpenAI API direct only | USD / 1M tokens | not publishedThe row exists on OpenAI’s own pricing table with every cell empty. That is the vendor’s presentation, not a failed fetch. | None published | None published | — | ||
gpt-5.5 | OpenAI API direct | USD / 1M tokens | $5.00 | $0.50 | $30.00 | None published | Batch and Flex $2.50 / $15.00 | Fast mode $12.50 / $75.00. Above 272,000 in: $10.00 / $45.00. |
gpt-5.5-pro | OpenAI API direct | USD / 1M tokens | $30.00 | not published | $180.00 | None published | Batch and Flex $15.00 / $90.00 | No fast mode row — no Pro-tier SKU has one. Above 272,000 in: $60.00 / $270.00. |
gpt-5.4 | OpenAI API direct | USD / 1M tokens | $2.50 | $0.25 | $15.00 | None published | Batch and Flex $1.25 / $7.50 | Fast mode $5.00 / $30.00. Above 272,000 in: $5.00 / $22.50. |
gpt-5.4-mini | OpenAI API direct | USD / 1M tokens | $0.75 | $0.075 | $4.50 | None published | Batch and Flex $0.375 / $2.25 | Fast mode $1.50 / $9.00. |
gpt-5.4-nano | OpenAI API direct | USD / 1M tokens | $0.20 | $0.02 | $1.25 | None published | Batch and Flex $0.10 / $0.625 | No fast mode row. |
gpt-5.3-codex | OpenAI API direct | USD / 1M tokens | $1.75 | $0.175 | $14.00 | None published | not published | Fast mode $3.50 / $28.00, exactly 2×. The sole surviving Codex variant — five siblings shut down on July 23, 2026. |
chat-latest | OpenAI API direct — “latest Instant model used in ChatGPT” | USD / 1M tokens | $5.00 | $0.50 | $30.00 | None published | None published | OpenAI publishes no mapping from the ChatGPT app’s “GPT Instant” and “GPT Reasoning” labels to any model ID. |
| GPT-5.6 Sol, Terra and Luna | Microsoft Foundry, version 2026-07-09 | USD / 1M tokens | page not locatedAzure’s model page defers rates to its own pricing page, which was not reached in this pass. Limits are verified; the money is not. | None published | Global Batch and Data Zone Batch are 50% off Global Standard. | Data Zone deployments are +10% on Global Standard. Azure Government adds a further premium. | ||
| OpenAI models on Amazon Bedrock | AWS Bedrock | USD / 1M tokens | page not locatedOpenAI’s own footnote is the only sourced statement: models on Bedrock “are billed through AWS and may differ from direct OpenAI pricing.” | — | Flex and Batch 50% off Standard. | Priority 1.75× Standard. | ||
| Google — Gemini Developer API, paid tier. Tier identity read from Google’s own tab markup. | ||||||||
| Gemini 3.7 Flash | Gemini Developer API and Vertex AI | USD / 1M tokens | $1.50 | $0.15 | $7.50 | $0.75 in · $0.075 cached · $3.75 outEnds December 31, 2026. Google publishes the post-promotional figure explicitly: $1.50 / $7.50 from January 1, 2027. Cache storage $0.50 per 1M tokens per hour promotional, $1.00 after. | Batch $0.375 / $1.875 · Flex $0.375 / $1.875 (both at the promotional level) | Priority $1.35 / $6.75 — 1.8× standard. GA and the newest model here; no long-context cliff on any Flash model. |
| Gemini 3.6 Flash | Gemini Developer API and Vertex AI | USD / 1M tokens | $1.50 | $0.15 | $7.50 | $0.75 / $3.75, ends December 31, 2026Google applied the 3.7 Flash introductory rate to 3.6 Flash as well. | Batch and Flex $0.375 / $1.875 | Priority $1.35 / $6.75. |
| Gemini 3.5 Flash | Gemini Developer API and Vertex AI | USD / 1M tokens | $1.50 | $0.15 | $9.00 | None published | Batch and Flex $0.75 / $4.50 | Priority $2.70 / $16.20. Note this older model’s output rate is 2.4× the newer 3.7 Flash promotional rate. |
| Gemini 3.5 Flash-Lite | Gemini Developer API and Vertex AI | USD / 1M tokens | $0.30 | $0.03 | $2.50 | None published | Batch and Flex $0.15 / $1.25 | Priority $0.54 / $4.50. Input is $0.30 flat across all modalities. |
| Gemini 3.1 Flash-Lite | Gemini Developer API and Vertex AI | USD / 1M tokens | $0.25 text, image, video$0.50 audio | $0.025 text, image, video · $0.05 audioFrom Vertex’s table; the Developer API page did not render this cell. | $1.50 | None published | Batch and Flex $0.125 ($0.25 audio) / $0.75 | Priority $0.45 ($0.90 audio) / $2.70. Its successor charges a flat $0.30, so 3.5 is dearer for text and cheaper for audio. |
| Gemini 3.1 Pro Preview | Gemini Developer API and Vertex AI | USD / 1M tokens | $2.00 at or below 200K in$4.00 above 200K | $0.20 at or below 200K · $0.40 aboveFrom Vertex’s table. | $12.00 at or below 200K in$18.00 above 200K | None published | Batch and Flex $1.00 / $6.00, and $2.00 / $9.00 in the upper band | Priority $3.60 / $21.60, and $7.20 / $32.40 in the upper band. Newest Pro model, still carrying a preview endpoint. |
| Gemini 3 Flash Preview | Gemini Developer API and Vertex AI | USD / 1M tokens | $0.50 text, image, video · $1.00 audio | $0.05 text, image, video · $0.10 audio (Vertex table) | $3.00 | None published | Batch and Flex $0.25 / $1.50 | Priority $0.90 / $5.40. |
| Gemini Omni Flash | Gemini Developer API | USD / 1M tokens | $1.50 | not published | $9.00 text$17.50 video | None published | not publishedOnly a Standard column exists; the batch, flex and priority tabs carry no rows for this model. | not published |
| Every Gemini model above | Vertex AI, non-global endpoint (data residency) | USD / 1M tokens | Exactly +10% on every published row. Gemini 3.7 Flash: $0.75 becomes $0.825 in, $3.75 becomes $4.125 out, cached $0.075 becomes $0.0825. | Promotional rate carries through, uplifted | Same tier structure as global. | Gemini 3.1 Pro Preview and Gemini 3 Flash Preview have no non-global row published. | ||
| CodeMender and AlphaEvolve | Vertex AI only | USD / 1M tokens, additive | An additive second price column, not a multiplier. AlphaEvolve on Gemini 3.1 Pro Preview adds $4.00 in and $24.00 out on top of the model’s $2.00 / $12.00, for $6.00 / $36.00 total. | Vendor tables disagree. Vertex’s CodeMender table prices 3.7 and 3.6 Flash at $1.50 / $7.50 — the post-promotional rate — while Vertex’s main table prices the same models at the promotional $0.75 / $3.75. Both are quoted as published; this index does not pick one. | Provisioned Throughput on 3.6 and 3.7 Flash carries a 50% monthly billing credit, August 13 to December 31, 2026, on Vertex only. | Neither agent exists on the Gemini Developer API. | ||
| DeepSeek — direct API. Peak is the base rate; off-peak is half of it. | ||||||||
deepseek-v4-pro | DeepSeek direct API | USD / 1M tokens | $1.32 peak | $0.044 peak | $3.96 peak | None — the clock is a tier, not a promotion | Off-peak $0.66 in · $0.022 cached · $1.98 out. Peak is 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday. | not publishedNo Batch API exists — there is no batch page in DeepSeek’s doc sidebar. |
deepseek-v4-flash and deepseek-v4-flash-vision-exp | DeepSeek direct API | USD / 1M tokens | $0.44 peak | $0.014 peak | $1.32 peak | None | Off-peak $0.22 · $0.007 · $0.66. V4-Pro is exactly 3× these rates on input and output; the cache-hit cells round to $0.007 and $0.022, which is nearer 3.14×. | Images on the vision model are converted to tokens by dimension and billed as input — no separate rate. |
deepseek-v4-pro, native CNY table | DeepSeek direct API, Chinese locale | CNY / 1M tokens | ¥9.0 peak | ¥0.30 peak | ¥27.0 peak | None | Off-peak ¥4.5 · ¥0.15 · ¥13.5 | Published natively in CNY, not converted from the USD table. |
deepseek-v4-flash, native CNY table | DeepSeek direct API, Chinese locale | CNY / 1M tokens | ¥3.0 peak | ¥0.10 peak | ¥9.0 peak | None | Off-peak ¥1.5 · ¥0.05 · ¥4.5 | The cache-hit row implies about 7.14 CNY per USD against the USD table, where the main rows imply about 6.82. Each locale is rounded to its own tidy numbers; do not convert between them. |
| Z.ai — GLM, direct API. All prices stated in USD by the vendor. | ||||||||
| GLM-5.3 | Z.ai direct API | USD / 1M tokens | $1.40 | $0.26 | $4.40 | Cached-input storage is “Limited-time Free” with no published end date — a promotion on a third billing axis. | not publishedNo Batch API on Z.ai’s pricing page. | Weights shipped August 28, 2026 under a bespoke glm-5.3 licence, not MIT. |
| GLM-5.3-Flash | Z.ai direct API | USD / 1M tokens | $0.15 | $0.03 | $0.50 | $0.075 in · $0.015 cached · $0.25 out50% off, and the vendor states it “ends at 24:00 on September 9, 2026 (UTC+8, Singapore time).” The timezone is part of the fact — that is 16:00 UTC on September 9. | not published | The only promotional rate in this index with an end date stated to the hour and the timezone. |
| GLM-5.2 and GLM-5.1 | Z.ai direct API | USD / 1M tokens | $1.40 | $0.26 | $4.40 | None on the token rate | None published | On the GLM Coding Plan, requests for these two are silently rerouted to GLM-5.3. |
| GLM-4.7 | Z.ai direct API | USD / 1M tokens | $0.60 | $0.11 | $2.20 | None on the token rate | None published | Rerouted to GLM-5.3-Flash on the coding plan. |
| GLM-4.7-Flash and GLM-4.6V-Flash | Z.ai direct API | USD / 1M tokens | Free | Free | Free | Not a promotion — listed as free, structurally | — | — |
| GLM-5.3 and GLM-5.3-Flash | GLM Coding Plan — a subscription surface | Credits, not tokens | No per-token rate. Credits are consumed at published multipliers: GLM-5.3 at 6.9 input, 1.7 cached input and 24 output; GLM-5.3-Flash at 2.3, 0.56 and 8. Divided by 10,000. | Off-peak is 50% of the standard credit rate. Peak is Monday to Friday, 14:00–18:00 Singapore time, which is 06:00–10:00 UTC. | Plan credits: Lite 2,000 per five hours and 10,000 weekly; Pro 12,000 and 60,000; Max 28,000 and 140,000. | page not locatedPer-tier monthly dollar prices sit behind the console. The docs say only “starting at just 18 USD per month.” | ||
| Moonshot, MiniMax, Alibaba and Tencent — direct APIs | ||||||||
kimi-k3 | Moonshot direct API | USD / 1M tokens, excluding applicable taxes | $3.00 cache miss | $0.30 cache hit | $15.00 | None published | Excluded from the batch tier. Moonshot prices batch at 60% of standard and names only kimi-k2.7-code, kimi-k2.6 and kimi-k2.5 as supported. | The only vendor in this index that flags tax treatment on its own price table. |
kimi-k2.7-code | Moonshot direct API | USD / 1M tokens, ex-tax | $0.95 cache miss | $0.19 cache hit | $4.00 | None published | Batch $0.57 miss · $0.114 hit · $2.40 out — 60% of standard | kimi-k2.7-code-highspeed is exactly 2× at $1.90 / $0.38 / $8.00. The vendor describes it as the same model with a faster output speed. Latency is the whole product. |
kimi-k2.6 | Moonshot direct API | USD / 1M tokens, ex-tax | $0.95 cache miss | $0.16 cache hit | $4.00 | None published | Batch $0.57 · $0.10 · $2.40 | Max output for every Kimi model below K3 is not published |
| MiniMax-M3, at or below 512K input | MiniMax direct API, Standard tier | USD / 1M tokens | $0.60 | $0.12 | $2.40 | $0.30 in · $0.06 cached · $1.20 outBadged “Permanent 50% off” — a discount the vendor calls permanent while still rendering a struck-through list price. No end date, by construction. | not publishedNo batch discount on the pay-as-you-go page. | Priority is 1.5× standard: $0.45 / $0.09 / $1.80 effective. M3 is the flagship; M2.5, M2.1 and M2 sit under MiniMax’s own “Legacy Models” accordion. |
| MiniMax-M3, above 512K input | MiniMax direct API, Standard tier | USD / 1M tokens | $1.20 | $0.24 | $4.80 | $0.60 · $0.12 · $2.40 effective | None published | Priority effective $0.90 / $0.18 / $3.60. The tier boundary is exactly 2× on every cell. |
| MiniMax-H3 and H3 Max | MiniMax direct API | USD per second of output video, not per token | H3 at $0.08 per second at 768P and $0.13 at 2K. H3 Max at $0.05 at 480P and $0.08 at 768P. Regeneration $0.05 per second. Only the output video is billed on H3 Max. | None published | — | MiniMax-H3-Context-IR is the exception and is priced per token: $0.90 in / $3.60 out. | ||
qwen3.8-max | Alibaba Cloud Model Studio, Singapore | USD / 1M tokens | $2.00 | $0.25 implicit · $0.17 explicit read · $2.50 explicit write | $6.00 | None published | Batch not flagged on this SKU. Alibaba’s rule: batch and cache discounts “cannot apply simultaneously.” | Tiered pricing bills every token in a request at the tier the request total falls into, not just the excess. |
qwen3.8-flash | Alibaba Cloud Model Studio, Singapore | USD / 1M tokens | $0.15 | not machine-readableQwen Cloud’s pricing block did not render on two attempts. The headline pair is vendor-primary from Alibaba’s billing page; the cache cells are not. | $0.47 | None published | Not flagged for batch on this SKU. | — |
qwen3.7-max | Alibaba Cloud Model Studio, Singapore | USD / 1M tokens | $2.50 | $0.50 implicit | $7.50 | $1.25 in · $0.25 cached · $3.75 outLabelled only “Limited-time 50% off” — no end date published. The promo applies to the floating alias; the pinned snapshots stay at list. | Caches discounted 50% alongside. | Same model, two prices, decided by which ID you call. |
qwen3.7-plus, at or below 256K input | Alibaba Cloud Model Studio, Singapore | USD / 1M tokens | $0.40 | $0.08 implicit | $1.60 | $0.32 · $0.064 · $1.28“Limited-time 20% off”, no end date published. | — | Above 256K input: $1.20 / $4.80 list, with the same 20% off. |
qwen3.7-flash, at or below 32K input | Alibaba Cloud Model Studio, Singapore | USD / 1M tokens | $0.030 | $0.006 implicit | $0.130 | None published | 50% batch discount, flagged | Above 32K: $0.100 / $0.400. Above 256K: $0.200 / $0.800. |
| Hy4 preview | Tencent, international release and TokenHub Singapore | USD / 1M tokens | $0.834 | $0.042 | $2.501 | Free on WorkBuddy and CodeBuddy for two weeks from the August 28, 2026 launch. Tencent publishes no calendar end date. | not published | Apache-2.0 open weights. Preview status, with the vendor saying an official release is still ahead. |
| Hy4 preview, Chinese release | Tencent, China | CNY / 1M tokens | ¥6 | “as low as ¥0.3”The Chinese text hedges the cache figure; the English text states $0.042 flat. Quote the language you are citing. | ¥18 | Same free-trial terms | — | Both currencies are vendor-published regional list prices. Neither is our conversion, and this index does not convert between them. |
| Hy3 | Tencent TokenHub Singapore | USD / 1M tokens | $0.132 | $0.033 | $0.528 | Free on WorkBuddy and CodeBuddy, extended until September 30 | not published | — |
| xAI, Meta, Mistral, Amazon and Cohere — direct APIs | ||||||||
grok-4.6 | xAI direct API | USD / 1M tokens | $2.00 | $0.50 | $6.00 | None — xAI publishes no promotional rates at all | “Batch API: Not supported.” The flagship loses a discount tier the older model has. | Prompts reaching 200k tokens are billed at $4.00 / $1.00 / $12.00 for every token in the request. Shipped August 12, 2026; xAI’s current flagship. |
grok-4.5 | xAI direct API | USD / 1M tokens | $2.00 | $0.30 | $6.00 | None | page not locatedIts per-model page was not fetched; the price table carries no batch column. | 200k and above: $4.00 / $0.60 / $12.00. |
grok-4.3 | xAI direct API | USD / 1M tokens | $1.25 | $0.20 | $2.50 | None | Batch at a 20% discount — not the 50% the rest of the industry uses | 200k and above: $2.50 / $0.40 / $5.00. Still live and still sold; no longer the flagship. |
grok-4.20 — reasoning, non-reasoning and multi-agent | xAI direct API | USD / 1M tokens | $1.25 | $0.20 | $2.50 | None | page not located | 200k and above: 2× on every cell. |
grok-build-0.1 | xAI direct API | USD / 1M tokens | $1.00 | $0.20 | $2.00 | None | page not located | 200k and above: 2× on every cell. |
muse-spark-1.2 and 1.1, Standard tier | Meta Model API direct | USD / 1M tokens | $1.25 | $0.15 | $4.25 | None published | not publishedNo batch endpoint exists. An absence by design. | Vendor’s own words: “There is no long-context premium: you pay the same rate whether your context window is mostly empty or almost full.” |
muse-spark-1.2-contributor | Meta Model API direct, Contributor tier | USD / 1M tokens, plus training permission | $0.10 | $0.002 | $0.20 | Not a promotion. Meta’s own framing: “Tier is a model attribute: it sets the price you pay and whether your data may be used to train future Meta models.” | None | Same checkpoint as Standard. Throttled to 100 requests per minute against Standard’s 3,000. See section 08. |
| Mistral Medium 3.5 | Mistral direct API | USD / 1M tokens | $1.50 | not publishedOnly a cross-cutting FAQ statement that cached input reduces input cost “by up to 90%.” | $7.50 | None published | Batch −50%, stated in the FAQ only | Priority Tier at 1.75× list, with a 99.5% uptime SLA. Standard and Batch carry no SLA. The vendor’s self-described frontier-class model costs 3× Large 3 on input. |
| Mistral Small 4 | Mistral direct API | USD / 1M tokens | $0.15 | not published | $0.60 | None published | Batch −50% | Priority Tier 1.75×. |
| Mistral Large 3 | Mistral direct API | USD / 1M tokens | $0.50 | not published | $1.50 | None published | Batch −50% | On AWS Bedrock in US regions this is $0.50 / $1.50 exactly — the reseller matches the vendor. |
| Z.ai GLM 5.2, hosted by Mistral | Mistral direct API, third-party model | USD / 1M tokens | $1.40 | $0.14 | $4.40 | None published | Batch −50% | Mistral publishes a max-output figure on this third-party card and omits the field entirely from all three of its own model cards. |
| Amazon Nova Pro and Nova Premier | AWS Bedrock | USD / 1M tokens | not machine-readableThe Nova rate tables are region-selector widgets that did not render. The section headings did, which confirms Nova is priced on the same Global versus Geo axis as everything else on Bedrock. | None published | Flex 50% off Standard. | Priority 1.75×. Nova Premier is marked Legacy with an end of life of September 14, 2026. | ||
| Cohere, current generative line | Cohere direct | USD per instance-hour, not per token | not publishedNo current frontier generative per-token card exists on Cohere’s pricing page. What it does publish is Model Vault per-instance pricing: $4.00 to $10.00 per hour depending on model and performance tier, or $2,500 to $6,500 per month. | None published | — | On Bedrock, Cohere appears only as Rerank 3.5 at $2.00 per 1,000 queries. One vendor, three billing units across two surfaces. | ||
Two structural observations the table makes that a per-vendor page cannot. First, newer is not reliably cheaper or dearer within a family. Gemini 3.7 and 3.6 Flash sit at a promotional $0.75 / $3.75 while the older 3.5 Flash is $1.50 / $9.00 — double the input and 2.4× the output. Claude Sonnet 5 undercuts both Sonnet 4.6 and 4.5 by a third on input. Grok 4.6 costs more than Grok 4.3 and loses batch support. Whatever rule you were using to guess a price from a version number, it does not hold.
Second, the discount tiers are not a shared vocabulary. Batch is 50% off at six vendors, 60% of standard at Moonshot, 20% off at xAI on exactly one model, and does not exist at all at DeepSeek, Z.ai, MiniMax, Meta or Tencent. Google runs four tiers where flex and batch are both exactly 50% off but are different products with different latency guarantees, and priority is a premium tier at 1.8× on every published row. Reading “batch” as a fixed discount across vendors will misprice a budget by a factor of two.
04 — PromotionsSeven vendors discounting, and three published end dates.
This is the table that makes the index falsifiable. Every promotional rate in the archive above appears here with whatever the vendor published about when it stops — including, in five rows, the fact that the vendor published no end date at all. A price index that folded these into a single rate column would be unfalsifiable for most of the vendors in it, because there would be no way to tell which cells were about to move.
| Vendor and model | Surface | List rate | Promotional rate | Published end date | What the vendor says follows |
|---|---|---|---|---|---|
| Google — Gemini 3.7 Flash and 3.6 Flash | Gemini Developer API and Vertex AI, global | $1.50 / $7.50 | $0.75 / $3.75 | December 31, 2026 | “Starting January 1, 2027, standard pricing of $1.5 / $7.5 per 1M tokens input / output will apply.” Both halves published, on two independent Google pages. |
| Google — 3.7 and 3.6 Flash cache read and storage | Gemini Developer API | $0.15 read · $1.00 per 1M tokens per hour storage | $0.075 read · $0.50 storage | December 31, 2026 | The post-promotional figures are published for both. |
| Google — 3.7 and 3.6 Flash Provisioned Throughput | Vertex AI only | No credit | 50% monthly billing credit on net eligible PT spend | December 31, 2026, effective August 13, 2026 | Credits issue on the 7th of the following month and expire 30 days after issuance — a discount with an expiry on the discount itself. |
| Z.ai — GLM-5.3-Flash | Z.ai direct API | $0.15 / $0.03 / $0.50 | $0.075 / $0.015 / $0.25 | 24:00 on September 9, 2026 (UTC+8, Singapore time) | The list prices are rendered as strikethrough beside the promo prices, so the post-promotional rate is on the same page. |
OpenAI — gpt-5.6-sol | OpenAI API direct, Standard tier | not published | $4.00 / $20.00 | A floor, not a date: “available at least through November 21, 2026” | No figure. The vendor describes the change as “a 20% reduction in input pricing and a 33% reduction in output pricing,” which back-solves to $5 / $30 — GPT-5.5’s rate. That is arithmetic on the vendor’s own words, not a published price, and it appears in no rate cell in this index. |
Alibaba — qwen3.7-max | Model Studio, Singapore | $2.50 / $7.50 | $1.25 / $3.75 | not publishedLabelled only “Limited-time 50% off.” | Nothing. The pinned snapshots -2026-06-08 and -2026-05-20 sit at full list, so the same model has two prices depending on which ID you call. |
Alibaba — qwen3.7-plus | Model Studio, Singapore | $0.40 / $1.60 | $0.32 / $1.28 | not published“Limited-time 20% off.” | Same alias-versus-snapshot split as qwen3.7-max. |
| MiniMax — MiniMax-M3, all four tier and service combinations | MiniMax direct API | $0.60 / $2.40 and up | $0.30 / $1.20 and up | not publishedThe badge reads “Permanent 50% off.” | A discount the vendor calls permanent while still rendering a struck-through list price. It is neither a list rate nor an expiring promotion, and it needs its own category. |
| Z.ai — cached-input storage, 16 SKUs | Z.ai direct API | Storage billed | Free | not published“Limited-time Free.” | A promotion on a third billing axis — storage, not per-token — which is easy to miss entirely. |
| Tencent — Hy4 preview | WorkBuddy and CodeBuddy | $0.834 / $2.501 | Free | “for two weeks” from the August 28, 2026 launch — no calendar date published | A product-surface free trial, not an API discount. Stated as the vendor states it rather than computed into a date. |
| Tencent — Hy3 | WorkBuddy and CodeBuddy | $0.132 / $0.528 | Free | September 30 — year not stated, 2026 from context | “Free access to Hy3 on both platforms has also been extended until September 30.” |
| Moonshot — file extraction and file storage | Moonshot direct API | not published | Free | not published“temporarily free.” | No underlying rate is published, so there is no way to model what happens when it stops. |
| Anthropic — Claude Sonnet 5 | Claude API direct | $2.00 / $10.00 | None. The promotion became the list price on August 10, 2026. | Not applicable | Included in this table deliberately, because it is the row a reader is most likely to arrive carrying a wrong number for. The scheduled $3 / $15 increase was cancelled and must not be cited as a future rate. |
Google publishes the post-promotional number as a hard figure and a hard date, on two independent pages, for the same commercial event that OpenAI documents with a soft floor and no rate at all. Same kind of promotion, opposite disclosure quality — and neither is more binding than the other. Anthropic published a future price with a date too, and then cancelled it. So Google’s January 1, 2027 figure is quotable as a plan, and citing it as what Gemini will cost in 2027 repeats exactly the mistake this page opens with.
05 — SurfacesThe same model, the same rate, a different bill.
The most useful thing this index does is refuse to give one number per model. A price is a rate, a unit, an endpoint and a tier — and three of those four change when you move surfaces without the rate moving at all.
The cleanest illustration is first-party and unambiguous. Anthropic bills through AWS Marketplace and Microsoft Foundry in Claude Consumption Units. Token usage is rated in USD at exactly the standard per-model rates in the index above, any negotiated discount is applied, and the result is converted at $0.01 per CCU — 100 CCU represents $1.00 — and reported hourly. Your AWS or Azure bill shows a single CCU line item. The rate is identical. The unit on the invoice is not. Anyone reconciling a cloud bill against a per-million-token model will find nothing that looks like a token.
The second illustration is a capability rather than a unit. Anthropic’s fast mode — $10 / $50 on Opus 5 and Opus 4.8, exactly 2× standard — is available on the Claude API first-party only. It is not available on Claude Platform on AWS or on partner-operated clouds. Same model, same headline rate, and one surface can buy a latency upgrade the others cannot.
| Model | xAI direct | AWS Bedrock | Google Vertex AI | Microsoft Foundry |
|---|---|---|---|---|
| Grok 4.6 | $2.00 / $6.00 | Global cross-region $2.00 / $6.00; in-region and geo cross-region $2.20 / $6.60, a +10% premium | $2.00 / $6.00, and $4.00 / $12.00 above 200K — byte-identical to the vendor, long-context rule included | Not offered |
| Grok 4.3 | $1.25 / $2.50 | In-region only, $1.25 / $2.50. GovCloud US-West $1.50 / $3.00, a +20% premium | $1.25 / $2.50, and $2.50 / $5.00 above 200K | Not offered |
| Grok 4.2 | Not listed | Not offered | Not offered | Global $1.25 / $2.50 · Data Zone $1.375 / $2.75, a +10% premium. Foundry’s newest Grok is four generations behind xAI’s. |
| Grok 4.1 Fast, reasoning and non-reasoning | Not listed — no longer on xAI’s model list | Not offered | $0.20 / $0.50, cache hit $0.05 — still sold | — |
| Grok-4 | Not listed | Not offered | Not offered | Global $3.00 / $15.00 · US Gov Global $3.75 / $18.75. On the reseller you can pay 2.5× on output for a model the vendor has dropped. |
| Batch support | 20% discount on Grok 4.3; “Not supported” on Grok 4.6 | Flex 0.5×, Priority 1.75× | Batch and Flex tiers published | Global Batch and Data Zone Batch at 50% off |
The regional dimension deserves its own line, because of all the surface premiums it is the one presented with no price vocabulary attached at all. AWS Bedrock’s regional price list is a premium AWS never labels as one. The only word on those tables is a region name. Measured against the US East and US West rows on AWS’s own table, Gemma 3 27B input runs $0.23 in US regions, $0.27 in Mumbai, Ireland and Milan — a 17.4% premium — and $0.36 in London, which is 56.5% above the US row. Six model families on that list land between +50% and +57% in London. Sydney is the opposite case at around +3%.
London, on AWS Bedrock
Gemma 3 27B input: $0.23 in US East and US West, $0.27 in Mumbai, Ireland and Milan, $0.36 in London. Six model families on the same AWS price list land between +50% and +57% in London. AWS's only label on those tables is a region name.
Five surfaces, one multiplier
OpenAI charges a 10% uplift on regional data-residency endpoints for models released on or after March 5, 2026. Vertex's non-global endpoint is exactly 1.10x global. Anthropic's inference_geo of us applies a 1.1x multiplier. Bedrock's in-region and geo cross-region options are +10% on global. Azure's Data Zone is +10% on Global Standard.
OpenRouter's per-token markup
The vendor's own words: it passes through the pricing of the underlying providers and there is no markup on inference pricing. The money is taken at the credit-purchase step instead — a 5.5% platform fee on pay-as-you-go, plus 5% on bring-your-own-key above a $25,000 per month list-price allowance.
Priority, at AWS and at Mistral
AWS Bedrock states priority tier pricing is at a 75% premium to standard and flex at a 50% discount. Mistral's Priority Tier is 1.75x standard list pricing, described as a 75% premium on input, output and cached tokens, with a 99.5% uptime SLA that the standard and batch tiers do not carry.
OpenRouter is worth stating precisely because the common assumption about it is wrong in a specific way. It takes no per-token markup — that is the vendor’s own claim, made twice on its own pages, and the rate rows in the index above bear it out: z-ai/glm-5.3, moonshotai/kimi-k2.6, qwen/qwen3.8-max and deepseek/deepseek-v4-pro-0813 all match their vendors to the cent. The fee is at the credit-purchase step, at 5.5% on pay-as-you-go, plus 5% on bring-your-own-key above a $25,000 per month list-price allowance. That allowance is measured in list-price inference cost, not in requests — a detail with real consequences for anyone modelling it as a request quota.
What OpenRouter does not guarantee is which provider serves a slug. With 80-plus providers and 500-plus models, the passthrough promise is per-provider, not per-model. Which is the setup for the next section.
06 — The sharpest specimenThree slugs where batch costs more than standard.
Everywhere else in this index, a batch tier is a discount. It is 50% off at six vendors, 60% of standard at Moonshot, 20% off at xAI on one model. On OpenRouter, three :batch slugs in the current catalogue cost more than the standard slug for the same model. Not a rounding difference — two of them are exactly double.
The mechanism is not a markup, and the passthrough claim in the last section survives intact. Each of the three is passing through a different real vendor rate than the standard slug is passing through. The suffix names an API mode; it does not name a price direction. That distinction is invisible until you check both slugs against the vendor’s own page, which is precisely the work an index is for.
| Slug | Standard slug | :batch slug | Multiple | What the batch price actually is |
|---|---|---|---|---|
z-ai/glm-5.3-flash | $0.075 / $0.25 | $0.15 / $0.50 | 2.00× on both | Exactly Z.ai’s post-promotional list price. The standard slug is passing through the 50%-off promotion that ends September 9; the batch slug is passing through the list rate. Z.ai publishes no batch API at all. |
deepseek/deepseek-v4-pro-0813 | $0.66 / $1.98 | $1.32 / $3.96 | 2.00× on both | Exactly DeepSeek’s peak rate against the standard slug’s off-peak rate. DeepSeek publishes no batch API either. The correct statement is that OpenRouter tracks DeepSeek’s clock, not that it displays a fixed column. |
qwen/qwen3.8-2.4t-a95b | $2.00 / $6.00 | $2.50 / $6.25 | 1.25× in · 1.042× out | A 25% premium on input and about 4% on output, against a standard slug that matches Alibaba’s hosted rate exactly. No vendor batch rate is flagged on this SKU. |
qwen/qwen3.5-9b | $0.10 / $0.15 | $0.17 / $0.25 | 1.70× in · 1.667× out | A fourth instance of the same shape on a smaller model. |
moonshotai/kimi-k3 | $3.00 / $15.00 | $3.00 / $15.00 | 1.00× — no discount | Not an inversion but the same lesson. Moonshot’s own batch tier is 60% of standard and excludes K3 entirely. A :batch slug exists for a model the vendor does not offer batch pricing for. |
minimax/minimax-m3 | $0.30 / $1.20 | $0.30 / $1.20 | 1.00× — no discount | Same rate, but the batch slug halves the advertised context to 524,288. The variant differs in something other than price. |
None of these six rows is a reseller overcharging. Every one is a reseller faithfully passing through a rate — and the rate it passes through is not the one the slug’s name implies. Two vendors here publish no batch API whatsoever, and both have :batch slugs on a reseller anyway.
The operational consequence is narrow and expensive: a cost model that routes non-urgent work to :batch slugs by string substitution, on the reasonable assumption that batch means cheaper, will double the bill on two of these models and add nothing but latency tolerance on two more.
07 — Context bandsOne number per model is materially incomplete.
Three of the vendors in this index charge a different rate for the same model depending on how long the prompt is, and in every case the higher rate applies to the whole request, not just the tokens above the threshold. A single input rate per model does not describe what any of them will charge you.
The thresholds differ, the multipliers differ, and one major vendor has no cliff at all on its most-used line. This is the reason the index carries a premium-tiers column rather than folding everything into one input figure — and the mechanics of these thresholds, including how to structure prompts around them, are the subject of our long-context pricing thresholds guide.
| Vendor and models | Threshold | Input multiplier | Output multiplier | Worked example |
|---|---|---|---|---|
| OpenAI — the GPT-5.6 family and later | Prompts above 272,000 input tokens | 2× | 1.5× | gpt-5.6-sol goes from $4.00 / $20.00 to $8.00 / $30.00 for the full request. Cache writes bill at 1.25× the uncached input rate. |
| Google — Gemini 3.1 Pro Preview and 2.5 Pro | Input at or above 200K tokens | 2× | 1.5× | Gemini 3.1 Pro Preview goes from $2.00 / $12.00 to $4.00 / $18.00. Same shape as OpenAI’s, at a lower threshold. |
| Google — every Flash and Flash-Lite model | None. Flat to the full 1,048,576-token input limit | 1× | 1× | The line most people actually use has no cliff, which is why “Gemini charges more for long context” is only true of the Pro models. |
| xAI — every Grok model | Prompts reaching 200k tokens | 2× | 2× | Grok 4.6 goes from $2.00 / $6.00 to $4.00 / $12.00. xAI is the only vendor here that doubles output as well as input. Vertex reproduces the rule; both AWS Bedrock Grok cards omit it and publish a single flat rate per inference option. |
| Anthropic — Claude 4.6 and later | None. The full 1M-token window at standard pricing | 1× | 1× | The vendor’s own gloss: a 900k-token request bills at the same per-token rate as a 9k-token request. Caching and batch discounts apply at standard rates across the full window. |
| Meta — Muse Spark | None, stated explicitly | 1× | 1× | “There is no long-context premium: you pay the same rate whether your context window is mostly empty or almost full.” The only vendor here that names the absence. |
Alibaba — qwen3.7-flash and qwen3.7-plus | Tiered by request total: 32K and 256K boundaries | Varies by tier | Varies by tier | A different mechanism with the same trap: qwen3.7-flash runs $0.030 / $0.130 up to 32K, $0.100 / $0.400 to 256K and $0.200 / $0.800 to 1M — and every token in a request bills at the tier the request total falls into. |
| MiniMax — MiniMax-M3 | Input above 512K | 2× | 2× | Effective rates go from $0.30 / $1.20 to $0.60 / $2.40 on the Standard tier, and the same doubling applies inside the Priority tier. |
One further trap sits underneath this table and is worth naming because it will otherwise silently corrupt a cost model built from it. OpenAI’s headline context window includes output. Its 1,050,000 figure is 922,000 maximum input plus 128,000 maximum output, exactly. Google’s 1,048,576 is an input limit with 65,536 of output on top. The two vendors’ “1M” numbers are not the same quantity and must not be compared directly. The ceilings themselves, by model and by surface, are censused separately in our context window and output limit census.
08 — A different currencyOne vendor charges in data, not dollars.
Every other row in this index is a price in money. Meta publishes one that is not. The Muse Spark 1.2 checkpoint is sold at two prices under two model IDs, and the cheaper one is bought with permission rather than cash.
Meta’s own framing is unusually direct about it: tier “sets the price you pay and whether your data may be used to train future Meta models,” and the Contributor ID carries “heavily discounted pricing in exchange for permission to use your prompts and completions to train future Meta models.” Measured against Meta’s own Standard rate for the same checkpoint, that discount is 92.0% on input, 95.3% on output and 98.7% on cached input.
| Billing category | Standard tier | Contributor tier | Discount | What is exchanged |
|---|---|---|---|---|
| Input | $1.25 | $0.10 | −92.0% | Permission to use your prompts and completions to train future Meta models. |
| Output | $4.25 | $0.20 | −95.3% | Same terms. |
| Cached input | $0.15 | $0.002 | −98.7% | Same terms. |
| Requests per minute | 3,000 | 100 | −96.7% | The discount tier is also throttled. Token-per-minute limits are closer: 4,000,000 against 3,000,000. Limits are per team, not per API key. |
No other vendor in this index prices on this axis, and it is worth being precise about why that is not the same as saying no other vendor trains on your data. Google runs the closest analogue and runs it the other way round: on the Gemini API free tier, content is used to improve Google’s products; on the paid tier it is not. Same model, same rate card, different data terms decided by the tier rather than by a model ID.
The difference matters for procurement. Google’s split is a consequence of not paying. Meta’s is a price: a published, quantified discount attached to a specific model ID, which means it can be compared against the market rate, put in a budget, and — the part that makes it a governance question rather than a finance one — chosen by a developer with a one-word change to a model string.
09 — ImplicationsHow to use this without getting burned.
The most common way to misuse a price index is to take one cell out of it. Every number above is true of a specific model, on a specific surface, in a specific tier, on August 30, 2026 — and at least one of those four qualifiers is doing real work in most rows.
Copy the end date with the rate, or copy neither
Two of the three headline rate changes this month were promotional, and both were reported without their footnotes. A rate in a spreadsheet with no expiry column will read as permanent to whoever inherits the spreadsheet. Model the post-promotional rate wherever the vendor publishes one, and model the uncertainty where it does not.
A price is not a number until you say where you bought it
The same Claude rate bills in Claude Consumption Units on two clouds and in dollars on a third. The same Grok model costs 10% more in-region on Bedrock and is not offered at all on Foundry. Anthropic's fast mode does not exist outside the first-party API. Write the surface into the line item, not just the model name.
Do not append :batch and assume a discount
Three OpenRouter slugs cost more with the batch suffix, two of them exactly double, because the suffix names an API mode rather than a price direction. Two of those vendors publish no batch API at all. Check both slugs against the vendor's own page before routing anything by name pattern.
Long prompts reprice the whole request
Above 272,000 input tokens OpenAI charges 2x input and 1.5x output for every token in the request, not just the excess. Google's Pro models do the same above 200K and its Flash models do not do it at all. If your median prompt sits near a threshold, your effective rate is neither of the two published numbers.
The wider point is about where the market has moved. Two years ago a price index could be a list of models and two numbers each. The direction of travel since is unmistakable: every large vendor has added tiers, thresholds, endpoints and units, and the differences between them are now the thing that decides a bill rather than a footnote to it. That is unlikely to reverse, because each of these dimensions is a real product with a real cost behind it — data residency is genuinely more expensive to serve, sheddable off-peak capacity genuinely is cheaper. What should be expected instead is more of it: more premium tiers, more units that are not tokens, and more prices that are a function of when and where you called rather than what you called.
The practical response is to stop treating price as an attribute of a model and start treating it as an attribute of a call. Our ledger of AI coding plan limit changes tracks the same volatility on the subscription side, where the ceiling moves rather than the rate. If you want senior help mapping what your workloads actually cost across these surfaces, and what a mid-quarter promotional expiry does to a committed budget, that is the kind of question our AI transformation engagements open with.
10 — ConclusionThe footnote is the price.
Two columns, three gap labels, and a URL that gets updated instead of replaced.
The design decision that makes this page worth citing is that it never merges a promotional rate into a list column, and never drops a row it could not verify. A reader who wants to argue that frontier inference is getting cheaper and a reader who wants to argue that the discounts are about to expire can both quote it, cell by cell, without either of them inheriting our judgement.
The two cases at the top are the reason the shape matters. Anthropic cancelled a scheduled increase five days after we published a post built on it, and OpenAI’s most-quoted rate is a promotion that most of the market — including our own internal anchors — recorded as a list price. Neither is anyone’s dishonesty. Both are what happens when a number travels without the sentence printed next to it.
This URL supersedes our four dated snapshots from March, April and August 2026, which stand as historical records of what the market published on their own dates. It is maintained by the Digital Applied Team, reviewed monthly and out of cycle when a frontier rate moves, and re-dated through its modified time rather than replaced. Cite the as-of date with any cell you take from it.