AI DevelopmentPricing Tracker37 min readPublished August 30, 2026

14 vendors · list and promotional rates in separate columns · data as of Aug 30, 2026

What Every AI Model Costs Per Million Tokens

A maintained index of what every current model API charges per million tokens, traced to each vendor’s own pricing page, as of August 30, 2026. Promotional rates sit in their own column with their published end dates. Rows we could not verify are marked and kept in.

DA
Digital Applied Team
Senior strategists · Published Aug 30, 2026
PublishedAug 30, 2026
Data as ofAug 30, 2026
Read time37 min
Vendors indexed
14
direct APIs, plus four reseller surfaces
Data as of
Aug 30
2026 · refreshed monthly at this URL
Snapshots superseded
4
March, April and August 2026
Scheduled increase cancelled
1
Sonnet 5 stays $2 / $10

This is a maintained AI model API price index: what every current model costs per million tokens on input, cached input and output, plus batch, flex, off-peak, priority and fast-mode rates wherever the vendor publishes them, organised by the surface you buy on. Data as of August 30, 2026. Every figure traces to the vendor’s own pricing page. Promotional rates never share a column with list rates.

The reason to build it this way is not tidiness. In the last three weeks two separate frontier rates moved for reasons that were printed in a footnote and dropped by nearly everyone who quoted the number — including, in one case, us. A reader looking at two confident secondary sources had no way to tell which one to believe, because neither carried the footnote that settled it.

So the index leads with those two cases, then gives you the tables: the full rate card, the promotions with their end dates, the surface layer where the same model bills differently depending on where you bought it, the batch slugs that cost more than standard, the long-context cliffs that make any one-number-per-model table incomplete, and the one vendor charging in data rather than money.

Key takeaways
  1. 01
    Claude Sonnet 5 is $2 / $10 and that is now the standard price.Anthropic announced the $2/$10 rate at launch as introductory pricing through August 31, 2026, then cancelled the scheduled September 1 increase to $3/$15 on August 10, 2026. Both the pricing docs and the Sonnet 5 launch post carry the reversal in writing. Any table still showing a September 1 cliff, or $3/$15 as a future Sonnet 5 rate, is wrong.
  2. 02
    GPT-5.6 Sol’s $4 / $20 is promotional, not a list price.OpenAI’s own model page states the promotional pricing is available at least through November 21, 2026. OpenAI publishes no figure for what follows. The vendor describes the cut as a 20% reduction in input and a 33% reduction in output, which back-solves to $5 / $30 — that is vendor arithmetic, not a published future price, and this index labels it as such.
  3. 03
    Promotional and list rates are in separate columns, always.Seven vendors here run a discount of some kind and only three publish a hard end date: Z.ai’s GLM-5.3-Flash promo ends at 24:00 on September 9, 2026 (UTC+8), Google’s Gemini 3.7 and 3.6 Flash promo ends December 31, 2026, and Tencent’s free Hy3 access is stated as running to September 30, year not printed on the page. Everything else is limited-time, temporarily, or — at MiniMax — permanently discounted.
  4. 04
    The same model bills differently depending on the surface.Anthropic bills through AWS Marketplace and Microsoft Foundry in Claude Consumption Units at $0.01 per CCU: same rate, different unit on the invoice. AWS Bedrock’s regional price list moves Gemma 3 27B input from $0.23 in US regions to $0.36 in London, a 56.5% premium AWS labels only as a region. Claude Fast mode is first-party only — not on Claude Platform on AWS or partner clouds.
  5. 05
    Rows we could not verify are marked and kept, never dropped.Three labels, three distinct styles: the vendor does not publish it, the page could not be located, or the page was located but could not be read. Cohere’s missing generative card is the first kind. Azure’s OpenAI per-token rates are the second. Amazon Nova’s region-selector tables are the third. They are not the same kind of gap.

01Why this existsTwo rates, two footnotes, and one of them was ours.

A price index earns the right to be cited when it can settle a disagreement that the reader cannot settle themselves. Two cases in August 2026 show exactly what that looks like, and they run in opposite directions from the same underlying mistake.

The first is Anthropic’s Claude Sonnet 5. When Sonnet 5 launched, $2 per million input tokens and $10 per million output tokens were published as introductory pricing through August 31, 2026, with a scheduled increase to $3 / $15 on September 1. Two secondary sources contradicted each other this month: one reported the increase had been cancelled, one reported it was landing. Only the vendor could settle it, and the vendor did. Anthropic’s pricing documentation now carries a callout stating that the $2 / $10 pricing “is now the standard price” and that “the previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur.” The Sonnet 5 launch post carries a dated edit note to the same effect, marked August 10, 2026.

Our own August 5 pricing tracker reported the expiry footnote accurately, quoting Anthropic’s own wording, and built a section on the September 1 step-up. Five days later the vendor withdrew the increase. The post was not careless; it was a correct snapshot of a rate whose footnote the vendor then deleted. That is the strongest argument for a maintained index we can make, and it is about our own work, which is why it belongs at the top of this page rather than buried in a correction note.

The primary, verbatim

From Anthropic’s pricing documentation, fetched August 30, 2026: “The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur.”

The same vendor’s launch post carries a dated edit: “Edit August 10, 2026: Sonnet 5’s introductory pricing of $2 per million input tokens and $10 per million output tokens is now permanent.” Anthropic also annotated its own cost-performance charts as stale, because they were plotted at $3 / $15. A vendor telling you which of its own published numbers is out of date is rarer than it should be.

The second is OpenAI’s GPT-5.6 Sol, and it is the same error running backwards. Sol’s current $4 / $20 was widely reported — and recorded in our own internal anchors — as a list price. It is not. OpenAI’s model page states plainly that “GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.” A promotional decrease was mistaken for a permanent one, where Sonnet 5 was a scheduled increase mistaken for an inevitable one. Both mistakes come from taking a rate without its footnote.

The two vendors also differ sharply in what they disclose about the other side of the promotion. Google publishes both halves for its Gemini 3.7 and 3.6 Flash promotion — the promotional rate, the end date, and the exact figure that applies afterwards. OpenAI publishes a soft floor and no post-promotional rate at all. What it does publish is a characterisation: the cut is “a 20% reduction in input pricing and a 33% reduction in output pricing.” Run that backwards and you land on $5 / $30, which is GPT-5.5’s rate. This index will report that only as vendor arithmetic, and never in a price cell, because OpenAI has not published it as a price.

Case one · Anthropic
A scheduled increase, cancelled
$2 / $10 · standard

Sonnet 5 launched with introductory pricing through August 31, 2026 and a scheduled step-up to $3 / $15. On August 10 Anthropic cancelled the increase and made $2 / $10 the standard price. Two secondary sources disagreed about this all month; the vendor's own docs and launch post settle it in two independent places.

Read the footnote
Case two · OpenAI
A promotional decrease, mislabelled
$4 / $20 · promotional

GPT-5.6 Sol's rate is promotional and carries a vendor-published floor of at least through November 21, 2026. No post-promotional figure is published. The $5 / $30 that the vendor's own percentages back-solve to is arithmetic, not a price, and appears nowhere in this index's rate columns.

Same error, opposite sign
The rule that follows
A rate without its footnote is not a rate
list ≠ promotional

Every promotional row in this index carries its published end date, or an explicit note that the vendor published none. Merging a promotional rate into a list column produces a table that cannot be falsified, which for four of the seven discounting vendors here is exactly what would happen.

The design decision

The habit of reading a rate together with its footnote is the same literacy our model catalogue literacy guide argues for across prices, dates and surfaces. This page is that argument turned into a dataset.

02MethodWhat was collected, what is missing, and who keeps it current.

The gaps do as much work here as the prices. A price index that silently drops the rows it could not verify looks more complete than it is, and quietly teaches the reader that everything absent from it does not exist. Every unverified row below is kept in and labelled, and the three labels mean three different things.

Methodology

What was collected. For every current model with a publicly published API price, the input rate, the cached-input rate, the output rate, and every discount or premium tier the vendor publishes — batch, flex, off-peak, priority, fast mode, long-context bands and regional endpoints. Rows are organised by surface: the vendor’s own direct API, the cloud resellers that meter the same model, and the coding plans that include it without publishing a token rate.

Sources. Every price traces to the vendor’s own pricing page, model page or documentation. Where a figure could only be obtained from a second vendor’s page — Google’s Vertex tables carry cache-read rates that the Gemini Developer API page does not render — the row says which page it came from. No figure in the rate columns comes from an aggregator or a price-tracker site.

Dates. Pages were fetched on August 30, 2026. The as-of date is August 30, 2026. Nothing dated after that appears as an event. Scheduled future prices — Google’s January 1, 2027 reversion is the clearest — are recorded as plans, not outcomes. The Sonnet 5 case above is the reason that distinction is enforced rather than assumed.

Currency. Every row states its currency and its unit. Where a vendor publishes a native non-USD price, both are reported as published and neither is converted. DeepSeek’s CNY table is not a conversion of its USD table — the implied cross-rate is about 6.82 CNY per USD on the main rows and about 7.14 on the V4-Flash cache-hit row, because each locale is rounded to its own tidy numbers. Tencent publishes $0.834 / $2.501 / $0.042 on its English release and ¥6 / ¥18 / “as low as ¥0.3” on its Chinese one; both are vendor-published regional list prices and neither is our arithmetic.

The three gap labels. Rows that could not be verified are kept in the tables with one of three marks, rendered three different ways so the distinction survives a screenshot: not published means the vendor does not publish this figure; page not located means the page that would carry it was not found; and not machine-readable means the page was located but its tables did not render to any fetch method used. The first is a fact about the vendor. The second and third are facts about this collection pass.

What is excluded. Non-token meters, except where a vendor prices a frontier product only that way — Cohere’s instance-hour Model Vault rates and MiniMax’s per-second video rates are in because their absence would misrepresent the vendor. Fine tuning, image, audio, embedding and tool-call pricing are out of the main table; they belong to a different index. Coding-plan subscriptions appear only as a surface with no published token rate, because that is the accurate description of them.

Refresh owner and cadence. Maintained by the Digital Applied Team, reviewed monthly and out of cycle when a frontier vendor moves a headline rate. The slug carries no date deliberately: each refresh updates the tables in place and re-dates the page through its modified time, so the URL stays stable and citable while the as-of date moves.

This page supersedes four dated snapshots

Our archive already holds four dated pricing snapshots. This URL is the one that gets updated from here on. Those four stand as historical records of what the market published on their own dates, and should be cited that way — not as current rates:

LLM API pricing index and cost tracker (March 26, 2026) · LLM API pricing index, Q2 2026 (April 12, 2026) · AI model API pricing tracker, Q2 2026 (April 23, 2026) · AI API pricing, August 2026: cuts, promos and traps (August 5, 2026).

Each of those four carries an update block pointing here, so the cross-link runs both ways. The August 5 post also carries a dated correction to its Sonnet 5 section, for the reason set out above.

Every rate column in this index resolves to one of the pages below. They are listed rather than footnoted so a reader can re-derive any cell, and so a later refresh has an explicit list to re-fetch.

Primary sources, all fetched August 30, 2026. Where a figure came from a second page belonging to the same vendor, the row in the index says so.
VendorWhat it suppliesPrimary page
AnthropicModel, batch, fast-mode and cache rates; CCU billing on AWS and Microsoft Foundry; data-residency multiplierplatform.claude.com pricing
OpenAIStandard, Batch, Flex and Fast mode tables; the Sol promotional footnote; the 272,000-token long-context ruleplatform.openai.com pricing
GoogleStandard, batch, flex and priority tabs; the promotional rates and their December 31, 2026 end dateai.google.dev pricing · Vertex AI pricing
DeepSeekPeak and off-peak tables in both USD and CNY; the peak-window definitionapi-docs.deepseek.com pricing
Z.aiGLM list and promotional rates with the September 9 end date; the coding-plan credit multipliersdocs.z.ai pricing
Moonshot AIPer-model cache-hit, cache-miss and output rates; the batch tier and its exclusionsplatform.kimi.ai pricing
MiniMaxM3 Standard and Priority tabs, the permanent-discount badge, and the per-second H3 video ratesplatform.minimax.io pricing
Alibaba CloudEvery Qwen SKU with its tiering, promotional labels, batch and cache rulesModel Studio billing
TencentThe Hy4 preview rates in both native currencies, on the vendor’s two language releasesHy4 preview release
xAIThe full Grok price table, the 200k long-context rule, and the per-model batch support statementsdocs.x.ai models
MetaStandard and Contributor tier rates, the tier definition, and the no-long-context-premium statementMeta Model API pricing
MistralPer-model cards and rates; the Priority Tier multiplier and its SLAdocs.mistral.ai priority tier
CohereModel Vault per-instance rates, and the absence of a current frontier per-token cardcohere.com pricing
Amazon Web ServicesThe Bedrock regional price list, the Priority and Flex multipliers, and the per-model cardsBedrock pricing
MicrosoftDeployment types, the Data Zone and Government premiums, and the long-context restatementFoundry deployment types
OpenRouterThe passthrough statement, the platform and BYOK fees, and the live slug rates including the batch inversionsopenrouter.ai pricing
Cite this
Digital Applied, “What Every AI Model Costs Per Million Tokens,” Digital Applied Blog, August 30, 2026, https://www.digitalapplied.com/blog/frontier-model-api-price-indexData as of August 30, 2026, which is distinct from the publication date and moves on every refresh. This is a maintained dataset at a stable, undated URL: refreshes update the tables in place and re-date the page through its modified time. Cite the as-of date alongside any cell, and carry the list-versus-promotional distinction with the number — they are separate claims and quoting one as the other is the failure this page exists to prevent.

03The IndexEvery current model, by surface.

Read the columns carefully. List holds the standard rate. Promotional holds a discounted rate together with its published end date, or the vendor’s own words where no date exists. They are never merged, and a blank promotional cell means the vendor publishes no promotion on that row — not that we did not look.

Discount tiers collects batch, flex and off-peak. Premium tiers collects fast mode, priority, peak rates and the long-context bands. Both are per-vendor vocabularies rather than a shared one: “batch” means a 50% discount at Anthropic, OpenAI, Google, Alibaba, Mistral and AWS; 60% of standard at Moonshot; a 20% discount at xAI on one model and nothing at all on its flagship; and, on three OpenRouter slugs, a price higher than standard. Section 06 takes that apart.

Frontier model API price index. Data as of August 30, 2026. Every figure is from the vendor’s own pricing page, model page or documentation. Promotional rates are in their own column with their published end dates and are never merged into the list columns. Cached input is the cache-read (hit) rate unless the cell says otherwise. Rows that could not be verified are kept in and carry one of three marks: not published, page not located, or not machine-readable.
ModelSurfaceCurrency and unitList inputList cached inputList outputPromotional rate and end dateDiscount tiersPremium tiers and bands
Anthropic — Claude API, first party
Claude Fable 5Claude API directUSD / 1M tokens$10.00$1.00$50.00None publishedBatch $5.00 / $25.00Cache write $12.50 for 5 minutes, $20.00 for 1 hour. No fast mode row.
Claude Mythos 5 (limited availability)Claude API directUSD / 1M tokens$10.00$1.00$50.00None publishedBatch $5.00 / $25.00On Bedrock, AWS states access “is gated and requires approval.” Capabilities are not documented on the pricing page and are not asserted here.
Claude Opus 5Claude API directUSD / 1M tokens$5.00$0.50$25.00None publishedBatch $2.50 / $12.50Fast mode $10.00 / $50.00 — exactly 2× standard, Claude API first party only. Cache write $6.25 / $10.00.
Claude Opus 4.8Claude API directUSD / 1M tokens$5.00$0.50$25.00None publishedBatch $2.50 / $12.50Fast mode $10.00 / $50.00, first party only. Applies across the full context window.
Claude Opus 4.7, 4.6 and 4.5Claude API directUSD / 1M tokens$5.00$0.50$25.00None publishedBatch $2.50 / $12.50No fast mode. On 4.7 a fast request returns an error; on 4.6 it runs at standard speed and bills at standard rates.
Claude Opus 4.1 (retired, except on Bedrock and Google Cloud)Claude API directUSD / 1M tokens$15.00$1.50$75.00None publishedBatch $7.50 / $37.50Retired on the first-party API; still sold on two reseller surfaces at these rates.
Claude Opus 4 (retired, except on Google Cloud)Claude API directUSD / 1M tokens$15.00$1.50$75.00None publishedBatch $7.50 / $37.50One surviving surface.
Claude Sonnet 5Claude API directUSD / 1M tokens$2.00$0.20$10.00None. The launch-era introductory pricing became the standard price on August 10, 2026, and the scheduled increase was cancelled. No expiry.Batch $1.00 / $5.00Cache write $2.50 / $4.00. No fast mode row.
Claude Sonnet 4.6 and 4.5Claude API directUSD / 1M tokens$3.00$0.30$15.00None publishedBatch $1.50 / $7.50Both older Sonnets cost 50% more on input than Sonnet 5 and 50% more on output.
Claude Sonnet 4 (retired, except on Bedrock and Google Cloud)Claude API directUSD / 1M tokens$3.00$0.30$15.00None publishedBatch $1.50 / $7.50
Claude Haiku 4.5Claude API directUSD / 1M tokens$1.00$0.10$5.00None publishedBatch $0.50 / $2.50Cache write $1.25 / $2.00.
Claude Haiku 3.5 (retired, except on Bedrock and Google Cloud)Claude API directUSD / 1M tokens$0.80$0.08$4.00None publishedBatch $0.40 / $2.00The cheapest Claude row anywhere, and it is retired on the first-party API.
Every Claude model aboveClaude Platform on AWS · Claude in Microsoft FoundryClaude Consumption Units, $0.01 per CCU (100 CCU = $1.00)Token usage is rated in USD at the standard per-model rates above, discounts applied, then converted to CCUs and reported hourly. The AWS or Azure bill shows one CCU line item.None publishedBatch discount applies at the token-rating step, before conversion.Fast mode is not available here. Data residency: inference_geo: "us" applies a 1.1× multiplier to every token category.
OpenAI — direct API, Standard tier, prompts at or below 272,000 input tokens
gpt-5.6-solOpenAI API directUSD / 1M tokensnot publishedOpenAI publishes no post-promotional figure for Sol.not publishednot published$4.00 in · $0.40 cached · $20.00 out“available at least through November 21, 2026.” The vendor calls the cut “a 20% reduction in input pricing and a 33% reduction in output pricing,” which back-solves to $5 / $30 — vendor arithmetic, not a published price.Batch $2.00 / $10.00 · Flex identical · cache write $2.50Fast mode $8.00 / $40.00. Above 272,000 input tokens: $8.00 in, $30.00 out, for the whole request.
gpt-5.6-terraOpenAI API directUSD / 1M tokens$2.00$0.20$12.00None — no promotional language on its model pageBatch and Flex $1.00 / $6.00 · cache write $1.25Fast mode $4.00 / $24.00. Above 272,000 in: $4.00 / $18.00.
gpt-5.6-lunaOpenAI API directUSD / 1M tokens$0.20$0.02$1.20None — no promotional language on its model pageBatch and Flex $0.10 / $0.60Fast mode $0.40 / $2.40. Above 272,000 in: $0.40 / $1.80.
gpt-5.6-cyber (Daybreak)OpenAI API direct only — absent from Azure’s model listUSD / 1M tokens$12.50$1.25$75.00None publishedNone. No batch row, no flex row, and v1/batch is listed as “Not supported.”No fast mode row. Cache write $15.625. The highest headline price in the index with no published way to reduce it.
gpt-5.5-cyber (Daybreak)OpenAI API direct onlyUSD / 1M tokens$12.50$1.25$75.00None publishedNone publishedRequires separate approval and provisioning.
gpt-5.4-cyber (Daybreak)OpenAI API direct onlyUSD / 1M tokensnot publishedThe row exists on OpenAI’s own pricing table with every cell empty. That is the vendor’s presentation, not a failed fetch.None publishedNone published
gpt-5.5OpenAI API directUSD / 1M tokens$5.00$0.50$30.00None publishedBatch and Flex $2.50 / $15.00Fast mode $12.50 / $75.00. Above 272,000 in: $10.00 / $45.00.
gpt-5.5-proOpenAI API directUSD / 1M tokens$30.00not published$180.00None publishedBatch and Flex $15.00 / $90.00No fast mode row — no Pro-tier SKU has one. Above 272,000 in: $60.00 / $270.00.
gpt-5.4OpenAI API directUSD / 1M tokens$2.50$0.25$15.00None publishedBatch and Flex $1.25 / $7.50Fast mode $5.00 / $30.00. Above 272,000 in: $5.00 / $22.50.
gpt-5.4-miniOpenAI API directUSD / 1M tokens$0.75$0.075$4.50None publishedBatch and Flex $0.375 / $2.25Fast mode $1.50 / $9.00.
gpt-5.4-nanoOpenAI API directUSD / 1M tokens$0.20$0.02$1.25None publishedBatch and Flex $0.10 / $0.625No fast mode row.
gpt-5.3-codexOpenAI API directUSD / 1M tokens$1.75$0.175$14.00None publishednot publishedFast mode $3.50 / $28.00, exactly 2×. The sole surviving Codex variant — five siblings shut down on July 23, 2026.
chat-latestOpenAI API direct — “latest Instant model used in ChatGPT”USD / 1M tokens$5.00$0.50$30.00None publishedNone publishedOpenAI publishes no mapping from the ChatGPT app’s “GPT Instant” and “GPT Reasoning” labels to any model ID.
GPT-5.6 Sol, Terra and LunaMicrosoft Foundry, version 2026-07-09USD / 1M tokenspage not locatedAzure’s model page defers rates to its own pricing page, which was not reached in this pass. Limits are verified; the money is not.None publishedGlobal Batch and Data Zone Batch are 50% off Global Standard.Data Zone deployments are +10% on Global Standard. Azure Government adds a further premium.
OpenAI models on Amazon BedrockAWS BedrockUSD / 1M tokenspage not locatedOpenAI’s own footnote is the only sourced statement: models on Bedrock “are billed through AWS and may differ from direct OpenAI pricing.”Flex and Batch 50% off Standard.Priority 1.75× Standard.
Google — Gemini Developer API, paid tier. Tier identity read from Google’s own tab markup.
Gemini 3.7 FlashGemini Developer API and Vertex AIUSD / 1M tokens$1.50$0.15$7.50$0.75 in · $0.075 cached · $3.75 outEnds December 31, 2026. Google publishes the post-promotional figure explicitly: $1.50 / $7.50 from January 1, 2027. Cache storage $0.50 per 1M tokens per hour promotional, $1.00 after.Batch $0.375 / $1.875 · Flex $0.375 / $1.875 (both at the promotional level)Priority $1.35 / $6.75 — 1.8× standard. GA and the newest model here; no long-context cliff on any Flash model.
Gemini 3.6 FlashGemini Developer API and Vertex AIUSD / 1M tokens$1.50$0.15$7.50$0.75 / $3.75, ends December 31, 2026Google applied the 3.7 Flash introductory rate to 3.6 Flash as well.Batch and Flex $0.375 / $1.875Priority $1.35 / $6.75.
Gemini 3.5 FlashGemini Developer API and Vertex AIUSD / 1M tokens$1.50$0.15$9.00None publishedBatch and Flex $0.75 / $4.50Priority $2.70 / $16.20. Note this older model’s output rate is 2.4× the newer 3.7 Flash promotional rate.
Gemini 3.5 Flash-LiteGemini Developer API and Vertex AIUSD / 1M tokens$0.30$0.03$2.50None publishedBatch and Flex $0.15 / $1.25Priority $0.54 / $4.50. Input is $0.30 flat across all modalities.
Gemini 3.1 Flash-LiteGemini Developer API and Vertex AIUSD / 1M tokens$0.25 text, image, video$0.50 audio$0.025 text, image, video · $0.05 audioFrom Vertex’s table; the Developer API page did not render this cell.$1.50None publishedBatch and Flex $0.125 ($0.25 audio) / $0.75Priority $0.45 ($0.90 audio) / $2.70. Its successor charges a flat $0.30, so 3.5 is dearer for text and cheaper for audio.
Gemini 3.1 Pro PreviewGemini Developer API and Vertex AIUSD / 1M tokens$2.00 at or below 200K in$4.00 above 200K$0.20 at or below 200K · $0.40 aboveFrom Vertex’s table.$12.00 at or below 200K in$18.00 above 200KNone publishedBatch and Flex $1.00 / $6.00, and $2.00 / $9.00 in the upper bandPriority $3.60 / $21.60, and $7.20 / $32.40 in the upper band. Newest Pro model, still carrying a preview endpoint.
Gemini 3 Flash PreviewGemini Developer API and Vertex AIUSD / 1M tokens$0.50 text, image, video · $1.00 audio$0.05 text, image, video · $0.10 audio (Vertex table)$3.00None publishedBatch and Flex $0.25 / $1.50Priority $0.90 / $5.40.
Gemini Omni FlashGemini Developer APIUSD / 1M tokens$1.50not published$9.00 text$17.50 videoNone publishednot publishedOnly a Standard column exists; the batch, flex and priority tabs carry no rows for this model.not published
Every Gemini model aboveVertex AI, non-global endpoint (data residency)USD / 1M tokensExactly +10% on every published row. Gemini 3.7 Flash: $0.75 becomes $0.825 in, $3.75 becomes $4.125 out, cached $0.075 becomes $0.0825.Promotional rate carries through, upliftedSame tier structure as global.Gemini 3.1 Pro Preview and Gemini 3 Flash Preview have no non-global row published.
CodeMender and AlphaEvolveVertex AI onlyUSD / 1M tokens, additiveAn additive second price column, not a multiplier. AlphaEvolve on Gemini 3.1 Pro Preview adds $4.00 in and $24.00 out on top of the model’s $2.00 / $12.00, for $6.00 / $36.00 total.Vendor tables disagree. Vertex’s CodeMender table prices 3.7 and 3.6 Flash at $1.50 / $7.50 — the post-promotional rate — while Vertex’s main table prices the same models at the promotional $0.75 / $3.75. Both are quoted as published; this index does not pick one.Provisioned Throughput on 3.6 and 3.7 Flash carries a 50% monthly billing credit, August 13 to December 31, 2026, on Vertex only.Neither agent exists on the Gemini Developer API.
DeepSeek — direct API. Peak is the base rate; off-peak is half of it.
deepseek-v4-proDeepSeek direct APIUSD / 1M tokens$1.32 peak$0.044 peak$3.96 peakNone — the clock is a tier, not a promotionOff-peak $0.66 in · $0.022 cached · $1.98 out. Peak is 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday.not publishedNo Batch API exists — there is no batch page in DeepSeek’s doc sidebar.
deepseek-v4-flash and deepseek-v4-flash-vision-expDeepSeek direct APIUSD / 1M tokens$0.44 peak$0.014 peak$1.32 peakNoneOff-peak $0.22 · $0.007 · $0.66. V4-Pro is exactly 3× these rates on input and output; the cache-hit cells round to $0.007 and $0.022, which is nearer 3.14×.Images on the vision model are converted to tokens by dimension and billed as input — no separate rate.
deepseek-v4-pro, native CNY tableDeepSeek direct API, Chinese localeCNY / 1M tokens¥9.0 peak¥0.30 peak¥27.0 peakNoneOff-peak ¥4.5 · ¥0.15 · ¥13.5Published natively in CNY, not converted from the USD table.
deepseek-v4-flash, native CNY tableDeepSeek direct API, Chinese localeCNY / 1M tokens¥3.0 peak¥0.10 peak¥9.0 peakNoneOff-peak ¥1.5 · ¥0.05 · ¥4.5The cache-hit row implies about 7.14 CNY per USD against the USD table, where the main rows imply about 6.82. Each locale is rounded to its own tidy numbers; do not convert between them.
Z.ai — GLM, direct API. All prices stated in USD by the vendor.
GLM-5.3Z.ai direct APIUSD / 1M tokens$1.40$0.26$4.40Cached-input storage is “Limited-time Free” with no published end date — a promotion on a third billing axis.not publishedNo Batch API on Z.ai’s pricing page.Weights shipped August 28, 2026 under a bespoke glm-5.3 licence, not MIT.
GLM-5.3-FlashZ.ai direct APIUSD / 1M tokens$0.15$0.03$0.50$0.075 in · $0.015 cached · $0.25 out50% off, and the vendor states it “ends at 24:00 on September 9, 2026 (UTC+8, Singapore time).” The timezone is part of the fact — that is 16:00 UTC on September 9.not publishedThe only promotional rate in this index with an end date stated to the hour and the timezone.
GLM-5.2 and GLM-5.1Z.ai direct APIUSD / 1M tokens$1.40$0.26$4.40None on the token rateNone publishedOn the GLM Coding Plan, requests for these two are silently rerouted to GLM-5.3.
GLM-4.7Z.ai direct APIUSD / 1M tokens$0.60$0.11$2.20None on the token rateNone publishedRerouted to GLM-5.3-Flash on the coding plan.
GLM-4.7-Flash and GLM-4.6V-FlashZ.ai direct APIUSD / 1M tokensFreeFreeFreeNot a promotion — listed as free, structurally
GLM-5.3 and GLM-5.3-FlashGLM Coding Plan — a subscription surfaceCredits, not tokensNo per-token rate. Credits are consumed at published multipliers: GLM-5.3 at 6.9 input, 1.7 cached input and 24 output; GLM-5.3-Flash at 2.3, 0.56 and 8. Divided by 10,000.Off-peak is 50% of the standard credit rate. Peak is Monday to Friday, 14:00–18:00 Singapore time, which is 06:00–10:00 UTC.Plan credits: Lite 2,000 per five hours and 10,000 weekly; Pro 12,000 and 60,000; Max 28,000 and 140,000.page not locatedPer-tier monthly dollar prices sit behind the console. The docs say only “starting at just 18 USD per month.”
Moonshot, MiniMax, Alibaba and Tencent — direct APIs
kimi-k3Moonshot direct APIUSD / 1M tokens, excluding applicable taxes$3.00 cache miss$0.30 cache hit$15.00None publishedExcluded from the batch tier. Moonshot prices batch at 60% of standard and names only kimi-k2.7-code, kimi-k2.6 and kimi-k2.5 as supported.The only vendor in this index that flags tax treatment on its own price table.
kimi-k2.7-codeMoonshot direct APIUSD / 1M tokens, ex-tax$0.95 cache miss$0.19 cache hit$4.00None publishedBatch $0.57 miss · $0.114 hit · $2.40 out — 60% of standardkimi-k2.7-code-highspeed is exactly 2× at $1.90 / $0.38 / $8.00. The vendor describes it as the same model with a faster output speed. Latency is the whole product.
kimi-k2.6Moonshot direct APIUSD / 1M tokens, ex-tax$0.95 cache miss$0.16 cache hit$4.00None publishedBatch $0.57 · $0.10 · $2.40Max output for every Kimi model below K3 is not published
MiniMax-M3, at or below 512K inputMiniMax direct API, Standard tierUSD / 1M tokens$0.60$0.12$2.40$0.30 in · $0.06 cached · $1.20 outBadged “Permanent 50% off” — a discount the vendor calls permanent while still rendering a struck-through list price. No end date, by construction.not publishedNo batch discount on the pay-as-you-go page.Priority is 1.5× standard: $0.45 / $0.09 / $1.80 effective. M3 is the flagship; M2.5, M2.1 and M2 sit under MiniMax’s own “Legacy Models” accordion.
MiniMax-M3, above 512K inputMiniMax direct API, Standard tierUSD / 1M tokens$1.20$0.24$4.80$0.60 · $0.12 · $2.40 effectiveNone publishedPriority effective $0.90 / $0.18 / $3.60. The tier boundary is exactly 2× on every cell.
MiniMax-H3 and H3 MaxMiniMax direct APIUSD per second of output video, not per tokenH3 at $0.08 per second at 768P and $0.13 at 2K. H3 Max at $0.05 at 480P and $0.08 at 768P. Regeneration $0.05 per second. Only the output video is billed on H3 Max.None publishedMiniMax-H3-Context-IR is the exception and is priced per token: $0.90 in / $3.60 out.
qwen3.8-maxAlibaba Cloud Model Studio, SingaporeUSD / 1M tokens$2.00$0.25 implicit · $0.17 explicit read · $2.50 explicit write$6.00None publishedBatch not flagged on this SKU. Alibaba’s rule: batch and cache discounts “cannot apply simultaneously.”Tiered pricing bills every token in a request at the tier the request total falls into, not just the excess.
qwen3.8-flashAlibaba Cloud Model Studio, SingaporeUSD / 1M tokens$0.15not machine-readableQwen Cloud’s pricing block did not render on two attempts. The headline pair is vendor-primary from Alibaba’s billing page; the cache cells are not.$0.47None publishedNot flagged for batch on this SKU.
qwen3.7-maxAlibaba Cloud Model Studio, SingaporeUSD / 1M tokens$2.50$0.50 implicit$7.50$1.25 in · $0.25 cached · $3.75 outLabelled only “Limited-time 50% off” — no end date published. The promo applies to the floating alias; the pinned snapshots stay at list.Caches discounted 50% alongside.Same model, two prices, decided by which ID you call.
qwen3.7-plus, at or below 256K inputAlibaba Cloud Model Studio, SingaporeUSD / 1M tokens$0.40$0.08 implicit$1.60$0.32 · $0.064 · $1.28“Limited-time 20% off”, no end date published.Above 256K input: $1.20 / $4.80 list, with the same 20% off.
qwen3.7-flash, at or below 32K inputAlibaba Cloud Model Studio, SingaporeUSD / 1M tokens$0.030$0.006 implicit$0.130None published50% batch discount, flaggedAbove 32K: $0.100 / $0.400. Above 256K: $0.200 / $0.800.
Hy4 previewTencent, international release and TokenHub SingaporeUSD / 1M tokens$0.834$0.042$2.501Free on WorkBuddy and CodeBuddy for two weeks from the August 28, 2026 launch. Tencent publishes no calendar end date.not publishedApache-2.0 open weights. Preview status, with the vendor saying an official release is still ahead.
Hy4 preview, Chinese releaseTencent, ChinaCNY / 1M tokens¥6“as low as ¥0.3”The Chinese text hedges the cache figure; the English text states $0.042 flat. Quote the language you are citing.¥18Same free-trial termsBoth currencies are vendor-published regional list prices. Neither is our conversion, and this index does not convert between them.
Hy3Tencent TokenHub SingaporeUSD / 1M tokens$0.132$0.033$0.528Free on WorkBuddy and CodeBuddy, extended until September 30not published
xAI, Meta, Mistral, Amazon and Cohere — direct APIs
grok-4.6xAI direct APIUSD / 1M tokens$2.00$0.50$6.00None — xAI publishes no promotional rates at all“Batch API: Not supported.” The flagship loses a discount tier the older model has.Prompts reaching 200k tokens are billed at $4.00 / $1.00 / $12.00 for every token in the request. Shipped August 12, 2026; xAI’s current flagship.
grok-4.5xAI direct APIUSD / 1M tokens$2.00$0.30$6.00Nonepage not locatedIts per-model page was not fetched; the price table carries no batch column.200k and above: $4.00 / $0.60 / $12.00.
grok-4.3xAI direct APIUSD / 1M tokens$1.25$0.20$2.50NoneBatch at a 20% discount — not the 50% the rest of the industry uses200k and above: $2.50 / $0.40 / $5.00. Still live and still sold; no longer the flagship.
grok-4.20 — reasoning, non-reasoning and multi-agentxAI direct APIUSD / 1M tokens$1.25$0.20$2.50Nonepage not located200k and above: 2× on every cell.
grok-build-0.1xAI direct APIUSD / 1M tokens$1.00$0.20$2.00Nonepage not located200k and above: 2× on every cell.
muse-spark-1.2 and 1.1, Standard tierMeta Model API directUSD / 1M tokens$1.25$0.15$4.25None publishednot publishedNo batch endpoint exists. An absence by design.Vendor’s own words: “There is no long-context premium: you pay the same rate whether your context window is mostly empty or almost full.”
muse-spark-1.2-contributorMeta Model API direct, Contributor tierUSD / 1M tokens, plus training permission$0.10$0.002$0.20Not a promotion. Meta’s own framing: “Tier is a model attribute: it sets the price you pay and whether your data may be used to train future Meta models.”NoneSame checkpoint as Standard. Throttled to 100 requests per minute against Standard’s 3,000. See section 08.
Mistral Medium 3.5Mistral direct APIUSD / 1M tokens$1.50not publishedOnly a cross-cutting FAQ statement that cached input reduces input cost “by up to 90%.”$7.50None publishedBatch −50%, stated in the FAQ onlyPriority Tier at 1.75× list, with a 99.5% uptime SLA. Standard and Batch carry no SLA. The vendor’s self-described frontier-class model costs 3× Large 3 on input.
Mistral Small 4Mistral direct APIUSD / 1M tokens$0.15not published$0.60None publishedBatch −50%Priority Tier 1.75×.
Mistral Large 3Mistral direct APIUSD / 1M tokens$0.50not published$1.50None publishedBatch −50%On AWS Bedrock in US regions this is $0.50 / $1.50 exactly — the reseller matches the vendor.
Z.ai GLM 5.2, hosted by MistralMistral direct API, third-party modelUSD / 1M tokens$1.40$0.14$4.40None publishedBatch −50%Mistral publishes a max-output figure on this third-party card and omits the field entirely from all three of its own model cards.
Amazon Nova Pro and Nova PremierAWS BedrockUSD / 1M tokensnot machine-readableThe Nova rate tables are region-selector widgets that did not render. The section headings did, which confirms Nova is priced on the same Global versus Geo axis as everything else on Bedrock.None publishedFlex 50% off Standard.Priority 1.75×. Nova Premier is marked Legacy with an end of life of September 14, 2026.
Cohere, current generative lineCohere directUSD per instance-hour, not per tokennot publishedNo current frontier generative per-token card exists on Cohere’s pricing page. What it does publish is Model Vault per-instance pricing: $4.00 to $10.00 per hour depending on model and performance tier, or $2,500 to $6,500 per month.None publishedOn Bedrock, Cohere appears only as Rerank 3.5 at $2.00 per 1,000 queries. One vendor, three billing units across two surfaces.

Two structural observations the table makes that a per-vendor page cannot. First, newer is not reliably cheaper or dearer within a family. Gemini 3.7 and 3.6 Flash sit at a promotional $0.75 / $3.75 while the older 3.5 Flash is $1.50 / $9.00 — double the input and 2.4× the output. Claude Sonnet 5 undercuts both Sonnet 4.6 and 4.5 by a third on input. Grok 4.6 costs more than Grok 4.3 and loses batch support. Whatever rule you were using to guess a price from a version number, it does not hold.

Second, the discount tiers are not a shared vocabulary. Batch is 50% off at six vendors, 60% of standard at Moonshot, 20% off at xAI on exactly one model, and does not exist at all at DeepSeek, Z.ai, MiniMax, Meta or Tencent. Google runs four tiers where flex and batch are both exactly 50% off but are different products with different latency guarantees, and priority is a premium tier at 1.8× on every published row. Reading “batch” as a fixed discount across vendors will misprice a budget by a factor of two.

List output price per million tokens for ten current models, data as of August 30, 2026Horizontal bars of published list output rates, largest to smallest. Claude Fable 5 $50.00, Claude Opus 5 $25.00, GPT-5.6 Terra $12.00, Claude Sonnet 5 $10.00, Gemini 3.7 Flash $7.50, Grok 4.6 $6.00, GLM-5.3 $4.40, DeepSeek V4-Pro peak $3.96, MiniMax-M3 $2.40, GPT-5.6 Luna $1.20. GPT-5.6 Sol is omitted because OpenAI publishes no list output rate for it.LIST OUTPUT · USD PER MILLION TOKENS · TEN CURRENT MODELSData as of 2026-08-30 · list column only · Sol omitted, no published listClaude Fable 5$50.00Claude Opus 5$25.00GPT-5.6 Terra$12.00Claude Sonnet 5$10.00Gemini 3.7 Flash$7.50Grok 4.6$6.00GLM-5.3$4.40DeepSeek V4-Pro peak$3.96MiniMax-M3$2.40GPT-5.6 Luna$1.20
List output rates from the table above, as of August 30, 2026. Promotional cells are excluded. GPT-5.6 Sol is omitted because OpenAI publishes no list output figure for it — only the $20 promotional rate.

04PromotionsSeven vendors discounting, and three published end dates.

This is the table that makes the index falsifiable. Every promotional rate in the archive above appears here with whatever the vendor published about when it stops — including, in five rows, the fact that the vendor published no end date at all. A price index that folded these into a single rate column would be unfalsifiable for most of the vendors in it, because there would be no way to tell which cells were about to move.

Every promotional or discounted rate in the index, with its published end date. As of August 30, 2026. A published future price is a plan, not an outcome — Anthropic’s cancellation of its own September 1 increase is the proof.
Vendor and modelSurfaceList ratePromotional ratePublished end dateWhat the vendor says follows
Google — Gemini 3.7 Flash and 3.6 FlashGemini Developer API and Vertex AI, global$1.50 / $7.50$0.75 / $3.75December 31, 2026“Starting January 1, 2027, standard pricing of $1.5 / $7.5 per 1M tokens input / output will apply.” Both halves published, on two independent Google pages.
Google — 3.7 and 3.6 Flash cache read and storageGemini Developer API$0.15 read · $1.00 per 1M tokens per hour storage$0.075 read · $0.50 storageDecember 31, 2026The post-promotional figures are published for both.
Google — 3.7 and 3.6 Flash Provisioned ThroughputVertex AI onlyNo credit50% monthly billing credit on net eligible PT spendDecember 31, 2026, effective August 13, 2026Credits issue on the 7th of the following month and expire 30 days after issuance — a discount with an expiry on the discount itself.
Z.ai — GLM-5.3-FlashZ.ai direct API$0.15 / $0.03 / $0.50$0.075 / $0.015 / $0.2524:00 on September 9, 2026 (UTC+8, Singapore time)The list prices are rendered as strikethrough beside the promo prices, so the post-promotional rate is on the same page.
OpenAI — gpt-5.6-solOpenAI API direct, Standard tiernot published$4.00 / $20.00A floor, not a date: “available at least through November 21, 2026”No figure. The vendor describes the change as “a 20% reduction in input pricing and a 33% reduction in output pricing,” which back-solves to $5 / $30 — GPT-5.5’s rate. That is arithmetic on the vendor’s own words, not a published price, and it appears in no rate cell in this index.
Alibaba — qwen3.7-maxModel Studio, Singapore$2.50 / $7.50$1.25 / $3.75not publishedLabelled only “Limited-time 50% off.”Nothing. The pinned snapshots -2026-06-08 and -2026-05-20 sit at full list, so the same model has two prices depending on which ID you call.
Alibaba — qwen3.7-plusModel Studio, Singapore$0.40 / $1.60$0.32 / $1.28not published“Limited-time 20% off.”Same alias-versus-snapshot split as qwen3.7-max.
MiniMax — MiniMax-M3, all four tier and service combinationsMiniMax direct API$0.60 / $2.40 and up$0.30 / $1.20 and upnot publishedThe badge reads “Permanent 50% off.”A discount the vendor calls permanent while still rendering a struck-through list price. It is neither a list rate nor an expiring promotion, and it needs its own category.
Z.ai — cached-input storage, 16 SKUsZ.ai direct APIStorage billedFreenot published“Limited-time Free.”A promotion on a third billing axis — storage, not per-token — which is easy to miss entirely.
Tencent — Hy4 previewWorkBuddy and CodeBuddy$0.834 / $2.501Free“for two weeks” from the August 28, 2026 launch — no calendar date publishedA product-surface free trial, not an API discount. Stated as the vendor states it rather than computed into a date.
Tencent — Hy3WorkBuddy and CodeBuddy$0.132 / $0.528FreeSeptember 30 — year not stated, 2026 from context“Free access to Hy3 on both platforms has also been extended until September 30.”
Moonshot — file extraction and file storageMoonshot direct APInot publishedFreenot published“temporarily free.”No underlying rate is published, so there is no way to model what happens when it stops.
Anthropic — Claude Sonnet 5Claude API direct$2.00 / $10.00None. The promotion became the list price on August 10, 2026.Not applicableIncluded in this table deliberately, because it is the row a reader is most likely to arrive carrying a wrong number for. The scheduled $3 / $15 increase was cancelled and must not be cited as a future rate.
The asymmetry worth naming

Google publishes the post-promotional number as a hard figure and a hard date, on two independent pages, for the same commercial event that OpenAI documents with a soft floor and no rate at all. Same kind of promotion, opposite disclosure quality — and neither is more binding than the other. Anthropic published a future price with a date too, and then cancelled it. So Google’s January 1, 2027 figure is quotable as a plan, and citing it as what Gemini will cost in 2027 repeats exactly the mistake this page opens with.

05SurfacesThe same model, the same rate, a different bill.

The most useful thing this index does is refuse to give one number per model. A price is a rate, a unit, an endpoint and a tier — and three of those four change when you move surfaces without the rate moving at all.

The cleanest illustration is first-party and unambiguous. Anthropic bills through AWS Marketplace and Microsoft Foundry in Claude Consumption Units. Token usage is rated in USD at exactly the standard per-model rates in the index above, any negotiated discount is applied, and the result is converted at $0.01 per CCU — 100 CCU represents $1.00 — and reported hourly. Your AWS or Azure bill shows a single CCU line item. The rate is identical. The unit on the invoice is not. Anyone reconciling a cloud bill against a per-million-token model will find nothing that looks like a token.

The second illustration is a capability rather than a unit. Anthropic’s fast mode — $10 / $50 on Opus 5 and Opus 4.8, exactly 2× standard — is available on the Claude API first-party only. It is not available on Claude Platform on AWS or on partner-operated clouds. Same model, same headline rate, and one surface can buy a latency upgrade the others cannot.

One vendor’s models across four surfaces, in USD per 1M input and output tokens. As of August 30, 2026. Grok is the clearest case because xAI, AWS, Google and Microsoft all publish rates for overlapping parts of the same catalogue — and disagree about which catalogue it is.
ModelxAI directAWS BedrockGoogle Vertex AIMicrosoft Foundry
Grok 4.6$2.00 / $6.00Global cross-region $2.00 / $6.00; in-region and geo cross-region $2.20 / $6.60, a +10% premium$2.00 / $6.00, and $4.00 / $12.00 above 200K — byte-identical to the vendor, long-context rule includedNot offered
Grok 4.3$1.25 / $2.50In-region only, $1.25 / $2.50. GovCloud US-West $1.50 / $3.00, a +20% premium$1.25 / $2.50, and $2.50 / $5.00 above 200KNot offered
Grok 4.2Not listedNot offeredNot offeredGlobal $1.25 / $2.50 · Data Zone $1.375 / $2.75, a +10% premium. Foundry’s newest Grok is four generations behind xAI’s.
Grok 4.1 Fast, reasoning and non-reasoningNot listed — no longer on xAI’s model listNot offered$0.20 / $0.50, cache hit $0.05 — still sold
Grok-4Not listedNot offeredNot offeredGlobal $3.00 / $15.00 · US Gov Global $3.75 / $18.75. On the reseller you can pay 2.5× on output for a model the vendor has dropped.
Batch support20% discount on Grok 4.3; “Not supported” on Grok 4.6Flex 0.5×, Priority 1.75×Batch and Flex tiers publishedGlobal Batch and Data Zone Batch at 50% off

The regional dimension deserves its own line, because of all the surface premiums it is the one presented with no price vocabulary attached at all. AWS Bedrock’s regional price list is a premium AWS never labels as one. The only word on those tables is a region name. Measured against the US East and US West rows on AWS’s own table, Gemma 3 27B input runs $0.23 in US regions, $0.27 in Mumbai, Ireland and Milan — a 17.4% premium — and $0.36 in London, which is 56.5% above the US row. Six model families on that list land between +50% and +57% in London. Sydney is the opposite case at around +3%.

Regional premium
London, on AWS Bedrock
+56.5%

Gemma 3 27B input: $0.23 in US East and US West, $0.27 in Mumbai, Ireland and Milan, $0.36 in London. Six model families on the same AWS price list land between +50% and +57% in London. AWS's only label on those tables is a region name.

Measured against AWS's own US row
Convergent design
Five surfaces, one multiplier
+10%

OpenAI charges a 10% uplift on regional data-residency endpoints for models released on or after March 5, 2026. Vertex's non-global endpoint is exactly 1.10x global. Anthropic's inference_geo of us applies a 1.1x multiplier. Bedrock's in-region and geo cross-region options are +10% on global. Azure's Data Zone is +10% on Global Standard.

Independently arrived at
Passthrough, verified
OpenRouter's per-token markup
0%

The vendor's own words: it passes through the pricing of the underlying providers and there is no markup on inference pricing. The money is taken at the credit-purchase step instead — a 5.5% platform fee on pay-as-you-go, plus 5% on bring-your-own-key above a $25,000 per month list-price allowance.

Corrects the common assumption
Same premium, two vendors
Priority, at AWS and at Mistral
1.75×

AWS Bedrock states priority tier pricing is at a 75% premium to standard and flex at a 50% discount. Mistral's Priority Tier is 1.75x standard list pricing, described as a 75% premium on input, output and cached tokens, with a 99.5% uptime SLA that the standard and batch tiers do not carry.

Two independent price sheets

OpenRouter is worth stating precisely because the common assumption about it is wrong in a specific way. It takes no per-token markup — that is the vendor’s own claim, made twice on its own pages, and the rate rows in the index above bear it out: z-ai/glm-5.3, moonshotai/kimi-k2.6, qwen/qwen3.8-max and deepseek/deepseek-v4-pro-0813 all match their vendors to the cent. The fee is at the credit-purchase step, at 5.5% on pay-as-you-go, plus 5% on bring-your-own-key above a $25,000 per month list-price allowance. That allowance is measured in list-price inference cost, not in requests — a detail with real consequences for anyone modelling it as a request quota.

What OpenRouter does not guarantee is which provider serves a slug. With 80-plus providers and 500-plus models, the passthrough promise is per-provider, not per-model. Which is the setup for the next section.

06The sharpest specimenThree slugs where batch costs more than standard.

Everywhere else in this index, a batch tier is a discount. It is 50% off at six vendors, 60% of standard at Moonshot, 20% off at xAI on one model. On OpenRouter, three :batch slugs in the current catalogue cost more than the standard slug for the same model. Not a rounding difference — two of them are exactly double.

The mechanism is not a markup, and the passthrough claim in the last section survives intact. Each of the three is passing through a different real vendor rate than the standard slug is passing through. The suffix names an API mode; it does not name a price direction. That distinction is invisible until you check both slugs against the vendor’s own page, which is precisely the work an index is for.

OpenRouter slugs where the :batch variant is priced above the standard variant, USD per 1M tokens, pulled August 30, 2026. The DeepSeek rows are clock-dependent: the pull was made on a Sunday, which is categorically off-peak under DeepSeek’s Monday-to-Friday peak window.
SlugStandard slug:batch slugMultipleWhat the batch price actually is
z-ai/glm-5.3-flash$0.075 / $0.25$0.15 / $0.502.00× on bothExactly Z.ai’s post-promotional list price. The standard slug is passing through the 50%-off promotion that ends September 9; the batch slug is passing through the list rate. Z.ai publishes no batch API at all.
deepseek/deepseek-v4-pro-0813$0.66 / $1.98$1.32 / $3.962.00× on bothExactly DeepSeek’s peak rate against the standard slug’s off-peak rate. DeepSeek publishes no batch API either. The correct statement is that OpenRouter tracks DeepSeek’s clock, not that it displays a fixed column.
qwen/qwen3.8-2.4t-a95b$2.00 / $6.00$2.50 / $6.251.25× in · 1.042× outA 25% premium on input and about 4% on output, against a standard slug that matches Alibaba’s hosted rate exactly. No vendor batch rate is flagged on this SKU.
qwen/qwen3.5-9b$0.10 / $0.15$0.17 / $0.251.70× in · 1.667× outA fourth instance of the same shape on a smaller model.
moonshotai/kimi-k3$3.00 / $15.00$3.00 / $15.001.00× — no discountNot an inversion but the same lesson. Moonshot’s own batch tier is 60% of standard and excludes K3 entirely. A :batch slug exists for a model the vendor does not offer batch pricing for.
minimax/minimax-m3$0.30 / $1.20$0.30 / $1.201.00× — no discountSame rate, but the batch slug halves the advertised context to 524,288. The variant differs in something other than price.
What this specimen is actually about

None of these six rows is a reseller overcharging. Every one is a reseller faithfully passing through a rate — and the rate it passes through is not the one the slug’s name implies. Two vendors here publish no batch API whatsoever, and both have :batch slugs on a reseller anyway.

The operational consequence is narrow and expensive: a cost model that routes non-urgent work to :batch slugs by string substitution, on the reasonable assumption that batch means cheaper, will double the bill on two of these models and add nothing but latency tolerance on two more.

07Context bandsOne number per model is materially incomplete.

Three of the vendors in this index charge a different rate for the same model depending on how long the prompt is, and in every case the higher rate applies to the whole request, not just the tokens above the threshold. A single input rate per model does not describe what any of them will charge you.

The thresholds differ, the multipliers differ, and one major vendor has no cliff at all on its most-used line. This is the reason the index carries a premium-tiers column rather than folding everything into one input figure — and the mechanics of these thresholds, including how to structure prompts around them, are the subject of our long-context pricing thresholds guide.

Long-context pricing bands, as of August 30, 2026. In every case the higher rate applies to all tokens in the request, not only the tokens above the threshold. Worked from each vendor’s own published pair of rates.
Vendor and modelsThresholdInput multiplierOutput multiplierWorked example
OpenAI — the GPT-5.6 family and laterPrompts above 272,000 input tokens1.5×gpt-5.6-sol goes from $4.00 / $20.00 to $8.00 / $30.00 for the full request. Cache writes bill at 1.25× the uncached input rate.
Google — Gemini 3.1 Pro Preview and 2.5 ProInput at or above 200K tokens1.5×Gemini 3.1 Pro Preview goes from $2.00 / $12.00 to $4.00 / $18.00. Same shape as OpenAI’s, at a lower threshold.
Google — every Flash and Flash-Lite modelNone. Flat to the full 1,048,576-token input limitThe line most people actually use has no cliff, which is why “Gemini charges more for long context” is only true of the Pro models.
xAI — every Grok modelPrompts reaching 200k tokensGrok 4.6 goes from $2.00 / $6.00 to $4.00 / $12.00. xAI is the only vendor here that doubles output as well as input. Vertex reproduces the rule; both AWS Bedrock Grok cards omit it and publish a single flat rate per inference option.
Anthropic — Claude 4.6 and laterNone. The full 1M-token window at standard pricingThe vendor’s own gloss: a 900k-token request bills at the same per-token rate as a 9k-token request. Caching and batch discounts apply at standard rates across the full window.
Meta — Muse SparkNone, stated explicitly“There is no long-context premium: you pay the same rate whether your context window is mostly empty or almost full.” The only vendor here that names the absence.
Alibaba — qwen3.7-flash and qwen3.7-plusTiered by request total: 32K and 256K boundariesVaries by tierVaries by tierA different mechanism with the same trap: qwen3.7-flash runs $0.030 / $0.130 up to 32K, $0.100 / $0.400 to 256K and $0.200 / $0.800 to 1M — and every token in a request bills at the tier the request total falls into.
MiniMax — MiniMax-M3Input above 512KEffective rates go from $0.30 / $1.20 to $0.60 / $2.40 on the Standard tier, and the same doubling applies inside the Priority tier.

One further trap sits underneath this table and is worth naming because it will otherwise silently corrupt a cost model built from it. OpenAI’s headline context window includes output. Its 1,050,000 figure is 922,000 maximum input plus 128,000 maximum output, exactly. Google’s 1,048,576 is an input limit with 65,536 of output on top. The two vendors’ “1M” numbers are not the same quantity and must not be compared directly. The ceilings themselves, by model and by surface, are censused separately in our context window and output limit census.

08A different currencyOne vendor charges in data, not dollars.

Every other row in this index is a price in money. Meta publishes one that is not. The Muse Spark 1.2 checkpoint is sold at two prices under two model IDs, and the cheaper one is bought with permission rather than cash.

Meta’s own framing is unusually direct about it: tier “sets the price you pay and whether your data may be used to train future Meta models,” and the Contributor ID carries “heavily discounted pricing in exchange for permission to use your prompts and completions to train future Meta models.” Measured against Meta’s own Standard rate for the same checkpoint, that discount is 92.0% on input, 95.3% on output and 98.7% on cached input.

Meta’s two tiers for the same Muse Spark 1.2 checkpoint, USD per 1M tokens, as of August 30, 2026. Percentages are computed against Meta’s own Standard rate for the same model, which is the stated denominator.
Billing categoryStandard tierContributor tierDiscountWhat is exchanged
Input$1.25$0.10−92.0%Permission to use your prompts and completions to train future Meta models.
Output$4.25$0.20−95.3%Same terms.
Cached input$0.15$0.002−98.7%Same terms.
Requests per minute3,000100−96.7%The discount tier is also throttled. Token-per-minute limits are closer: 4,000,000 against 3,000,000. Limits are per team, not per API key.

No other vendor in this index prices on this axis, and it is worth being precise about why that is not the same as saying no other vendor trains on your data. Google runs the closest analogue and runs it the other way round: on the Gemini API free tier, content is used to improve Google’s products; on the paid tier it is not. Same model, same rate card, different data terms decided by the tier rather than by a model ID.

The difference matters for procurement. Google’s split is a consequence of not paying. Meta’s is a price: a published, quantified discount attached to a specific model ID, which means it can be compared against the market rate, put in a budget, and — the part that makes it a governance question rather than a finance one — chosen by a developer with a one-word change to a model string.

09ImplicationsHow to use this without getting burned.

The most common way to misuse a price index is to take one cell out of it. Every number above is true of a specific model, on a specific surface, in a specific tier, on August 30, 2026 — and at least one of those four qualifiers is doing real work in most rows.

Carry the footnote
Copy the end date with the rate, or copy neither

Two of the three headline rate changes this month were promotional, and both were reported without their footnotes. A rate in a spreadsheet with no expiry column will read as permanent to whoever inherits the spreadsheet. Model the post-promotional rate wherever the vendor publishes one, and model the uncertainty where it does not.

Rate plus expiry, always
Name the surface
A price is not a number until you say where you bought it

The same Claude rate bills in Claude Consumption Units on two clouds and in dollars on a third. The same Grok model costs 10% more in-region on Bedrock and is not offered at all on Foundry. Anthropic's fast mode does not exist outside the first-party API. Write the surface into the line item, not just the model name.

Model plus surface plus tier
Never substitute by string
Do not append :batch and assume a discount

Three OpenRouter slugs cost more with the batch suffix, two of them exactly double, because the suffix names an API mode rather than a price direction. Two of those vendors publish no batch API at all. Check both slugs against the vendor's own page before routing anything by name pattern.

Verify each slug
Budget the cliff
Long prompts reprice the whole request

Above 272,000 input tokens OpenAI charges 2x input and 1.5x output for every token in the request, not just the excess. Google's Pro models do the same above 200K and its Flash models do not do it at all. If your median prompt sits near a threshold, your effective rate is neither of the two published numbers.

Model the distribution

The wider point is about where the market has moved. Two years ago a price index could be a list of models and two numbers each. The direction of travel since is unmistakable: every large vendor has added tiers, thresholds, endpoints and units, and the differences between them are now the thing that decides a bill rather than a footnote to it. That is unlikely to reverse, because each of these dimensions is a real product with a real cost behind it — data residency is genuinely more expensive to serve, sheddable off-peak capacity genuinely is cheaper. What should be expected instead is more of it: more premium tiers, more units that are not tokens, and more prices that are a function of when and where you called rather than what you called.

The practical response is to stop treating price as an attribute of a model and start treating it as an attribute of a call. Our ledger of AI coding plan limit changes tracks the same volatility on the subscription side, where the ceiling moves rather than the rate. If you want senior help mapping what your workloads actually cost across these surfaces, and what a mid-quarter promotional expiry does to a committed budget, that is the kind of question our AI transformation engagements open with.

10ConclusionThe footnote is the price.

The index, as of August 30, 2026

Two columns, three gap labels, and a URL that gets updated instead of replaced.

The design decision that makes this page worth citing is that it never merges a promotional rate into a list column, and never drops a row it could not verify. A reader who wants to argue that frontier inference is getting cheaper and a reader who wants to argue that the discounts are about to expire can both quote it, cell by cell, without either of them inheriting our judgement.

The two cases at the top are the reason the shape matters. Anthropic cancelled a scheduled increase five days after we published a post built on it, and OpenAI’s most-quoted rate is a promotion that most of the market — including our own internal anchors — recorded as a list price. Neither is anyone’s dishonesty. Both are what happens when a number travels without the sentence printed next to it.

This URL supersedes our four dated snapshots from March, April and August 2026, which stand as historical records of what the market published on their own dates. It is maintained by the Digital Applied Team, reviewed monthly and out of cycle when a frontier rate moves, and re-dated through its modified time rather than replaced. Cite the as-of date with any cell you take from it.

Budget your model spend properly

The vendor sets the rate. The footnote decides what you pay.

Our team maps what your AI workloads actually cost across direct APIs, cloud resellers and coding plans — including the tier, threshold and residency dimensions that a single per-token figure hides.

Free consultationExpert guidanceTailored solutions
What we work on

AI cost engagements

  • Model and surface selection against real workloads
  • Prompt-length distribution against long-context thresholds
  • Cache and batch strategy where the vendor actually offers one
  • Promotional-expiry exposure in committed budgets
  • Reseller versus direct API procurement reviews
FAQ · AI model API pricing

What this index gets asked.

No. That increase was cancelled. Sonnet 5 launched with $2 per million input tokens and $10 per million output tokens described as introductory pricing through August 31, 2026, with a scheduled step-up to $3 / $15 on September 1. On August 10, 2026 Anthropic made the introductory rate permanent and withdrew the scheduled increase. Two first-party pages carry it: the pricing documentation states the $2 / $10 rate is now the standard price and that the increase will not occur, and the Sonnet 5 launch post carries a dated edit note to the same effect. Any table still showing a September 1 cliff, or citing $3 / $15 as a future Sonnet 5 rate, is out of date. Our own August 5 pricing tracker reported the expiry footnote accurately on the day it published and now carries a dated correction.
Related dispatches

Continue exploring AI model costs.