AI DevelopmentPricing Tracker7 min readPublished September 25, 2026

18 tiers · 9 vendors · read September 25 · every speed figure is the vendor's

What Paying for a Faster AI Model Actually Costs You

A census of 18 speed and priority tiers where one vendor sells the same model faster: the multiple, the vendor's speed claim, and what else changes.

DA
Digital Applied Team
Research and practical guidance
PublishedSeptember 25, 2026
Rows18 tiers, 2 contrast rows

The same model, served faster, is now a product with its own price. In the week to September 25, 2026 Xiaomi, Z.ai and Anthropic each put a faster version of an existing model on a price list, and Alibaba's Prime tiers surfaced on its pricing pages, joining the fast and priority tiers OpenAI, Google, xAI, Moonshot and Mistral already sell. The multiple is not standard. Most tiers cost exactly twice the base. Google's cost 1.8 times. One costs ten times. And we found no measurement of any of them by anyone other than the vendor.

This census puts every priced tier we could find on one table: the multiple, the vendor's speed claim in the vendor's own unit, and what else changes when you flip the flag, because a faster tier usually loses something too, such as batch pricing, a cache, a region or a rate limit. Then a decision rule for when a multiple is worth paying and when a smaller model is the cheaper answer to the same latency problem.

Key takeaways
  1. 01
    Ten of the 18 tiers cost exactly 2× the base; Google's cost 1.8×, Mistral's 1.75×, Xiaomi's 10×.OpenAI's older models vary per model from 1.67× to 2.5×, Z.ai's FlashX is 2.47× / 2.5×, and Alibaba's Wan video Prime is 1.36× to 1.40× against list. Anthropic's Priority Tier publishes no multiple at all and is no longer sold.
  2. 02
    The speed claims do not share a unit, and none is independently measured.Anthropic promises output tokens per second and says not time to first token; Z.ai and Moonshot give absolute rates; Xiaomi says up to 20× with no unit; Google, xAI and Mistral promise queue position. Artificial Analysis lists no fast or priority tier for any model.
  3. 03
    The premium usually costs you something besides money.Batch pricing drops out at Anthropic, Xiaomi and xAI. Anthropic's fast and standard requests do not share a cache. Gemini Priority's default rate limit is 0.3× standard. Opus fast mode is Claude API only; Grok 4.7 Fast is not on xAI's own API.
  4. 04
    Four vendors will serve your request at standard speed and bill standard without telling you, except in a response field.OpenAI's ramp limit, Gemini's downgrade, xAI's priority flag and Mistral's over-limit fallback all fall back silently. Read the service-tier field on every response if you are paying for the fast one.

01 — The findingThree shapes of premium, and no independent number for any of them

A
The common multiple
2.0×

Both Anthropic fast-mode rows, OpenAI Fast on GPT-6 and on GPT-5.6 Sol, xAI Priority, Grok 4.7 Fast under 200K tokens, Moonshot HighSpeed and all three Alibaba Prime text SKUs. Ten rows.

10 of 18
B
Google's number
1.8×

Every Gemini Priority price on the pricing page is exactly 1.8× standard, including caching and the two new TTS models, although the docs say 75 to 100% more.

Gemini API
C
The outlier
10×

MiMo-V2.6-Pro-UltraSpeed is the same Pro checkpoint at ten times the price on input, output and cache hits, for a claim of up to 20× speed in no stated unit.

Xiaomi
D
Independent measurements
0

No fast, priority, prime, FlashX, HighSpeed or UltraSpeed slug exists on Artificial Analysis. Aggregator throughput fields for the Xiaomi routes were empty when read.

Measured by nobody

The scope rule for the table is strict: a row qualifies only when one vendor sells the same model at two service levels for two prices. A different model whose name contains "Fast" or "Turbo" is not a speed tier, it is a different model, and is excluded. Cheaper-and-slower tiers such as OpenAI Flex and Gemini Flex are listed below the table for contrast, because on one model they show the full spread: Gemini 3.8 Flash costs 0.5× on Flex and 1.8× on Priority, so the fastest synchronous tier costs 3.6 times the cheapest.

02 — The censusEighteen tiers, read September 25, 2026

Multiples are the tier price divided by the base price at the same unit and context band, computed from each vendor's pricing page. Speed claims are quoted as the vendor wrote them. The fourth column is what else the vendor's documentation says changes.

Vendor pricing pages and service-tier documentation, read September 25, 2026. Prices in US dollars per million tokens unless marked ¥ (CNY) or per second of video.
Tier, over baseInput / output multipleVendor speed claimWhat else changes
MiMo-V2.6-Pro-UltraSpeed (Xiaomi), over MiMo-V2.6-Pro at $0.435 / $0.8710.00× / 10.00×"up to 20x inference speed", unit not statedNo Batch API. Cache-hit price also 10×. Same context and output limits as Pro on the aggregator route.
GLM-5.3-FlashX (Z.ai), over GLM-5.3-Flash at $0.15 / $0.502.47× / 2.50×"inference speeds of 200 tokens/s"Not on the GLM Coding Plan. Cached input 2.5×. No FlashX batch route seen. Documented on the Flash model page with no separate model card. On the aggregator, Z.ai's own FlashX and Flash endpoints carry the same 1,048,576 context and 131,072 max output.
Claude Opus 5.5 fast mode (Anthropic), over $4 / $202.00× / 2.00×"up to 2.5x higher output tokens per second"; the docs say the target is output rate, not time to first tokenResearch preview by waitlist. Claude API only: not Bedrock, Google Cloud or Microsoft Foundry. Own rate limit. No Batch API. Fast and standard do not share cached prefixes, so a fallback is a cache miss. Not combinable with a Priority Tier commitment.
Claude Opus 5 and Opus 4.8 fast mode (Anthropic), over $5 / $252.00× / 2.00×Same "up to 2.5x" output-rate statementAs above. Opus 4.7 returns an error; Opus 4.6 accepts the fast flag and runs at standard speed and standard price.
Grok 4.7 Fast (xAI), over Grok 4.7 at $2 / $6 under 200K prompt tokens2.00× / 2.00×; 1.50× above 200K"twice the output speed at twice the price"; docs call it the same model on faster infrastructureNot on the public xAI API: Cursor and Grok Build only. Cached input 2×. No batch discount exists for Grok 4.7 at any tier.
xAI Priority Processing (service_tier priority), over standard rates2.00× / 2.00×"typically results in lower TTFT and faster inter-token latency, especially during periods of high demand"; no multiple claimedBilled at priority only when the response says so; otherwise standard. Not for batch, image or video. Cache discount applied before the 2×.
OpenAI Fast mode, GPT-6 Astra, Sol and Luna (service_tier fast), over $10 / $50, $2 / $10 and $0.10 / $0.502.00× / 2.00× at every context bandPage-level: "up to 2.5× faster speeds and more consistent latency"; the only model-specific figure is for GPT-5.6 SolShares the standard rate limit. Traffic that ramps too fast is downgraded to standard and billed standard. No latency SLA on GPT-6 Astra. EU data residency is standard-only for GPT-6. Cached input also 2×. Not for fine-tuned models or embeddings.
OpenAI Fast mode, GPT-5.6 Sol, over $4 / $20 short context2.00× / 2.00×"up to 2.5× faster than Standard processing" for this model specificallyAs above, plus Scale Tier SLA treatment on GPT-5.6 and earlier. Promotional price runs at least through November 21, 2026.
OpenAI Fast mode, older models (gpt-5.5, gpt-4.1, gpt-4o, gpt-5-mini, gpt-4o-mini)2.50×, 1.75×, 1.70×, 1.80×, 1.67×Same page-level claim; no per-model figureAs above. The multiple is not uniform across the older lineup; it is set per model on the pricing page.
Gemini API Priority inference (service_tier priority), every priced text model1.80× / 1.80×Latency "Low (Seconds)" against standard "Seconds to minutes"; docs say "75-100% more than Standard", every price is exactly 1.8×Preview. Default rate limit is 0.3× the standard limit. Over-limit traffic is downgraded to standard and billed standard; a response header records which tier served it. Context caching also 1.8×.
Gemini 3.8 Flash TTS and Flash-Lite TTS, Priority1.80× / 1.80×None specific to speechPriced on the pricing page but absent from the Priority page's supported-models list. Promotional prices double on January 1, 2027 for every tier, so the ratio holds.
Qwen3.8-Max-Prime (Alibaba Model Studio), over Qwen3.8-Max at ¥12 / ¥36 in Beijing2.00× / 2.00×Output rate "increased to 1.5 to 2 times that of the standard API"Listed for the Beijing region only, with no free quota. The base row carries batch and caching discount tags; the Prime row carries neither. Rate limiting is waived when spare capacity exists.
GLM-5.3-Prime, sold by Alibaba Model Studio, over its GLM-5.3 at ¥8 / ¥282.00× / 2.00×Same 1.5 to 2× output-rate rangeListed only on the Chinese Prime page as of September 25; Z.ai's own documentation has no Prime tier anywhere. The aggregator route's single provider is Alibaba.
GLM-5.2-fast-preview (Alibaba Model Studio), over GLM-5.2 at ¥8 / ¥282.00× / 2.00×Same 1.5 to 2× rangeNo free quota against the base model's 1M tokens.
Wan3.0-Video-Prime (Alibaba), over Wan3.0-Video per second of video1.36× to 1.40× against Singapore list (1.50× in Beijing); about 2× against the current 30%-off promotionNo video-specific figure; the Prime page's rate line is written for tokensSame free quota as base. The English Prime page and the pricing page print different Beijing prices; name the comparator when quoting a multiple.
Kimi K2.7 Code HighSpeed (Moonshot), over Kimi K2.7 Code at $0.95 / $4.002.00× / 2.00×"approximately 180 Tokens/s" and "up to 260 Tokens/s in short context scenarios"Vendor warns capacity is limited and the experience may fluctuate. Cache hit 2×. Same 256K context. Batch not stated.
Mistral Priority Tier (service_tier auto), over standard list1.75× / 1.75×Latency "Seconds" against "Seconds to minutes"; priority queue under loadEnterprise entitlement with custom per-model limits; falls back to standard when over. Carries a 99.5% uptime SLA that standard does not. Cached tokens 1.75×.
Anthropic Priority Tier (committed throughput)Not published"targets 99.5% uptime with prioritized computational resources"; no speed claimNo longer available for purchase; existing commitments run to contract end. Not supported on Opus 5.5, Opus 5, Fable 5.1, Mythos or Sonnet 5.

Output-price multiple of the fast or priority tier over its base

Vendor pricing pages, read September 25, 2026. Bars scaled to Xiaomi's 10×. Grok 4.7 Fast shown at its under-200K multiple.
MiMo-V2.6-Pro-UltraSpeedXiaomi
10.0×
GLM-5.3-FlashXZ.ai
2.5×
Claude Opus 5.5 fast modeAnthropic
2.0×
OpenAI Fast, GPT-6 familyOpenAI
2.0×
Kimi K2.7 Code HighSpeedMoonshot
2.0×
Qwen3.8-Max-PrimeAlibaba
2.0×
Gemini PriorityGoogle
1.8×
Mistral Priority TierMistral
1.75×

Two contrast rows sit outside the census because they run the other way. OpenAI Flex serves GPT-6 Astra at $5 and $25, half of standard, with the documentation warning of slower responses and occasional unavailability that returns an uncharged error. Gemini Flex, in preview, is also 0.5× with a stated target of one to fifteen minutes and no automatic upgrade to standard. The batch tiers across vendors are a separate, mostly 0.5×, market that our batch pricing landscape covers.

The Xiaomi row is the one people will quote, and it needs its context. The base model is cheap: MiMo-V2.6-Pro is $0.435 input and $0.87 output, so ten times that is $4.35 and $8.70, in the same band as a frontier model's standard tier. Our post on the MiMo-V2.6 release covers what Xiaomi published about the model itself.

03 — The mechanismWhy the same weights behave differently through a different route

The rows split into two kinds of product, and vendors use the same word, priority, for both. The first kind is a faster serving configuration: Anthropic describes fast mode as the same model weights on a faster inference configuration, xAI calls Grok 4.7 Fast the same model on faster infrastructure, and Z.ai and Moonshot quote absolute token rates. The second kind is queue position: Gemini's priority inference, xAI's priority processing and Mistral's priority tier promise that your request is served ahead of standard traffic under load, which lowers time to first token when the queue is long and does nothing when it is empty. OpenAI's Fast mode describes both faster speeds and more consistent latency.

That distinction decides whether a benchmark on a quiet afternoon tells you anything. A serving-configuration tier should show its speedup at any hour. A queue-priority tier shows nothing until the vendor is busy, which is also when you cannot reproduce the test. Our latency benchmark reference explains why time to first token and output rate must be reported separately, and the table above shows that vendors already choose which one to promise.

Five other things change with the flag, and each is in a row above. Batch drops out at Anthropic, Xiaomi and xAI, and Alibaba's Prime rows lose the batch and caching tags their base rows carry. The cache can reset: Anthropic's fast and standard requests do not share cached prefixes, so a fallback to standard on a long agent session pays the cache write again. Rate limits move in both directions: Anthropic gives fast mode its own limit, Gemini Priority defaults to 0.3× standard, OpenAI shares the standard limit and downgrades traffic that ramps too quickly. Availability narrows, sometimes to a single region or a pair of clients. And the premium can silently not apply: OpenAI, Google, xAI and Mistral all serve over-limit or ramping traffic at standard and bill it standard, recorded only in a response field. Anthropic's Opus 4.6 goes further and accepts the fast flag while running at standard speed and price, with no error.

Read the response, not the request

If you pay for a fast or priority tier, log the tier the vendor says it served: OpenAI's service_tier field, Anthropic's usage.speed, Gemini's x-gemini-service-tier header, xAI's and Mistral's service_tier in the response. A dashboard that reports the tier you asked for is reporting your intent, not your bill.

04 — The decisionWhen a multiple is worth paying, and when a smaller model is the cheaper answer

A speed tier is a way to buy latency without changing the answer. A smaller model is a way to buy latency by changing it. The right choice depends on which part of the latency you are paying for and whether the task can tolerate a weaker model at all.

A person is waiting, output is long, and quality cannot drop
The 2× tiers are built for this: an agent writing code or a document while someone watches. Pay the multiple on the fraction of traffic that is interactive and route the rest to standard or batch. Check the vendor's claim is output rate, not queue position, or the speedup vanishes when the vendor is quiet.
Fast tier
The bottleneck is the wait before the first token under load
That is queue priority, not a faster configuration. Gemini, xAI and Mistral sell it as such. Measure your time to first token at your busiest hour before buying; if it is fine at 3 p.m., priority buys nothing at 3 p.m.
Priority tier
The task is classification, extraction or short answers
A smaller model at standard price is usually faster than the large model at 2× and costs a tenth. Our active-parameters reference shows how much of the speed difference is architecture. Test quality on your own data first; the saving is only real if the answers hold.
Smaller model
Nobody is waiting
Batch or Flex at 0.5×. On Gemini 3.8 Flash the spread between Flex and Priority is 3.6×, which is the price of misclassifying a nightly job as interactive.
Batch or Flex
A 10× tier is on the table
Ask what unit the speed claim is in. Xiaomi's is unstated. At $4.35 and $8.70 the UltraSpeed tier costs what a frontier standard tier costs, so the comparison is not against MiMo-V2.6-Pro but against every other model at that price.
Compare across vendors

The arithmetic that settles most of these is simple. If a tier costs 2× and the vendor claims up to 2.5× output speed, then at best you get 25% more output speed per dollar and at worst, when the claim is a ceiling, you pay more. A smaller model at a fifth of the price that is twice as fast wins on both axes whenever its answers are good enough. Our active-parameters reference has the architecture side of that comparison and our price index the standard-tier prices.

05 — The gapsWhat no vendor publishes, and what we could not verify

No vendor in the table publishes a measurement method for its speed claim: not the prompt length, the output length, the batch size, the region or the hour. Anthropic comes closest by naming the metric it targets and the one it does not. Google's documentation gives a latency band rather than a number. Xiaomi gives a multiple with no unit. Until one of them publishes a method, the right way to quote any of these figures is the way this table does: in quotation marks, with the vendor's name.

Not verified and therefore not in the census: Vertex AI's own provisioned-throughput pricing; a third-party host's fast endpoint for GLM-5.3 that appears on an aggregator at 1.5× but has no vendor documentation we could read; any DeepSeek speed tier, because none exists on its pricing page; Kimi K3, which has no high-speed SKU; and any fast mode for Claude Fable 5.1 or Sonnet 5, which Anthropic's fast-mode page does not list. Alibaba's English and Chinese Prime pages disagree on which models are included, and the English Prime page and the pricing page disagree on one regional video price, so the Prime rows cite the page each figure came from.

Methodology

A census of vendor-documented service tiers with every multiple computed from published prices and every speed claim quoted as written.

What was collected
Every case found where one vendor sells the same model at two service levels for two prices: fast modes, priority tiers, high-speed SKUs. For each: base and tier prices, the multiple, the vendor's speed claim and its unit, limits that differ, batch and cache treatment, and availability. Cheaper-and-slower tiers recorded separately for contrast.
Sources
Vendor pricing pages and service-tier documentation for Anthropic, OpenAI, Google, xAI, Z.ai, Alibaba Cloud Model Studio (English and Chinese pages), Moonshot, Xiaomi and Mistral. Aggregator endpoint data was used only to confirm a tier's existence and provider, never as a price source, and is labelled where it appears.
As-of date
All pages read on September 25, 2026.
Arithmetic
Multiple = tier price ÷ base price at the same unit and the same context band, rounded to two decimals in the table. Alibaba multiples computed in CNY on the Beijing pricing page; Wan video multiples given against both list and promotional base prices.
Exclusions and limitations
Different models with "Fast" or "Turbo" in the name; third-party hosts serving another vendor's model; committed-throughput products with no published price, except Anthropic's Priority Tier, listed for completeness. No speed figure was measured by us or by any independent party we could find.
Refresh
Refreshed in place when a listed multiple changes, a tier is withdrawn, or an independent measurement of any tier is published. Gemini promotional prices double on January 1, 2027 with multiples unchanged.

06 — ConclusionThe multiple is published; the speed is a claim

What to do

Pay a speed multiple only on interactive traffic, only after measuring your own latency at your busiest hour, and only while logging which tier actually served each request

Ten of eighteen tiers cost twice the base. What that buys is written in the vendor's own unit and measured by nobody else, and it usually costs a cache, a batch discount or a region as well. Our AI transformation practice runs the latency measurement before a client buys the tier.

Digital Applied

Buy speed where a person is waiting, and nowhere else.

We measure your time to first token and output rate at your busiest hour, split traffic into interactive, standard and batch, and log the tier each request was actually served on.

Latency measurementTier routingCost reconciliation
Your next project

A speed budget that matches the bill

  • →Interactive traffic identified and routed
  • →Served tier logged per request
  • →Batch and Flex for everything unattended
Questions and answers

The questions we get about fast and priority tiers

Twice the standard price: $8 input and $40 output per million tokens against $4 and $20, as of September 25, 2026. Anthropic claims up to 2.5× higher output tokens per second. It is a research preview on the Claude API only, excludes the Batch API, and fast and standard requests do not share cached prefixes.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

AI Development

Voice AI Benchmarks and Listener Votes Disagree: 14 Models

Artificial Analysis scores realtime voice systems two ways. The arena system with the best reasoning and turn-taking scores ranks last on listener preference.

September 25, 2026 · 6 minRead
AI Development

DeepSeek V4 Flash Vision: Images, Same Price, Two Clocks

DeepSeek shipped a separate experimental V4 Flash vision model on August 21 at text prices, on a peak and off-peak clock that doubles the rate.

August 21, 2026 · 14 minRead
AI Development

Claude Sonnet 5: Near-Opus Agentic Coding at Sonnet Price

Anthropic shipped Claude Sonnet 5 on June 30, 2026. Its most agentic Sonnet yet lands near Opus 4.8 on coding and knowledge work at roughly half the price.

June 30, 2026 · 9 minRead
AI Development

How Much VRAM to Run an LLM? KV-Cache & Context Math

Calculate exactly how much VRAM an LLM needs: weight bytes by quantization plus the KV-cache formula, so you know the real context limits per GPU.

June 28, 2026 · 12 minRead
AI Development

Google AI Plans: Free vs Plus vs Pro vs Ultra 2026

Google's AI subscription tiers after I/O 2026 — AI Plus $7.99, AI Pro $19.99, AI Ultra $100 (new), AI Ultra $200 (was $250). Feature matrix and decision tree.

May 23, 2026 · 14 minRead
AI Development

Computer-Use Agents: Microsoft vs Anthropic vs Google

Microsoft GA, Anthropic public beta, and Google Gemini preview — OSWorld scores now 78% across frontier models above the ~72% human baseline. Routing guide.

May 22, 2026 · 16 minRead