AI DevelopmentPricing Tracker18 min readPublished Aug 5, 2026

Two cuts landed Jul 30 · one promo expires Aug 31 · four surfaces per model

AI API Pricing, August 2026: Cuts, Promos, and Traps

OpenAI cut GPT-5.6 Luna by 80% and Terra by 20% on July 30. Claude Sonnet 5’s introductory rate expires August 31. DeepSeek has announced an increase with no date attached. Underneath all of it, the same model carries up to four different prices depending on which surface you buy it from — and marketplace listings do not reliably show the vendor’s standard rate.

DA
Digital Applied Team
Senior strategists · Published Aug 5, 2026
PublishedAug 5, 2026
Read time18 min
SourcesVendor rate cards + marketplace API
GPT-5.6 Luna list cut
80%
effective July 30, 2026
Terra cut 20%
Sonnet 5 promo cliff
Aug 31
$2/$10 becomes $3/$15
+50% both ways
Pricing surfaces per model
4
standard · batch/flex · fast · marketplace
DeepSeek V4 Flash gap
37%
marketplace base alias under vendor list
vendor-direct is not the floor

AI API pricing in August 2026 is not a list of numbers — it is a list of numbers plus the surface each one belongs to. On July 30 OpenAI cut GPT-5.6 Luna by 80% and Terra by 20%. Claude Sonnet 5 is running an introductory rate that expires on August 31. DeepSeek has warned of an increase without naming a date. And the same model, on the same day, can be quoted at four different prices.

That last point is where budgets actually break. A model has a standard rate, a batch or flex rate at half of it, a fast rate at double it, and then a third-party marketplace listing that may equal any of those — or undercut all of them. We pulled the vendor rate cards and a marketplace models API side by side at the time of writing and found that the marketplace listing for one GPT-5.6 model matches OpenAI’s standard rate exactly while the listings for the other two match the discounted batch tier. Check one model and you will conclude the marketplace is honest. Check three and you will find the asymmetry.

This tracker covers what changed, what expires, and where the reconciliation traps sit across four model families: GPT-5.6 Sol, Terra and Luna, Claude Sonnet 5, DeepSeek V4 Flash and V4 Pro, and the Qwen3.7 tier. Every price below is labelled with the surface it came from. Where a vendor list price could not be located, we say so rather than borrowing a marketplace number and calling it a list price.

Key takeaways
  1. 01
    Two cuts landed on July 30 and they are permanent.OpenAI reduced GPT-5.6 Luna’s price by 80% and Terra’s by 20%, effective July 30, 2026. These are list changes, not promotions — there is no expiry date to plan around, only a new baseline to rebase forecasts onto.
  2. 02
    One promo expires on August 31 with a 50% step-up.Claude Sonnet 5’s introductory $2/$10 per million reverts to $3/$15 on September 1, per Anthropic’s own pricing footnote. That is a uniform 50% increase on both input and output, and the predecessor Sonnet 4.6 already sits at the reverted level.
  3. 03
    A marketplace listing is not automatically the list price.At the time of writing, the GPT-5.6 Sol listing on OpenRouter equals OpenAI’s standard rate, while the Terra and Luna listings equal the batch and flex rate instead. The discrepancy is per-model, so a single spot-check proves nothing.
  4. 04
    Vendor-direct is not always the cheapest route.DeepSeek’s own page lists V4 Flash 0731 at $0.14/$0.28 cache-miss, while OpenRouter’s provider rates for the same model read $0.09/$0.18 on the dated snapshot and $0.0882/$0.1764 on the base alias. The same inversion does not apply to V4 Pro.
  5. 05
    The cheapest tiers carry conditions, not just low rates.Qwen3.7 Flash’s headline marketplace rate applies only below 32K prompt tokens and rises to about 6.7× the headline input rate at the 256K boundary. DeepSeek has announced a significant increase with no date. Read the boundary before you build the forecast.

01The CutJuly 30: 80% off Luna, 20% off Terra.

OpenAI’s GPT-5.6 model page carries a dated update banner recording the change: on July 30, 2026, the company reduced the price of GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%. The accompanying price-performance announcement frames it the same way. Both are primary sources, and the developer platform’s rate card reflects the post-cut numbers, so the change is verifiable on two separate OpenAI-published surfaces rather than inferred from a single blog post.

The resulting standard list, per million tokens, reads $5.00 input / $30.00 output for Sol, $2.00 / $12.00 for Terra, and $0.20 / $1.20 for Luna. Cached input runs $0.50, $0.20 and $0.02 respectively; cache writes run $6.25, $2.50 and $0.25. The context window is 1,050,000 tokens across all three. If you last priced GPT-5.6 during its public GA, the Luna and Terra lines on your spreadsheet are now wrong.

Cut 80%
GPT-5.6 Luna
$0.20 / $1.20 standard · per 1M in / out

The cheap end of the family and the model that moved most. Cached input $0.02, cache write $0.25. Batch and flex land at $0.10 / $0.60; Fast mode at $0.40 / $2.40. OpenAI describes Luna as delivering performance comparable to models that were frontier-class a year ago at roughly 6 cents on the dollar per task — a vendor claim, on the vendor’s own framing.

Standard list · after July 30, 2026
Cut 20%
GPT-5.6 Terra
$2.00 / $12.00 standard · per 1M in / out

The balanced tier. Cached input $0.20, cache write $2.50. Batch and flex at $1.00 / $6.00; Fast at $4.00 / $24.00. On input, Terra’s standard rate now sits exactly level with Claude Sonnet 5’s promotional rate — a coincidence that stops being one on September 1.

Standard list · after July 30, 2026
Not named in the cut
GPT-5.6 Sol
$5.00 / $30.00 standard · per 1M in / out

OpenAI’s July 30 note names only Luna and Terra. Sol’s standard list reads $5.00 / $30.00 at the time of writing, with cached input $0.50 and cache write $6.25. Batch and flex at $2.50 / $15.00; Fast at $10.00 / $60.00.

Standard list · at the time of writing
“Starting today, GPT-5.6 Luna, our fastest and most affordable model, will cost 80% less, while GPT-5.6 Terra, our balanced model for everyday work, will cost 20% less.”— OpenAI, price-performance announcement for GPT-5.6, published on or about July 30, 2026

One structural detail travelled with the cut and is easy to miss: OpenAI’s Fast mode replaced Priority Processing on the same date, and requests tagged priority now auto-route to Fast mode. Fast is priced at twice the standard list. If a service in your stack still sets a priority flag from the old regime, it is now buying the 2× tier by default rather than the tier it was configured for. That is a silent doubling, not an error you will see in a log.

The subtler pattern in the post-cut card is how clean the ladder became. Luna, Terra and Sol now sit at $0.20, $2.00 and $5.00 on input and $1.20, $12.00 and $30.00 on output — a uniform 1 : 10 : 25 ratio in both directions. Terra is exactly 10× Luna and Sol is exactly 2.5× Terra, on input and on output alike. That means model substitution inside the GPT-5.6 family has a single multiplier regardless of how input-heavy or output-heavy your workload is, which is unusual and genuinely useful for capacity planning.

One more line item on the OpenAI card
Regional-processing endpoints — the data-residency option for OpenAI models released on or after March 5, 2026 — carry a 10% uplift over the standard table. It is a small number that compounds quietly, and it is the kind of line that never makes it into a comparison spreadsheet built from headline rates.

02Promo CliffsWhat expires, what already landed, and what has no date.

Cuts and promotions look identical on a rate card and behave completely differently in a forecast. A cut is a new baseline. A promotion is a temporary rate with a reversion attached, and the reversion is usually printed in a footnote rather than the table. Anthropic prints its footnote plainly, which makes Sonnet 5 the cleanest cliff on the board right now.

The Sonnet 5 footnote, verbatim
“Introductory pricing of $2/$10 per million input/output tokens through August 31, 2026; $3/$15 standard pricing thereafter.” Anthropic’s public pricing page, read at the time of writing.

Counting August 5 and August 31 themselves, that leaves twenty-seven days of promotional pricing. The step-up is uniform: $2.00 to $3.00 on input and $10.00 to $15.00 on output is a 50% increase in both directions. Two supporting details matter. First, Sonnet 4.6 — the prior-generation Sonnet — is already listed at $3.00 / $15.00, so the reverted level is not speculative; it is the level the predecessor already occupies. Second, the promotional prompt-cache rates of $2.50 per million on writes and $0.20 per million on reads carry the same promotional asterisk as the base rate, and the page does not print what they become afterwards. Do not assume they hold. If cache economics carry your Sonnet 5 workload, that is an unpriced variable in September.

Anthropic also advertises 50% savings for batch processing as a blanket line beside the model cards. Applied to the promotional rate, that puts batch Sonnet 5 at roughly $1.00 / $5.00 — which is exactly what the marketplace batch variant lists at, so the two surfaces agree. We covered the model itself when Sonnet 5 launched at near-Opus capability on a mid-tier price; the pricing question then was whether the intro rate was the real rate. The footnote answers it.

Promo-cliff calendar assembled from OpenAI, Anthropic and DeepSeek pricing pages, grouped into changes that already landed, changes with a dated cliff, and changes that are announced but undated or unconfirmed, with the budget action implied by each.
ChangeWhere it standsRate before, rate afterBudget action
Already landed — permanent list changes, no expiry
GPT-5.6 Luna standard listEffective July 30, 2026Cut by 80%; the list now reads $0.20 / $1.20 per 1MRebase the forecast. There is no reversion date to model.
GPT-5.6 Terra standard listEffective July 30, 2026Cut by 20%; the list now reads $2.00 / $12.00 per 1MRebase the forecast. Re-test Terra against Sonnet 5 after August 31.
Priority Processing becomes Fast modeReplaced July 30, 2026Fast is 2× the standard list; requests tagged priority auto-route to itAudit for stray priority tags before they silently double a line item.
Dated cliff — a reversion you can put in a calendar
Claude Sonnet 5 introductory pricingRuns through August 31, 2026 — twenty-seven days remain, counting August 5 and August 31$2.00 / $10.00 becomes $3.00 / $15.00 per 1M — a uniform 50% increaseModel the September number now; Sonnet 4.6 already sits at $3.00 / $15.00.
Sonnet 5 prompt-cache ratesCarry the same promotional asterisk as the base rate$2.50 write / $0.20 read per 1M; no post-promotion figure is publishedTreat September cache economics as unpriced rather than unchanged.
Undated or unconfirmed — watch items, not planning inputs
DeepSeek API pricingAnnounced on the vendor pricing page; no date givenCurrent vendor list, then a significant increase described in general termsTreat the current list as provisional and keep a second route warm.
Marketplace promotional banner on Terra and LunaObserved promotional language; we could not re-confirm it in the static page source at the time of writingProvider rates read $1.00 / $6.00 and $0.10 / $0.60 — the batch and flex figures, unchangedDo not stack a banner discount on top of the batch/flex gap.

DeepSeek is the loosest entry on that calendar, and deliberately so. Its own pricing page states that it plans to raise overall pricing for its API services in the near future, that a significant increase is expected, and that the specific plan will be subject to official notice. No date accompanies the statement. An announced-but-undated increase is worse for planning than a dated one, because it cannot be modelled and cannot be ignored. The only sane response is a routing fallback rather than a forecast line.

03Four SurfacesOne model, four prices, all of them real.

The single most common budgeting mistake in 2026 is treating a model name as a price. It is not. A model name plus a surface is a price. Across the vendors in this tracker there are four surfaces in regular use, and the spread between the cheapest and most expensive of them, for the same model on the same day, reaches 4×.

Surface one
Standard
1.0×

The synchronous, published list rate — the number on the vendor’s pricing page and the number most comparisons quote. It is the reference point for every other surface, and it is also the one most likely to be quietly out of date on a third-party page.

The baseline everything else multiplies
Surface two
Batch and flex
0.5×

GPT-5.6’s batch and flex tiers are both 50% of the standard list across Sol, Terra and Luna. Anthropic advertises 50% batch savings as a blanket line. The trade is latency and scheduling flexibility, not capability — which is why it is the right default for anything asynchronous.

Half of standard, same model
Surface three
Fast
2.0×

OpenAI’s Fast mode replaced Priority Processing on July 30 and prices at twice the standard list. The vendor describes Fast on Sol as delivering up to 2.5× faster speeds than standard processing at twice the price, with no change in intelligence. It is a latency purchase, not a quality one.

Double standard · 4× the batch tier
Surface four
Marketplace provider rate
?

A gateway or marketplace lists its own provider rate. Sometimes it equals the vendor standard, sometimes it equals the vendor batch tier, and sometimes it undercuts the vendor list entirely. It is the only surface with no predictable relationship to the others — which is exactly what makes it the trap.

No fixed relationship to the vendor card

There is a fifth axis that cuts across all four: prompt caching. And here something quietly interesting has happened. OpenAI prices cached input on GPT-5.6 at one tenth of the input rate and cache writes at 1.25× the input rate — uniformly across Sol, Terra and Luna. Anthropic prices Sonnet 5’s promotional cache reads at $0.20 against a $2.00 input rate, and cache writes at $2.50. That is also one tenth on reads and 1.25× on writes. Two vendors, independently, landed on identical cache multipliers. Pricing structure is converging into a shared grammar even while headline rates diverge — which is good news for anyone modelling across vendors, because the cache maths now ports.

If you are routing across several of these surfaces programmatically rather than picking one, the reconciliation problem becomes an infrastructure problem, and the shape of it is covered in our LLM gateway architecture reference.

04ReconciliationThe grid no single vendor page will ever show you.

The table below puts the vendor standard rate, the vendor batch or flex rate, and the marketplace provider rate for the same model on the same row, then names what the marketplace figure actually corresponds to. Vendor cells come from OpenAI’s developer-platform pricing page, Anthropic’s public pricing page and DeepSeek’s API pricing page; marketplace cells come from OpenRouter’s public models API. Every one was read at the time of writing. Where a vendor list price could not be located, the cell says so rather than borrowing a marketplace figure.

Pricing-surface reconciliation grid for GPT-5.6 Sol, Terra and Luna, Claude Sonnet 5 and Sonnet 4.6, DeepSeek V4 Flash and V4 Pro, and the Qwen3.7 family, showing vendor standard rates, vendor batch or flex rates, OpenRouter provider rates, and what each marketplace figure corresponds to.
ModelVendor standard (in / out per 1M)Vendor batch or flexMarketplace provider rateWhat the marketplace figure is
OpenAI — GPT-5.6, standard list after the July 30 cut
GPT-5.6 Sol$5.00 / $30.00$2.50 / $15.00$5.00 / $30.00Equals the vendor standard rate exactly
GPT-5.6 Terra$2.00 / $12.00$1.00 / $6.00$1.00 / $6.00Equals the batch and flex rate — not the standard rate
GPT-5.6 Luna$0.20 / $1.20$0.10 / $0.60$0.10 / $0.60Equals the batch and flex rate — not the standard rate
Anthropic — Claude
Claude Sonnet 5$2.00 / $10.00 through Aug 31; $3.00 / $15.00 thereafter$1.00 / $5.00 during the promo (50% batch line)$2.00 / $10.00Equals the vendor promotional rate, not a separate discount
Claude Sonnet 4.6$3.00 / $15.00not checked this passnot checked this passReference row — the level Sonnet 5 reverts to on September 1
DeepSeek — vendor list quoted at the cache-miss rate
DeepSeek V4 Flash, 0731 snapshot$0.14 / $0.28 cache-miss; $0.0028 cache-hit inputnot published$0.09 / $0.18 on the -0731 and -latest snapshots; $0.0882 / $0.1764 on the base aliasUndercuts the vendor list by about 36% on the dated snapshots and 37% on the base alias
DeepSeek V4 Pro$0.435 / $0.87 cache-miss; $0.003625 cache-hit inputnot published$0.435 / $0.87Equals the vendor list exactly — the inversion is Flash-only
Alibaba Qwen3.7 — no vendor list price located this pass
Qwen3.7 Flashnot locatednot located$0.03 / $0.13 under 32K prompt; $0.10 / $0.40 at 32K and above; $0.20 / $0.80 at 256K and aboveMarketplace provider rate only, tiered by prompt length
Qwen3.7 Plusnot locatednot located$0.32 / $1.28 under 256K prompt; $0.96 / $3.84 at 256K and aboveMarketplace provider rate only, tiered by prompt length
Qwen3.7 Maxnot locatednot located$1.475 / $4.425Marketplace provider rate only, no tiering shown

Read the OpenAI block first, because it contains the finding. The Sol listing matches OpenAI’s standard rate to the cent. The Terra and Luna listings match the batch and flex rate instead. Both facts are true simultaneously, on the same marketplace, on the same day. The practical consequence is that a spot-check does not generalise: verify Sol, conclude the marketplace mirrors vendor list, and you will budget Luna at $0.10 per million input when a direct standard call costs $0.20 — a 2× miss on the line you probably run most.

Note also what this is not. It is not evidence that the marketplace is cheaper, and it is not evidence that it is misleading. A gateway listing a provider rate is stating what it charges, not what the upstream vendor charges. The error is on the reader’s side, and it is entirely avoidable by writing the surface into the cell next to the number. That single habit is what this whole table is arguing for.

What we could not confirm
Promotional banner language offering a discount on Terra and Luna has been observed on the marketplace surface, but we could not re-confirm it in the static page source at the time of writing — it is likely client-rendered. The underlying provider rates we did read are unchanged batch and flex figures. Treat any such banner as describing the existing gap between the batch tier and the standard list, not as a second discount stacked on top of it, unless you can verify it live in your own account.

05Vendor-Direct TrapDeepSeek: when the vendor is not the cheapest route.

The assumption that buying direct from the model’s creator is the floor price is intuitive, widely held, and — for DeepSeek V4 Flash at the time of writing — wrong. DeepSeek’s own API pricing page lists a single Flash model, versioned DeepSeek-V4-Flash-0731, at $0.14 per million input on a cache miss and $0.28 per million output, with a cache-hit input rate of $0.0028. OpenRouter’s provider rates for the same model version read lower on every alias: $0.09 / $0.18 on both the -0731 and -latest snapshots, and $0.0882 / $0.1764 on the base alias.

Recomputed from those two sourced pairs, the dated snapshots sit about 36% below the vendor cache-miss list on both input and output, and the base alias about 37% below. We are not going to speculate about why — DeepSeek has published no statement on the difference, and third-party inference providers on a marketplace set their own economics. The reportable fact is the gap and the fact that it is model-specific, because V4 Pro’s marketplace rate of $0.435 / $0.87 matches DeepSeek’s own list exactly.

DeepSeek V4 Flash · input rate per 1M tokens, by surface

Sources: DeepSeek API pricing page; OpenRouter public models API — input rates read at the time of writing
Vendor list · cache missDeepSeek’s own pricing page, V4-Flash-0731
$0.14
Marketplace · -0731 and -latestProvider rate for the dated snapshot aliases
$0.09
Marketplace · base aliasProvider rate for deepseek-v4-flash without a snapshot
$0.0882
Vendor list · cache hitDeepSeek’s own pricing page, cached input only
$0.0028

The cache-hit bar deserves its own note, because it is the largest single lever on that chart and it is not a discount you negotiate — it is one you engineer. DeepSeek’s cache-hit input rate for V4 Flash is exactly one fiftieth of its cache-miss rate. For V4 Pro the ratio is steeper still: $0.003625 against $0.435 is exactly one hundred and twentieth. A workload with high prompt-prefix reuse changes tier without changing vendor. If you are running V4 Flash in production, the release detail behind that 0731 version string is in our write-up of the 0731 checkpoint’s official release.

Layer the announced increase on top and DeepSeek becomes a routing decision rather than a default. The vendor has said an increase is coming and has not said when or by how much. That is not a reason to move off it — the rates are genuinely low — but it is a reason to keep the second route configured and tested rather than theoretical. Undated risk is best answered with switchable infrastructure, not with a scenario in a spreadsheet.

06Tiered RatesQwen3.7: the headline rate has a boundary.

Two cautions before any Qwen3.7 number. First, no standalone vendor list price for the Qwen3.7 family was locatable on Alibaba’s public model-endpoint listing during this pass — the page carries endpoint references without a pricing table. Every Qwen3.7 figure below is therefore a marketplace provider rate and nothing more, and it should never be quoted as a vendor list price. Second, unlike every other model in this tracker, Qwen3.7 rates are tiered by prompt length, so the headline number is a floor that applies only inside a bound.

Prompts under 32K
Qwen3.7 Flash base tier
0.03/1M in

Output at $0.13 per million. Cache read $0.006 and cache write $0.038 at this tier. The context window is 1M tokens, but the base rate stops applying long before you fill it — the boundary is the prompt, not the window.

Marketplace provider rate
Prompts 32K and above
First escalation
0.10/1M in

Output rises to $0.40 per million. That is 3.3× the base input rate and 3.1× the base output rate — crossed automatically by any moderately sized document, transcript or retrieved context bundle, with no change on your side.

Marketplace provider rate
Prompts 256K and above
Second escalation
0.20/1M in

Output rises to $0.80 per million. Input is now 6.7× the headline rate and output 6.2×. A long-context workload priced off the $0.03 figure will land materially over budget, and nothing in the model name signals it.

Marketplace provider rate

The same shape appears one tier up. Qwen3.7 Plus lists at $0.32 / $1.28 below 256K prompt tokens and $0.96 / $3.84 above it — an exact 3× step on both input and output. Qwen3.7 Max is the outlier: $1.475 / $4.425 flat, with no tiering shown. The pattern across the family is that the cheaper the model, the more aggressively its price depends on how you use it, which is a reasonable design and a terrible thing to discover after a month of traffic.

The broader point is that the sub-$0.30 tier now has internal structure. If your interest is what that tier actually buys you at volume rather than how it is quoted, we worked the bulk economics separately in what the sub-$0.30 tier unlocks at bulk volume. This tracker stays on the quoting problem.

07Worked CostOne job, priced on every surface.

Rates per million are hard to feel. So here is a single fixed reference job — 2,000,000 input tokens and 200,000 output tokens, assuming no cache hits — priced on every surface each model exposes. The arithmetic per cell is simply input tokens multiplied by the input rate plus output tokens multiplied by the output rate, at per-million list rates. No quality weighting, no vendor efficiency claims folded in. Totals under a dollar are shown to three decimals.

Cost of a fixed reference job of two million input tokens and two hundred thousand output tokens, recomputed from published rates for each pricing surface of GPT-5.6 Sol, Terra and Luna, Claude Sonnet 5, DeepSeek V4 Flash and Qwen3.7 Flash, with each surface expressed as a multiple of that model’s cheapest listed surface.
ModelSurfaceRate (in / out per 1M)Job costMultiple of that model’s cheapest surface
OpenAI GPT-5.6 — standard list after the July 30 cut
SolBatch or flex$2.50 / $15.00$8.001.0×
SolStandard$5.00 / $30.00$16.002.0×
SolFast$10.00 / $60.00$32.004.0×
TerraBatch or flex$1.00 / $6.00$3.201.0×
TerraStandard$2.00 / $12.00$6.402.0×
TerraFast$4.00 / $24.00$12.804.0×
LunaBatch or flex$0.10 / $0.60$0.3201.0×
LunaStandard$0.20 / $1.20$0.6402.0×
LunaFast$0.40 / $2.40$1.284.0×
Anthropic Claude Sonnet 5 — promo through Aug 31, reversion from Sep 1
Sonnet 5Batch, promo window$1.00 / $5.00$3.001.0×
Sonnet 5Standard, through Aug 31$2.00 / $10.00$6.002.0×
Sonnet 5Standard, from Sep 1$3.00 / $15.00$9.003.0×
DeepSeek V4 Flash, 0731 snapshot — marketplace provider rates and the vendor cache-miss list
V4 FlashMarketplace, base alias$0.0882 / $0.1764$0.2121.00×
V4 FlashMarketplace, -0731 and -latest$0.09 / $0.18$0.2161.02×
V4 FlashVendor list, cache miss$0.14 / $0.28$0.3361.59×
Alibaba Qwen3.7 Flash — marketplace provider rate, tiered by prompt length
Qwen3.7 FlashPrompts under 32K$0.03 / $0.13$0.0861.00×
Qwen3.7 FlashPrompts 32K and above$0.10 / $0.40$0.2803.26×
Qwen3.7 FlashPrompts 256K and above$0.20 / $0.80$0.5606.51×

Three things fall out of the arithmetic that no vendor page states. First, the GPT-5.6 spread is exactly 4× from batch to Fast on all three models, because the two multipliers are uniform — so the surface decision is worth more than the model decision inside a single tier band. Second, Sonnet 5 currently undercuts Terra’s standard rate on this job mix by 6.25% ($6.00 against $6.40), and on September 1 the same comparison flips to Sonnet 5 costing about 41% more ($9.00 against $6.40). One calendar date reverses the ranking of two models with no change to either product.

Third, and least expected: OpenAI’s batch tier for Luna now sits marginally below DeepSeek’s own vendor list for V4 Flash on this job — $0.320 against $0.336, about 5% cheaper. The cheap end of the Western frontier stack has caught the cheap end of the open-weight stack at the vendor’s own quoted price. It has not caught the marketplace provider rates for the same DeepSeek model, which come in at $0.212 and are still about 34% below Luna’s batch tier. Which route is cheapest now depends entirely on which surface you are allowed to buy from.

Project that forward and the direction of travel is fairly clear. If a permanent 80% cut on a frontier vendor’s cheapest model lands it within a few cents of an open-weight vendor’s list price, then price alone stops being the reason to run open weights, and sovereignty, latency and control take over as the deciding factors. Meanwhile the gap between a vendor’s list and a marketplace’s provider rate is becoming the number that actually determines your bill — which is a procurement problem, not a model-selection problem, and most teams are still staffing it as the latter.

08Operating RulesFour habits that make a pricing sheet survive a month.

None of this requires new tooling. It requires four disciplines applied to the document you already keep, plus a willingness to write “not located” in a cell instead of filling it with a number from a different surface.

Budget lines
Price by surface, never by model name

Every rate in your sheet gets a surface label and a date: standard, batch or flex, fast, or marketplace provider rate. An unlabelled number is not a price, it is a rumour. This one change would have caught every trap in this article.

Label the surface
Marketplace routing
Reconcile every listing, not a sample

The GPT-5.6 asymmetry proves a sample is worthless: Sol matches vendor standard while Terra and Luna match the batch tier. Reconcile each model you actually call, and re-run it whenever a vendor announces a change — a cut on the vendor side does not guarantee the marketplace listing moved with it.

Check per model
Promo exposure
Model the cliff before it arrives

Sonnet 5 steps up 50% on September 1 and its cache rates have no published post-promotion figure. Put the reverted rate in the forecast now, flag the cache line as unpriced, and decide in August whether the September number still wins on your workload rather than discovering it in an invoice.

Forecast the reversion
Cheap-tier conditions
Read the boundary, not the headline

Qwen3.7 Flash escalates about 6.7× on input past 256K prompt tokens, and DeepSeek has announced an undated increase. Cheap tiers are cheap under conditions. Bound your prompt sizes deliberately, and keep a tested fallback route for any vendor that has told you a rise is coming.

Bound the inputs

If you want a longer baseline than one month, the Q2 2026 edition of this tracker holds the earlier data points — read as a historical snapshot, not as current rates, since most of those figures have since moved. Teams that want the reconciliation and forecasting built into their reporting rather than maintained by hand usually start with our analytics and measurement work, and the routing side sits inside AI and digital transformation engagements.

09ConclusionThe number is fine. The label is what is missing.

Where AI API pricing stands, August 2026

A model does not have a price. A model plus a surface has a price.

August 2026 opened with one large permanent cut, one promotion on a countdown, and one undated warning. GPT-5.6 Luna is 80% cheaper and Terra 20% cheaper than they were on July 29, and neither change expires. Claude Sonnet 5 costs $2 and $10 per million through August 31 and then $3 and $15. DeepSeek has said an increase is coming and has not said when. Those three facts are the easy part.

The hard part is that the same model can be quoted at four prices, and the marketplace listing has no fixed relationship to the vendor card. One GPT-5.6 model’s listing matches the vendor standard rate exactly; the other two match the discounted batch tier. The marketplace provider rates for DeepSeek V4 Flash sit about a third below DeepSeek’s own list for the identical model version, while its V4 Pro list matches to the cent. Qwen3.7’s cheapest rate applies only under 32K prompt tokens. None of that is hidden — it is just never assembled in one place.

So the discipline is small and the payoff is not. Write the surface and the date next to every rate you record. Reconcile each model you actually call rather than a sample. Put the reversion in the forecast before the cliff arrives. Do that and a pricing sheet survives a month like July 30; skip it and the cheapest-looking number on the page is the one most likely to be wrong.

Get your AI spend under control

Most AI overspend is not the model choice. It is the surface nobody labelled.

We reconcile vendor rate cards against gateway listings, model promo reversions before they land, and route workloads to the surface that actually wins — so the number in your forecast is the number on your invoice.

Free consultationExpert guidanceTailored solutions
What we work on

AI cost and routing engagements

  • Surface-labelled rate cards across every vendor you call
  • Gateway-versus-vendor reconciliation on the models you use
  • Promo-cliff forecasting and reversion modelling
  • Batch, flex and cache routing for asynchronous workloads
  • Fallback routes for vendors with announced increases
FAQ · AI API pricing, August 2026

The pricing questions we get every week.

Yes. OpenAI reduced the price of GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%, effective July 30, 2026. The change is recorded as a dated update banner on OpenAI's GPT-5.6 model page and restated in the company's price-performance announcement, and the developer platform's rate card reflects the post-cut figures — so it is verifiable on two separate OpenAI-published surfaces. The resulting standard list, per million tokens, reads $5.00 input and $30.00 output for Sol, $2.00 and $12.00 for Terra, and $0.20 and $1.20 for Luna. These are permanent list changes rather than promotions, so there is no expiry date to plan around. OpenAI's July 30 note names only Luna and Terra; Sol's standard rate reads $5.00 / $30.00 at the time of writing.
Related dispatches

Continue exploring AI cost strategy.

AI Development

LLM Gateway Architecture: 2026 Engineering Reference

The LLM gateway is now critical AI infrastructure. Compare LiteLLM, Portkey, Cloudflare, Vercel, and OpenRouter on caching, routing, and build-vs-buy economics.

June 3, 2026 · 11 minRead
AI Development

OpenRouter Fusion: Multi-Model AI Response Synthesis

OpenRouter Fusion sends queries to multiple AI models, analyzes outputs, and fuses optimal results. Deep Research agents preferred Fusion to their own outputs.

April 1, 2026 · 14 minRead
AI Development

DeepSeek V4 Flash: A Bulk AI Workload Playbook for 2026

Worked cost math for classification, extraction, dedup and translation at million-row scale, plus the benchmark evidence on where cheap models should not go.

August 1, 2026 · 15 minRead
AI Development

DeepSeek V4 Flash 0731: Official Release, Agent Benchmarks

DeepSeek V4 Flash exits preview into public beta as the 0731 checkpoint: same 284B architecture, vendor-stated agent benchmarks, no weights posted yet.

July 31, 2026 · 10 minRead
AI Development

AI Video Generation 2026: Omni vs Sora vs Veo 3 Compared

Gemini Omni, OpenAI Sora 2, and Google Veo 3.1 compared for video — quality, per-second cost spread of 17x, and the September 24 Sora API sunset clock.

May 22, 2026 · 15 minRead
AI Development

Google Intelligent Eyewear: Gemini AI Glasses Fall 2026

Google announces Gemini-powered smart glasses with Samsung, Gentle Monster, and Warby Parker at I/O. Audio glasses ship fall 2026; display tier TBD.

May 20, 2026 · 18 minRead