AI API pricing in August 2026 is not a list of numbers — it is a list of numbers plus the surface each one belongs to. On July 30 OpenAI cut GPT-5.6 Luna by 80% and Terra by 20%. Claude Sonnet 5 is running an introductory rate that expires on August 31. DeepSeek has warned of an increase without naming a date. And the same model, on the same day, can be quoted at four different prices.
That last point is where budgets actually break. A model has a standard rate, a batch or flex rate at half of it, a fast rate at double it, and then a third-party marketplace listing that may equal any of those — or undercut all of them. We pulled the vendor rate cards and a marketplace models API side by side at the time of writing and found that the marketplace listing for one GPT-5.6 model matches OpenAI’s standard rate exactly while the listings for the other two match the discounted batch tier. Check one model and you will conclude the marketplace is honest. Check three and you will find the asymmetry.
This tracker covers what changed, what expires, and where the reconciliation traps sit across four model families: GPT-5.6 Sol, Terra and Luna, Claude Sonnet 5, DeepSeek V4 Flash and V4 Pro, and the Qwen3.7 tier. Every price below is labelled with the surface it came from. Where a vendor list price could not be located, we say so rather than borrowing a marketplace number and calling it a list price.
- 01Two cuts landed on July 30 and they are permanent.OpenAI reduced GPT-5.6 Luna’s price by 80% and Terra’s by 20%, effective July 30, 2026. These are list changes, not promotions — there is no expiry date to plan around, only a new baseline to rebase forecasts onto.
- 02One promo expires on August 31 with a 50% step-up.Claude Sonnet 5’s introductory $2/$10 per million reverts to $3/$15 on September 1, per Anthropic’s own pricing footnote. That is a uniform 50% increase on both input and output, and the predecessor Sonnet 4.6 already sits at the reverted level.
- 03A marketplace listing is not automatically the list price.At the time of writing, the GPT-5.6 Sol listing on OpenRouter equals OpenAI’s standard rate, while the Terra and Luna listings equal the batch and flex rate instead. The discrepancy is per-model, so a single spot-check proves nothing.
- 04Vendor-direct is not always the cheapest route.DeepSeek’s own page lists V4 Flash 0731 at $0.14/$0.28 cache-miss, while OpenRouter’s provider rates for the same model read $0.09/$0.18 on the dated snapshot and $0.0882/$0.1764 on the base alias. The same inversion does not apply to V4 Pro.
- 05The cheapest tiers carry conditions, not just low rates.Qwen3.7 Flash’s headline marketplace rate applies only below 32K prompt tokens and rises to about 6.7× the headline input rate at the 256K boundary. DeepSeek has announced a significant increase with no date. Read the boundary before you build the forecast.
01 — The CutJuly 30: 80% off Luna, 20% off Terra.
OpenAI’s GPT-5.6 model page carries a dated update banner recording the change: on July 30, 2026, the company reduced the price of GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%. The accompanying price-performance announcement frames it the same way. Both are primary sources, and the developer platform’s rate card reflects the post-cut numbers, so the change is verifiable on two separate OpenAI-published surfaces rather than inferred from a single blog post.
The resulting standard list, per million tokens, reads $5.00 input / $30.00 output for Sol, $2.00 / $12.00 for Terra, and $0.20 / $1.20 for Luna. Cached input runs $0.50, $0.20 and $0.02 respectively; cache writes run $6.25, $2.50 and $0.25. The context window is 1,050,000 tokens across all three. If you last priced GPT-5.6 during its public GA, the Luna and Terra lines on your spreadsheet are now wrong.
GPT-5.6 Luna
The cheap end of the family and the model that moved most. Cached input $0.02, cache write $0.25. Batch and flex land at $0.10 / $0.60; Fast mode at $0.40 / $2.40. OpenAI describes Luna as delivering performance comparable to models that were frontier-class a year ago at roughly 6 cents on the dollar per task — a vendor claim, on the vendor’s own framing.
GPT-5.6 Terra
The balanced tier. Cached input $0.20, cache write $2.50. Batch and flex at $1.00 / $6.00; Fast at $4.00 / $24.00. On input, Terra’s standard rate now sits exactly level with Claude Sonnet 5’s promotional rate — a coincidence that stops being one on September 1.
GPT-5.6 Sol
OpenAI’s July 30 note names only Luna and Terra. Sol’s standard list reads $5.00 / $30.00 at the time of writing, with cached input $0.50 and cache write $6.25. Batch and flex at $2.50 / $15.00; Fast at $10.00 / $60.00.
“Starting today, GPT-5.6 Luna, our fastest and most affordable model, will cost 80% less, while GPT-5.6 Terra, our balanced model for everyday work, will cost 20% less.”— OpenAI, price-performance announcement for GPT-5.6, published on or about July 30, 2026
One structural detail travelled with the cut and is easy to miss: OpenAI’s Fast mode replaced Priority Processing on the same date, and requests tagged priority now auto-route to Fast mode. Fast is priced at twice the standard list. If a service in your stack still sets a priority flag from the old regime, it is now buying the 2× tier by default rather than the tier it was configured for. That is a silent doubling, not an error you will see in a log.
The subtler pattern in the post-cut card is how clean the ladder became. Luna, Terra and Sol now sit at $0.20, $2.00 and $5.00 on input and $1.20, $12.00 and $30.00 on output — a uniform 1 : 10 : 25 ratio in both directions. Terra is exactly 10× Luna and Sol is exactly 2.5× Terra, on input and on output alike. That means model substitution inside the GPT-5.6 family has a single multiplier regardless of how input-heavy or output-heavy your workload is, which is unusual and genuinely useful for capacity planning.
02 — Promo CliffsWhat expires, what already landed, and what has no date.
Cuts and promotions look identical on a rate card and behave completely differently in a forecast. A cut is a new baseline. A promotion is a temporary rate with a reversion attached, and the reversion is usually printed in a footnote rather than the table. Anthropic prints its footnote plainly, which makes Sonnet 5 the cleanest cliff on the board right now.
Counting August 5 and August 31 themselves, that leaves twenty-seven days of promotional pricing. The step-up is uniform: $2.00 to $3.00 on input and $10.00 to $15.00 on output is a 50% increase in both directions. Two supporting details matter. First, Sonnet 4.6 — the prior-generation Sonnet — is already listed at $3.00 / $15.00, so the reverted level is not speculative; it is the level the predecessor already occupies. Second, the promotional prompt-cache rates of $2.50 per million on writes and $0.20 per million on reads carry the same promotional asterisk as the base rate, and the page does not print what they become afterwards. Do not assume they hold. If cache economics carry your Sonnet 5 workload, that is an unpriced variable in September.
Anthropic also advertises 50% savings for batch processing as a blanket line beside the model cards. Applied to the promotional rate, that puts batch Sonnet 5 at roughly $1.00 / $5.00 — which is exactly what the marketplace batch variant lists at, so the two surfaces agree. We covered the model itself when Sonnet 5 launched at near-Opus capability on a mid-tier price; the pricing question then was whether the intro rate was the real rate. The footnote answers it.
| Change | Where it stands | Rate before, rate after | Budget action |
|---|---|---|---|
| Already landed — permanent list changes, no expiry | |||
| GPT-5.6 Luna standard list | Effective July 30, 2026 | Cut by 80%; the list now reads $0.20 / $1.20 per 1M | Rebase the forecast. There is no reversion date to model. |
| GPT-5.6 Terra standard list | Effective July 30, 2026 | Cut by 20%; the list now reads $2.00 / $12.00 per 1M | Rebase the forecast. Re-test Terra against Sonnet 5 after August 31. |
| Priority Processing becomes Fast mode | Replaced July 30, 2026 | Fast is 2× the standard list; requests tagged priority auto-route to it | Audit for stray priority tags before they silently double a line item. |
| Dated cliff — a reversion you can put in a calendar | |||
| Claude Sonnet 5 introductory pricing | Runs through August 31, 2026 — twenty-seven days remain, counting August 5 and August 31 | $2.00 / $10.00 becomes $3.00 / $15.00 per 1M — a uniform 50% increase | Model the September number now; Sonnet 4.6 already sits at $3.00 / $15.00. |
| Sonnet 5 prompt-cache rates | Carry the same promotional asterisk as the base rate | $2.50 write / $0.20 read per 1M; no post-promotion figure is published | Treat September cache economics as unpriced rather than unchanged. |
| Undated or unconfirmed — watch items, not planning inputs | |||
| DeepSeek API pricing | Announced on the vendor pricing page; no date given | Current vendor list, then a significant increase described in general terms | Treat the current list as provisional and keep a second route warm. |
| Marketplace promotional banner on Terra and Luna | Observed promotional language; we could not re-confirm it in the static page source at the time of writing | Provider rates read $1.00 / $6.00 and $0.10 / $0.60 — the batch and flex figures, unchanged | Do not stack a banner discount on top of the batch/flex gap. |
DeepSeek is the loosest entry on that calendar, and deliberately so. Its own pricing page states that it plans to raise overall pricing for its API services in the near future, that a significant increase is expected, and that the specific plan will be subject to official notice. No date accompanies the statement. An announced-but-undated increase is worse for planning than a dated one, because it cannot be modelled and cannot be ignored. The only sane response is a routing fallback rather than a forecast line.
03 — Four SurfacesOne model, four prices, all of them real.
The single most common budgeting mistake in 2026 is treating a model name as a price. It is not. A model name plus a surface is a price. Across the vendors in this tracker there are four surfaces in regular use, and the spread between the cheapest and most expensive of them, for the same model on the same day, reaches 4×.
Standard
The synchronous, published list rate — the number on the vendor’s pricing page and the number most comparisons quote. It is the reference point for every other surface, and it is also the one most likely to be quietly out of date on a third-party page.
Batch and flex
GPT-5.6’s batch and flex tiers are both 50% of the standard list across Sol, Terra and Luna. Anthropic advertises 50% batch savings as a blanket line. The trade is latency and scheduling flexibility, not capability — which is why it is the right default for anything asynchronous.
Fast
OpenAI’s Fast mode replaced Priority Processing on July 30 and prices at twice the standard list. The vendor describes Fast on Sol as delivering up to 2.5× faster speeds than standard processing at twice the price, with no change in intelligence. It is a latency purchase, not a quality one.
Marketplace provider rate
A gateway or marketplace lists its own provider rate. Sometimes it equals the vendor standard, sometimes it equals the vendor batch tier, and sometimes it undercuts the vendor list entirely. It is the only surface with no predictable relationship to the others — which is exactly what makes it the trap.
There is a fifth axis that cuts across all four: prompt caching. And here something quietly interesting has happened. OpenAI prices cached input on GPT-5.6 at one tenth of the input rate and cache writes at 1.25× the input rate — uniformly across Sol, Terra and Luna. Anthropic prices Sonnet 5’s promotional cache reads at $0.20 against a $2.00 input rate, and cache writes at $2.50. That is also one tenth on reads and 1.25× on writes. Two vendors, independently, landed on identical cache multipliers. Pricing structure is converging into a shared grammar even while headline rates diverge — which is good news for anyone modelling across vendors, because the cache maths now ports.
If you are routing across several of these surfaces programmatically rather than picking one, the reconciliation problem becomes an infrastructure problem, and the shape of it is covered in our LLM gateway architecture reference.
04 — ReconciliationThe grid no single vendor page will ever show you.
The table below puts the vendor standard rate, the vendor batch or flex rate, and the marketplace provider rate for the same model on the same row, then names what the marketplace figure actually corresponds to. Vendor cells come from OpenAI’s developer-platform pricing page, Anthropic’s public pricing page and DeepSeek’s API pricing page; marketplace cells come from OpenRouter’s public models API. Every one was read at the time of writing. Where a vendor list price could not be located, the cell says so rather than borrowing a marketplace figure.
| Model | Vendor standard (in / out per 1M) | Vendor batch or flex | Marketplace provider rate | What the marketplace figure is |
|---|---|---|---|---|
| OpenAI — GPT-5.6, standard list after the July 30 cut | ||||
| GPT-5.6 Sol | $5.00 / $30.00 | $2.50 / $15.00 | $5.00 / $30.00 | Equals the vendor standard rate exactly |
| GPT-5.6 Terra | $2.00 / $12.00 | $1.00 / $6.00 | $1.00 / $6.00 | Equals the batch and flex rate — not the standard rate |
| GPT-5.6 Luna | $0.20 / $1.20 | $0.10 / $0.60 | $0.10 / $0.60 | Equals the batch and flex rate — not the standard rate |
| Anthropic — Claude | ||||
| Claude Sonnet 5 | $2.00 / $10.00 through Aug 31; $3.00 / $15.00 thereafter | $1.00 / $5.00 during the promo (50% batch line) | $2.00 / $10.00 | Equals the vendor promotional rate, not a separate discount |
| Claude Sonnet 4.6 | $3.00 / $15.00 | not checked this pass | not checked this pass | Reference row — the level Sonnet 5 reverts to on September 1 |
| DeepSeek — vendor list quoted at the cache-miss rate | ||||
| DeepSeek V4 Flash, 0731 snapshot | $0.14 / $0.28 cache-miss; $0.0028 cache-hit input | not published | $0.09 / $0.18 on the -0731 and -latest snapshots; $0.0882 / $0.1764 on the base alias | Undercuts the vendor list by about 36% on the dated snapshots and 37% on the base alias |
| DeepSeek V4 Pro | $0.435 / $0.87 cache-miss; $0.003625 cache-hit input | not published | $0.435 / $0.87 | Equals the vendor list exactly — the inversion is Flash-only |
| Alibaba Qwen3.7 — no vendor list price located this pass | ||||
| Qwen3.7 Flash | not located | not located | $0.03 / $0.13 under 32K prompt; $0.10 / $0.40 at 32K and above; $0.20 / $0.80 at 256K and above | Marketplace provider rate only, tiered by prompt length |
| Qwen3.7 Plus | not located | not located | $0.32 / $1.28 under 256K prompt; $0.96 / $3.84 at 256K and above | Marketplace provider rate only, tiered by prompt length |
| Qwen3.7 Max | not located | not located | $1.475 / $4.425 | Marketplace provider rate only, no tiering shown |
Read the OpenAI block first, because it contains the finding. The Sol listing matches OpenAI’s standard rate to the cent. The Terra and Luna listings match the batch and flex rate instead. Both facts are true simultaneously, on the same marketplace, on the same day. The practical consequence is that a spot-check does not generalise: verify Sol, conclude the marketplace mirrors vendor list, and you will budget Luna at $0.10 per million input when a direct standard call costs $0.20 — a 2× miss on the line you probably run most.
Note also what this is not. It is not evidence that the marketplace is cheaper, and it is not evidence that it is misleading. A gateway listing a provider rate is stating what it charges, not what the upstream vendor charges. The error is on the reader’s side, and it is entirely avoidable by writing the surface into the cell next to the number. That single habit is what this whole table is arguing for.
05 — Vendor-Direct TrapDeepSeek: when the vendor is not the cheapest route.
The assumption that buying direct from the model’s creator is the floor price is intuitive, widely held, and — for DeepSeek V4 Flash at the time of writing — wrong. DeepSeek’s own API pricing page lists a single Flash model, versioned DeepSeek-V4-Flash-0731, at $0.14 per million input on a cache miss and $0.28 per million output, with a cache-hit input rate of $0.0028. OpenRouter’s provider rates for the same model version read lower on every alias: $0.09 / $0.18 on both the -0731 and -latest snapshots, and $0.0882 / $0.1764 on the base alias.
Recomputed from those two sourced pairs, the dated snapshots sit about 36% below the vendor cache-miss list on both input and output, and the base alias about 37% below. We are not going to speculate about why — DeepSeek has published no statement on the difference, and third-party inference providers on a marketplace set their own economics. The reportable fact is the gap and the fact that it is model-specific, because V4 Pro’s marketplace rate of $0.435 / $0.87 matches DeepSeek’s own list exactly.
DeepSeek V4 Flash · input rate per 1M tokens, by surface
Sources: DeepSeek API pricing page; OpenRouter public models API — input rates read at the time of writingThe cache-hit bar deserves its own note, because it is the largest single lever on that chart and it is not a discount you negotiate — it is one you engineer. DeepSeek’s cache-hit input rate for V4 Flash is exactly one fiftieth of its cache-miss rate. For V4 Pro the ratio is steeper still: $0.003625 against $0.435 is exactly one hundred and twentieth. A workload with high prompt-prefix reuse changes tier without changing vendor. If you are running V4 Flash in production, the release detail behind that 0731 version string is in our write-up of the 0731 checkpoint’s official release.
Layer the announced increase on top and DeepSeek becomes a routing decision rather than a default. The vendor has said an increase is coming and has not said when or by how much. That is not a reason to move off it — the rates are genuinely low — but it is a reason to keep the second route configured and tested rather than theoretical. Undated risk is best answered with switchable infrastructure, not with a scenario in a spreadsheet.
06 — Tiered RatesQwen3.7: the headline rate has a boundary.
Two cautions before any Qwen3.7 number. First, no standalone vendor list price for the Qwen3.7 family was locatable on Alibaba’s public model-endpoint listing during this pass — the page carries endpoint references without a pricing table. Every Qwen3.7 figure below is therefore a marketplace provider rate and nothing more, and it should never be quoted as a vendor list price. Second, unlike every other model in this tracker, Qwen3.7 rates are tiered by prompt length, so the headline number is a floor that applies only inside a bound.
Qwen3.7 Flash base tier
Output at $0.13 per million. Cache read $0.006 and cache write $0.038 at this tier. The context window is 1M tokens, but the base rate stops applying long before you fill it — the boundary is the prompt, not the window.
First escalation
Output rises to $0.40 per million. That is 3.3× the base input rate and 3.1× the base output rate — crossed automatically by any moderately sized document, transcript or retrieved context bundle, with no change on your side.
Second escalation
Output rises to $0.80 per million. Input is now 6.7× the headline rate and output 6.2×. A long-context workload priced off the $0.03 figure will land materially over budget, and nothing in the model name signals it.
The same shape appears one tier up. Qwen3.7 Plus lists at $0.32 / $1.28 below 256K prompt tokens and $0.96 / $3.84 above it — an exact 3× step on both input and output. Qwen3.7 Max is the outlier: $1.475 / $4.425 flat, with no tiering shown. The pattern across the family is that the cheaper the model, the more aggressively its price depends on how you use it, which is a reasonable design and a terrible thing to discover after a month of traffic.
The broader point is that the sub-$0.30 tier now has internal structure. If your interest is what that tier actually buys you at volume rather than how it is quoted, we worked the bulk economics separately in what the sub-$0.30 tier unlocks at bulk volume. This tracker stays on the quoting problem.
07 — Worked CostOne job, priced on every surface.
Rates per million are hard to feel. So here is a single fixed reference job — 2,000,000 input tokens and 200,000 output tokens, assuming no cache hits — priced on every surface each model exposes. The arithmetic per cell is simply input tokens multiplied by the input rate plus output tokens multiplied by the output rate, at per-million list rates. No quality weighting, no vendor efficiency claims folded in. Totals under a dollar are shown to three decimals.
| Model | Surface | Rate (in / out per 1M) | Job cost | Multiple of that model’s cheapest surface |
|---|---|---|---|---|
| OpenAI GPT-5.6 — standard list after the July 30 cut | ||||
| Sol | Batch or flex | $2.50 / $15.00 | $8.00 | 1.0× |
| Sol | Standard | $5.00 / $30.00 | $16.00 | 2.0× |
| Sol | Fast | $10.00 / $60.00 | $32.00 | 4.0× |
| Terra | Batch or flex | $1.00 / $6.00 | $3.20 | 1.0× |
| Terra | Standard | $2.00 / $12.00 | $6.40 | 2.0× |
| Terra | Fast | $4.00 / $24.00 | $12.80 | 4.0× |
| Luna | Batch or flex | $0.10 / $0.60 | $0.320 | 1.0× |
| Luna | Standard | $0.20 / $1.20 | $0.640 | 2.0× |
| Luna | Fast | $0.40 / $2.40 | $1.28 | 4.0× |
| Anthropic Claude Sonnet 5 — promo through Aug 31, reversion from Sep 1 | ||||
| Sonnet 5 | Batch, promo window | $1.00 / $5.00 | $3.00 | 1.0× |
| Sonnet 5 | Standard, through Aug 31 | $2.00 / $10.00 | $6.00 | 2.0× |
| Sonnet 5 | Standard, from Sep 1 | $3.00 / $15.00 | $9.00 | 3.0× |
| DeepSeek V4 Flash, 0731 snapshot — marketplace provider rates and the vendor cache-miss list | ||||
| V4 Flash | Marketplace, base alias | $0.0882 / $0.1764 | $0.212 | 1.00× |
| V4 Flash | Marketplace, -0731 and -latest | $0.09 / $0.18 | $0.216 | 1.02× |
| V4 Flash | Vendor list, cache miss | $0.14 / $0.28 | $0.336 | 1.59× |
| Alibaba Qwen3.7 Flash — marketplace provider rate, tiered by prompt length | ||||
| Qwen3.7 Flash | Prompts under 32K | $0.03 / $0.13 | $0.086 | 1.00× |
| Qwen3.7 Flash | Prompts 32K and above | $0.10 / $0.40 | $0.280 | 3.26× |
| Qwen3.7 Flash | Prompts 256K and above | $0.20 / $0.80 | $0.560 | 6.51× |
Three things fall out of the arithmetic that no vendor page states. First, the GPT-5.6 spread is exactly 4× from batch to Fast on all three models, because the two multipliers are uniform — so the surface decision is worth more than the model decision inside a single tier band. Second, Sonnet 5 currently undercuts Terra’s standard rate on this job mix by 6.25% ($6.00 against $6.40), and on September 1 the same comparison flips to Sonnet 5 costing about 41% more ($9.00 against $6.40). One calendar date reverses the ranking of two models with no change to either product.
Third, and least expected: OpenAI’s batch tier for Luna now sits marginally below DeepSeek’s own vendor list for V4 Flash on this job — $0.320 against $0.336, about 5% cheaper. The cheap end of the Western frontier stack has caught the cheap end of the open-weight stack at the vendor’s own quoted price. It has not caught the marketplace provider rates for the same DeepSeek model, which come in at $0.212 and are still about 34% below Luna’s batch tier. Which route is cheapest now depends entirely on which surface you are allowed to buy from.
Project that forward and the direction of travel is fairly clear. If a permanent 80% cut on a frontier vendor’s cheapest model lands it within a few cents of an open-weight vendor’s list price, then price alone stops being the reason to run open weights, and sovereignty, latency and control take over as the deciding factors. Meanwhile the gap between a vendor’s list and a marketplace’s provider rate is becoming the number that actually determines your bill — which is a procurement problem, not a model-selection problem, and most teams are still staffing it as the latter.
08 — Operating RulesFour habits that make a pricing sheet survive a month.
None of this requires new tooling. It requires four disciplines applied to the document you already keep, plus a willingness to write “not located” in a cell instead of filling it with a number from a different surface.
Price by surface, never by model name
Every rate in your sheet gets a surface label and a date: standard, batch or flex, fast, or marketplace provider rate. An unlabelled number is not a price, it is a rumour. This one change would have caught every trap in this article.
Reconcile every listing, not a sample
The GPT-5.6 asymmetry proves a sample is worthless: Sol matches vendor standard while Terra and Luna match the batch tier. Reconcile each model you actually call, and re-run it whenever a vendor announces a change — a cut on the vendor side does not guarantee the marketplace listing moved with it.
Model the cliff before it arrives
Sonnet 5 steps up 50% on September 1 and its cache rates have no published post-promotion figure. Put the reverted rate in the forecast now, flag the cache line as unpriced, and decide in August whether the September number still wins on your workload rather than discovering it in an invoice.
Read the boundary, not the headline
Qwen3.7 Flash escalates about 6.7× on input past 256K prompt tokens, and DeepSeek has announced an undated increase. Cheap tiers are cheap under conditions. Bound your prompt sizes deliberately, and keep a tested fallback route for any vendor that has told you a rise is coming.
If you want a longer baseline than one month, the Q2 2026 edition of this tracker holds the earlier data points — read as a historical snapshot, not as current rates, since most of those figures have since moved. Teams that want the reconciliation and forecasting built into their reporting rather than maintained by hand usually start with our analytics and measurement work, and the routing side sits inside AI and digital transformation engagements.
09 — ConclusionThe number is fine. The label is what is missing.
A model does not have a price. A model plus a surface has a price.
August 2026 opened with one large permanent cut, one promotion on a countdown, and one undated warning. GPT-5.6 Luna is 80% cheaper and Terra 20% cheaper than they were on July 29, and neither change expires. Claude Sonnet 5 costs $2 and $10 per million through August 31 and then $3 and $15. DeepSeek has said an increase is coming and has not said when. Those three facts are the easy part.
The hard part is that the same model can be quoted at four prices, and the marketplace listing has no fixed relationship to the vendor card. One GPT-5.6 model’s listing matches the vendor standard rate exactly; the other two match the discounted batch tier. The marketplace provider rates for DeepSeek V4 Flash sit about a third below DeepSeek’s own list for the identical model version, while its V4 Pro list matches to the cent. Qwen3.7’s cheapest rate applies only under 32K prompt tokens. None of that is hidden — it is just never assembled in one place.
So the discipline is small and the payoff is not. Write the surface and the date next to every rate you record. Reconcile each model you actually call rather than a sample. Put the reversion in the forecast before the cliff arrives. Do that and a pricing sheet survives a month like July 30; skip it and the cheapest-looking number on the page is the one most likely to be wrong.