Off-peak LLM pricing has arrived — and the first thing to understand about it is that a cheaper hour is not the same as a cheaper bill. DeepSeek's pricing page now posts exact peak windows and a full two-tier rate table that takes effect on Sunday, August 16, 2026 at 16:00 UTC, and every cell in that table — peak and off-peak alike — sits above the flat rate in force today.
The off-peak tier is real: it is precisely half of the peak tier, on every line item, exactly as the vendor states. But half of peak is not half of today. Recompute each cell against the flat rate still live at the time of writing and the off-peak floor alone runs roughly 1.5× to 6.1× current prices depending on the line item. This is not a discount scheme. It is a repricing in which off-peak is the smaller of two increases.
The same week, Z.ai shipped GLM-5.3 and restructured its GLM Coding Plan around a points system with its own peak clock — weekday afternoons in China time, everything else off-peak. This guide lays out both schedules converted to US and EU business hours, the recomputed price deltas the vendor pages don't print, and a practical playbook for deciding which workloads to move into the cheap windows and which to leave alone.
- 01This is a repricing, not a discount.On all six DeepSeek line items (cache-hit input, cache-miss input, output — for both V4-Flash and V4-Pro), the announced off-peak rate is higher than today's flat rate: roughly 1.5× to 6.1× depending on the cell. Peak runs higher still.
- 02DeepSeek's windows are exact and UTC-denominated.Peak is 01:00–04:00 and 06:00–10:00 UTC — 7 hours a day — with all other hours off-peak at half the peak rate. The switch lands Sunday, August 16 at 16:00 UTC, which is midnight Beijing time going into August 17.
- 03GLM's peak clock is narrow — don't invert it.On the restructured GLM Coding Plan ($18 / $80 / $168 per month), peak is Monday–Friday 14:00–18:00 UTC+8 only. Every other hour, including entire weekends, bills at 50% of standard points.
- 04US business hours are already off-peak on both vendors.Converted to August 2026 offsets, neither vendor's peak window touches a US 9-to-5 in Eastern or Pacific time. EU teams are the opposite case: both vendors' peaks overlap the 08:00–12:00 CEST business morning.
- 05Budget at peak; treat off-peak capture as upside.The windows and the rates are the vendor's to change — DeepSeek went from a vague footnote to a full dated table within days. Forecast spend at peak rates, then let deliberate scheduling deliver the difference.
01 — The AnnouncementWhat lands on Sunday.
DeepSeek announced the change in a changelog entry dated Thursday, August 13, 2026 — the same entry that took DeepSeek-V4-Pro to general availability on app, web, and API. This post doesn't retell the GA story; it covers the pricing structure that ships alongside it. The essentials, all from DeepSeek's own pages:
- Effective moment: 16:00 UTC on Sunday, August 16, 2026. Stated another way — the same instant — that is 00:00 Beijing time going into Monday, August 17. Until then, the flat rates remain in force; nothing described here is live at the time of writing.
- Peak windows: the pricing page states them exactly — "01:00 - 04:00 and 06:00 - 10:00 UTC (all other hours are off-peak)." That is 7 peak hours per day; the remaining 17 of 24 hours (≈71% of each day) are off-peak.
- The ratio: off-peak is precisely half of peak on every line item — confirmed cell-by-cell on the posted table, and stated in prose in both the changelog and the accompanying news post.
Worth recording for anyone tracking how vendors communicate price changes: until this announcement, the same pricing page carried only an undated footnote warning that an overall increase was planned. As fetched on August 14, that vague line is gone, replaced by the concrete two-tier table. The vendor moved from a soft warning to exact numbers and a dated switch inside a single news cycle — which is itself the strongest argument in this post for budgeting off posted schedules, not current rates.
02 — The ArithmeticBoth tiers rise against today.
DeepSeek's news post states, accurately: "Off-peak rates are 50% lower than peak, enabling more flexible workload scheduling." What neither the news post nor the pricing page states is how either tier compares to the flat rate you are paying today. Both the current flat rates and the new two-tier table appear on the same live pricing page, so the comparison takes nothing more than division. We ran it for all six line items — the deltas below are our arithmetic on the vendor's posted numbers, not vendor-stated figures.
| Line item · per 1M tokens | Today (flat) | Off-peak (Aug 16+) | Peak (Aug 16+) | Off-peak vs today | Peak vs today |
|---|---|---|---|---|---|
| deepseek-v4-flash | |||||
| Cache-hit input | $0.0028 | $0.007 | $0.014 | +150% | +400% |
| Cache-miss input | $0.14 | $0.22 | $0.44 | +57% | +214% |
| Output | $0.28 | $0.66 | $1.32 | +136% | +371% |
| deepseek-v4-pro | |||||
| Cache-hit input | $0.003625 | $0.022 | $0.044 | +507% | +1,114% |
| Cache-miss input | $0.435 | $0.66 | $1.32 | +52% | +203% |
| Output | $0.87 | $1.98 | $3.96 | +128% | +355% |
Rate columns are from DeepSeek's own pricing page as fetched August 14, 2026; the two delta columns are percentages of today's flat rate, computed by us. Three readings matter. The most extreme relative jump is V4-Pro cache-hit input, where peak lands at roughly 12.1× today's rate — but cache-hit is also the cheapest absolute cell, so the dollar impact per token is small. The line item most workloads actually spend on is output, and there the picture is stark enough: V4-Flash output goes from $0.28 flat to $0.66 off-peak and $1.32 peak; V4-Pro output goes from $0.87 to $1.98 and $3.96.
The off-peak floor · cheapest new tier as a multiple of today's flat rate
Source: DeepSeek pricing page rates, Aug 14, 2026 · multiples computed by Digital Applied03 — Timezone MathThe peak windows on your clock.
Both vendors publish their windows in their own denominations — DeepSeek in UTC, Z.ai in UTC+8. What a scheduling decision needs is those windows in your team's local time. The grid below is our conversion at August 2026 offsets: US Eastern on daylight time (UTC−4), US Pacific on daylight time (UTC−7), Central Europe on CEST (UTC+2), and China Standard Time (UTC+8, no daylight saving).
| Peak window | Vendor-stated | US Eastern (EDT) | US Pacific (PDT) | Central Europe (CEST) | China (CST) |
|---|---|---|---|---|---|
| DeepSeek · block 1 · daily | 01:00–04:00 UTC | 21:00–00:00 (prev. day) | 18:00–21:00 (prev. day) | 03:00–06:00 | 09:00–12:00 |
| DeepSeek · block 2 · daily | 06:00–10:00 UTC | 02:00–06:00 | 23:00 (prev. day)–03:00 | 08:00–12:00 | 14:00–18:00 |
| GLM Coding Plan · Mon–Fri only | 14:00–18:00 UTC+8 | 02:00–06:00 | 23:00 (prev. day)–03:00 | 08:00–12:00 | 14:00–18:00 |
Three conclusions fall straight out of the grid. First, a US-based team is already off-peak on both vendors without lifting a finger — neither DeepSeek block nor the GLM window touches a 9-to-5 in Eastern or Pacific time. Second, EU teams are the squeezed case: both vendors' peaks overlap the 08:00–12:00 CEST business morning, so an EU team running bulk or agent workloads before lunch is paying peak by default and needs to actively defer that work to the afternoon. Third, GLM's single weekday window converts to exactly the same UTC hours as DeepSeek's second block (06:00–10:00 UTC) — so one scheduling rule covers both vendors' worst overlap.
One framing note, because the history matters. Converted to China Standard Time, DeepSeek's two blocks land on 09:00–12:00 and 14:00–18:00 — ordinary Beijing business hours with an off-peak lunch gap. Before August 13 there was no primary source describing any such scheme, and unsourced claims circulating earlier were rightly not credited — including by us. The mapping above is straightforward arithmetic on the UTC hours DeepSeek has now actually published, not vindication of anything that circulated before the announcement existed.
04 — The Comparison CaseGLM's points clock runs narrower.
Z.ai shipped GLM-5.3 on Friday, August 14, 2026 and restructured the GLM Coding Plan around a credits system in the same launch — we cover the model itself separately; this section is only about its clock. Where DeepSeek defines 7 peak hours every day, GLM's peak is a single four-hour weekday window — and it is worth being precise here, because the scheme is easy to state backwards.
The billing mechanics differ from DeepSeek's in kind, not just in window shape. Coding Plan usage is not billed per token in dollars; it draws down credits by a published formula — input tokens times an input multiplier, plus cached input times its multiplier, plus output times its multiplier, divided by 10,000. For GLM-5.3 those multipliers are 6.9 / 1.7 / 24 (input / cached-input / output). Off-peak halves the credit draw, not a dollar price. No standalone per-token dollar price for GLM-5.3 appears in the launch post or the developer-docs overview we checked — Coding Plan credits are the only published pricing for it, an absence worth knowing before you try to build a cross-vendor $/token comparison.
Weekly credits
2,000 credits per rolling 5-hour window, 10,000 per week. Z.ai's own estimate: roughly 43–87M GLM-5.3 tokens weekly at its assumed 90.9% cache-hit rate — the low end all-peak, the high end all-off-peak.
Weekly credits
12,000 per 5-hour window, 60,000 weekly — 6× Lite usage, which works out to $13.33 per Lite-equivalent unit. Vendor-estimated 263–526M tokens weekly on the same cache-hit assumption.
Weekly credits
28,000 per 5-hour window, 140,000 weekly — 14× Lite usage, or $12.00 per Lite-equivalent unit, the cheapest per-unit rung. Vendor-estimated 613M–1,226M tokens weekly.
Two caveats on those tiers. The token-allowance ranges are the vendor's own estimates, resting on its stated assumption of a 90.9% cache-hit rate as "the average level for coding workloads" — your mix will differ, and a low-cache workload lands well under the printed range. And these are launch-day numbers from the live subscribe page, fetched August 14; the plan structure itself was reworked as part of this launch, so treat the figures as a snapshot, not a commitment. One operational detail from the same docs: requests to the previous GLM-5.2 and GLM-5.1 models are now routed automatically to GLM-5.3 on all Coding Plan tiers, so the 5.3 multipliers above are effectively the plan's economics for everyone.
05 — Scheduling PlaybookWhich workloads move, which don't.
Time-of-day pricing only pays you if some of your token spend is genuinely movable. The test is simple: does anything break if this job runs eight hours later? Workloads split cleanly on that question.
Batch enrichment, classification, embeddings refresh
Millions of rows, no human waiting. These jobs already run on a cron somewhere — moving the cron into an off-peak window is a one-line change that halves the rate versus peak. The archetypal winner.
Evals, regression sweeps, nightly QA
Eval suites and LLM-judge passes tolerate latency by design and often run nightly already. Pin them inside off-peak hours and they never pay the peak tier at all.
Overnight agent runs & research sweeps
Multi-hour autonomous runs care when they start, not when you read the results. Kick them off after the peak window closes in your timezone; an overnight US run naturally clears both vendors' peaks.
Interactive, customer-facing, in-workday tooling
Chat surfaces, support flows, and anything a person is waiting on cannot chase a clock — the request happens when the user shows up. Budget these at peak and don't contort the product to save on tokens.
Two disciplines make the movable column actually work. First, engineering: a deferred million-row job needs idempotency, checkpointing, and a QA gate far more than it needs a cheap hour — the engineering discipline for a million-row job doesn't change because the clock does. Second, don't confuse the two half-price lanes now on offer. Batch APIs give an always-on ~50% lane in exchange for submit-and-wait semantics; the batch-API half-price lane trades latency for money at any hour. Off-peak pricing is a time-of-day lane: full interactivity, but only inside the vendor's cheap window. A workload that can tolerate hours of delay generally belongs in a batch lane where one exists; off-peak windows are for work that needs synchronous APIs but not a particular hour.
For US teams the playbook is almost embarrassingly simple: your business day is already off-peak on both vendors, so the only real action is making sure overnight crons don't accidentally land inside 21:00–00:00 or 02:00–06:00 Eastern, DeepSeek's two converted blocks. EU teams have the genuine scheduling problem — the 08:00–12:00 CEST overlap on both vendors means morning bulk work is peak-priced by default, and the fix is moving batch and agent workloads to afternoons and evenings. If your stack routes meaningful spend through these vendors, this is exactly the kind of cost-structure decision our AI transformation engagements model before it shows up on an invoice.
06 — Cost GovernanceReasoning about a schedule you don't control.
The deeper issue this week raises is not the specific windows — it's that a growing share of AI cost structure is now set by vendor-controlled schedules: time-of-day tiers, dated switches, promo cliffs. Three rules keep a budget honest under that regime.
- Budget at peak, bank off-peak. Forecast every DeepSeek line item at the peak rate and treat scheduled off-peak capture as realized savings, not as the plan. Any forecast built on the off-peak floor assumes your scheduling discipline never slips and the vendor never moves the window — two assumptions you don't control.
- Name the denominator on every "saving." Moving a job off-peak saves 50% versus peak — and still costs roughly 1.5× to 6.1× what the same job costs today, depending on line item. Both statements are true; only the pair of them is honest. The same discipline applies to intro pricing elsewhere — the case for budgeting at rates already posted to double is the same argument with a calendar instead of a clock.
- Treat announced schedules as changeable. DeepSeek replaced a vague warning footnote with a full dated rate table inside a week; Z.ai restructured its plan tiers on launch day. What a vendor actually owes you before a price change is a contract question — we walk the notice-period fine print separately — but the operational posture is the same either way: re-fetch the pricing page before every budget cycle.
Where does this go next? Time-of-day pricing is standard practice in electricity markets precisely because demand is spiky and capacity is fixed — and inference has the same shape. DeepSeek's own framing ("allocate resources more reasonably") reads like a utility smoothing its load curve, and GLM's points clock does the same for subscription capacity. It would not be surprising to see other capacity-constrained providers test similar windows, nor to see today's windows widen, narrow, or reprice as the vendors watch how much load actually moves. The durable skill is not knowing this week's hours — it's having your movable workloads already separated from your immovable ones, so the next schedule change is a config edit rather than an architecture project.
07 — ConclusionCheap hours are real. Cheap is relative.
Read every discount against the rate you pay today, not the tier printed beside it.
DeepSeek's switch is best understood as two moves in one announcement: a substantial across-the-board price rise, and a genuine 50% time-of-day spread laid on top of it. Both are real. The vendor pages document the spread; the comparison to today's flat rates — off-peak alone at roughly 1.5× to 6.1× current prices — you have to compute yourself, which is why we did.
The scheduling opportunity is also real, and unevenly distributed. US teams inherit off-peak status on both vendors by timezone accident; EU teams have to earn it by pushing bulk work out of the 08:00–12:00 morning; and every team should hold the line that interactive, customer-facing workloads budget at peak rather than bending product behavior around a vendor's clock.
Most of all, this week is a procurement lesson in how fast soft signals harden. A vague footnote became an exact, dated, two-tier rate table in days. The teams that absorb that without drama are the ones that budget at the posted ceiling, separate movable from immovable spend, and re-read the pricing page as routinely as they read their invoice.