China's open-frontier August put three major model events on one board: DeepSeek moved V4-Pro to general availability on August 13, Z.ai launched GLM-5.3 on August 14, and both landed while Moonshot's Kimi K3 — weights shipped July 27 — had been the newest downloadable open checkpoint.
The three stories look similar from a distance and diverge sharply up close. One model's weights are a download you can start today under MIT terms. One model's weights are a dated promise. One model's weights have been available for weeks — under a bespoke licence that is not Apache 2.0 and that most "open source" headlines glossed over. On pricing, one vendor published a schedule that raises every rate starting Sunday, August 16; one publishes no per-token API price at all; one's headline figures come from a reseller listing rather than a directly captured vendor page.
This scoreboard covers each model's availability, pricing surface, benchmark posture, and true open-weight status — and closes with the most transferable finding of the week: two vendors' own published benchmark tables disagree with each other on the same third-party model, on the same-named benchmark. If you compare vendor tables for a living, that section is the one to keep.
- 01Open-weight means three different things on this board.DeepSeek's GA checkpoint is on Hugging Face under MIT. GLM-5.3's weights are a stated two-week promise. Kimi K3's weights shipped July 27 under a bespoke kimi-k3 licence with revenue and branding thresholds.
- 02DeepSeek's announced schedule raises every rate.The peak/off-peak pricing effective August 16 at 16:00 UTC is higher than today's flat list on every line item — off-peak is the smaller of two increases, not a discount from today.
- 03GLM-5.3 has no per-token API price.The launch post publishes no input/output rate; the restructured GLM Coding Plan ($18 / $80 / $168 list) with a credits system is the only priced surface, and no GLM-5.3 row existed on OpenRouter as of the August 14 pull.
- 04Kimi K3 is open-weight, not unconditionally open.The licence grants broad free commercial use but requires a separate agreement once model-as-a-service revenue passes $20M over any consecutive 12 months, plus a branding condition at 100M MAU or $20M/month revenue.
- 05Vendor benchmark tables need version pinning.DeepSeek's and Z.ai's own tables diverge by up to 3.2 points on the same model and benchmark name — and AutomationBench differs by harness version, not disagreement. Pin version, harness, and effort before comparing.
01 — The ScoreboardThe week on one page.
Everything below is drawn from the vendors' own pages — DeepSeek's API changelog and pricing page, the Hugging Face repositories and licence files, Z.ai's launch post and subscribe page, and OpenRouter's public model listing — all retrieved August 14, 2026. The table is the post in miniature; the sections that follow expand each row. For where these providers stood before this week, see our Q2 2026 market-share snapshot of the same providers.
| Model | Weights today | Licence | API price surface | OpenRouter listing |
|---|---|---|---|---|
| DeepSeek V4-Pro-0813 | Shipped — Hugging Face repo created August 13, ~893 GB across 67 safetensors shards | MIT | Published by the vendor; a new peak/off-peak schedule takes effect August 16, 16:00 UTC, raising every line | Listed; the listing date reads August 12 while the changelog header reads August 13 — we report both |
| GLM-5.3 | Not shipped — Z.ai states weights come "in two weeks" from the August 14 launch; no zai-org repo existed on Hugging Face as of the August 14 pull | Unstated until weights ship | No per-token rate published; the priced surface is the restructured GLM Coding Plan (credits-based) | Not listed as of the August 14 pull — the newest z-ai row was still GLM-5.2 |
| Kimi K3 | Shipped July 27 — ~1.56 TB across 118 files, with image-text-to-text tags confirming native vision at the weights level | Bespoke "kimi-k3" licence — free commercial use with revenue and branding thresholds | $3 in / $15 out per 1M from OpenRouter's listing; the $0.30 cache-hit input figure comes from third-party trackers, which OpenRouter does not expose as a field — no directly captured vendor table either way | Listed July 16 |
Read the second column twice. In procurement conversations all three models get filed under "open Chinese frontier," but only two have weights you can download at the time of writing, and only one of those carries a standard permissive licence. That distinction — download versus promise versus conditional grant — decides whether your on-prem plan is executable this week, next month, or only after a licence review.
02 — DeepSeekV4-Pro goes GA — MIT weights, higher prices scheduled.
DeepSeek's API changelog carries an entry dated August 13, 2026 announcing that the GA release of DeepSeek-V4-Pro "has been rolled out on the APP, Web, and API" — the model alias deepseek-v4-pro now resolves to DeepSeek-V4-Pro-0813. Archive captures narrow the entry's appearance to a roughly 32-hour window between August 12 and August 13 UTC, so we won't claim a precise publish hour. The full announcement story is in our DeepSeek V4-Pro GA coverage.
The open-weight question has a clean answer: the deepseek-ai/DeepSeek-V4-Pro-0813 repository exists on Hugging Face, created August 13, under an MIT licence, at roughly 893 GB across 67 safetensors shards. One honest caveat travels with it — the 0813 model card does not restate parameter counts, and the widely repeated figures for this family trace only to the April Preview disclosure, so we attach no parameter claims to the GA checkpoint.
On benchmarks, the changelog entry includes a ten-metric table — Terminal Bench 2.1 at 87.9, Toolathlon-Verified at 74.1, AutomationBench (Public) at 31.8 among them — evaluated, in DeepSeek's own label, with "the minimal mode of DeepSeek Harness" at max reasoning effort. Those are vendor-run numbers on a vendor-built harness; hold that caveat, because it becomes the centerpiece of Section 05. DeepSeek documents a low/high/max effort ladder for both V4 models and does not state which rung is the default, so neither will we.
deepseek-v4-pro. On Sunday, August 16 at 16:00 UTC a peak/off-peak schedule takes effect: peak hours are 01:00–04:00 and 06:00–10:00 UTC, all other hours (weekends included) are off-peak at half the peak rate. Both tiers are increases. V4-Pro output goes from $0.87 today to $1.98 off-peak and $3.96 peak; cache-miss input goes from $0.435 to $0.66 off-peak and $1.32 peak. V4-Flash rises on every line as well ($0.14/$0.28 today; $0.22/$0.66 off-peak; $0.44/$1.32 peak, cache-miss/output).V4-Pro output per 1M tokens · today vs the announced August 16 schedule
Source: api-docs.deepseek.com pricing page, retrieved August 14, 2026The framing matters: "off-peak discount" is technically true relative to peak and misleading relative to today. Off-peak is the smaller of two increases — 52% on V4-Pro cache-miss input against today's $0.435, 128% on output against today's $0.87. Teams that can shift batch workloads into the off-peak windows will blunt the rise, not escape it. The scheduling mechanics and a workload playbook are in our companion piece on off-peak LLM pricing, published today.
03 — Z.aiGLM-5.3 — strong launch, promised weights.
Z.ai launched GLM-5.3 on August 14 with an unusually direct opening claim: "Today we are releasing GLM-5.3. It uses the same base model as GLM-5.2 — every gain comes from post-training." On availability, the model is live on Z.ai's own surfaces. On weights, the launch post is explicit that nothing has shipped yet: "We will release the weights in two weeks after launch, once safety evaluation and hardening are complete." We verified the gap between promise and delivery directly — no GLM-5.3 repository existed under the zai-org Hugging Face organization as of the August 14 API pull, and OpenRouter's model list still showed GLM-5.2 as the newest z-ai row in the same pull. Until the weights land, GLM-5.3 is a hosted model with an open-weight commitment, and should be planned as exactly that.
The API surface changed in a breaking way: the reasoning-effort ladder is now low/high/max — there is no medium rung — and thinking can no longer be disabled. Z.ai's migration note warns that requests still sending thinking.type: "disabled" will fail outright unless switched to enabled with effort set to low before the model ID is updated. Where each vendor's ladder sits, and which rungs converge across the Chinese labs, is mapped in our cross-vendor effort-ladder field guide, also published today.
Pricing is the odd one out on this scoreboard: the launch post publishes no per-token API rate anywhere on the page — we checked the full text, not a summary. The priced surface is the GLM Coding Plan, restructured on launch day to a credits system with off-peak point pricing: calls outside peak hours consume 50% of standard points, and peak is narrowly defined as Monday–Friday, 14:00–18:00 UTC+8 — every other hour, weekends included, is off-peak.
10,000 credits per week
The entry tier of the restructured plan. Point usage is metered separately for input, cached input, and output tokens; off-peak calls consume half the points.
6× Lite usage (list)
Restructured at the GLM-5.3 launch — the plan moved from a quota-multiplier model to credits, and the Pro tier's list price and multiplier both changed with it.
14× Lite usage (list)
The top tier of the credits system. Point metering and the 50% off-peak point rate work the same way here as on Lite and Pro.
On benchmarks, Z.ai publishes its own cross-vendor table, footnoted as evaluated "in Claude Code 2.1.207" with most rows at max reasoning effort and per-benchmark settings individually footnoted. The table is not a sweep, and Z.ai does not pretend it is: GPT-5.6 Sol takes Terminal Bench 2.1 (88.8 vs GLM-5.3's 88.2), Kimi K3 takes Toolathlon Verified (76.5 vs 73.0), while GLM-5.3 posts the table's best AutomationBench v1.0.6 score at 48.2. The launch also leans hard into security research — a vulnerability-disclosure ledger whose totals (2,436 findings across 269 open-source projects, 1,097 of those 2,436 rated critical or high) are entirely Z.ai's own count of Z.ai's own program, with no independent reporting found as of the August 14 research pass. We unpack that program, defensively framed, in our GLM-5.3 disclosure-ledger analysis, and the full launch in the GLM-5.3 launch breakdown.
04 — Moonshot AIKimi K3 — shipped weights, bespoke terms.
Kimi K3 is the veteran of this scoreboard: announced by Moonshot in mid-July, with full weights landing on Hugging Face on July 27 — the deadline the launch blog set, which promised the weights by that date rather than on it. The repository is substantial (~1.56 TB across 118 files) and its image-text-to-text tags confirm native vision at the weights level. But the licence is where headlines went wrong. Hugging Face tags it license: other with the name "kimi-k3" — it is not Apache 2.0, and calling K3 "open source" without qualification is inaccurate. The licence text grants broad free use, including commercial use, modification, and redistribution, with two conditions: a model-as-a-service business built on the weights that passes $20 million in aggregate revenue over any consecutive 12 months needs a separate agreement with Moonshot before continuing, and any commercial product on the weights exceeding 100 million monthly active users or $20 million per month in revenue must display "Kimi K3" prominently in its UI. Our clause-by-clause read is in our full analysis of K3's bespoke licence.
On pricing, label the surface: the commonly cited $3.00 fresh input / $0.30 cache-hit input / $15.00 output per 1M tokens, flat across the full 1M-token context, comes from OpenRouter's listing (which matches Moonshot's list price on the fields it exposes) plus third-party trackers for the cache-hit figure specifically. Moonshot's own pricing table renders via JavaScript and was not captured directly in this research pass — so treat these as corroborated reseller-surface figures, not a vendor-page citation. K3 was listed on OpenRouter on July 16, a day before Moonshot's own launch framing — the same listing-before-announcement pattern DeepSeek's 0813 checkpoint repeated in August.
One more hedge the spec sheets skip: at launch, Moonshot stated that "Kimi K3 will use max thinking effort by default, with low- and high-effort modes to be introduced in subsequent updates." We found no Moonshot changelog entry confirming those tiers had shipped as of this writing, so treat K3's effort ladder as max-only until the vendor documents otherwise. Moonshot's launch post is also unusually candid on positioning, stating that K3's overall performance "still trails the most powerful proprietary models" while claiming frontier-level results across its own evaluation suite — a vendor-run suite, with the same caveat as every table in this post.
05 — The Real FindingSame benchmark, two tables — a version-pinning lesson.
Here is the part of this week that will still matter in six months. DeepSeek's model card and Z.ai's launch post each publish a cross-vendor benchmark table, and both score the same third-party models on several of the same benchmarks. We diffed the two tables cell by cell. Mostly, they agree exactly — five Terminal Bench 2.1 entries and four DeepSWE entries match to the decimal across two independently published vendor tables, which is itself evidence nobody is inventing numbers. But three cells diverge for the same third-party model on the same benchmark name, and one benchmark splits on harness version outright.
| Benchmark | Model scored | DeepSeek's table | Z.ai's table | Gap (pts) |
|---|---|---|---|---|
| Exact agreement — same number in both vendors' tables | ||||
| Terminal Bench 2.1 | Kimi K3 | 88.3 | 88.3 | 0.0 |
| Terminal Bench 2.1 | DeepSeek-V4-Pro-0813 | 87.9 | 87.9 | 0.0 |
| Terminal Bench 2.1 | GLM-5.2 | 81.0 | 81.0 | 0.0 |
| DeepSWE | Opus 4.8 | 58.0 | 58.0 | 0.0 |
| Divergence — same model, same benchmark name, different score | ||||
| Toolathlon-Verified | Fable 5 (w/ fallback) | 77.9 | 74.7 | 3.2 |
| CyberGym | Fable 5 (w/ fallback) | 83.1 | 83.8 | 0.7 |
| DeepSWE | Fable 5 (w/ fallback) | 70.0 | 69.7 | 0.3 |
| Version mismatch — same name, different harness, not comparable | ||||
| AutomationBench | DeepSeek-V4-Pro-0813 | 31.8 ("Public") | 43.2 (v1.0.6) | n/a |
| AutomationBench | Kimi K3 | 30.8 ("Public") | 46.7 (v1.0.6) | n/a |
| AutomationBench | GLM-5.2 | 12.9 ("Public") | 26.2 (v1.0.6) | n/a |
The divergence rows carry no accusation. A 0.3-to-3.2-point spread on the same closed model plausibly reflects different evaluation dates, minor version drift in the model being scored, or rounding — neither vendor says, so neither will we. The operational rule is simpler: do not average the two figures, and do not treat either as more authoritative. Each number is only meaningful inside its own harness — DeepSeek's on "the minimal mode of DeepSeek Harness," Z.ai's "in Claude Code 2.1.207."
AutomationBench is the sharper case because it is not a disagreement at all. DeepSeek scored the "Public" version; Z.ai scored v1.0.6, explicitly footnoted as incorporating a null-type handling fix from an upstream pull request — and every shared model scores far higher on the fixed harness. Citing "AutomationBench: X%" without a version number is close to meaningless for cross-vendor comparison. Z.ai's own page even carries a smaller version of the same lesson internally: its benchmark table labels one column "Fable 5 (w/ fallback)" while the prose attributes the identical figures to "Mythos 5," and the page never reconciles the two names. We record that inconsistency; we do not resolve it.
06 — Buyer's PlaybookWhat buyers should do with this week.
The trend beneath the news: the Chinese frontier labs now ship on a weekly cadence, but their openness models are diverging rather than converging. DeepSeek doubled down on unconditional MIT weights while raising API prices; Z.ai decoupled launch from weights release and priced only its subscription surface; Moonshot invented licence terms that are free for nearly everyone and conditional for exactly the hyperscale competitors most likely to commercialize the weights. "Open" has become a spectrum with contract language attached, and each vendor is choosing a different point on it.
Sovereignty-bound deployment
DeepSeek V4-Pro-0813 is the only frontier checkpoint on this board that is downloadable today under a standard permissive licence. MIT terms, ~893 GB. Budget serious hardware and verify the licence file in the repo, not the headline.
Promise-tracking, not planning
A dated commitment is a calendar entry, not a dependency. Do not architect around GLM-5.3 weights until the repo and its licence text exist — the licence terms are unstated until the weights ship, and terms are where K3 surprised people.
Licence-first due diligence
For most teams the kimi-k3 conditions never bite — the thresholds are $20M MaaS revenue over 12 months and 100M MAU or $20M/month for branding. But read them against your growth path before the weights enter production, not after.
Version-pinned evals only
Record benchmark version, harness, effort setting, and date next to every score you cite — and rerun the handful of benchmarks you actually route traffic on, in your own harness, on your own workloads.
Looking forward, two dates on this board are procurement events in their own right: DeepSeek's August 16 switch and Z.ai's roughly two-week weights window. The forward risk is that time-of-day pricing and launch-then-weights sequencing become templates other labs copy — which would make "what notice does my vendor owe me before a price change" a standing contract question rather than a one-off. That question gets its own treatment in our guide to AI price schedules and soft commitments, published today. And if your team is deciding which of these models — or none of them — belongs in a production routing mix, our AI transformation engagements start with exactly this kind of licence-and-eval due diligence on your own workloads.
07 — ConclusionThree ships, three definitions of open.
Check the repo, read the licence, pin the harness.
The surface story is velocity: three frontier-class releases from three Chinese labs, two of them in the week before publication. The useful story is precision. DeepSeek V4-Pro's GA build is verifiably open — MIT weights on Hugging Face, same day as the changelog entry — while its API prices rise on every line from August 16. GLM-5.3 is a strong hosted launch whose open-weight status is a two-week promise with unstated licence terms. Kimi K3 has been genuinely downloadable since July 27, under bespoke conditions that make "open source" the wrong word without a footnote.
The benchmark lesson outlives all three news cycles. Two vendors published overlapping tables this week, and the cells that matched exactly are as instructive as the ones that diverged: vendor numbers are mostly honest and never interchangeable. A score without its benchmark version, harness, effort setting, and date is not evidence — it is decoration. Pin those four things on every number you cite, and rerun the ones you route real traffic on.
Treat this scoreboard as dated the day it published. The weights promise has a clock on it, the pricing schedule has a switch date, and the next divergent benchmark cell is one launch post away.