A stealth model reveal is not reassurance — it is the start of a re-verification job. On August 26, 2026, OpenRouter's anonymous stealth/ox-alpha listing resolved to Z.ai's GLM-5.3-Flash, with MIT-licensed weights published the same day. Every team that put the free anonymous endpoint into something real now has a day of work to do, and most of it only became possible this morning.
The scale of that "something real" is documented. OpenRouter itself posted on August 24 that "Ox Alpha is on track to hit nearly 6 trillion tokens today" — a dated, single-day run rate, two days before anyone knew whose model it was. (Cumulative totals for the stealth period circulate too, but they disagree with each other by roughly a factor of five and none traces to a primary source, so we publish none.) A lot of production traffic was riding on an endpoint with no vendor name attached.
This checklist covers the four checks that only become possible, or only become urgent, the moment anonymity ends: the pinned slug, the price ladder, the serving configuration, and the data terms. It closes with an honest scorecard of what the reveal resolves and what it does not.
- 01A reveal starts a re-verification job.Nothing about a name change confirms the endpoint is unchanged. OpenRouter's stealth terms say nothing about reveal-day mechanics — the transition is silent by design, so the burden of verification is yours.
- 02Your before state is archived — use it.A Wayback Machine capture from August 24 preserves the stealth listing at $0 with its own no-training line. That dated snapshot is exactly what makes the after state checkable.
- 03Prices are surface- and date-specific.The standard surface shows a promo of $0.075/$0.25 per 1M through September 9; list is $0.15/$0.50 after. The :batch surface already prints the list numbers today — on this dateline, batch is not cheaper.
- 04Four things must be verified, not assumed.Whether ox-alpha carried supplemental terms, whether serving config differs across the 21 providers, what a pinned stealth slug now returns, and the exact text of Alibaba's Section 5.2 — none is established. Each is a check, not an asserted risk.
- 05One unknowable survives the reveal.A named vendor gives you someone to ask about training data — it does not retroactively certify that your evals against the anonymous endpoint were contamination-free. A name is an address for questions, not a corrected score.
01 — Why This Post ExistsThe identity arrived. The work arrives with it.
Yesterday we published a guide to evaluating a model whose identity you cannot know. Its stated spine was an inventory of the unknowables that no amount of testing reveals about an anonymous endpoint: training-data provenance, retention, deprecation risk, licence, and liability. This post is the deliberate sequel. A reveal inverts part of that inventory — unknowables become checks you can now run, and the honest ones come with a warning about which items a name still does not settle.
Two boundaries, so you reach for the right checklist. First, this is not an open-weights adoption checklist: our Kimi K3 and Qwen3.8 checklists trigger on "weights are coming, prepare hosting," and there is no self-hosting or hardware-prep content here at all — this post is for the hosted endpoint you already had in production. Second, this is not the signing-stage checklist: the procurement questions to ask before you sign cover the contract you have not yet entered. A reveal is the already-in-production case — the dependency exists, and the question is what changed underneath it. (If the reveal has you rethinking how much rides on one model at all, that is the single-model dependency checklist's territory, not this one's.)
The trap this post is built to avoid is the generic checklist. "Re-run your evals" and "read the terms" apply to any Tuesday. Every item below has to pass a harder test: it is only possible, or only urgent, because anonymity ended today. Anything that would apply equally to a routine version bump was cut.
02 — The Before StateYour baseline is archived. Pull it before you check anything.
The single most useful artifact for reveal-day work is a dated snapshot of what you were operating under before the name landed. For ox-alpha, one exists: a Wayback Machine capture from August 24 preserves the stealth listing priced at $0, with its own terms stating that prompts "are retained by the provider and are not used for training; all other use is governed by the Stealth Model Terms." That is the before state — dated, archived, and comparable. Every check in this post is a diff against it.
The general framework behind that listing is OpenRouter's Stealth Program EULA, and it is worth re-reading on reveal day for what it does not say: no clause covers endpoint persistence, pricing continuity, or serving-config carryover once a provider is named. The reveal transition is silent by design. The EULA's only removal clause is unconditional:
"Stealth Models may be removed from our Stealth Program at any time upon request of the Stealth Provider or at OpenRouter's sole discretion, with or without notice to you."— OpenRouter Stealth Program EULA, retrieved August 26, 2026
The EULA's training posture during the anonymous period is equally blunt — opt-out by abstention: "If you do not want your User Content to be provided to Stealth Providers for Stealth Model training, then you should refrain from accessing or using the Stealth Models." There is no per-request or account-level toggle. Individual stealth listings can override that default — ox-alpha's archived listing carried its own no-training line — which is precisely why the archived per-model text matters more than the general EULA.
Archived listing price
The Wayback capture of August 24 preserves the listing at zero cost with its own no-training line. Screenshot-grade evidence of the terms you actually operated under.
Guaranteed notice
The Stealth Program EULA permits removal at any time, at the provider's request or OpenRouter's sole discretion, with or without notice. No notice is owed to you at all.
Verify this yourself
OpenRouter publishes provider-specific supplemental terms for individual stealth models. Whether ox-alpha carried any could not be resolved from archived text — check the page for any stealth model you still use.
That third card is the first of this post's four verify-yourself items. OpenRouter maintains a separate Stealth Program Supplemental Terms page listing additional, provider-specific terms attached to individual stealth models. Our research could not resolve whether ox-alpha carried supplemental terms during its stealth window — the page returns an index, not archived per-model text. So the item reads as an instruction, not a risk: if you used any stealth listing, check the supplemental page for it, and archive what you find with a date on it.
03 — Check 1 · EndpointFind every config that pins the dead slug.
This check was impossible yesterday for the simplest reason: there was no named successor to migrate to. Today there is. A direct fetch of the named model's OpenRouter page on August 26 states the stealth slug was revealed to be this model and shows no continuing or parallel stealth/ox-alpha route. Treat "the free slug is gone" as the working assumption — it is one fetch's read of a live page, so re-verify — and act on it: grep every config, environment variable, and routing rule for the stealth slug, today, while the people who added it still remember why.
The documented pattern is that the stealth slug does not survive as itself. Per our census of resolved stealth listings, every other Alpha-named slot has been retired or renamed out of the live API: Quasar Alpha and Optimus Alpha went when OpenRouter confirmed them as GPT-4.1 prereleases, and the two Sonoma listings went under a reported but never-confirmed Grok 4 Fast identity. A stealth slug is scaffolding; the reveal is when it comes down. That is a structurally worse version of the pinning problem a floating alias creates — an alias at least keeps resolving to something.
stealth/ox-alpha reference returns after retirement — an error, a redirect, or a silent failure — is not documented anywhere we could find; we could not locate an OpenRouter support page or changelog that states it. Do not assume any of the three: send one canary request against the pinned slug from a staging environment, log exactly what comes back, and build your migration around the observed behavior rather than a guess.04 — Check 2 · PricingRead the price ladder by surface and date, not by headline.
"Free" was never a price — it was step one of a ladder, and the reveal is when the ladder becomes visible. Our own stealth-model census sketched the shape as anonymous-then-list, and we got half of it wrong: the free slug does appear to be gone, but the named model did not land at list price. It landed at a promo.
| Period · surface | Input / output per 1M | What triggered it | Where to verify |
|---|---|---|---|
| Anonymous · stealth listing | $0 / $0 | Stealth Program listing; archived at $0 on August 24 | Wayback Machine capture of the listing, 2026-08-24 |
| Promo · standard surface (Aug 26 – Sep 9) | $0.075 / $0.25 ($0.015 cached in) | Reveal-day launch promo — exactly 50% off list on all three rates | Z.ai pricing docs; OpenRouter provider table |
| List · standard surface (after Sep 9) | $0.15 / $0.50 ($0.03 cached in) | Promo expiry, vendor-stated as September 9 at 24:00 UTC+8; no grandfathering clause found | Z.ai pricing docs — re-check before Sep 9 |
| Batch surface · today (Aug 26) | $0.15 / $0.50 | Separate surface pricing — already prints the post-promo list numbers | OpenRouter :batch listing for the named model |
Three disciplines fall out of that table. First, never quote a price without naming its surface and its date. On August 26, "GLM-5.3-Flash costs $0.075 per million in" is true of the standard surface and false of the batch surface; "batch is cheaper" is simply false on this dateline, because the :batch surface already prints $0.15/$0.50 while the standard surface holds the promo. Second, the promo is exactly a 50% discount — $0.075 doubles to $0.15, $0.25 to $0.50, $0.015 to $0.03 — confirmed on Z.ai's own pricing page and independently by OpenRouter's provider table, so every cost model built on today's rates doubles after September 9 (the precise 24:00 UTC+8 cutoff is vendor-stated via secondary coverage — confirm it on the primary before you rely on the hour). Third, Z.ai's pricing page states no policy for in-flight work at the deadline — no lock-in, no grandfathering, no proration. That is a genuine gap, not an inference: assume no grandfathering unless you find a clause saying otherwise, and re-run the cost model at list price before September 9.
The reveal also exposed something a single anonymous listing hides: the same named model is not one price. Twenty-one providers served z-ai/glm-5.3-flash on OpenRouter on reveal day, at three distinct price tiers.
One named model, three prices · 21 OpenRouter providers, reveal day
Bar length = tier's input price as a share of the $0.15/1M list input rate. Source: OpenRouter z-ai/glm-5.3-flash provider table, retrieved Aug 26, 202605 — Check 3 · Serving ConfigThe named docs publish defaults you could not read yesterday.
The strongest now-answerable check the reveal enables is also the one with an immediate invoice attached: the named endpoint's own documentation shows reasoning effort defaulting to its maximum setting, with no way to disable thinking. While the model was anonymous, there were no vendor docs to read — you could observe verbose outputs, but you could not know they were a documented, non-optional default. Now you can, and the cost consequence for high-volume workloads is direct. Our reveal-day analysis of GLM-5.3-Flash covers the mechanics; the checklist action is to read the named model's parameter defaults today and re-estimate output-token spend under them.
The other serving-config item is a verify-yourself entry. With 21 providers serving one named slug, it is reasonable to ask whether quantization or serving configuration differs across them. Our research confirmed only price variance — the provider table did not surface a quantization column for this model, so provider-level config variance is plausible but not established. Do not treat it as a fact; treat it as a check: before you let a benchmark run against one provider stand for "the model," confirm which provider served it and what that provider discloses about its serving setup.
06 — Check 4 · Data TermsRead the terms that now apply — for the surface you actually landed on.
During anonymity, the applicable data terms were a structural unknowable: the general stealth EULA's train-by-default posture, modified per listing, with abstention as the only opt-out. The reveal changes the kind of question this is. The terms that govern you going forward are the named vendor's own published terms — a document with an author you can name in a contract, and one that can differ meaningfully by product tier within the same vendor. That last clause is where "read the licence" collapses into a single line item and hides the real work.
Alibaba's Qwen line supplies a second, independent instance of the problem. Qwen3.8-Flash-Next and Qwen3.8-Flash are two different artifacts, not one product with two names. Which one your team "adopted" determines which terms bind you:
Qwen3.8-Flash-Next
Downloadable weights under a community licence. The licence text — not the family name — is what binds you. What hosting them takes is a different post's territory; here the action is only: read it.
Qwen3.8-Flash
A hosted service under Alibaba Cloud Model Studio terms. General API access carries a stated no-training policy; the Coding Plan subscription carries an explicit training authorization. Same vendor, different posture per tier.
The Alibaba example is worth spelling out because the tier, not the vendor name, is what sets your training-data posture. Alibaba Cloud states that for standard Model Studio API usage it "will never use your data for model training." Its Coding Plan subscription is the named exception: "By using the Coding Plan, you authorize use of your model inputs and generated content to improve the service and optimize models" (Alibaba Cloud Model Studio — Coding Plan). And the revocation is forward-only: discontinuing the plan revokes future authorization but "does not apply to data already authorized." For GLM-5.3-Flash the same read-the-named-terms action applies on Z.ai's side, and the reveal came bundled with an MIT licence on the Hugging Face repo — which covers the weights, not the hosted API's data handling. Our census of free- and promo-tier data terms maps this landscape vendor by vendor.
07 — ScorecardWhat the reveal resolves — and what it doesn't.
"We know who it is now" quietly overstates what a name delivers. The table below takes dimensions that were unknowable while the model was anonymous — not a complete inventory of everything an anonymous endpoint hides — and checks each against what August 26 actually made knowable.
| Unknowable while anonymous | Resolved? | What you can now check | Still open |
|---|---|---|---|
| Resolved — the reveal answered it | |||
| Vendor identity and accountable entity | Yes | Z.ai, named on its own announcement — a counterparty you can put in a contract and a jurisdiction you can look up | — |
| Licence for downstream use | Yes | MIT, verified on the Hugging Face repo — the text exists to be read, today | The weights licence does not govern the hosted API's data handling |
| Documented cost defaults | Yes | Published parameter defaults — reasoning effort at maximum, thinking not disableable — plus a dated price ladder | How in-flight work is treated at the September 9 promo deadline is unpublished |
| Partially resolved — a check exists, an answer does not | |||
| Applicable data terms | Partial | The named vendor's published terms govern you going forward, tier by tier | Whether ox-alpha carried supplemental terms during the stealth window — archived per-model text unresolved |
| Serving configuration | Partial | Price variance is visible — 21 providers, 3 tiers, one day | Whether quantization or serving config differs per provider — verify before generalizing any benchmark |
| Not resolved — the name does not settle it | |||
| Benchmark contamination risk | No | You can now ask a named vendor about training data and read its disclosures | Nothing retroactively certifies that your stealth-era eval runs were contamination-free — a name is someone to ask, not a corrected score |
One more entry belongs in the scorecard: the shape of the transition itself. The named model landed at a 50%-off promo, not at list — which means naive "reveal equals list price" logic fails in both directions: it overestimates what you pay today and underestimates the jump coming on September 9. The general lesson projects forward to the next reveal, whoever ships it: expect the transition to be a ladder with a promotional middle rung and a dated cliff, not a single step, and diary the cliff the day you learn its date. This whole re-verification pass — endpoint, ladder, config, terms — is a few focused hours for one team; it is also exactly the kind of dependency audit we run inside our AI transformation engagements when the model surface under a production system shifts.
08 — ConclusionTreat the name as a work order.
A name is an address for your questions, not an answer to them.
The comfortable reading of a reveal — "now we know what we were using" — is exactly wrong. What you knew yesterday was an endpoint's behavior; what you learned today is an identity. The gap between those two is where the four checks live: whether your pinned slug still resolves, which surface and date your price quote belongs to, whether the config you benchmarked is the config you now call, and which named tier's data terms govern you.
Run the checks in cost order. The pinned-slug grep is minutes and prevents a silent failure. The price-ladder diary is minutes and prevents a doubled invoice after September 9. The eval re-run is an afternoon and is the only empirical link between the endpoint you tested and the one you now pay for. The terms read is the slowest and the most durable — archive dated copies, because this post's entire method depended on someone having archived the before state.
And keep the humility the scorecard forces: one unknowable survives every reveal. Knowing who trained the model gives you someone to ask about contamination; it does not clean your old scores. The teams that handle reveals well are not the ones that celebrate the name — they are the ones that treat it as the day the real verification work finally became possible, and do it.