Evaluating an anonymous AI model is two different jobs that most teams collapse into one. The first job — measuring what the model does — works fine without a name: latency, tool-call reliability, refusal behavior, and output stability are all observable from your side of the API. The second job — establishing who stands behind the model — does not work at all, because there is no one to ask.
The distinction stopped being academic the moment stealth listings became a normal way to ship a frontier-class model. When ox-alpha first appeared on OpenRouter, teams started routing real work through an endpoint whose maker will not identify itself. Any theory about which lab built it is community fingerprinting, not a fact.
This guide is the due-diligence half of the eval problem. Our standard metric toolkit covers how to measure quality; this post covers what an identity buys you that no benchmark can — and how to run honest empirical tests on the parts that remain measurable. The spine is an explicit inventory of the unknowables: training-data provenance, retention, deprecation risk, license, and liability.
- 01A benchmark score and an identity are different products.A score tells you nothing a vendor couldn't have trained toward; an identity tells you who retains your data and who is liable when the model is wrong. No amount of the first substitutes for the second.
- 02The stealth default is “train, evaluate, and improve” — with no per-user opt-out.OpenRouter's Stealth Program EULA grants stealth providers the right to train on submitted content, and states the only opt-out is not using the stealth models at all. Listings can be removed at any time, with or without notice.
- 03Named vendors publish exactly what a stealth listing cannot.OpenAI (30-day default retention, opt-in training), Anthropic (7-day log retention, express-permission training), and Google (paid-tier no-training commitment) each name a legal entity behind each promise.
- 04Plenty is still measurable without a name.Latency, context behavior under load, tool-call reliability, refusal surface, and rerun stability are all testable client-side with your own logs and repeated calls — none of it requires, or produces, any information about provenance or liability.
- 05Private eval suites are an aspiration, not armor.Contamination-resistant benchmarks are an open research problem, and contamination detectors substantially undercount for reasoning models. Build held-out suites — and treat your own “clean” results with vendor-benchmark skepticism.
01 — The DistinctionA score is not an identity.
Every model evaluation you can run produces evidence about behavior: how the model answers your prompts, today, under your conditions. That evidence is genuinely useful and the second half of this post is about collecting it honestly. But behavior is also the one thing a motivated vendor can shape — anything that has ever appeared in a public benchmark could have been trained toward, which is precisely why an anonymous model arriving with strong self-reported numbers proves so little.
An identity produces a different kind of evidence entirely: commitments. A named vendor publishes a retention window, a training-use default, and a deprecation policy — and attaches a legal entity to each one. When the model misbehaves, that entity owns a changelog, a status page, and a support queue. When its output infringes someone's copyright, that entity is the counterparty your contract points at. None of this is observable in a response payload. It exists only in published terms, signed by someone.
That is the frame for everything below. The question is never "is the anonymous model good?" — you can test that. The question is "what am I giving up by not knowing who made it?" — and the honest answer is a specific, enumerable list, most of which cannot be recovered by any amount of clever testing.
02 — Stealth TermsWhat the stealth terms actually say.
The best thing about OpenRouter's Stealth Program is that its rules are written down. The Stealth Program EULA (last updated July 6, 2026) defines "Stealth Provider," "Stealth Model," and "User Content" as formal terms, and then says three things every team should read before sending a single production prompt to an anonymous endpoint.
Train, evaluate, improve
Stealth providers “may train, evaluate, and improve” listed models using submitted User Content — and “may not further sublicense or otherwise use User Content for any other purpose.” That second clause is a limit on resale, not an opt-out from training.
Refrain from use
The EULA's stated remedy for users who don't want their content used for stealth training is not using stealth models at all. Content passes to the provider under a hashed identifier, so the individual user is not directly identifiable to them.
Any time, no notice
Stealth listings can be removed “at any time upon request of the Stealth Provider or at OpenRouter's sole discretion, with or without notice to you.” A listing being live today carries no forward guarantee of any kind.
“If you do not want your User Content to be provided to Stealth Providers for Stealth Model training, then you should refrain from accessing or using the Stealth Models.”— OpenRouter Stealth Program EULA, last updated July 6, 2026
The contrast with the rest of the platform is instructive. OpenRouter's own account settings expose prompt-logging, chat-logging, zero-data-retention, and model-training toggles at the organization level — these are user-controllable dimensions for named providers generally. A stealth listing's provider-side policy is not controllable the same way. And the platform's main Terms of Service (updated July 29, 2026) add two disclaimers that matter here: OpenRouter states it "strives to accurately represent the status of prompt logging and training for each Model on our Site" while disclaiming liability for errors in Model Terms — so even the listing page's description of a stealth model's data policy is not warranted — and it is "not responsible for any incorrect location reporting to the Models," which undercuts any residency assumption you might be tempted to make about an anonymous endpoint.
03 — The ContrastWhat a named vendor discloses.
To see the shape of the hole an anonymous listing leaves, look at what three named API vendors publish — OpenAI's enterprise privacy commitments, Anthropic's API data-retention documentation, and Google's Gemini API data-logging policy. This is not a tutorial on their data practices — it's a contrast instrument. Each of the commitments below is a published document with a legal entity's name on it, which is exactly the artifact a stealth listing cannot produce.
Default retention cap
API inputs and outputs are retained up to 30 days for abuse monitoring, then deleted unless legally required otherwise; Zero Data Retention is available for eligible endpoints. API data (after March 1, 2023) is not used for training by default — opt-in only.
Standard log retention
Reduced from 30 days effective September 14, 2025. Retained data is “never used for model training without your express permission.” Zero Data Retention is available for qualifying enterprise API customers.
Paid vs unpaid split
On paid usage, Google states it does not use prompts or responses to improve its products and processes them under a Data Processing Addendum. On unpaid usage, data may be used to improve model quality. Both halves of that split matter.
The structural point is what these documents provide: a training-use default attached to a named legal entity — and, for OpenAI and Anthropic, a published retention window. Neither of those attributes exists for an anonymous provider: no named retention SLA, and no auditable training default beyond what the aggregator chooses to disclose on the listing's behalf.
04 — Identity LedgerThe Identity Ledger: what a name buys you, row by row.
The table below is this post's core deliverable — the inventory, not a checklist. Each row is one thing procurement or engineering eventually asks about a model. The third column is the honest kicker: for most rows, there is no test you can run.
| Dimension | Named vendor (API pattern) | Anonymous stealth listing | How you'd find out yourself |
|---|---|---|---|
| Training-data provenance | Undisclosed in detail, but a named entity you can put contractual representations to | No counterparty; any lab-identity theory is community fingerprinting, not fact | Cannot be established client-side |
| Training on your prompts | OpenAI: opt-in only. Anthropic: express permission required. Google: not used on paid tier | Default is "may train, evaluate, and improve"; only opt-out is not using the model | Not verifiable from outside — even the aggregator disclaims warranty for its own listing accuracy |
| Retention window | Published: up to 30 days (OpenAI, page updated Jan 8, 2026); 7 days (Anthropic, effective Sep 14, 2025); ZDR options exist | No named window; content reaches the provider under a hashed identifier | Untestable — retention happens on someone else's infrastructure |
| Deprecation / withdrawal notice | An accountable party; OpenRouter itself recommends fallback-chain engineering, citing 70+ models pulled or deprecated (no stated window) | Removable "at any time... with or without notice to you" | You find out when requests fail; fallback chains are mitigation, not information |
| IP / output indemnity | A named party that can sign one; the terms are whatever its own commercial agreement says | No counterparty to indemnify against — the clause cannot exist | Cannot be established; there is no entity to sign |
| Liability counterparty | A named legal entity | OpenRouter disclaims all liability for a provider's acts or omissions; the provider is unreachable by design | Read the aggregator's ToS — that document is the entire record |
| Incident / changelog trail | Changelog entry, status-page incident, or support ticket with an accountable party | A GitHub issue full of speculation — no vendor changelog or postmortem is possible | Your own request/response logs are the only durable record |
05 — Measurable AnywayWhat you can test without a name.
None of the above means empirical evaluation is pointless — it means being precise about what it produces. Four dimensions are fully measurable client-side, against any endpoint, using nothing but your own request/response logs and repeated calls. No cooperation from the provider is required, and no identity is revealed by the results.
Response behavior under load
Time-to-first-token, throughput, and context-window behavior at increasing fill are directly observable per request. Log every call; the distribution is the deliverable, not the best run.
The observed weak spot
Send known tool-call payloads and assert non-empty structured output. A live GitHub issue documents an anonymous stealth endpoint returning silent empty responses on certain tool-call payloads — exactly the failure a harness catches and a listing page never mentions.
Policy behavior mapping
A fixed adversarial and edge-case prompt set, run and logged, maps what the model declines and how. This is the held-out-twin pattern applied to policy behavior rather than accuracy.
Rerun variance
Identical prompts, repeated runs, scored deltas. Stability is a property you measure, not one you assume — section 07 puts observed numbers on how much scores move between identical runs.
The tool-call case deserves the extra attention because it's concrete: the issue filed against a third-party agent framework documents stealth/ox-alpha returning silent empty, zero-token responses on certain tool-call payloads — a dated, reproducible divergence from a named model's expected tool-call contract, with no vendor changelog or postmortem possible because there is no named vendor to publish one. If your workload is agentic, build the harness before you route traffic; a full agent-eval pipeline is the systematic version of this.
And hold on to the boundary: none of these tests require — or produce — any information about training-data provenance, retention practice, or legal liability. They describe what the model does, not what was done to build it or who stands behind it. A team that runs all four beautifully has still learned nothing from the Identity Ledger's rows.
06 — Private SuitesPrivate suites the vendor can't have trained on.
The standard defense against benchmark gaming is a private task suite — items the model's builders have never seen. Against an anonymous model this is not optional: since you cannot know the training cutoff, the lab, or the data pipeline, a public benchmark score is unfalsifiable as evidence of generalization. Three construction patterns from the current literature apply.
The held-out twin. Build a second benchmark matching an existing one's distribution, difficulty, and scoring rubric — but composed of items that did not exist when a candidate model's training data could have been frozen. Comparing scores on the public split versus the twin is the practical contamination-detection technique named across this literature.
Disposable synthetic suites. Work in the S3Eval strand generates large numbers of synthetic tasks specifically because they postdate any given model's training cutoff and are cheap to keep regenerating. The design goal is disposability — a suite you can throw away and remint — rather than a one-time held-out set that ages into the next training crawl.
Contamination-resistant design. A May 2026 arXiv position paper ("LLM Benchmark Datasets Should Be Contamination-Resistant", Al-Lawati, Lucas, Lee, Wang) argues public benchmark datasets should be engineered to be contamination-resistant — unlearnable while still supporting inference. Note the framing: this is a still-open research direction, not a shipped methodology you can adopt today.
07 — Run DisciplineRerun discipline: the worst repeated result is the real one.
Suite construction is half the job; run discipline is the other half, and it's where single-run scores quietly lie. Applied eval-engineering guidance from Braintrust (Jess Wang, March 6, 2026) recommends starting a private suite with as few as five representative examples chosen to hit different layout and behavior challenges rather than attempting comprehensive coverage up front — the cited case study then ran 70+ experiments across 11 prompt versions over two days.
The same case study measured concrete rerun variance: at temperature 0.7, three identical runs of one task showed a 10% swing in one scored dimension and 13% in another; at temperature 0.0, output was deterministic. The stated practical rule: run a candidate configuration three times before trusting a score, and take the worst repeated result as the real number — "If v10 scored 50% once and 40% twice, the real performance is 40%."
Minimum viable private suite
Applied guidance, not a statistical minimum: five representative examples chosen for coverage of distinct behaviors, then expand. The point is to start iterating, not to certify.
Before trusting any score
Run each candidate config three times; take the worst repeated result as the real number. A single strong run is indistinguishable from variance.
Swing at temperature 0.7
Three identical runs of one task: 10% swing on one scored dimension, 13% on another. At temperature 0.0 the same task was deterministic — variance is a setting, not fate.
For an anonymous endpoint this discipline does double duty — repeated runs are also how you detect behavior shifts on a model that will never announce them. And because reruns cost real tokens, score against cost-per-successful-task rather than raw accuracy — the metric that survives a model being cheap, chatty, and wrong.
08 — Longevity RiskThe listing can vanish mid-sprint.
Model disappearance is routine platform behavior, not an edge case. OpenRouter's own tutorial content states that more than 70 models have been pulled or deprecated by providers "in the last few years" — an unscoped count with no stated population or fixed window, so resist turning it into a rate — and recommends fallback-chain engineering as the mitigation. That's for the platform at large, named models included. A stealth listing adds the EULA's explicit removal clause on top: "Stealth Models may be removed from our Stealth Program at any time upon request of the Stealth Provider or at OpenRouter's sole discretion, with or without notice to you."
Historically, stealth previews resolve in one of two ways: they graduate to a named, priced model, or they are withdrawn. The binary itself is the risk signal — "still listed today" carries no forward guarantee, and the record of how often stealth listings turn out to be who you'd guess is a separate question from whether yours will exist on Friday. Plan for the endpoint's disappearance the way you plan for a spot instance's: checkpoint your prompts, keep a fallback chain warm, and never let an anonymous listing become a single point of failure.
The forward projection follows from the incentives. Stealth listings are cheap distribution for labs that want frontier-scale feedback without frontier-scale accountability, and the EULA's train-by-default terms are what make free access economically rational — so expect more anonymous endpoints, not fewer. The teams that benefit will be the ones that treat them as what they are: a free trial with data-terms pricing, useful for scoped, non-sensitive workloads, and never a substitute for a counterparty. If you're formalizing that boundary for your own organization — which workloads may touch an anonymous endpoint, which require a named counterparty — our AI transformation engagements build exactly that kind of model-governance policy alongside the eval harness.
09 — ConclusionEvaluate the behavior, price the silence.
A benchmark tells you what a model does. Only an identity tells you who answers for it.
The empirical half of anonymous-model evaluation is genuinely tractable: latency, tool-call reliability, refusal surface, and rerun stability are all measurable from your own logs, and a held-out private suite — built with the twin pattern, run three times, scored by the worst repeated result — is the honest way to read capability claims no one will put their name to.
The other half is not tractable, and pretending otherwise is the real failure mode. Training-data provenance, training-on-prompts defaults, retention windows, withdrawal notice, and liability are all things a name discloses and a stealth listing structurally cannot. No harness recovers what only a counterparty can sign.
So run both ledgers. Test what is testable, price what is silent, and route accordingly: anonymous endpoints for scoped experiments where the data terms are acceptable and disappearance is survivable; named vendors for everything a client, regulator, or court might one day ask about.