AI DevelopmentFramework11 min readPublished August 25, 2026

Benchmarks measure behavior · identity assigns liability · the stealth terms put the difference in writing

Evaluating an Anonymous Model: What You Can Actually Know

An anonymous endpoint can be benchmarked all day and still tell you nothing about who trains on your prompts, how long they keep them, or who answers when it breaks. This framework separates what you can establish empirically from what stays structurally unknowable without a name — sourced to the published terms on both sides of that line.

DA
Digital Applied Team
Senior strategists · Published Aug 25, 2026
PublishedAug 25, 2026
Read time11 min
Sources10 primary
Stealth training opt-outs
0
per-user, under the Stealth EULA
only exit: don't use it
OpenAI API retention
30d
default cap, then deletion unless legally required
Anthropic API log retention
7d
reduced from 30 · Sep 14, 2025
−23 days
Models pulled or deprecated
70+
per OpenRouter · no stated window

Evaluating an anonymous AI model is two different jobs that most teams collapse into one. The first job — measuring what the model does — works fine without a name: latency, tool-call reliability, refusal behavior, and output stability are all observable from your side of the API. The second job — establishing who stands behind the model — does not work at all, because there is no one to ask.

The distinction stopped being academic the moment stealth listings became a normal way to ship a frontier-class model. When ox-alpha first appeared on OpenRouter, teams started routing real work through an endpoint whose maker will not identify itself. Any theory about which lab built it is community fingerprinting, not a fact.

This guide is the due-diligence half of the eval problem. Our standard metric toolkit covers how to measure quality; this post covers what an identity buys you that no benchmark can — and how to run honest empirical tests on the parts that remain measurable. The spine is an explicit inventory of the unknowables: training-data provenance, retention, deprecation risk, license, and liability.

Key takeaways
  1. 01
    A benchmark score and an identity are different products.A score tells you nothing a vendor couldn't have trained toward; an identity tells you who retains your data and who is liable when the model is wrong. No amount of the first substitutes for the second.
  2. 02
    The stealth default is “train, evaluate, and improve” — with no per-user opt-out.OpenRouter's Stealth Program EULA grants stealth providers the right to train on submitted content, and states the only opt-out is not using the stealth models at all. Listings can be removed at any time, with or without notice.
  3. 03
    Named vendors publish exactly what a stealth listing cannot.OpenAI (30-day default retention, opt-in training), Anthropic (7-day log retention, express-permission training), and Google (paid-tier no-training commitment) each name a legal entity behind each promise.
  4. 04
    Plenty is still measurable without a name.Latency, context behavior under load, tool-call reliability, refusal surface, and rerun stability are all testable client-side with your own logs and repeated calls — none of it requires, or produces, any information about provenance or liability.
  5. 05
    Private eval suites are an aspiration, not armor.Contamination-resistant benchmarks are an open research problem, and contamination detectors substantially undercount for reasoning models. Build held-out suites — and treat your own “clean” results with vendor-benchmark skepticism.

01The DistinctionA score is not an identity.

Every model evaluation you can run produces evidence about behavior: how the model answers your prompts, today, under your conditions. That evidence is genuinely useful and the second half of this post is about collecting it honestly. But behavior is also the one thing a motivated vendor can shape — anything that has ever appeared in a public benchmark could have been trained toward, which is precisely why an anonymous model arriving with strong self-reported numbers proves so little.

An identity produces a different kind of evidence entirely: commitments. A named vendor publishes a retention window, a training-use default, and a deprecation policy — and attaches a legal entity to each one. When the model misbehaves, that entity owns a changelog, a status page, and a support queue. When its output infringes someone's copyright, that entity is the counterparty your contract points at. None of this is observable in a response payload. It exists only in published terms, signed by someone.

That is the frame for everything below. The question is never "is the anonymous model good?" — you can test that. The question is "what am I giving up by not knowing who made it?" — and the honest answer is a specific, enumerable list, most of which cannot be recovered by any amount of clever testing.

02Stealth TermsWhat the stealth terms actually say.

The best thing about OpenRouter's Stealth Program is that its rules are written down. The Stealth Program EULA (last updated July 6, 2026) defines "Stealth Provider," "Stealth Model," and "User Content" as formal terms, and then says three things every team should read before sending a single production prompt to an anonymous endpoint.

Data default
Train, evaluate, improve
the EULA's stated default

Stealth providers “may train, evaluate, and improve” listed models using submitted User Content — and “may not further sublicense or otherwise use User Content for any other purpose.” That second clause is a limit on resale, not an opt-out from training.

EULA · updated Jul 6, 2026
The only opt-out
Refrain from use
no per-user toggle

The EULA's stated remedy for users who don't want their content used for stealth training is not using stealth models at all. Content passes to the provider under a hashed identifier, so the individual user is not directly identifiable to them.

No per-user opt-out
Withdrawal
Any time, no notice
provider request or platform discretion

Stealth listings can be removed “at any time upon request of the Stealth Provider or at OpenRouter's sole discretion, with or without notice to you.” A listing being live today carries no forward guarantee of any kind.

Removal clause
“If you do not want your User Content to be provided to Stealth Providers for Stealth Model training, then you should refrain from accessing or using the Stealth Models.”— OpenRouter Stealth Program EULA, last updated July 6, 2026

The contrast with the rest of the platform is instructive. OpenRouter's own account settings expose prompt-logging, chat-logging, zero-data-retention, and model-training toggles at the organization level — these are user-controllable dimensions for named providers generally. A stealth listing's provider-side policy is not controllable the same way. And the platform's main Terms of Service (updated July 29, 2026) add two disclaimers that matter here: OpenRouter states it "strives to accurately represent the status of prompt logging and training for each Model on our Site" while disclaiming liability for errors in Model Terms — so even the listing page's description of a stealth model's data policy is not warranted — and it is "not responsible for any incorrect location reporting to the Models," which undercuts any residency assumption you might be tempted to make about an anonymous endpoint.

03The ContrastWhat a named vendor discloses.

To see the shape of the hole an anonymous listing leaves, look at what three named API vendors publish — OpenAI's enterprise privacy commitments, Anthropic's API data-retention documentation, and Google's Gemini API data-logging policy. This is not a tutorial on their data practices — it's a contrast instrument. Each of the commitments below is a published document with a legal entity's name on it, which is exactly the artifact a stealth listing cannot produce.

OpenAI API
Default retention cap
30days

API inputs and outputs are retained up to 30 days for abuse monitoring, then deleted unless legally required otherwise; Zero Data Retention is available for eligible endpoints. API data (after March 1, 2023) is not used for training by default — opt-in only.

Page updated Jan 8, 2026
Anthropic API
Standard log retention
7days

Reduced from 30 days effective September 14, 2025. Retained data is “never used for model training without your express permission.” Zero Data Retention is available for qualifying enterprise API customers.

Effective Sep 14, 2025
Google Gemini API
Paid vs unpaid split
2tiers

On paid usage, Google states it does not use prompts or responses to improve its products and processes them under a Data Processing Addendum. On unpaid usage, data may be used to improve model quality. Both halves of that split matter.

Per Gemini API terms

The structural point is what these documents provide: a training-use default attached to a named legal entity — and, for OpenAI and Anthropic, a published retention window. Neither of those attributes exists for an anonymous provider: no named retention SLA, and no auditable training default beyond what the aggregator chooses to disclose on the listing's behalf.

04Identity LedgerThe Identity Ledger: what a name buys you, row by row.

The table below is this post's core deliverable — the inventory, not a checklist. Each row is one thing procurement or engineering eventually asks about a model. The third column is the honest kicker: for most rows, there is no test you can run.

The Identity Ledger: what a named API vendor discloses versus what an anonymous stealth listing discloses, and whether each dimension can be established empirically without an identity
DimensionNamed vendor (API pattern)Anonymous stealth listingHow you'd find out yourself
Training-data provenanceUndisclosed in detail, but a named entity you can put contractual representations toNo counterparty; any lab-identity theory is community fingerprinting, not factCannot be established client-side
Training on your promptsOpenAI: opt-in only. Anthropic: express permission required. Google: not used on paid tierDefault is "may train, evaluate, and improve"; only opt-out is not using the modelNot verifiable from outside — even the aggregator disclaims warranty for its own listing accuracy
Retention windowPublished: up to 30 days (OpenAI, page updated Jan 8, 2026); 7 days (Anthropic, effective Sep 14, 2025); ZDR options existNo named window; content reaches the provider under a hashed identifierUntestable — retention happens on someone else's infrastructure
Deprecation / withdrawal noticeAn accountable party; OpenRouter itself recommends fallback-chain engineering, citing 70+ models pulled or deprecated (no stated window)Removable "at any time... with or without notice to you"You find out when requests fail; fallback chains are mitigation, not information
IP / output indemnityA named party that can sign one; the terms are whatever its own commercial agreement saysNo counterparty to indemnify against — the clause cannot existCannot be established; there is no entity to sign
Liability counterpartyA named legal entityOpenRouter disclaims all liability for a provider's acts or omissions; the provider is unreachable by designRead the aggregator's ToS — that document is the entire record
Incident / changelog trailChangelog entry, status-page incident, or support ticket with an accountable partyA GitHub issue full of speculation — no vendor changelog or postmortem is possibleYour own request/response logs are the only durable record
The liability row, in the platform's own words
The sharpest row is the least discussed. OpenRouter's Terms of Service state: "OpenRouter disclaims all liability for any suspension, restriction, disabling, termination, removal, unavailability, degradation or modification of any Model arising from or related to Model Terms of the acts or omissions of any Model Provider." That clause applies to every listed model — but for a named model, there is still a provider with a name behind the aggregator. For an anonymous one, the disclaimer is where the chain of accountability ends.

05Measurable AnywayWhat you can test without a name.

None of the above means empirical evaluation is pointless — it means being precise about what it produces. Four dimensions are fully measurable client-side, against any endpoint, using nothing but your own request/response logs and repeated calls. No cooperation from the provider is required, and no identity is revealed by the results.

Latency & context
Response behavior under load

Time-to-first-token, throughput, and context-window behavior at increasing fill are directly observable per request. Log every call; the distribution is the deliverable, not the best run.

Testable client-side
Tool-call reliability
The observed weak spot

Send known tool-call payloads and assert non-empty structured output. A live GitHub issue documents an anonymous stealth endpoint returning silent empty responses on certain tool-call payloads — exactly the failure a harness catches and a listing page never mentions.

Testable — and where anonymous endpoints have failed
Refusal surface
Policy behavior mapping

A fixed adversarial and edge-case prompt set, run and logged, maps what the model declines and how. This is the held-out-twin pattern applied to policy behavior rather than accuracy.

Testable client-side
Output stability
Rerun variance

Identical prompts, repeated runs, scored deltas. Stability is a property you measure, not one you assume — section 07 puts observed numbers on how much scores move between identical runs.

Testable client-side

The tool-call case deserves the extra attention because it's concrete: the issue filed against a third-party agent framework documents stealth/ox-alpha returning silent empty, zero-token responses on certain tool-call payloads — a dated, reproducible divergence from a named model's expected tool-call contract, with no vendor changelog or postmortem possible because there is no named vendor to publish one. If your workload is agentic, build the harness before you route traffic; a full agent-eval pipeline is the systematic version of this.

And hold on to the boundary: none of these tests require — or produce — any information about training-data provenance, retention practice, or legal liability. They describe what the model does, not what was done to build it or who stands behind it. A team that runs all four beautifully has still learned nothing from the Identity Ledger's rows.

06Private SuitesPrivate suites the vendor can't have trained on.

The standard defense against benchmark gaming is a private task suite — items the model's builders have never seen. Against an anonymous model this is not optional: since you cannot know the training cutoff, the lab, or the data pipeline, a public benchmark score is unfalsifiable as evidence of generalization. Three construction patterns from the current literature apply.

The held-out twin. Build a second benchmark matching an existing one's distribution, difficulty, and scoring rubric — but composed of items that did not exist when a candidate model's training data could have been frozen. Comparing scores on the public split versus the twin is the practical contamination-detection technique named across this literature.

Disposable synthetic suites. Work in the S3Eval strand generates large numbers of synthetic tasks specifically because they postdate any given model's training cutoff and are cheap to keep regenerating. The design goal is disposability — a suite you can throw away and remint — rather than a one-time held-out set that ages into the next training crawl.

Contamination-resistant design. A May 2026 arXiv position paper ("LLM Benchmark Datasets Should Be Contamination-Resistant", Al-Lawati, Lucas, Lee, Wang) argues public benchmark datasets should be engineered to be contamination-resistant — unlearnable while still supporting inference. Note the framing: this is a still-open research direction, not a shipped methodology you can adopt today.

A clean detector result proves less than you think
Contamination detection itself is fragile for reasoning models. A 2025-cycle paper found that because reasoning models train on chain-of-thought traces, not just question/answer pairs, standard contamination detectors — which only see the final Q/A pair — substantially undercount contamination. A "clean" detector result does not prove an anonymous model wasn't trained on your eval set. Treat "held-out eval" claims from any source — including your own — with the same skepticism you'd apply to a vendor benchmark.

07Run DisciplineRerun discipline: the worst repeated result is the real one.

Suite construction is half the job; run discipline is the other half, and it's where single-run scores quietly lie. Applied eval-engineering guidance from Braintrust (Jess Wang, March 6, 2026) recommends starting a private suite with as few as five representative examples chosen to hit different layout and behavior challenges rather than attempting comprehensive coverage up front — the cited case study then ran 70+ experiments across 11 prompt versions over two days.

The same case study measured concrete rerun variance: at temperature 0.7, three identical runs of one task showed a 10% swing in one scored dimension and 13% in another; at temperature 0.0, output was deterministic. The stated practical rule: run a candidate configuration three times before trusting a score, and take the worst repeated result as the real number — "If v10 scored 50% once and 40% twice, the real performance is 40%."

Starting size
Minimum viable private suite
5examples

Applied guidance, not a statistical minimum: five representative examples chosen for coverage of distinct behaviors, then expand. The point is to start iterating, not to certify.

Braintrust · Mar 6, 2026
Rerun rule
Before trusting any score
3runs

Run each candidate config three times; take the worst repeated result as the real number. A single strong run is indistinguishable from variance.

Worst repeated result wins
Observed variance
Swing at temperature 0.7
13%

Three identical runs of one task: 10% swing on one scored dimension, 13% on another. At temperature 0.0 the same task was deterministic — variance is a setting, not fate.

Same task, same prompt

For an anonymous endpoint this discipline does double duty — repeated runs are also how you detect behavior shifts on a model that will never announce them. And because reruns cost real tokens, score against cost-per-successful-task rather than raw accuracy — the metric that survives a model being cheap, chatty, and wrong.

08Longevity RiskThe listing can vanish mid-sprint.

Model disappearance is routine platform behavior, not an edge case. OpenRouter's own tutorial content states that more than 70 models have been pulled or deprecated by providers "in the last few years" — an unscoped count with no stated population or fixed window, so resist turning it into a rate — and recommends fallback-chain engineering as the mitigation. That's for the platform at large, named models included. A stealth listing adds the EULA's explicit removal clause on top: "Stealth Models may be removed from our Stealth Program at any time upon request of the Stealth Provider or at OpenRouter's sole discretion, with or without notice to you."

Historically, stealth previews resolve in one of two ways: they graduate to a named, priced model, or they are withdrawn. The binary itself is the risk signal — "still listed today" carries no forward guarantee, and the record of how often stealth listings turn out to be who you'd guess is a separate question from whether yours will exist on Friday. Plan for the endpoint's disappearance the way you plan for a spot instance's: checkpoint your prompts, keep a fallback chain warm, and never let an anonymous listing become a single point of failure.

The forward projection follows from the incentives. Stealth listings are cheap distribution for labs that want frontier-scale feedback without frontier-scale accountability, and the EULA's train-by-default terms are what make free access economically rational — so expect more anonymous endpoints, not fewer. The teams that benefit will be the ones that treat them as what they are: a free trial with data-terms pricing, useful for scoped, non-sensitive workloads, and never a substitute for a counterparty. If you're formalizing that boundary for your own organization — which workloads may touch an anonymous endpoint, which require a named counterparty — our AI transformation engagements build exactly that kind of model-governance policy alongside the eval harness.

09ConclusionEvaluate the behavior, price the silence.

The frame to keep

A benchmark tells you what a model does. Only an identity tells you who answers for it.

The empirical half of anonymous-model evaluation is genuinely tractable: latency, tool-call reliability, refusal surface, and rerun stability are all measurable from your own logs, and a held-out private suite — built with the twin pattern, run three times, scored by the worst repeated result — is the honest way to read capability claims no one will put their name to.

The other half is not tractable, and pretending otherwise is the real failure mode. Training-data provenance, training-on-prompts defaults, retention windows, withdrawal notice, and liability are all things a name discloses and a stealth listing structurally cannot. No harness recovers what only a counterparty can sign.

So run both ledgers. Test what is testable, price what is silent, and route accordingly: anonymous endpoints for scoped experiments where the data terms are acceptable and disappearance is survivable; named vendors for everything a client, regulator, or court might one day ask about.

Put a due-diligence floor under your model choices

Know what the model does — and what stays unknowable.

We help teams build eval harnesses, model-governance policies, and vendor-selection frameworks that hold up when the model's maker won't say who they are — delivered in days, not quarters.

Free consultationExpert guidanceTailored solutions
What we work on

Model evaluation & governance engagements

  • Private eval suites your vendors can't have trained on
  • Tool-call reliability harnesses for agentic workloads
  • Model-governance policies: which workloads touch which endpoints
  • Fallback-chain engineering for deprecation resilience
  • Named-vendor due diligence: retention, indemnity, liability
FAQ · Anonymous model evaluation

The questions we get every week.

A stealth listing is a model published on an aggregator — OpenRouter's Stealth Program is the formalized version — without disclosing which lab built it. The provider is bound by the aggregator's stealth agreement rather than by its own published terms, and the listing typically offers free or cheap access. For labs, the appeal is frontier-scale real-world feedback without attaching their name to a pre-release model; the Stealth Program EULA's default — providers “may train, evaluate, and improve” their models on submitted content — is what makes the free access economically rational. Historically these listings either graduate to a named, priced model or are withdrawn.
Related dispatches

Continue exploring model evaluation.