"Independently evaluated" appears on model pages, in procurement answers and in policy submissions, and it can mean a dozen different things. On September 18, 2026 Anthropic announced that Faculty, Accenture's AI unit, will place evaluators inside the company with access comparable to an employee's, and that Anthropic will fund the work directly. In the same post Anthropic said there are as yet no standards for what such evaluators should see or how they should report.
That is the right moment to lay the arrangements side by side. This census reads the primary documents for 12 of them, across Anthropic, OpenAI and Google DeepMind, the nonprofits METR and Apollo Research, and the UK and US government institutes, and answers five questions for each: who evaluates, who pays, what access they get, whether results are published, and whether the evaluator can delay a release. Every cell is what the document states; where it is silent the cell says unknown. This blog runs on Anthropic's models, and Anthropic's rows are held to the same rule as the rest.
- 01No document in the census gives an outside evaluator the power to delay a release.Where a document addresses the question, the mechanism either reviews a report, feeds an internal decision body, or is periodic and decoupled from launches; two rows leave it unaddressed and are marked unknown rather than no. That is an observed absence across these 12 documents, not a claim about arrangements we did not read.
- 02Where the payer is stated, it splits three ways: the lab, the government, or nobody.Anthropic funds Accenture directly and OpenAI offers compensation to all its assessors. UK AISI and US CAISI are government-funded. METR states it has taken no payment for company-identifying assessments and declines lab donations.
- 03Publication rights range from unrestricted to approval-gated to unstated.Anthropic's Risk Report reviewers may publish anything but confidential information. OpenAI's contract excerpt requires written approval before publication. The Accenture post states no publication mechanism at all.
- 04Access is described in words, rarely in items.Only two documents list access as items: OpenAI's account of the UK AISI engagement and METR's Frontier Risk Report. 'Comparable to an employee's' and 'shared access to data and ideas' are the norm; weights are never explicitly granted in any document here.
01 — The hookThe September 18 announcement
Anthropic's post frames the partnership as a step toward a commitment its CEO made in an essay to embed evaluators within the company. The scope is evaluating and red-teaming models, alignment assessments and safeguard testing. Three terms are stated plainly and matter for a buyer. Anthropic funds Accenture's work directly. Each company expects to invest at least $1 billion over five years in building capacity in this area, which is an investment horizon and not a contract term; no term is stated. And the partnership is non-exclusive, with other evaluators to be announced and Accenture free to work with other developers.
Two things the post does not say are as important. It does not say whether Faculty's findings will be published or on what approval. And it does not give the evaluators any power over releases; it says the opposite, that the safety of the models remains Anthropic's responsibility. The post also says Anthropic is in dialogue with METR and other nonprofits to pilot elements of embedded evaluation using their own funding, and that long-term funding should come from pooled or government sources. Our post on what the CEO's essay promised covers the origin of the commitment.
There are, as yet, no standards for what information embedded evaluators should have access to, or how they should report what they find. There is also no settled system for funding independent evaluation.Anthropic, embedded-evaluation partnership announcement, September 18, 2026
02 — The censusTable 1: who evaluates, who pays, what access
Twelve arrangements. An arrangement is one documented relationship between an evaluator and a lab, or one policy mechanism that defines such a relationship. "Unknown" means the primary documents consulted do not say. Where the payer is a government, the row states that and notes whether any lab payment appears.
| Arrangement | Who evaluates | Who pays | Access granted |
|---|---|---|---|
| Anthropic × Accenture embedded evaluation | Faculty, Accenture's specialist AI business | Anthropic, directly, per Anthropic's post | Described as comparable to an employee's: watching training, following build and deploy decisions, speaking to staff. No item-level list of weights, checkpoints or logs. |
| Anthropic Risk Report external review (RSP v3.4 §3.6) | One or more third parties, unnamed in the policy; Board consulted, Long-Term Benefit Trust approves | Not stated in the policy | The private Risk Report, unredacted or minimally redacted, within one week of Board submission; follow-up questions answered. |
| Anthropic annual procedural compliance review | An unnamed third party commissioned by Anthropic | Anthropic (it commissions the review) | Not stated |
| METR review of Anthropic's Opus 4.6 Sabotage Risk Report | METR, a nonprofit | Not Anthropic: METR's conflict-of-interest policy says it has received no payment for company-identifying risk assessments to date | An unredacted version of the Sabotage Risk Report and other materials, per METR. |
| METR Frontier Risk Report (Feb–Mar 2026) | METR, covering Anthropic, Google, Meta and OpenAI | Not the participants; METR cites broad independent funders | Each participant's most capable internal models at the time, including raw chains of thought, plus non-public capability and monitoring information. |
| OpenAI third-party assessors | Named examples in OpenAI's post: METR, SecureBio, Apollo Research, Irregular | OpenAI offers compensation to all assessors; some decline; no payment contingent on results | Early checkpoints, select evaluation results, models with fewer mitigations, direct chain-of-thought access for some. Weights not stated. |
| UK AI Security Institute × OpenAI biosecurity red-teaming | UK AISI, a government research organisation | UK government; no OpenAI payment stated on either side | Five stated forms including helpful-only variants, safeguard prototypes, monitor chain of thought and selective mitigation disabling. |
| US CAISI agreements with OpenAI and Anthropic | US Center for AI Standards and Innovation, at NIST; the US AI Safety Institute when the 2024 agreements were signed | US government; no lab payment stated | Major new models before and after public release, per the 2024 NIST announcement. Weights, logs, variants: unknown. |
| UK AISI × Google DeepMind research MOU | UK AISI jointly with Google DeepMind | UK government funds AISI; cost-sharing unknown | Shared access to data and ideas and joint publications, per AISI. Model access terms: unknown. |
| Google DeepMind Frontier Safety Framework v3.1 | No third-party evaluator is named; external parties involved where required or appropriate | Unknown | Unknown; the framework states no access terms for external parties |
| Apollo Research × OpenAI | Apollo Research, a public benefit corporation as of 2026 | Unknown individually; covered by OpenAI's blanket compensation statement | Pre-deployment testing of o1 in 2024 and anti-scheming work in 2025, per Apollo; specific access forms unknown |
| UK AISI × Anthropic cyber capability testing | UK AISI | UK government; no Anthropic payment stated | A pre-release Claude Mythos 5 running without cyber safeguards and with internet access, per Anthropic's incident account. Precise terms unknown. |
03 — The censusTable 2: publication and release power
The same 12 rows on the two questions a buyer or policy reader most needs answered: does the public ever see the evaluator's findings, and can the evaluator stop or delay a model shipping.
| Arrangement | Results published? | Can delay a release? | Source |
|---|---|---|---|
| Anthropic × Accenture | Unknown; evaluators 'can also report incidents', no publication right or cadence stated | No such power stated; Anthropic says safety remains its responsibility | Anthropic, Sep 18, 2026 |
| Anthropic Risk Report review | Yes: a written report made public, commentary within 30 days, no restriction beyond confidentiality | No; the mechanism reviews reports, not releases | RSP v3.4, Jul 8, 2026 |
| Anthropic compliance review | Unknown; the policy commits to commissioning, not publishing | No; scope is procedural compliance, not outcomes | RSP v3.4, Jul 8, 2026 |
| METR × Anthropic Opus 4.6 | Yes: two review documents with METR's disagreements | No; the model had been deployed for weeks at review time | METR, Mar 12, 2026 |
| METR Frontier Risk Report | Yes, with participants approving what non-public information could be disclosed | No; entity-based and decoupled from releases by design | METR, May 19, 2026 |
| OpenAI third-party assessors | Conditionally: summaries in system cards; assessor publications need OpenAI's written approval | Not stated; deployment decisions sit with OpenAI's Safety Advisory Group | OpenAI, Nov 19, 2025; Preparedness Framework v2, Apr 15, 2025 |
| UK AISI × OpenAI biosecurity | Partially, by OpenAI; no standalone AISI report cited | No; described as ongoing rather than tied to a launch | OpenAI, Sep 12, 2025 |
| US CAISI agreements | Partially: CAISI publishes evaluation news items; agreement text not public | Unknown; not addressed | NIST, Aug 29, 2024; CAISI site |
| UK AISI × Google DeepMind MOU | Joint publications intended; MOU text not public | No such power stated; a research MOU | UK AISI, Dec 11, 2025 |
| Google DeepMind FSF v3.1 | Not committed; information shared with governments if a critical capability level is reached | Internal governance bodies only | FSF v3.1, effective Apr 17, 2026 |
| Apollo × OpenAI | Yes in named cases, after OpenAI's confidentiality and accuracy review | No such power stated | Apollo site; OpenAI, Nov 19, 2025 |
| UK AISI × Anthropic cyber | AISI reported an incident Aug 4, 2026; Anthropic published its own account Aug 31 | Unknown; not stated | Anthropic, Aug 31, 2026 |
Anthropic's Responsible Scaling Policy version 3.4 is the only document here that excludes a class of evaluator by financial dependence: reviewers must not be teams whose revenue, reputation and success depend entirely on Anthropic and similar companies. METR's conflict-of-interest policy, last updated August 28, 2026, is the only evaluator document that states its payment position outright.
04 — The typesFour kinds of arrangement
The 12 rows sort into four types, and each type can tell a buyer something different. None is better in the abstract; the question is what evidence each can produce.
Commercial embedded evaluator
Deepest access on paper and the lab pays. Can tell you how decisions were made inside; cannot yet tell you anything in public, because no publication mechanism is stated and no access or reporting standard exists.
Contracted independent lab
Pre-deployment or report-level testing under contract or on the evaluator's own funds. Can produce a published assessment; what appears in public depends on the contract's approval clause, which ranges from none to written approval.
Government institute
Publicly funded, with access described by the lab rather than by a published agreement. Can tell you a state body tested the model; cannot show you the agreement, and neither institute's page claims power over releases.
Published external review
A review of the lab's own risk documentation, published with the reviewer's disagreements. Can tell you whether the lab's reasoning holds up; cannot test the model itself, and runs after or alongside deployment.
05 — The questionsThree questions for any vendor
Put these to any vendor whose page says independently evaluated, and match the answers to a row above.
- Who paid the evaluator, and was any payment contingent on results?OpenAI answers the second half in writing; METR answers the first; most documents answer neither
- Payer
- What exactly did the evaluator get: weights, checkpoints, helpful-only variants, logs?Only the UK AISI × OpenAI account and METR's report list items; ask for the list, not the adjective
- Access
- Where can I read the evaluator's own report, and who approved it before publication?Unrestricted, approval-gated, or nowhere: all three appear in the census
- Publication
How a lab measures oversight once a model is deployed is a separate question, covered in our post on agent oversight metrics; what gets disclosed when something goes wrong is in our incident-disclosure post. If your organisation is writing an AI procurement policy, our AI transformation service uses this census as the starting checklist.
06 — The gapWhat no arrangement here does
None of the 12 documents gives an outside party the power to stop a release. Anthropic's external review covers Risk Reports, and the policy only says the company will try to complete it before Board review. OpenAI's Preparedness Framework gives the deployment decision to its Safety Advisory Group, with external input as one part of the evidence. Google DeepMind's framework names internal governance bodies for pre-deployment review and no external evaluator at all. The government institutes describe testing, not gating; UK AISI's own about page describes testing systems before public release and collaborating with companies to improve safety. And none publishes a full agreement text; every government arrangement here is announced but not published. A buyer who wants a veto in the chain will not find one in these documents.
07 — How we built itMethodology
A census of stated terms, not an assessment of any evaluator's independence or rigour, and not a ranking.
- Population
- Twelve arrangements involving Anthropic, OpenAI or Google DeepMind and an outside evaluator, where a primary document from the lab, the evaluator or the government body describes the terms. Meta appears only as a METR participant, and xAI only as a company whose public models were available to that exercise; no Meta or xAI safety framework was consulted. Academic evaluators are absent because no primary document establishing one's access, funding or publication rights was located.
- Inclusion rule
- A row needs a first-party document: a lab policy or announcement, an evaluator's own report or policy, or a government body's own page. Trade-press figures were not used for any cell; the $1 billion figures are Anthropic's own.
- Classification rule
- "Unknown" means the documents consulted do not address the question. "No such power stated" means the document describes the arrangement without granting it. Neither is a finding that the power or the fact does not exist elsewhere.
- Dates
- Each row cites its document's own date. The Anthropic policy is version 3.4, posted July 8, 2026. The OpenAI Preparedness Framework is version 2 of April 15, 2025, which OpenAI's May 28, 2026 governance framework states it does not supersede. The Google framework is version 3.1, effective April 17, 2026.
- What was excluded
- An automated extraction that returned a newer Preparedness Framework version number, which did not survive verification. Gemini model cards, which were not read. Any characterisation of an arrangement as independent or not independent beyond the document's own words.
- As-of date
- This page belongs to the September 20, 2026 batch; its sources were collected on September 22, 2026. Undated pages, including the UK AISI and CAISI sites and the METR and Apollo about pages, are cited as published at the time of writing.
- Known limitations
- Twelve rows from three labs is a census of the documented, not of the field. Access described in words cannot be verified against what was actually shared. This blog runs on Anthropic's models; Anthropic's rows are classified by the same rule as the rest.
- Refresh
- Extended when a lab or evaluator publishes an arrangement with stated terms, when Anthropic names the further evaluators its September 18 post promised, or when any government agreement text is published.
08 — Next stepIndependent describes the evaluator, not the arrangement
Ask for the row, not the adjective, before you sign
When a vendor says independently evaluated, ask which of the four types it was, who paid, what the evaluator saw and where the report is. Write the answers into the procurement record beside the vendor's claim. If the answer to any of the three questions is the same word the marketing page used, you do not yet have an answer.