A contamination claim should be no stronger than the evidence behind it. Publicly available answers establish a possible route of exposure. Recovering a hidden solution demonstrates something more specific. Showing how much that access changed a score requires another test again.
This reference gives writers and AI buyers a way to separate those claims. It is an original evidence-classification method, reviewed September 12, 2026, with primary-source examples. It does not declare a new contamination incident or certify any benchmark as clean.
- 01Identify the mechanism.Training exposure, runtime answer access and defective grading are different problems.
- 02Separate existence from effect.Showing a leakage route does not quantify its contribution to a score.
- 03Verify repairs at the same boundary.A patched environment needs both leakage checks and valid task behavior.
01 — Practical decisionName the failure before judging the score
Training contamination concerns information encountered before evaluation. Runtime leakage concerns answers or evaluation material available during the run. A grading defect concerns whether the test accepts the right behavior. All can distort a benchmark, but they require different evidence and different repairs.
OpenAI's February 23 SWE-bench Verified analysis reports recoverable problem details or gold patches in tested models and separately examines flawed tasks. Its hard-problem audit is a selected subset. A percentage from that subset should not be presented as the prevalence across every benchmark problem.
For a claim you plan to publish, write the proposed mechanism in one sentence. If you cannot distinguish exposure from retrieval or grading, the finding is not ready for a precise headline. Our benchmark methodology guide covers broader score interpretation; this reference is specifically about the strength of an allegation and its repair evidence.
02 — Practical decisionMatch the observation to the supported claim
These twelve rows are an editorial evidence ladder, not a numerical scoring system. The first group concerns exposure and retrieval, the second concerns effect, and the third concerns repairs. Evidence can skip or combine rows; do not add the rows into a confidence percentage.
The final column names what remains unproven. Keeping that column is essential: a writer should be able to state both the finding and its limit. A negative result belongs in the record too, including a model that does not reproduce an answer or a repair that fails to preserve a legitimate solution.
| Observation | Claim it can support | What remains unproven |
|---|---|---|
| Answers were publicly available | Exposure was possible | Whether this model encountered them. |
| A dataset overlap was located | Matching material exists in the inspected data | Training inclusion and influence without further evidence. |
| A model reproduces a distinctive answer | The tested prompt elicited answer-specific information | The learning route and effect on the full benchmark. |
| A run reads hidden evaluation material | Runtime answer access occurred in that trace | How often it occurred or whether it caused success. |
| A suspicious solution passes | The grader accepted that attempt | Semantic correctness or the source of the solution. |
| A controlled access restriction changes outcomes | The restricted channel affected the tested setup | Generalization beyond the tested tasks and settings. |
| Scores differ on a replacement task set | The sets produce different observed results | A causal contamination estimate if difficulty also changed. |
| Independent reviewers reproduce the finding | The observed mechanism is reproducible within that scope | Its prevalence in untested models or environments. |
| A repair removes the identified channel | The stated channel was changed | Whether equivalent access paths remain. |
| A regression check no longer reproduces access | That exploit check failed under the repaired setup | Absence of every other exploit. |
| Valid solutions still pass after repair | The repair preserves the checked legitimate behavior | Fairness across the complete task distribution. |
| A versioned independent rerun is published | The repaired result has inspectable external evidence | Permanent immunity to later data or environment changes. |
03 — Practical decisionKeep task-quality audits distinct from leakage evidence
OpenAI's July 8 coding-evaluation audit describes problems such as underspecified prompts, overly strict tests and incomplete coverage in SWE-Bench Pro. Its method combines agent-assisted investigation with human review. Those findings concern whether tasks measure the intended behavior; they do not establish that every passing agent used leaked answers.
This distinction changes the remedy. Removing a hidden solution from an environment does not fix a test that rejects a valid implementation. Repairing an underspecified prompt does not remove an answer embedded elsewhere. An evaluation can need both changes, but the evidence for each should stay separate.
Avoid using a single word such as broken as a substitute for the mechanism. A buyer deciding whether to trust a score needs to know whether success might be inflated, failure might be overstated or both. That can change which tasks remain useful for their own trial.
04 — Practical decisionTreat a proposed repair as a claim to inspect
The preprint SWE-Bench Pro Verified, submitted September 8, 2026, distinguishes leakage of gold solutions or hidden evaluation information from task-quality issues. Its abstract describes safeguards and task refinement. This article uses that bounded description; it does not reproduce the paper's experiments or claim independent confirmation.
A useful repair record identifies the original channel, the changed environment or task, the version tested and the outcomes before and after the change. It should also show that legitimate solutions still work. Otherwise a repair could lower the score simply by making the task impossible or removing an allowed capability.
Ask whether the score belongs to the original benchmark, a patched harness or a revised task set. Those labels matter for comparison. The research proof reference separates reported methods from replicated findings; a paper title containing verified does not itself supply a new independent verification.
05 — Practical decisionDo not infer a causal effect from an unmatched score drop
Suppose a model scores lower on a newly curated set. The tasks may be harder, the repository distribution may differ, or the harness may impose another budget. That observation can motivate investigation, but it cannot isolate the effect of contamination without controlling relevant differences.
A stronger proposed check holds the model, task set, settings and harness constant while changing the suspected information channel. Retain the attempts, including failures, and document what the intervention removed. If the intervention changes useful task information too, the result needs a narrower interpretation.
Even a controlled check has a population boundary. A reproduced exploit on selected tasks is not a measured prevalence across all models. Keep the task selection method beside the result. Our source independence guide also helps distinguish an independent rerun from several articles repeating the same investigation.
Illustrative working record
Illustrative wording: an investigator reproduced access to a hidden answer in the specified environment. The report does not yet establish the frequency of that access across the benchmark or the size of its effect on the published score.
This example shows calibrated language. It is not a report that Digital Applied reproduced an exploit.
06 — Practical decisionWrite a finding that another reader can check
Record the benchmark and version, affected model or agent, mechanism, evidence artifact, test conditions and unresolved alternatives. State who performed the investigation and whether anyone independently repeated it. If a claim depends on a preprint abstract alone, keep the description at that level.
For a repaired result, retain the original and revised identifiers so later readers can tell which number a chart used. Avoid overwriting the history with a clean label. An honest correction should explain the changed evidence and the limits of the new conclusion.
The worksheet below leaves observed status and evidence blank. Fill it with a concrete report before using it to support procurement or publication. An AI evaluation engagement should likewise separate documented findings from planned experiments and leave uncertain mechanisms visible.
Download the blank acceptance worksheet. The reference rows are filled; observed status, evidence and checked date are deliberately empty. Record pass, fail, unknown or not applicable only after inspecting the relevant artifact or claim.
Evidence and scope
- Dates
- Editorial allocation: September 11, 2026. Sources retrieved and article reviewed September 12, 2026. Event dates are stated separately.
- Sources
- OpenAI analyses dated February 23 and July 8, 2026; SWE-Bench Pro Verified v1 abstract and submission history dated September 8. No secondary benchmark figures adopted.
- Method
- Twelve original observation-to-claim rows across exposure, effect and repair. SVG group counts derive from the table. This is a definitional reference, not a census of contaminated models.
- Limits
- No exploit, benchmark or independent replication was run. Unknown means evidence is missing; a failed reproduction does not prove absence. No confidence percentage or contamination prevalence is calculated.
07 — Next stepPublish the supported claim and its limit
Publish the supported claim and its limit
Name the mechanism, retain the evidence and state what it does not establish. A narrower, reproducible finding is more useful than a sweeping contamination verdict that collapses exposure, causation and repair into one label.