An AI hallucination rate is meaningful only when you know what was asked, what counted as an error, and which answers entered the denominator. A short factual question, a sourced research summary, and a library import exercise test different capabilities. Their scores cannot be combined into a universal model reliability ranking.
Correction, October 4, 2026: The original version of this article claimed that Digital Applied ran a 5,000-prompt study. We could not substantiate that study with run-level records in the repository and publication history reviewed. Its original measurement claims are unsupported. We withdraw its model rankings, error rates, confidence intervals, and mitigation-effect estimates. The guidance below replaces those results; it does not report a new experiment.
This guide keeps the original topic: factual recall, citation accuracy, code references, reasoning, and mitigation. The public benchmark descriptions are dated to their original publications and were checked on October 4, 2026. They explain evaluation methods, not a current leaderboard of frontier models. The proposed workflow and fictional calculation are clearly distinguished from published research.
- 01Match the benchmark to the task.Short-answer correctness, document grounding, citation support, and executable code are different outcomes. Name the tested population before applying a score to production.
- 02Report abstention alongside errors.An incorrect-answer rate over all questions can fall when a system answers less often. Show answer coverage and the error rate among attempted answers together.
- 03A citation must support the claim.A real paper or working URL establishes existence. Checking the source passage establishes whether the cited claim is supported.
- 04Treat mitigation benefits as measurements to make.Compare reasoning, retrieval, and verification on the same frozen workload. Preserve responses, evidence, settings, and grader decisions before publishing an effect size.
01 — DenominatorsDefine the rate before comparing it.
Start with the unit of analysis. An answer-level measure asks whether a response contains an error. A claim-level measure divides a long answer into factual statements and grades each one. A citation-level measure asks whether references exist and support their associated claims. These units answer different questions, even when every result is printed with a percent sign.
Then define the labels. For a proposed short-answer evaluation, use correct, incorrect, and not attempted, with a separate unresolved bucket when the reference or grading decision cannot be established. Do not quietly treat an unresolved reference as a model failure or a success. Resolve it before the headline calculation, or disclose how many cases remain outside the scored population.
Fictional example, not benchmark data: imagine 100 answerable questions producing 70 correct answers, 10 incorrect answers, and 20 abstentions. The incorrect-answer rate over all questions is 10/100 = 10%. Coverage is 80/100 = 80%, and errors among attempted answers are 10/80 = 12.5%. Reporting only the first figure hides the unanswered workload. Correctness among attempts is 70/80 = 87.5%; it is not the overall success rate of 70%.
Choose the denominator before examining competing systems. If one returns long explanations and another returns only a short answer, claim counts and opportunities for error differ. Preserve the full output, disclose response-length constraints, and report results within task families. Our benchmark methodology guide gives broader context on interpreting evaluation conditions.
Response outcome
Use for a question with a reference answer. Publish coverage alongside correctness so abstention cannot disappear inside the headline rate.
Claim support
Record the evidence for each factual statement. Define how incomplete evidence differs from a demonstrated contradiction.
Existence and entailment
Check bibliographic identity separately from the relationship between the source passage and the generated assertion.
02 — Published methodsWhat factual recall benchmarks establish.
OpenAI introduced SimpleQA on October 30, 2024. Its original release contains 4,326 short, fact-seeking questions. Responses are classified as correct, incorrect, or not attempted. Questions were designed to have stable, verifiable answers and to challenge the models used during construction.
That is a deliberately selected question set, not a random sample of customer interactions. The launch article explicitly limits the benchmark to short answers and leaves the relationship with factual long-form writing unresolved. Treat a published SimpleQA result as evidence about that setup. It does not establish a production failure rate for research reports, customer support, or code generation.
The practical lesson is to describe your own sample just as carefully. Document where questions came from, the date of the reference material, inclusion rules, exclusions, and how you separated development examples from evaluation examples. If difficult cases are deliberately overrepresented, label the set as a stress test. If you want an estimate for ordinary traffic, explain how the sample reflects that traffic and which segments it misses.
Keep answerable and unanswerable questions distinguishable. Declining an impossible request can be the desired result; declining an answerable customer question creates an unresolved task. A single refusal rate cannot distinguish those cases. Set the expected behavior for each class before running the test, and preserve evidence for why the reference is considered answerable.
03 — Evidence chainsSeparate citation accuracy from document grounding.
Google DeepMind introduced FACTS Grounding on December 17, 2024. Its original dataset has 1,719 document-based examples. The evaluation checks whether the response addresses the request and whether it is grounded in the supplied document. The launch methodology combines multiple model judges.
This tests a different condition from recalling a fact without a supplied document. It also differs from verifying sources found on the open web. A document-grounded answer can faithfully repeat an inaccurate document. Conversely, a true statement from outside the document may violate a task that requires using only the provided source. Grounding and truth require distinct questions.
For long-form factuality, Min and colleagues’ FActScore paper, published at EMNLP 2023, decomposes text into atomic facts and measures the share supported by a reliable knowledge source. Its human evaluation examined generated biographies. FActScore is a separate research method from FACTS Grounding; the similar names do not make their scores interchangeable.
For a proposed citation test, preserve the exact claim, the reference returned by the model, the resolved source, and the supporting passage. Check title, author, date, and identifier where applicable. Then ask whether the passage supports the whole claim, including its population, time period, units, and qualifiers. A real study about one population cannot substantiate a broader claim merely because it discusses the same subject.
Include an unavailable-source state. If the paper is behind access controls or the URL fails, you have an access limitation until another authoritative copy resolves it. That alone does not demonstrate an invented paper. Distinguish a fabricated reference, a real but irrelevant reference, an overextended interpretation, and an inaccessible reference. The remedy depends on which failure actually occurred.
Resolve the reference
Match the returned identifier and bibliographic details to the source. Preserve mismatches rather than silently correcting the citation before scoring it.
Read the relevant passage
Check whether the source establishes the generated claim, including its date, sample, metric, and qualifications. A working link is only the start.
Keep uncertainty visible
Record unavailable full text or conflicting authoritative versions. Route unresolved references to review instead of claiming they are false.
04 — Executable checksTest code references against a pinned environment.
A code reference evaluation needs the package version, runtime, language, import path, and expected behavior. An API may exist in one release and be absent in another. Without a pinned environment, a result can confuse a version mismatch with an invented symbol. Preserve the dependency manifest and lockfile with the prompt and generated answer.
Use several distinct checks. Does the referenced package exist? Does the installed version export the symbol? Does the call satisfy the actual signature? Does it behave as the task requires? Record each outcome separately. A successful type check does not demonstrate semantic correctness, and a failed test does not automatically mean the model hallucinated an API.
The TypeScript handbook illustrates how type checking can identify incompatible values and absent properties under the documented checking rules. The Python unittest documentation describes assertions and test cases for checking expected behavior. These are verification tools; neither provides a model hallucination rate.
Run unfamiliar generated code in an isolated test environment with no production credentials or live side effects. Keep setup errors, dependency failures, timeouts, and model-output defects in separate categories. Otherwise, a broken evaluation environment can make a correct answer look wrong. Include negative cases such as a requested function that does not exist, so the desired behavior includes identifying an invalid premise.
05 — Controlled comparisonsMeasure extended thinking on the same workload.
This correction does not retain the original article’s claimed reductions from extended thinking. To establish an effect for your workload, compare configurations using the same frozen questions, references, scoring rules, and tool permissions. Record the exact reasoning setting the provider supports rather than treating similarly named settings from different providers as equivalent amounts of computation.
Change one factor at a time where possible. If a higher reasoning setting also receives retrieval access or more output space, the result measures the combined system change. That can be a useful deployment comparison, but it cannot isolate reasoning as the cause. State whether you are comparing complete systems or conducting an ablation of a particular feature.
Score answer coverage, errors, successful task completion, latency, and cost together. A configuration that improves correctness among attempted answers may still leave more work for a human. A longer response may introduce additional claims. Record these differences instead of using a single accuracy number as a complete business case.
Use repeated attempts when randomness matters, retaining every attempt and its identity. Repeats of the same question are not independent samples of new user needs. If you calculate uncertainty, use a method that respects the sampling design and report the numerator, denominator, and interval method. Do not append generic confidence bands to a leaderboard without the observations needed to calculate them.
06 — Failure taxonomyClassify errors so each one has a remedy.
The categories below are a proposed diagnostic rubric, not a measured distribution of model errors. Assign them from the observable output and reference evidence. Avoid inferring an internal mental process from a fluent explanation or an exposed reasoning summary. A plausible story about why an answer failed is not a substitute for the failure record.
Allow a response to have more than one defect if the rubric permits it. A citation can have both an incorrect date and an unsupported inference. If you later publish category percentages, state whether categories are mutually exclusive and which unit is being counted. Overlapping labels should not be presented as shares of a partition.
Invented or mismatched entity
The emitted identifier or entity does not match the authoritative reference. Preserve the lookup evidence and distinguish confirmed mismatch from missing access.
A source says less than the answer
The source is real, but the answer generalizes beyond it or changes its meaning. Compare the exact source passage with the generated assertion.
A fact belongs to another version
The answer uses information from the wrong time or software environment. Check against the date or version requested, not whichever source is easiest to find.
Keep grader disagreements visible. Have reviewers work from the same rubric and evidence, record their initial labels, and document the resolution. If automated scoring assists the process, save its prompt, model identifier, and output as well. A headline score is reproducible only when someone can trace it back through the judgments that produced it.
07 — InterventionsEvaluate retrieval, tools, and verification separately.
Retrieval is a candidate intervention when the needed evidence exists in an accessible source collection. Test whether the retriever finds the right material before assessing how well the model uses it. Save retrieved passages and their order. A failure to retrieve relevant evidence and a failure to follow available evidence are different problems with different fixes.
Tool grounding can make an evaluation more inspectable by returning a database record, package definition, or source passage that a reviewer can check. It also changes the system under test. Record tool versions, permissions, results, and failures. Do not attribute a tool-assisted result to the unaided model or assume that a successful tool call guarantees a faithful final answer.
Repeated generation and majority voting need an external correctness check. Agreement can show consistency while every candidate repeats the same error. If disagreement triggers abstention, report the resulting coverage and fallback workload. Review the disputed cases rather than treating consensus as ground truth.
Prompt instructions such as acknowledging uncertainty are worth testing against a baseline, with outcomes measured in the same way. A stricter prompt could reduce unsupported answers by declining more requests. Whether that is useful depends on the cost of a wrong answer, the cost of review, and the availability of a fallback. The original article’s universal improvement percentages have been withdrawn; this guide supplies no replacement effect sizes.
Before release, choose acceptance conditions for the actual workflow and document who handles unresolved outputs. A research assistant may route uncertain citations to an editor; a code assistant may require a successful test; a customer-facing answer may need escalation. Use an evaluation set that includes those transitions, and inspect whether the system completes the handoff rather than merely announcing it.
For broader instruction design, see our agent instruction-file audit. Keep the instructions used in an evaluation with the run records so a later prompt edit cannot silently change the meaning of a comparison.
08 — Deployment choiceTurn a benchmark into an accountable decision.
Choose an evidence standard before choosing a model.
Use public benchmarks to understand the capability being tested and the limits of the measurement. Then evaluate the system you intend to deploy against its real sources, permissions, output requirements, and fallback path. Preserve the evidence needed to revisit the decision when the workload or model changes.
This page no longer supplies Digital Applied original-study results or a ranking of models by hallucination rate. Its useful conclusion is methodological: distinguish incorrect answers from unanswered work, verify that citations support claims, and test interventions under comparable conditions. See our evaluation harness guide for further implementation context, or discuss evaluation design through our AI transformation services.