AI DevelopmentFramework15 min readPublished July 27, 2026

Accuracy 33% → 46% · hallucination rate 39% → 51% · scores as of July 27, 2026

Kimi K3 Benchmarks: The Hallucination Number Nobody Charted

Kimi K3 is the strongest open-weights model on Artificial Analysis’s Intelligence Index. It is also, on the same organisation’s knowledge benchmark, meaningfully more likely to answer confidently when it doesn’t know. Both things are true, and only one of them made the launch coverage.

DA
Digital Applied Team
Senior strategists · Published Jul 27, 2026
PublishedJul 27, 2026
Read time15 min
Sources10 primary & independent
AA-Omniscience hallucination
51%
K3 · was 39% on K2.6
+12 pts
AA-Omniscience accuracy
46%
K3 · was 33% on K2.6
+13 pts
AA Intelligence Index
57
4th overall, Jul 27, 2026
#1 open weights
Terminal-Bench 2.1 spread
7.4pts
vendor vs independent harness

The Kimi K3 benchmarks that made the launch coverage tell one story: a 57 on Artificial Analysis’s Intelligence Index, the leading open-weights model on the board, and 4th overall as of July 27, 2026. The number that didn’t make the charts tells another. On the same organisation’s knowledge and hallucination benchmark, K3’s hallucination rate rose from 39% on Kimi K2.6 to 51% — while its accuracy rose from 33% to 46%.

Read those two movements together and you get an uncomfortable characterisation of the model: K3 is simultaneously more often right and more often confidently wrong. Its predecessor abstained more. K3 answers more, gets more of them correct, and — of the questions it still doesn’t know — guesses rather than declines at a noticeably higher rate. For a chat assistant, that trade may be fine. For an agent writing to a CRM, drafting client-facing copy, or summarising a contract, it is a materially different risk profile than the headline Index score implies.

This guide is not a general survey of hallucination benchmarks — we’ve covered how hallucination benchmarks are measured across models separately. It is about one specific, checkable case: what K3’s own numbers say, what Moonshot chose not to measure, how far the same benchmark moves when someone else runs it, and the eval-before-adopt gate that turns all of that into a decision rather than a vibe.

Key takeaways
  1. 01
    Accuracy and hallucination both went up.Artificial Analysis reports K3's AA-Omniscience accuracy rising from 33% to 46% versus K2.6, and its hallucination rate rising from 39% to 51%. Both figures are from Artificial Analysis's own launch article, published July 17, 2026.
  2. 02
    The composite Index hid the regression.K3's AA-Omniscience Index still improved from +6 to +18, because the Index rewards correct answers more than it penalises confident wrong ones. A single headline number can improve while the failure mode you actually care about gets worse.
  3. 03
    Moonshot publishes no factuality metric at all.The Kimi K3 GitHub README carries 40-plus named benchmarks across reasoning, coding, agentic and vision categories. None of them measures hallucination or factual reliability. The number came from a third party, not the vendor.
  4. 04
    The same benchmark returns different scores.On Terminal-Bench 2.1, Moonshot's own harness reports 88.3% while Vals AI's independently hosted run scores K3 at 80.9% — a 7.4-point spread on one benchmark, one model, one version.
  5. 05
    Run your K2.6 harness against K3 before switching.If you already qualified K2.6 on a private workload, the aggregate delta is not your delta. Re-run the harness you already own, add an abstention-rate metric, and gate the swap on that — not on a leaderboard rank.

01The SplitMore often right, more often confidently wrong.

Artificial Analysis published its K3 assessment on July 17, 2026, the day the hosted API went live and ten days before the open weights shipped on July 27. Its own article states that K3’s AA-Omniscience accuracy rate rose from 33% to 46% against K2.6, that the hallucination rate rose from 39% to 51%, and that the composite AA-Omniscience Index moved from +6 to +18. Those three numbers come from the evaluating organisation’s own write-up, not from a press relay, and independent tech press reported the same movement the day before publication.

Hallucination rate
Up 12 points from K2.6
51%

Of everything that wasn't a clean correct answer, this is the share that was a confident wrong guess rather than an honest abstention. K2.6 sat at 39%. Higher is worse.

AA-Omniscience · Jul 17, 2026
Accuracy
Up 13 points from K2.6
46%

Share of the 6,000-question set answered correctly, regardless of whether the model chose to answer. K2.6 sat at 33%. Higher is better — and this is the number that carried the coverage.

AA-Omniscience · Jul 17, 2026
Omniscience Index
Up from +6 on K2.6
+18

The bounded composite that rewards correct answers, penalises hallucinated wrong ones, and applies no penalty for declining to answer. It improved — which is precisely the problem.

Range −100 to +100

The two movements are close in magnitude — accuracy up 13 points, hallucination rate up 12 — but they are not equivalent in consequence. An accuracy gain is realised by the user who asks a question the model happens to know. A hallucination-rate increase is realised by the user who asks a question the model doesn’t know, and who now receives an assertion instead of a hedge. Those are different populations of query, and in production work they are rarely the same people.

"Its hallucination rate climbed from 39 percent to 51 percent, meaning K3 fabricates more answers even as it gets more questions right."— Matthias Bastian, the-decoder.com, July 16, 2026

None of this makes K3 a bad model. It is the leading open-weights model on Artificial Analysis’s open-source board as of July 27, 2026 — a 57 against 51 for the next-placed GLM-5.2, with MiniMax-M3 and DeepSeek V4 Pro tied at 44 behind that. It sits 4th on the overall Intelligence Index, behind Claude Opus 5 at 61, Claude Fable 5 at 60, and GPT-5.6 Sol at 59. (Artificial Analysis’s own July 17 article ranked it 3rd; Opus 5 launched in the interim and moved the rank without moving K3’s score. Date-stamp every one of these figures — the index is live and it moves.)

What it makes K3 is a model whose adoption decision cannot be read off a rank. That is the entire argument of this piece, and the reason we build a standing eval harness to qualify new models before any of them touch client work.

02Benchmark DesignThree metrics, and only one of them travelled.

AA-Omniscience is built on 6,000 questions spanning 42 topics across six domains — Business, Health, Law, Software Engineering, Humanities & Social Sciences, and Science, Engineering & Mathematics — derived, per the methodology page, from authoritative academic and industry sources. The question set and an accompanying paper are published openly, which is what makes the numbers checkable rather than merely quotable.

The design decision that matters here is that it reports three separate metrics rather than collapsing to one. Most coverage carried one of the three.

Metric 01
Accuracy
% of 6,000 questions correct

How much the model actually knows, measured regardless of whether it chose to attempt the question. This is the metric that behaves like a normal benchmark and the one that dominated K3's launch coverage.

K3: 46% · K2.6: 33%
Metric 02
Hallucination rate
incorrect ÷ (incorrect + partial + not-attempted)

Of everything that wasn't a clean correct answer, what share was a confident wrong guess rather than an abstention. It is a measure of behaviour under uncertainty, not of knowledge.

K3: 51% · K2.6: 39%
Metric 03
Omniscience Index
bounded −100 to +100

The composite. Rewards correct answers, penalises hallucinated wrong answers, applies no penalty for refusing to answer. Zero means as many correct as incorrect.

K3: +18 · K2.6: +6
Why the denominator matters
The hallucination rate is not “percentage of answers that were wrong.” It is wrong answers as a share of everything that wasn’t cleanly correct — so abstaining pushes the number down and guessing pushes it up. A model that says “I don’t know” more often scores better on this metric without knowing anything more. That is a deliberate design choice by Artificial Analysis, and it is the reason this number captures something an accuracy score structurally cannot.

For context on how hard the metric is, the benchmark’s original write-up in November 2025 found the leading model of the day topping the Index at just +4.8, one of only three models to score above zero at all. Most frontier models, in other words, produced more confident wrong answers than correct ones on net when refusal was available to them. K3’s +18 is a real improvement against that backdrop. It is also, as the next section shows, an improvement that partly conceals the thing it’s measuring.

03Index MathThe composite improved. The failure mode got worse.

K3’s Omniscience Index tripled, from +6 to +18. That is the number a procurement deck would carry. It rose because the Index weights correct-answer volume against hallucinated wrong answers, and a 13-point accuracy gain outweighs a 12-point hallucination-rate increase in that arithmetic. The composite is doing exactly what it was designed to do. The problem is what a reader infers from it.

AA-Omniscience · Kimi K2.6 vs Kimi K3

Source: Artificial Analysis launch assessment, July 17, 2026 · scores as of July 27, 2026
K2.6 accuracyShare of 6,000 questions answered correctly
33%
K3 accuracy+13 points, or roughly +39% relative to K2.6
46%
Improved
K2.6 hallucination rateConfident wrong guesses as a share of non-correct outcomes
39%
K3 hallucination rate+12 points, or roughly +31% relative to K2.6
51%
Regressed
Higher is betterHigher is worse

There is a further derivation worth doing yourself, because it changes the tone of the finding. Artificial Analysis defines the hallucination rate as incorrect answers divided by the sum of incorrect, partial and not-attempted answers. If those four outcome categories — correct, incorrect, partial, not-attempted — are exhaustive, then the denominator is simply everything that wasn’t correct, and you can recover the share of the whole question set that drew an outright wrong answer.

Our derivation, not a published figure
Applying that formula to the published numbers: on K2.6, non-correct outcomes were 67% of the set, of which 39% were confident wrong answers — roughly 26% of all questions. On K3, non-correct outcomes were 54%, of which 51% were confident wrong answers — roughly 27.5%. The absolute volume of outright wrong answers barely moved, about 1.4 points of the full set. What collapsed was the abstention cushion: K3 has far less “didn’t attempt” territory left, so nearly the same quantity of wrong answers now occupies a much larger share of its non-correct behaviour. This is our arithmetic applied to Artificial Analysis’s stated formula and published percentages, not a figure either organisation reports.

That reframing is more useful than the headline. K3 did not dramatically start fabricating more; it dramatically stopped declining. For anyone building agentic workflows, the operational translation is direct: the guardrail you need is not a smarter model, it is an explicit abstention path — a way for the model to return “insufficient information” that your orchestration layer treats as a valid, routable outcome rather than a failure.

One mechanical confound is worth naming and then not overclaiming. Artificial Analysis also reports that K3 consumed 21% fewer output tokens than K2.6 across the same Intelligence Index run — 132 million against 166 million. Terser models have less room for hedging language and caveats, and it is plausible that some of the hallucination-rate movement is a byproduct of that concision rather than a change in epistemic behaviour. That is our hypothesis about a correlation, not a finding either organisation has published, and it is exactly the kind of thing your own harness can test in an afternoon and a leaderboard cannot.

04What Moonshot MeasuresForty benchmarks, and zero of them about truth.

The Kimi K3 GitHub repository is not shy about evaluation. Its README organises results into four categories. Reasoning & Knowledge carries GPQA Diamond, CritPt, AA-LCR and HLE-Full. Coding carries DeepSWE, ProgramBench, Terminal-Bench 2.1, FrontierSWE, SWE-Marathon, SciCode and more. Agentic runs to 22 entries — BrowseComp, DeepSearchQA, GDPval-AA v2, Toolathlon-Verified, OSWorld-Verified, τ³-Banking, Legal Research Bench, and others. Vision adds another ten. A full technical report sits in the same repository.

Across all forty-plus named benchmarks, none measures hallucination or factual reliability. The metric that shifted most sharply between generations is the one the vendor’s own materials are silent on — not suppressed, simply never in scope. This is not unique to Moonshot; it is the industry norm. It is also precisely why a buyer reading only vendor tables gets a systematically incomplete picture.

What the vendor does disclose
Moonshot is unusually candid about limitations elsewhere. Its launch post states plainly: “Despite being a highly competitive model overall, K3 nonetheless exhibits a noticeable gap in user experience compared with Claude Fable 5 and GPT 5.6 Sol.” It also documents two behaviours that matter for agent builders: a sensitivity to how thinking history is handled — “If the agent harness fails to pass back all the historical thinking content as required...generation quality may become highly unstable” — and a tendency toward excessive proactiveness, where “When it encounters minor issues or ambiguous user intent during task execution, it may make unexpected decisions on the user’s behalf.” Moonshot’s own advice is to impose more explicit behavioural constraints if your application needs the agent to stay in narrow bounds.

Read those two disclosed behaviours next to the hallucination movement and a coherent profile emerges. A model that answers rather than abstains, and that makes unexpected decisions on the user’s behalf under ambiguity, is the same disposition expressed at two different layers of the stack. Neither is disqualifying. Both argue for tighter scaffolding than you would give a more reticent model — narrower tool permissions, mandatory confirmation steps on writes, and an eval suite that specifically probes ambiguous inputs rather than well-formed ones.

05Harness VarianceOne benchmark, one model, three different scores.

If the hallucination split is the reason to run your own factuality evals, Terminal-Bench 2.1 is the reason to distrust single-number coding claims generally. Three organisations have reported a K3 score on that benchmark. They do not agree, and the disagreement is larger than the gaps that decide most model-selection arguments.

Kimi K3 Terminal-Bench 2.1 scores as reported by Moonshot AI, Artificial Analysis via a secondary outlet, and Vals AI, with the harness used and the sourcing posture of each figure.
EvaluatorK3 scoreHarness / methodSourcing posture
Vendor self-report
Moonshot AI88.3%Moonshot’s own KimiCode harness, as published in the K3 README and technical reportVendor primary; independently echoed in press aggregation of launch coverage
Third-party runs
Artificial Analysis85% — single-sourced, treat as indicativeIndependently run harness, figure relayed by an eval-methodology outlet rather than read off a machine-readable pageReported by emergent.sh, July 23, 2026; not re-derived from a static Artificial Analysis page
Vals AI80.9%Vals AI’s own hosted variant of the benchmark, run independently of the vendorIndependent evaluator, primary — published on its own Terminal-Bench 2.1 leaderboard
Derived spread
Vendor minus lowest independent7.4 pts88.3% − 80.9%, the two hardest-sourced numbers in the tableOur arithmetic on the two published figures
Narrowest adjacent gap3.3 pts88.3% − 85%, if the indicative middle figure holdsOur arithmetic; inherits the middle row’s single-source caveat

The two figures we would actually stake a decision on are the vendor self-report at 88.3% and Vals AI’s independent replication at 80.9%. That is a 7.4-point gap on a single benchmark, at a single version, for a single model. Vals AI’s broader composite tells a consistent story — it scores K3 at 74.7 on its own index, with 71.6 on CorpFin v2 and 48.9 on MedCode, a set of numbers noticeably more sober than the launch-week framing.

Nothing here implies bad faith. Vendor harnesses are tuned for the vendor’s model: prompt formats, retry policy, tool-call plumbing, timeout budgets and scaffold assumptions all move scores, and a lab that has spent months optimising its own agent loop will naturally extract more from its own weights than a standardised third-party rig does. Harness variance of this size is also not unique to Moonshot — comparable vendor-versus-standardised gaps have been reported for closed frontier models by the same eval-methodology outlet, though we have not re-verified those figures against a first-party source and would not print them as facts.

The practical consequence is simple. If your model-selection argument turns on a gap of under eight points on a coding benchmark, you do not have an argument — you have a harness artefact. That threshold is a useful thing to write into your own cost-aware model routing playbook as a decision rule rather than rediscovering it per release.

06Across The FieldThis isn’t a Moonshot problem. It’s an evaluation problem.

The tempting reading of the K3 numbers is that Moonshot traded reliability for scores. The comparator data doesn’t support that reading. The same profile — leading accuracy sitting alongside a hallucination rate that is anything but low — shows up at the top of the same leaderboard today, in a closed frontier model from a lab with very different incentives. To be explicit about what that is and isn’t: it is a single-date snapshot, not a generational comparison. We have no prior-generation AA-Omniscience baseline for Fable 5 and are not claiming its hallucination rate moved in either direction. What the snapshot does show is that an accuracy-led profile is where the top of this board currently sits, which makes the question an evaluation-discipline one rather than a Moonshot-specific flaw.

AA-Omniscience accuracy, hallucination rate and Index for Kimi K2.6, Kimi K3 and Claude Fable 5, with the generational delta and a one-line reading for each row.
ModelAccuracyHallucination rateIndexWhat it means
Moonshot, generation over generation
Kimi K2.633%39%+6Knew less, declined more. Barely positive on the composite.
Kimi K346%51%+18Knows more, declines much less. Composite triples.
Delta, K2.6 → K3+13 pts (+39% rel.)+12 pts (+31% rel.)+12Both metrics move up together — the trade the headline hides.
Closed-frontier comparator
Claude Fable 561%Reported above K3’s by outlets citing Artificial Analysis data; we could not re-derive the figure from a machine-readable page, so we don’t print it+40Tops the accuracy metric. Its Index lead is accuracy-driven, not hallucination-driven.

The Fable 5 row is the load-bearing one, and it is worth being precise about what we can and can’t assert. Artificial Analysis’s own pages confirm that Fable 5 posts the highest AA-Omniscience accuracy on the board at 61% and an Index of 40, and its own commentary attributes that leading position to accuracy rather than to low hallucination. Its per-model hallucination percentage renders inside a client-side chart that we could not read as text, and while several independent outlets relay a specific figure sourced to Artificial Analysis, we are not printing a number we could not verify at source. Directionally: the leader on this benchmark is not the model that guesses least.

That is the finding, and it is a snapshot claim rather than a trend claim. Read the board on one date and the best-scoring knowledge model and the best-scoring open-weights model both present accuracy-led profiles — accuracy is where the training effort went, and abstention discipline is not what the composite rewards. The only generational movement we can actually evidence here is Moonshot’s own, K2.6 to K3; if you were expecting that kind of release-over-release progress to monotonically reduce confident wrong answers, that one comparison already says otherwise. We deliberately aren’t relitigating the broader capability comparison here; our pre-release Fable 5 comparison covers that ground, and our release write-up covers what actually shipped.

Projecting forward: if the composite indices that drive procurement keep rewarding attempt-rate, the rational training response is to attempt more. We would expect the next twelve months to produce models that score better on knowledge composites while getting harder to deploy unsupervised — unless buyers start asking for the sub-metric rather than the index. The correction, if it comes, will come from procurement discipline rather than from model releases.

07Eval Before AdoptFour gates before a new model touches client work.

None of the preceding sections requires a research budget to act on. The gates below are what we run when a model like K3 lands, and they are deliberately cheap — the point is that they are run at all, and that a model can fail one of them without failing the others.

Gate 01
Re-run the harness you already own
Same prompts, same scaffold, new weights

If you qualified K2.6 on a private workload, that harness is your best comparator — better than any public delta. Swap the model, change nothing else, and read the difference on your own tasks. The aggregate leaderboard delta is not your delta.

Half a day
Gate 02
Measure abstention, not just accuracy
Seed unanswerable questions into the set

Add items your corpus genuinely cannot answer and score three outcomes, not two: correct, wrong, declined. A model that never declines will look fine on accuracy and behave badly in production. This is the metric K3's numbers argue you're missing.

The gate most teams skip
Gate 03
Test ambiguity, not well-formed input
Underspecified prompts, conflicting instructions

Moonshot itself flags that K3 may make unexpected decisions on the user's behalf under ambiguous intent. Benchmarks use clean inputs; clients do not. Score whether the model asks or assumes.

Agentic workloads only
Gate 04
Price the run, not the token
Blended cost per completed task

K3's hosted API lists at $3.00 per million input tokens, $0.30 on a cache hit, and $15.00 per million output. Blended cost on the Intelligence Index run came to $0.94 per task. Measure your own tasks — cache hit rate and reasoning length dominate the answer.

Per task, not per 1M

Gate four deserves a note on where K3 actually sits commercially, because open weights invite an assumption of cheapness that the hosted numbers don’t support. At $0.94 blended per task on Artificial Analysis’s Index run, K3 came in below GPT-5.6 Sol at $1.04 and roughly half of Claude Opus 4.8 at $1.80 — but roughly three times GLM-5.2 at $0.32 and more than twenty times DeepSeek V4 Pro at $0.04. It is mid-pack, not the budget option. Throughput reinforces the point: Artificial Analysis measured K3’s output at around 33 tokens per second against roughly 72 for Claude Fable 5 in the same side-by-side, so latency-sensitive surfaces will feel the difference before the invoice does.

The cache-hit economics are the lever worth engineering around. A cache hit costs a tenth of a miss on input tokens, which means prompt architecture — stable system prefixes, deterministic tool schemas, batched context — moves your bill more than model choice does at these volumes. That is the same discipline we apply across production content pipelines regardless of which model sits underneath.

08Where K3 FitsWhere K3 earns a slot — and where it doesn’t.

Model selection is per-workload, not per-model. K3’s profile — strong agentic and coding numbers, open weights, mid-pack cost, an elevated tendency to answer rather than abstain — maps cleanly onto some jobs and poorly onto others.

Agentic engineering
Tool-using coding agents

This is K3's strongest territory. It leads its predecessor decisively on agentic knowledge work — its GDPval-AA v2 Elo of 1,668 against K2.6's 1,190 is one of the larger generational jumps on that board. Verify on your own repos, and pass full thinking history back through the harness as Moonshot instructs.

Strong candidate
Self-hosted / sovereign
On-prem open weights

K3 is the leading open-weights model on Artificial Analysis's open-source board as of July 27, 2026, at 57 against 51 for the next placed. If weights-on-your-hardware is a hard requirement, it's the strongest option available — subject to the infrastructure cost of running a model this size.

Best open option
Unsupervised factual output
Client-facing claims & citations

A 51% hallucination rate on AA-Omniscience is the wrong profile for anything that ships to a client without a human read. Use retrieval grounding, require citations, and treat unsourced assertions as a validation failure rather than trusting the model's confidence.

Not without guardrails
Latency-sensitive UX
Interactive assistants

Measured output speed of roughly 33 tokens per second is around half of Claude Fable 5's in the same comparison. For streaming chat surfaces where perceived responsiveness is the product, that gap will be felt regardless of benchmark rank.

Check throughput first

The through-line is that none of these calls came from the Intelligence Index. They came from sub-metrics, vendor disclosures and independent replications — the layer underneath the rank. Any team adopting frontier models at pace needs that layer as standing infrastructure rather than a per-launch scramble, which is the starting point of most of our AI transformation engagements.

09ConclusionThe number that isn’t on the chart is the one to check.

Eval discipline, July 2026

A model can improve on every published number and still be riskier to deploy.

Kimi K3 is a genuinely strong release. It leads the open-weights field, sits 4th on Artificial Analysis’s overall Intelligence Index as of July 27, 2026, and posts agentic numbers that justify serious evaluation. It also answers more questions it doesn’t know the answer to than its predecessor did, on a benchmark its own vendor doesn’t run, at a rate that never appeared in a launch chart.

Both of those statements come from the same organisation’s data, published the same week. The reason only one of them circulated is not conspiracy — it is that composite indices are legible and sub-metrics are not. The Omniscience Index tripled, so the Omniscience Index is what travelled. The hallucination rate sat one level down, moving the wrong way, in the same article.

The defensible position isn’t to avoid K3. It is to stop treating any single published score as a procurement input. Run the harness you already own against the new weights. Score abstention alongside accuracy. Assume an eight-point coding-benchmark gap is harness noise until you reproduce it yourself. Date-stamp every figure, because the boards are live and volatile. That discipline costs a day per model and is the only thing standing between a leaderboard rank and a production incident.

Qualify models before they touch client work

Leaderboard ranks don’t survive contact with your workload. Your own eval harness does.

We build standing eval harnesses that qualify new frontier models against your own workloads — accuracy, abstention, cost per completed task, and harness-variance checks — so model swaps are a decision, not a gamble.

Free consultationExpert guidanceTailored solutions
What we work on

Model evaluation engagements

  • Private eval harnesses on your own tasks and corpus
  • Abstention and hallucination scoring, not just accuracy
  • Harness-variance checks against vendor benchmark claims
  • Blended cost-per-task modelling and cache-hit engineering
  • Multi-model routing rules with explicit swap gates
FAQ · Kimi K3 evaluation

The questions buyers ask every launch week.

Artificial Analysis reports Kimi K3 at a 51% hallucination rate on its AA-Omniscience benchmark, up from 39% on the previous Kimi K2.6, in an assessment published July 17, 2026. Over the same generation, accuracy rose from 33% to 46%. It's important to read the metric correctly: AA-Omniscience defines the hallucination rate as incorrect answers divided by the sum of incorrect, partial and not-attempted answers — so it measures how often a model guesses rather than abstains when it doesn't know, not the raw share of all answers that were wrong. A model that declines more often scores better on this metric without knowing anything more. Scores are as of July 27, 2026 on a leaderboard the organisation states updates continuously.
Related dispatches

Continue exploring model evaluation.