The Kimi K3 benchmarks that made the launch coverage tell one story: a 57 on Artificial Analysis’s Intelligence Index, the leading open-weights model on the board, and 4th overall as of July 27, 2026. The number that didn’t make the charts tells another. On the same organisation’s knowledge and hallucination benchmark, K3’s hallucination rate rose from 39% on Kimi K2.6 to 51% — while its accuracy rose from 33% to 46%.
Read those two movements together and you get an uncomfortable characterisation of the model: K3 is simultaneously more often right and more often confidently wrong. Its predecessor abstained more. K3 answers more, gets more of them correct, and — of the questions it still doesn’t know — guesses rather than declines at a noticeably higher rate. For a chat assistant, that trade may be fine. For an agent writing to a CRM, drafting client-facing copy, or summarising a contract, it is a materially different risk profile than the headline Index score implies.
This guide is not a general survey of hallucination benchmarks — we’ve covered how hallucination benchmarks are measured across models separately. It is about one specific, checkable case: what K3’s own numbers say, what Moonshot chose not to measure, how far the same benchmark moves when someone else runs it, and the eval-before-adopt gate that turns all of that into a decision rather than a vibe.
- 01Accuracy and hallucination both went up.Artificial Analysis reports K3's AA-Omniscience accuracy rising from 33% to 46% versus K2.6, and its hallucination rate rising from 39% to 51%. Both figures are from Artificial Analysis's own launch article, published July 17, 2026.
- 02The composite Index hid the regression.K3's AA-Omniscience Index still improved from +6 to +18, because the Index rewards correct answers more than it penalises confident wrong ones. A single headline number can improve while the failure mode you actually care about gets worse.
- 03Moonshot publishes no factuality metric at all.The Kimi K3 GitHub README carries 40-plus named benchmarks across reasoning, coding, agentic and vision categories. None of them measures hallucination or factual reliability. The number came from a third party, not the vendor.
- 04The same benchmark returns different scores.On Terminal-Bench 2.1, Moonshot's own harness reports 88.3% while Vals AI's independently hosted run scores K3 at 80.9% — a 7.4-point spread on one benchmark, one model, one version.
- 05Run your K2.6 harness against K3 before switching.If you already qualified K2.6 on a private workload, the aggregate delta is not your delta. Re-run the harness you already own, add an abstention-rate metric, and gate the swap on that — not on a leaderboard rank.
01 — The SplitMore often right, more often confidently wrong.
Artificial Analysis published its K3 assessment on July 17, 2026, the day the hosted API went live and ten days before the open weights shipped on July 27. Its own article states that K3’s AA-Omniscience accuracy rate rose from 33% to 46% against K2.6, that the hallucination rate rose from 39% to 51%, and that the composite AA-Omniscience Index moved from +6 to +18. Those three numbers come from the evaluating organisation’s own write-up, not from a press relay, and independent tech press reported the same movement the day before publication.
Up 12 points from K2.6
Of everything that wasn't a clean correct answer, this is the share that was a confident wrong guess rather than an honest abstention. K2.6 sat at 39%. Higher is worse.
Up 13 points from K2.6
Share of the 6,000-question set answered correctly, regardless of whether the model chose to answer. K2.6 sat at 33%. Higher is better — and this is the number that carried the coverage.
Up from +6 on K2.6
The bounded composite that rewards correct answers, penalises hallucinated wrong ones, and applies no penalty for declining to answer. It improved — which is precisely the problem.
The two movements are close in magnitude — accuracy up 13 points, hallucination rate up 12 — but they are not equivalent in consequence. An accuracy gain is realised by the user who asks a question the model happens to know. A hallucination-rate increase is realised by the user who asks a question the model doesn’t know, and who now receives an assertion instead of a hedge. Those are different populations of query, and in production work they are rarely the same people.
"Its hallucination rate climbed from 39 percent to 51 percent, meaning K3 fabricates more answers even as it gets more questions right."— Matthias Bastian, the-decoder.com, July 16, 2026
None of this makes K3 a bad model. It is the leading open-weights model on Artificial Analysis’s open-source board as of July 27, 2026 — a 57 against 51 for the next-placed GLM-5.2, with MiniMax-M3 and DeepSeek V4 Pro tied at 44 behind that. It sits 4th on the overall Intelligence Index, behind Claude Opus 5 at 61, Claude Fable 5 at 60, and GPT-5.6 Sol at 59. (Artificial Analysis’s own July 17 article ranked it 3rd; Opus 5 launched in the interim and moved the rank without moving K3’s score. Date-stamp every one of these figures — the index is live and it moves.)
What it makes K3 is a model whose adoption decision cannot be read off a rank. That is the entire argument of this piece, and the reason we build a standing eval harness to qualify new models before any of them touch client work.
02 — Benchmark DesignThree metrics, and only one of them travelled.
AA-Omniscience is built on 6,000 questions spanning 42 topics across six domains — Business, Health, Law, Software Engineering, Humanities & Social Sciences, and Science, Engineering & Mathematics — derived, per the methodology page, from authoritative academic and industry sources. The question set and an accompanying paper are published openly, which is what makes the numbers checkable rather than merely quotable.
The design decision that matters here is that it reports three separate metrics rather than collapsing to one. Most coverage carried one of the three.
Accuracy
How much the model actually knows, measured regardless of whether it chose to attempt the question. This is the metric that behaves like a normal benchmark and the one that dominated K3's launch coverage.
Hallucination rate
Of everything that wasn't a clean correct answer, what share was a confident wrong guess rather than an abstention. It is a measure of behaviour under uncertainty, not of knowledge.
Omniscience Index
The composite. Rewards correct answers, penalises hallucinated wrong answers, applies no penalty for refusing to answer. Zero means as many correct as incorrect.
For context on how hard the metric is, the benchmark’s original write-up in November 2025 found the leading model of the day topping the Index at just +4.8, one of only three models to score above zero at all. Most frontier models, in other words, produced more confident wrong answers than correct ones on net when refusal was available to them. K3’s +18 is a real improvement against that backdrop. It is also, as the next section shows, an improvement that partly conceals the thing it’s measuring.
03 — Index MathThe composite improved. The failure mode got worse.
K3’s Omniscience Index tripled, from +6 to +18. That is the number a procurement deck would carry. It rose because the Index weights correct-answer volume against hallucinated wrong answers, and a 13-point accuracy gain outweighs a 12-point hallucination-rate increase in that arithmetic. The composite is doing exactly what it was designed to do. The problem is what a reader infers from it.
AA-Omniscience · Kimi K2.6 vs Kimi K3
Source: Artificial Analysis launch assessment, July 17, 2026 · scores as of July 27, 2026There is a further derivation worth doing yourself, because it changes the tone of the finding. Artificial Analysis defines the hallucination rate as incorrect answers divided by the sum of incorrect, partial and not-attempted answers. If those four outcome categories — correct, incorrect, partial, not-attempted — are exhaustive, then the denominator is simply everything that wasn’t correct, and you can recover the share of the whole question set that drew an outright wrong answer.
That reframing is more useful than the headline. K3 did not dramatically start fabricating more; it dramatically stopped declining. For anyone building agentic workflows, the operational translation is direct: the guardrail you need is not a smarter model, it is an explicit abstention path — a way for the model to return “insufficient information” that your orchestration layer treats as a valid, routable outcome rather than a failure.
One mechanical confound is worth naming and then not overclaiming. Artificial Analysis also reports that K3 consumed 21% fewer output tokens than K2.6 across the same Intelligence Index run — 132 million against 166 million. Terser models have less room for hedging language and caveats, and it is plausible that some of the hallucination-rate movement is a byproduct of that concision rather than a change in epistemic behaviour. That is our hypothesis about a correlation, not a finding either organisation has published, and it is exactly the kind of thing your own harness can test in an afternoon and a leaderboard cannot.
04 — What Moonshot MeasuresForty benchmarks, and zero of them about truth.
The Kimi K3 GitHub repository is not shy about evaluation. Its README organises results into four categories. Reasoning & Knowledge carries GPQA Diamond, CritPt, AA-LCR and HLE-Full. Coding carries DeepSWE, ProgramBench, Terminal-Bench 2.1, FrontierSWE, SWE-Marathon, SciCode and more. Agentic runs to 22 entries — BrowseComp, DeepSearchQA, GDPval-AA v2, Toolathlon-Verified, OSWorld-Verified, τ³-Banking, Legal Research Bench, and others. Vision adds another ten. A full technical report sits in the same repository.
Across all forty-plus named benchmarks, none measures hallucination or factual reliability. The metric that shifted most sharply between generations is the one the vendor’s own materials are silent on — not suppressed, simply never in scope. This is not unique to Moonshot; it is the industry norm. It is also precisely why a buyer reading only vendor tables gets a systematically incomplete picture.
Read those two disclosed behaviours next to the hallucination movement and a coherent profile emerges. A model that answers rather than abstains, and that makes unexpected decisions on the user’s behalf under ambiguity, is the same disposition expressed at two different layers of the stack. Neither is disqualifying. Both argue for tighter scaffolding than you would give a more reticent model — narrower tool permissions, mandatory confirmation steps on writes, and an eval suite that specifically probes ambiguous inputs rather than well-formed ones.
05 — Harness VarianceOne benchmark, one model, three different scores.
If the hallucination split is the reason to run your own factuality evals, Terminal-Bench 2.1 is the reason to distrust single-number coding claims generally. Three organisations have reported a K3 score on that benchmark. They do not agree, and the disagreement is larger than the gaps that decide most model-selection arguments.
| Evaluator | K3 score | Harness / method | Sourcing posture |
|---|---|---|---|
| Vendor self-report | |||
| Moonshot AI | 88.3% | Moonshot’s own KimiCode harness, as published in the K3 README and technical report | Vendor primary; independently echoed in press aggregation of launch coverage |
| Third-party runs | |||
| Artificial Analysis | 85% — single-sourced, treat as indicative | Independently run harness, figure relayed by an eval-methodology outlet rather than read off a machine-readable page | Reported by emergent.sh, July 23, 2026; not re-derived from a static Artificial Analysis page |
| Vals AI | 80.9% | Vals AI’s own hosted variant of the benchmark, run independently of the vendor | Independent evaluator, primary — published on its own Terminal-Bench 2.1 leaderboard |
| Derived spread | |||
| Vendor minus lowest independent | 7.4 pts | 88.3% − 80.9%, the two hardest-sourced numbers in the table | Our arithmetic on the two published figures |
| Narrowest adjacent gap | 3.3 pts | 88.3% − 85%, if the indicative middle figure holds | Our arithmetic; inherits the middle row’s single-source caveat |
The two figures we would actually stake a decision on are the vendor self-report at 88.3% and Vals AI’s independent replication at 80.9%. That is a 7.4-point gap on a single benchmark, at a single version, for a single model. Vals AI’s broader composite tells a consistent story — it scores K3 at 74.7 on its own index, with 71.6 on CorpFin v2 and 48.9 on MedCode, a set of numbers noticeably more sober than the launch-week framing.
Nothing here implies bad faith. Vendor harnesses are tuned for the vendor’s model: prompt formats, retry policy, tool-call plumbing, timeout budgets and scaffold assumptions all move scores, and a lab that has spent months optimising its own agent loop will naturally extract more from its own weights than a standardised third-party rig does. Harness variance of this size is also not unique to Moonshot — comparable vendor-versus-standardised gaps have been reported for closed frontier models by the same eval-methodology outlet, though we have not re-verified those figures against a first-party source and would not print them as facts.
The practical consequence is simple. If your model-selection argument turns on a gap of under eight points on a coding benchmark, you do not have an argument — you have a harness artefact. That threshold is a useful thing to write into your own cost-aware model routing playbook as a decision rule rather than rediscovering it per release.
06 — Across The FieldThis isn’t a Moonshot problem. It’s an evaluation problem.
The tempting reading of the K3 numbers is that Moonshot traded reliability for scores. The comparator data doesn’t support that reading. The same profile — leading accuracy sitting alongside a hallucination rate that is anything but low — shows up at the top of the same leaderboard today, in a closed frontier model from a lab with very different incentives. To be explicit about what that is and isn’t: it is a single-date snapshot, not a generational comparison. We have no prior-generation AA-Omniscience baseline for Fable 5 and are not claiming its hallucination rate moved in either direction. What the snapshot does show is that an accuracy-led profile is where the top of this board currently sits, which makes the question an evaluation-discipline one rather than a Moonshot-specific flaw.
| Model | Accuracy | Hallucination rate | Index | What it means |
|---|---|---|---|---|
| Moonshot, generation over generation | ||||
| Kimi K2.6 | 33% | 39% | +6 | Knew less, declined more. Barely positive on the composite. |
| Kimi K3 | 46% | 51% | +18 | Knows more, declines much less. Composite triples. |
| Delta, K2.6 → K3 | +13 pts (+39% rel.) | +12 pts (+31% rel.) | +12 | Both metrics move up together — the trade the headline hides. |
| Closed-frontier comparator | ||||
| Claude Fable 5 | 61% | Reported above K3’s by outlets citing Artificial Analysis data; we could not re-derive the figure from a machine-readable page, so we don’t print it | +40 | Tops the accuracy metric. Its Index lead is accuracy-driven, not hallucination-driven. |
The Fable 5 row is the load-bearing one, and it is worth being precise about what we can and can’t assert. Artificial Analysis’s own pages confirm that Fable 5 posts the highest AA-Omniscience accuracy on the board at 61% and an Index of 40, and its own commentary attributes that leading position to accuracy rather than to low hallucination. Its per-model hallucination percentage renders inside a client-side chart that we could not read as text, and while several independent outlets relay a specific figure sourced to Artificial Analysis, we are not printing a number we could not verify at source. Directionally: the leader on this benchmark is not the model that guesses least.
That is the finding, and it is a snapshot claim rather than a trend claim. Read the board on one date and the best-scoring knowledge model and the best-scoring open-weights model both present accuracy-led profiles — accuracy is where the training effort went, and abstention discipline is not what the composite rewards. The only generational movement we can actually evidence here is Moonshot’s own, K2.6 to K3; if you were expecting that kind of release-over-release progress to monotonically reduce confident wrong answers, that one comparison already says otherwise. We deliberately aren’t relitigating the broader capability comparison here; our pre-release Fable 5 comparison covers that ground, and our release write-up covers what actually shipped.
Projecting forward: if the composite indices that drive procurement keep rewarding attempt-rate, the rational training response is to attempt more. We would expect the next twelve months to produce models that score better on knowledge composites while getting harder to deploy unsupervised — unless buyers start asking for the sub-metric rather than the index. The correction, if it comes, will come from procurement discipline rather than from model releases.
07 — Eval Before AdoptFour gates before a new model touches client work.
None of the preceding sections requires a research budget to act on. The gates below are what we run when a model like K3 lands, and they are deliberately cheap — the point is that they are run at all, and that a model can fail one of them without failing the others.
Re-run the harness you already own
If you qualified K2.6 on a private workload, that harness is your best comparator — better than any public delta. Swap the model, change nothing else, and read the difference on your own tasks. The aggregate leaderboard delta is not your delta.
Measure abstention, not just accuracy
Add items your corpus genuinely cannot answer and score three outcomes, not two: correct, wrong, declined. A model that never declines will look fine on accuracy and behave badly in production. This is the metric K3's numbers argue you're missing.
Test ambiguity, not well-formed input
Moonshot itself flags that K3 may make unexpected decisions on the user's behalf under ambiguous intent. Benchmarks use clean inputs; clients do not. Score whether the model asks or assumes.
Price the run, not the token
K3's hosted API lists at $3.00 per million input tokens, $0.30 on a cache hit, and $15.00 per million output. Blended cost on the Intelligence Index run came to $0.94 per task. Measure your own tasks — cache hit rate and reasoning length dominate the answer.
Gate four deserves a note on where K3 actually sits commercially, because open weights invite an assumption of cheapness that the hosted numbers don’t support. At $0.94 blended per task on Artificial Analysis’s Index run, K3 came in below GPT-5.6 Sol at $1.04 and roughly half of Claude Opus 4.8 at $1.80 — but roughly three times GLM-5.2 at $0.32 and more than twenty times DeepSeek V4 Pro at $0.04. It is mid-pack, not the budget option. Throughput reinforces the point: Artificial Analysis measured K3’s output at around 33 tokens per second against roughly 72 for Claude Fable 5 in the same side-by-side, so latency-sensitive surfaces will feel the difference before the invoice does.
The cache-hit economics are the lever worth engineering around. A cache hit costs a tenth of a miss on input tokens, which means prompt architecture — stable system prefixes, deterministic tool schemas, batched context — moves your bill more than model choice does at these volumes. That is the same discipline we apply across production content pipelines regardless of which model sits underneath.
08 — Where K3 FitsWhere K3 earns a slot — and where it doesn’t.
Model selection is per-workload, not per-model. K3’s profile — strong agentic and coding numbers, open weights, mid-pack cost, an elevated tendency to answer rather than abstain — maps cleanly onto some jobs and poorly onto others.
Tool-using coding agents
This is K3's strongest territory. It leads its predecessor decisively on agentic knowledge work — its GDPval-AA v2 Elo of 1,668 against K2.6's 1,190 is one of the larger generational jumps on that board. Verify on your own repos, and pass full thinking history back through the harness as Moonshot instructs.
On-prem open weights
K3 is the leading open-weights model on Artificial Analysis's open-source board as of July 27, 2026, at 57 against 51 for the next placed. If weights-on-your-hardware is a hard requirement, it's the strongest option available — subject to the infrastructure cost of running a model this size.
Client-facing claims & citations
A 51% hallucination rate on AA-Omniscience is the wrong profile for anything that ships to a client without a human read. Use retrieval grounding, require citations, and treat unsourced assertions as a validation failure rather than trusting the model's confidence.
Interactive assistants
Measured output speed of roughly 33 tokens per second is around half of Claude Fable 5's in the same comparison. For streaming chat surfaces where perceived responsiveness is the product, that gap will be felt regardless of benchmark rank.
The through-line is that none of these calls came from the Intelligence Index. They came from sub-metrics, vendor disclosures and independent replications — the layer underneath the rank. Any team adopting frontier models at pace needs that layer as standing infrastructure rather than a per-launch scramble, which is the starting point of most of our AI transformation engagements.
09 — ConclusionThe number that isn’t on the chart is the one to check.
A model can improve on every published number and still be riskier to deploy.
Kimi K3 is a genuinely strong release. It leads the open-weights field, sits 4th on Artificial Analysis’s overall Intelligence Index as of July 27, 2026, and posts agentic numbers that justify serious evaluation. It also answers more questions it doesn’t know the answer to than its predecessor did, on a benchmark its own vendor doesn’t run, at a rate that never appeared in a launch chart.
Both of those statements come from the same organisation’s data, published the same week. The reason only one of them circulated is not conspiracy — it is that composite indices are legible and sub-metrics are not. The Omniscience Index tripled, so the Omniscience Index is what travelled. The hallucination rate sat one level down, moving the wrong way, in the same article.
The defensible position isn’t to avoid K3. It is to stop treating any single published score as a procurement input. Run the harness you already own against the new weights. Score abstention alongside accuracy. Assume an eight-point coding-benchmark gap is harness noise until you reproduce it yourself. Date-stamp every figure, because the boards are live and volatile. That discipline costs a day per model and is the only thing standing between a leaderboard rank and a production incident.