AI DevelopmentMethodology12 min readPublished August 31, 2026

Same models, same nominal task, inverted answers — and no combined ranking exists

Which AI Model Actually Writes Best? The Boards Disagree

Two published creative-writing leaderboards — EQ-Bench Creative Writing v3 and arena.ai’s Creative Writing category — rank the same models in close to opposite order. This page puts their published rankings side by side, explains the instruments behind them, and refuses to average two scales that cannot be averaged.

DA
Digital Applied Team
Senior strategists · Published Aug 31, 2026
PublishedAug 31, 2026
Boards compared2
Read time12 min
claude-fable-5
#1 / 6th
arena.ai rank / EQ-Bench rank, same nominal task
kimi-k3
2nd / #23
EQ-Bench / arena.ai — the inversion runs both ways
Open-weight gap
38 · 45.3
points behind #1 on each board’s own, separate scale
Human votes
1.2M
arena.ai Creative Writing category, stated Aug 27, 2026

If you have just read a “best AI for writing” article, the honest answer to whether you should believe it is: not on the strength of either leaderboard alone. EQ-Bench Creative Writing v3 and arena.ai’s Creative Writing category — whose page states an August 27, 2026 snapshot — do not merely re-rank the same models. They invert. claude-fable-5 is first on arena.ai and sixth on EQ-Bench; kimi-k3 is second on EQ-Bench and twenty-third on arena.ai.

The disagreement is measurable, and — more usefully — explainable. One board is a human blind-voting arena; the other is an LLM-judged rubric-plus-Elo pipeline. The full side-by-side table is directly below, and the rest of the page explains why two carefully built instruments pointed at the same task return opposite answers.

Key takeaways
  1. 01
    The boards invert, not just disagree.claude-fable-5: #1 on arena.ai, 6th on EQ-Bench, 183 Elo behind claude-opus-5. kimi-k3: 2nd on EQ-Bench, #23 on arena.ai. claude-opus-5: 1st on EQ-Bench, 10th on arena.ai.
  2. 02
    They are different instruments, on different scales.arena.ai aggregates human blind pairwise votes; EQ-Bench uses two Claude judges for a rubric score and a pairwise Elo. The scales are anchored differently and can never be averaged or converted.
  3. 03
    They agree on exactly one thing.Each board puts the highest-ranked open-weight model shown here just behind its #1 — 38 points on arena.ai, 45.3 Elo on EQ-Bench — two separate observations on two incompatible scales.
  4. 04
    They even disagree on a checkable fact.arena.ai labels glm-5.3-max as MIT-licensed. The zai-org/GLM-5.3 model card says license: other with a bespoke licence name. A leaderboard’s metadata columns deserve the same skepticism as its scores.

01The findingSame models, opposite answers.

Two definitions before the table, because both carry weight in it. Creative Writing is the narrow category both boards publish: fiction, storytelling, and poetry written to a prompt. Leaderboard platforms track it separately from general “Writing” tasks like emails, essays, and editing, so a rank here is a creative-writing rank, nothing broader. An Elo score, as leaderboards use the term, is a relative rating inferred from many pairwise comparisons: winning against a strong opponent moves a model up more than winning against a weak one. It is meaningful only within one board’s pool, solver, and anchor points — which is why the two Elo columns below run on visibly different number lines and must be read as two separate rankings, never one.

Same model family, two published ranks. Scores and ranks are from each board’s own published table. The tier suffixes (-high, -max) are arena.ai’s naming; pairings are deliberate family-level matches, not assertions that the identical deployment was measured. The two score columns use different, incomparable Elo scales.
Model familyarena.ai rank · scoreEQ-Bench rank · EloPlaces moved
claude-fable-5same name on both boards#1 · 1505±96th · 1933.2Down 5 places, board to board
claude-opus-5arena.ai lists claude-opus-5-high#10 · 1471±91st · 2116.1Up 9 places — #1 on the other board
kimi-k3arena.ai lists kimi-k3-max#23 · 1458±112nd · 2070.8Up 21 places — the sharpest inversion
GLM-5.3arena.ai lists glm-5.3-max, possibly a hosted-only tier#13 · 1467±183rd · 2062.4Up 10 places
claude-opus-4-7arena.ai lists claude-opus-4-7-high#4 · 1489±78th · 1906.4Down 4 places
muse-sparkarena.ai carries no version suffix; EQ-Bench lists muse-spark-1.1 — a family match, not an exact one#20 · 1464±147th · 1915.2Up 13 places

Read the first three rows again. The model arena.ai’s voters rank first sits 183 Elo behind the EQ-Bench leader on EQ-Bench’s own scale. The model EQ-Bench ranks second is twenty-third with human voters. The model EQ-Bench ranks first does not crack arena.ai’s top nine. This is not noise around a shared consensus — the two boards return structurally different orderings of the same vendor lineups on the same nominal task. Neither board is the writing leaderboard, and any article that cites exactly one of them as settled truth has made a choice it probably has not told you about.

Two boards that agree on almost no individual ranking each place their top open-weight row just behind their top closed one — 38 points on one scale, 45.3 Elo on the other. Two separate readings that happen to point the same way.Digital Applied analysis, August 31, 2026

02MethodWhat was collected, and how the boards were matched.

This is a comparison of two published datasets, not a benchmark we ran. That makes the matching rules — which names were paired, which scales were kept apart, which board stamps its own snapshot — the part a citing reader needs most.

Methodology

Both boards were pulled first-party; no figure comes from an aggregator or a secondary write-up.

What was collected
Each board’s published table, read from the board itself: EQ-Bench’s rank, Elo, rubric, slop and length columns; arena.ai’s rank, score with confidence interval, vote counts, and vendor and licence labels. Plus each board’s own methodology documentation. Rows reproduced on this page are a selection.
Collected
2026-09-01, in a single pass per board. This page is dated August 31, 2026; the figures are as collected, one day later.
Board-stated dates
arena.ai states Aug 27, 2026 on the category page, with 1,214,472 votes across 393 models in this category. EQ-Bench publishes no as-of stamp anywhere on its Creative Writing v3 page; its only dated signal is a March 1, 2026 changelog note about a judge switch.
Sources
The EQ-Bench Creative Writing v3 leaderboard and its methodology page; the arena.ai Creative Writing leaderboard (lmarena.ai now redirects there); LMSYS’s first-party Style Control post; and one peer-reviewed study of LLM-judge bias, “Judging the Judges” (TMLR 2026).
Units and scales
Both boards publish an Elo-style score, on different scales: arena.ai’s Creative Writing scores cluster around 1500, EQ-Bench’s top out above 2100 with fixed anchors (DeepSeek-R1 at 1500, ministral-3b at 200). The two are never averaged, converted, or combined into one ranking anywhere on this page.
Name matching
arena.ai lists tier suffixes (claude-opus-5-high, kimi-k3-max, glm-5.3-max); EQ-Bench lists bare model names. Pairings in the inversion table are deliberate family-level matches, flagged per-row. muse-spark carries no version suffix on arena.ai and was matched to EQ-Bench’s muse-spark-1.1 as a family, not an exact deployment.
Exclusions and limitations
Models present on only one board are excluded from the side-by-side table. The numeric effect of arena.ai’s Style Control toggle on this category’s ranking could not be captured; only the toggle’s existence and documented design are reported. Our own post archive is not used as a comparison dataset.

03The dataThe two boards, side by side.

First, EQ-Bench Creative Writing v3’s top ten. Note the rubric column alongside the Elo — the same board publishes two scores per model, and section 05 is about the rows where they tell different stories. Length is the average output length the board reports per model, and it does real analytical work here.

EQ-Bench Creative Writing v3, top 10 by Elo, from the board’s own published table. The page carries no as-of stamp (see methodology). Slop is the board’s word-frequency measure of overused language-model phrasing; lower is better.
RankModelEloRubricSlopLength (chars)
1claude-opus-52116.185.350.96,003
2kimi-k32070.884.251.35,488
3GLM-5.32062.485.201.25,913
4gpt-5.6-sol1964.183.901.68,548
5ox-alpha1959.784.451.46,112
6claude-fable-51933.284.051.45,887
7muse-spark-1.11915.282.701.77,551
8claude-opus-4-71906.482.851.55,692
9gpt-5.6-terra1850.082.801.710,271
10gpt-5.51844.085.051.812,945

Now the arena.ai rows for the same families, from a page that states its own snapshot date — Aug 27, 2026 — and its own sample: 1,214,472 human votes across 393 models in this category. The confidence intervals matter: glm-5.3-max, qwen3.8-max, muse-spark, and kimi-k3-max sit within a few points of one another, with overlapping intervals, on 1,286 to 3,208 votes each.

arena.ai Creative Writing, selected rows, from the board’s own table as stated Aug 27, 2026. Ranks are the board’s own; this table shows selected rows, not the full ranking. The licence label on the glm-5.3-max row is the board’s own text and is disputed — see the correction below the table.
RankModelScoreVotesVendor · licence label
1claude-fable-51505±95,160Anthropic
2claude-opus-4-6-high1500±712,682Anthropic
4claude-opus-4-7-high1489±710,548Anthropic
5gemini-3-pro1483±86,244Google · Proprietary
10claude-opus-5-high1471±96,810Anthropic
13glm-5.3-max1467±181,286Z.ai · “MIT” — disputed by the model card; see below
16qwen3.8-max1464±142,244Alibaba · Proprietary
20muse-spark1464±141,950Meta
23kimi-k3-max1458±113,208Moonshot · Kimi K3 licence

One row needs a correction the board has not made. arena.ai’s licence label reads “Z.ai · MIT” for glm-5.3-max. The Hugging Face model card for zai-org/GLM-5.3 states license: other with license_name: glm-5.3 — a bespoke licence, not MIT. We documented that licence in detail in our August 28 post on the GLM-5.3 weights. To be precise about what was checked: the correction rests on the zai-org/GLM-5.3 model card, and glm-5.3-max may be a hosted-only tier — but a filter column that lets a reader sort by “MIT” and returns this model is disagreeing with the model’s own card. The two boards, in other words, diverge on a checkable fact as well as on scores.

04The instrumentsTwo instruments, not two opinions.

The inversion stops being mysterious once you look at how each number is produced. arena.ai is human pairwise voting: two anonymous outputs shown blind, side by side, one vote for the better one, aggregated through a Bradley-Terry model — the standard statistical method for turning many pairwise win-loss records into one rating per model. EQ-Bench is LLM-judged, twice over: outputs are first graded against a scoring rubric by one Claude judge, then run through pairwise matchups against neighbouring models — judged by a second, different Claude judge, with win margins feeding a modified Glicko/Trueskill-style solver — a rating system of the same broad family as Elo, extended to weigh how decisively each matchup was won — run until ranks stabilise. Per the site’s March 1, 2026 note, the Elo judge is Claude Sonnet 4.6 while the rubric judge remains Claude Sonnet 4: two judges, two columns, one leaderboard.

What each board actually measures, from each board’s own methodology documentation.
DimensionEQ-Bench Creative Writing v3arena.ai Creative Writing
RaterTwo LLM judges: Claude Sonnet 4 for the rubric, Claude Sonnet 4.6 for pairwise Elo (per the site’s March 1, 2026 note)Anonymous humans voting blind on side-by-side outputs
Unit of scoringA rubric score per output, plus pairwise matchups with graded win margins feeding a Glicko/Trueskill-style Elo solverOne vote per pairwise comparison, aggregated by a Bradley-Terry model into a score with a confidence interval
Length handlingPairwise judging truncates outputs to a standardised length; rubric scoring keeps the full output — two policies on one board, by designA Style Control toggle models length as an independent variable in the Bradley-Terry regression
Style handlingA user-set Vocab Control slider penalising “overly complex vocab usage”, and a word-frequency Slop score built on a published word listStyle Control also covers markdown headers, bold elements, and lists, alongside length
Self-preference controlNone — the site’s own words: “We do not control for the judge possibly preferring its own outputs.”Not applicable in the same sense — raters are human; the documented risk is voters’ style and length preferences instead
Scale anchoringFixed anchors: DeepSeek-R1 at 1500, ministral-3b at 200 — top scores run above 2100Scores in this category cluster around 1500
Snapshot stampNone published; the page’s only dated signal is a March 1, 2026 changelog noteStates Aug 27, 2026 on the page
Sample behind the numbersDeliberately adversarial prompts; the board states scoring a model costs around $10 in API fees1,214,472 votes across 393 models in this category

Two design details deserve their own sentences. First, EQ-Bench is adversarial by design — its prompts are chosen to be “challenging for weaker models and therefore highly discriminative”, and the site says plainly that “the purpose of the evaluation is not to help models write their best. Instead, we are deliberately exposing weaknesses.” arena.ai measures preference under whatever its voters bring. An instrument built to expose weaknesses and an instrument built to aggregate preference are answering different questions even when the category name matches. Second, arena.ai’s Style Control exists precisely because human votes carry style bias: LMSYS’s own launch post explains, “We explicitly model style as an independent variable in our Bradley-Terry regression. For example, we added length as a feature — just like each model, the length difference has its own Arena Score!” The toggle is live on the Creative Writing page today; what its activation does numerically to the 2026 gaps in this post could not be captured, so this page claims no narrowing or widening from it.

The bias is measured, not hypothetical

A 2026 peer-reviewed study in Transactions on Machine Learning Research, “Judging the Judges”, found style bias is the dominant LLM-judge bias — 0.10 to 0.76 across judge models, favouring markdown over plain prose, against at most 0.04 for position bias. It also found length preference is not universal: Claude-family judges in the study preferred concise answers (−0.12). Both of EQ-Bench’s judges are Claude models — which cuts against any simple “LLM judges always reward length” reading of the board, and makes the divergences in the next section more interesting, not less.

05The nested disagreementRubric and Elo disagree inside one board, too.

You do not need two leaderboards to see the disagreement — EQ-Bench publishes it in two columns of its own table. gpt-5.5 scores 85.05 on the rubric against claude-opus-5’s 85.35 — a 0.3 gap on graded quality — while sitting 272 Elo lower in the pairwise ranking, at 2.2× the output length. Further down the same table, horizon-beta posts a rubric of 83.30 — level with claude-opus-4-8 — at 14,202 characters, with an Elo of 1624.5. The board publishes both columns and they do not tell the same story.

Rubric vs Elo divergence inside EQ-Bench Creative Writing v3, from the board’s own table. The rubric judge sees full outputs; the Elo judge sees outputs truncated to a standardised length.
ModelRubricEloLength (chars)
claude-opus-585.352116.16,003
gpt-5.585.051844.012,945
horizon-beta83.301624.514,202
The board’s own answer, verbatim

EQ-Bench does not hide this — its methodology page addresses it directly: “Pairwise matchups allow the judge to be more discriminative than scoring a single item in isolation... The scores may also differ because we use different criteria in the judging prompts between rubric & pairwise. The judge will also be subject to different biases depending on the evaluation method.” And on which column to believe: “Why do these scores disagree? ... Which one is right? Well, both and neither.”

One more pattern from the same table, because it breaks a default assumption buyers carry: newer is not better at writing. On EQ-Bench, claude-opus-4-8 (1835.3) ranks below claude-opus-4-7 (1906.4), and claude-sonnet-5 (1787.6) ranks below claude-sonnet-4-6 (1804.5) — same-vendor version regressions, on this benchmark, published in the vendor-neutral place vendors do not control. If a leaderboard only ever confirmed release-note ordering, it would not be measuring anything.

06The convergenceThe one thing both boards agree on.

Here is the finding that survives everything above. The two boards disagree about nearly every individual rank — and still agree about the shape of the market: on each board, measured on its own scale, the highest-ranked open-weight model in the rows above sits a few dozen points behind that board’s #1 — 38 points on arena.ai, 45.3 Elo on EQ-Bench. These are two separate observations on two incompatible scales, and they concern different models. That is what makes the agreement striking rather than circular.

arena.ai · its own scale
Top open-weight model shown, behind #1
38pts

glm-5.3-max at 1467 sits 38 points behind claude-fable-5's 1505 on the human-vote board, as stated Aug 27, 2026 — the highest-ranked open-weight family among the rows shown above.

Human votes
EQ-Bench · its own scale
Top open-weight model shown, behind #1
45.3Elo

kimi-k3 at 2070.8 sits 45.3 Elo behind claude-opus-5's 2116.1 on the LLM-judged board. A different open model, a different scale, the same shape.

LLM-judged
The caveat that keeps it honest
Separate observations, never one number
2scales

The boards' Elo scales are anchored differently — one clusters around 1500, the other tops out above 2100. The two gaps cannot be averaged, converted, or combined into a single open-vs-closed figure.

Do not convert

Neither board states this finding — each publishes only its own ranking, so the convergence only becomes visible when the two are laid side by side. It is also the most defensible takeaway for a buyer: whichever instrument you trust, the gap between the top open-weight models on these boards and the frontier is real but narrow, and the models occupying it are cheap to try.

07ImplicationsHow to act when the boards cancel out.

So which board should you act on? Neither, on its own — and this post will not hand you a winner, because a single verdict is exactly the artefact the data argues against. We have published a head-to-head frontier-model verdict where the evidence supported one; here the evidence supports a refusal. What the two boards give you instead is a pair of directional signals with known, documented biases — and EQ-Bench itself says the quiet part on its own methodology page: “The scores and rankings should only ever be interpreted as a rough guide of writing ability”, and “it’s good to be skeptical of benchmark numbers by default.”

Match the instrument to your reader
Pick the board whose rater resembles your audience

arena.ai aggregates untrained human preference under blind comparison — closer to how a newsletter subscriber or a client experiences copy. EQ-Bench is a graded, adversarial examination by Claude judges. If your writing is judged by people skimming, the human-vote signal is the nearer proxy; if it is judged on craft under scrutiny, the rubric-and-Elo signal is.

Rater ≈ audience
Treat single-board citations as a flag
One board quoted alone is a choice, not a fact

Any article declaring a best writing model from one leaderboard has silently picked an instrument whose biases it probably has not read. Ask which board, which category, and what the other board says about the same model before repeating the claim.

Always ask: which board?
Shortlist from the agreement, test on your work
Use the one convergent finding, then run your own briefs

Both boards place a top-ranked open-weight model close behind their own #1, on their own scales. That makes a two-tier shortlist cheap: one frontier model, one open model, your actual prompts, your judgement on the outputs. Ten of your own briefs beat either board for your use case.

Shortlist, then test

For reading any leaderboard — contamination, category design, cherry-picking — our benchmark methodology guide covers the general literacy this post applies to one specific disagreement. If the model is only half the equation for you — the other half being prompts, style constraints, and editorial scaffolding — the free skills that measurably improve AI writing matter more than a five-place rank difference ever will. And if the real question is standing up an editorial pipeline where model choice, briefs, and quality control are designed together rather than argued from leaderboards, that is what our Content Engine service builds.

08ConclusionThe disagreement is the finding.

Two boards, one task

Neither board alone, both boards together, and one number they agree on.

These two creative-writing leaderboards invert each other on the same nominal task: first place on one is sixth on the other, second on one is twenty-third on the other. That is not a scandal and not noise — it is what happens when a human blind-voting arena and a two-judge LLM pipeline, each with documented and partly self-admitted biases, are pointed at something as contested as writing quality.

What survives is precise: the inversion itself, the rubric-vs-Elo split inside EQ-Bench, and the one convergent finding — a top open-weight model sitting close behind #1 on each board, on two scales that must never be merged. A reader who carries those three things can evaluate any “best AI for writing” headline in about ten seconds.

The practical move is unchanged from the data: shortlist one frontier and one open model, run your own briefs, and trust your reading of the outputs over either board’s ordering. Both boards, to their credit, would tell you the same.

Choose writing models on evidence

The leaderboards disagree. Your briefs settle it.

Our team designs editorial pipelines where model selection is tested on your briefs and your audience — not inherited from whichever leaderboard an article happened to cite.

Free consultationExpert guidanceTailored solutions
What we work on

AI writing engagements

  • Model shortlists tested on your actual briefs
  • Editorial pipelines with quality gates, not vibes
  • Open-weight vs frontier cost and quality trade-offs
  • Style and voice control that survives model swaps
  • Benchmark literacy for content and marketing teams
FAQ · AI writing leaderboards

The questions we get about writing leaderboards.

The data supports a shortlist, not a verdict: one frontier model and one open-weight model — both boards place a top open-weight model close behind their own #1, on their own separate scales. Run your real briefs through both and judge the outputs. Your audience resembles one board’s rater more than the other’s; weight that board accordingly.
Related dispatches

Continue exploring AI model evaluation.