AI DevelopmentDecision Matrix6 min readPublished September 22, 2026

Leads 4 of 4 coding and computer-use rows · 40% of the token price · Fable 5.1 becomes the escalation, not the default

Is Claude Fable 5.1 Still Worth It for Agentic Coding?

Opus 5.5 outscores Fable 5.1 on every coding benchmark Anthropic published, at 40% of the token price. When Fable 5.1 still earns its place.

DA
Digital Applied Team
Research and practical guidance
Opus 5.5 releasedSeptember 22, 2026
Fable 5.1 releasedSeptember 1, 2026

For most agentic coding, no. Claude Opus 5.5, released on September 22, 2026, scores higher than Claude Fable 5.1 on all four coding and computer-use rows in Anthropic’s own comparison, and costs 40% as much per input and output token. On Artificial Analysis’s independent runs, Opus 5.5 at high effort beat Fable 5.1’s best Terminal-Bench 4.0 score at under a third of the cost per task.

Anthropic’s own guidance now says the same. Its model selection guide opens with “Most workloads start with Claude Opus 5.5” and tells teams to move to Fable 5.1 only “if your evals at xhigh or max effort still fall short on demanding reasoning or long-horizon agentic work.” Fable 5.1 still has a job. It is now the model you escalate to, not the one you start on.

Key takeaways
  1. 01
    Opus 5.5 leads Fable 5.1 on every coding row Anthropic published.Terminal-Bench 4.0 66.4% vs 55.8%, CursorBench 4.0 57.8% vs 51.8%, FrontierCode 54.4% vs 50.3%. Computer use is within noise (81.8% vs 80.7%).
  2. 02
    Independently measured, Opus 5.5 is roughly 3× to 4× cheaper for the same result.Artificial Analysis: Opus 5.5 at medium matched Fable 5.1 at high on Terminal-Bench 4.0 (52.5% vs 52.0%) at $1.34 against $3.91 per task.
  3. 03
    Fable 5.1 still has a job: the hardest long-horizon work, and as an escalation.Anthropic lists it for agent sessions that run for hours, deep research and finished documents. A conversation can move from Opus 5.5 up to Fable 5.1 and keep its reasoning.
  4. 04
    Two non-benchmark factors favour Opus 5.5.Zero data retention is available on Opus 5.5 but not on Fable 5.1 by default, and only Opus 5.5 offers fast mode.

01Vendor numbersAnthropic’s coding numbers

These are the four rows from Anthropic’s Opus 5.5 announcement that bear on agentic coding. Both models were run by Anthropic, mostly at max effort; Opus 5.5’s Terminal-Bench 4.0 score is its xhigh result, its best on that test.

Source: Anthropic, Introducing Claude Opus 5.5, September 22, 2026. Gap in percentage points, our arithmetic.
BenchmarkOpus 5.5Fable 5.1Gap
Terminal-Bench 4.066.4%55.8%+10.6
CursorBench 4.057.8%51.8%+6.0
FrontierCode v1.1 Main54.4%50.3%+4.1
OSWorld 2.0, partial (computer use)81.8%80.7%+1.1

The Terminal-Bench lead is the one that clears the noise. Anthropic gives a standard error of ±2.6 points for Opus 5.5 on that test, and the gap is 10.6. CursorBench and FrontierCode are smaller but consistent. Other vendors’ runs of Fable 5.1 land close to Anthropic’s: OpenAI scored it at 55.8% on Terminal-Bench 4.0 and 50.9% on FrontierCode Main, and xAI at 51.8% on CursorBench 4.0.

Two cautions. Anthropic’s own Fable 5.1 figures moved between its two announcements: Humanity’s Last Exam with tools went from 65.0% on September 1 to 65.6% on September 22, and OSWorld 2.0 partial from 77.9% to 80.7%. And Anthropic says it thinks the table overstates the difference:

In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.Anthropic, Introducing Claude Opus 5.5, September 22, 2026

That cuts both ways. It means Opus 5.5 is probably not dramatically better than Fable 5.1 in real use. It also means the two are close enough that price should decide, and on price the answer is clear. Anthropic’s headline describes Opus 5.5 as performing “at the level of Claude Fable 5.1 on most work,” which is a claim of parity, not superiority.

02Independent dataThe independent check, effort by effort

Artificial Analysis ran both models at all five effort levels on its own Terminal-Bench 4.0 harness (model pages for Opus 5.5 and Fable 5.1). The table shows each score beside the model’s cost per task across the firm’s full Intelligence Index (v4.3.2, read on September 22), which includes Terminal-Bench and nine other evaluations.

Terminal-Bench 4.0 score · cost per Intelligence Index task. Source: Artificial Analysis v4.3.2, read September 22, 2026. Both Claude models run with fallback on.
EffortOpus 5.5Fable 5.1
low31.3% · $0.5540.4% · $2.37
medium52.5% · $1.3444.9% · $2.98
high56.6% · $1.8252.0% · $3.91
xhigh59.6% · $3.4655.1% · $5.98
max59.6% · $5.9852.0% · $7.63

Three things stand out. Opus 5.5 at high (56.6%) already beats Fable 5.1’s best result at any effort (55.1% at xhigh). Fable 5.1 scores lower at max than at xhigh on this test, so its most expensive setting buys nothing here. And Fable 5.1 has the higher floor: at low effort it scores 40.4% against Opus 5.5’s 31.3%. But Fable at low still costs more per task ($2.37) than Opus 5.5 at medium ($1.34), which scores 52.5%.

Cost per Intelligence Index task, with Terminal-Bench 4.0 score

Artificial Analysis Intelligence Index v4.3.2, read September 22, 2026
Opus 5.5 · mediumTerminal-Bench 52.5%
$1.34
Opus 5.5 · highTerminal-Bench 56.6%
$1.82
Opus 5.5 · xhighTerminal-Bench 59.6%
$3.46
Fable 5.1 · highTerminal-Bench 52.0%
$3.91
Fable 5.1 · xhighTerminal-Bench 55.1%
$5.98
Fable 5.1 · maxTerminal-Bench 52.0%
$7.63

A second independent evaluator, Vals AI, published its Opus 5.5 results the same day, and its longest tasks point the same way. On its RSI Index, where agents get 24 hours to improve things like a language-model training run or a compressor, Opus 5.5 ranked first and Fable 5.1 second. On ProgramBench, which asks a model to rebuild a command-line program from its binary and counts a task only when every hidden test passes, Opus 5.5 fully solved 37 of 200 tasks to Fable 5.1’s 14.

Source: Vals AI model, RSI Index and ProgramBench pages, September 22, 2026. Opus 5.5 at max effort; RSI tasks run in Claude Code.
Vals AI benchmarkOpus 5.5Fable 5.1
Vals RSI Index (24-hour AI research tasks)37.13% · #1 of 1935.03% · #2
ProgramBench, fully resolved (of 200)37 · 18.5%14 · 7.0%
ProgramBench, raw test pass rate87.0%82.7%
LM Training, bits per byte (lower is better)0.8310.854
Vals Index (broad, all domains)66.16% · #4 of 6068.83% · #1

On LM Training, Opus 5.5’s small model reached 0.831 bits per byte, below the 0.85 published human reference; Vals notes that its setup is not a matched reproduction of the reference paper. Two caveats keep this in proportion. Vals ran Opus 5.5 with Opus 5 and Opus 4.8 as fallbacks for refused requests; counting those tasks as failures drops its Terminal-Bench 4.0 score from 61.62% to 53.54%. And on Vals’ broad index across legal, finance, medical and coding work, Fable 5.1 still ranks first and Opus 5.5 fourth.

03The savingHow much cheaper Opus 5.5 is for the same result

Matching the two models at similar scores gives the clearest answer to how much more cost-effective Opus 5.5 is. All figures are from Artificial Analysis; the ratios are our arithmetic.

  • Same Terminal-Bench score: Opus 5.5 at medium (52.5%) against Fable 5.1 at high (52.0%) is $1.34 against $3.91 per task, about 2.9× cheaper. On the Terminal-Bench portion of that cost alone, the ratio is also about 2.9×.
  • Fable’s best coding result: Opus 5.5 at high beats Fable 5.1 at xhigh (56.6% against 55.1%) for $1.82 against $5.98, about 3.3× cheaper.
  • Same overall index score: Opus 5.5 at high (53.6) against Fable 5.1 at max (53.4) is $1.82 against $7.63, about 4.2× cheaper.

Anthropic’s own long coding test points the same way. Asked to translate the HAProxy load balancer from C into Rust, both models passed nearly all of its regression tests. Opus 5.5 finished in 9.5 hours against Fable 5.1’s 12, at 51% lower cost.

Per token, the gap is smaller on one line only. Fable 5.1’s cache reads are cheap relative to its own input price, at $0.25 against Opus 5.5’s $0.20. That was the argument for Fable 5.1 over Opus 5 in our September 1 Fable 5.1 post. Against Opus 5.5, Fable 5.1 costs more on every line.

Per million tokens. Source: Anthropic pricing page, September 22, 2026. Last column is Opus 5.5’s price as a share of Fable 5.1’s.
LineOpus 5.5Fable 5.1Opus share
Input$4$1040%
Output$20$5040%
Cache read$0.20$0.2580%
Cache write, 5 min / 1 hr$5 / $8$12.50 / $2040%
Batch, input / output$2 / $10$5 / $2540%

04The exceptionsWhere Fable 5.1 still earns its place

Anthropic still calls Fable 5.1 its “most capable widely released model,” and its selection matrix lists it for “the highest available capability”: agent sessions that run for hours, multistep deep research, and analysis carried through to a finished document, spreadsheet or deck. Opus 5.5 is the matrix’s pick for complex agentic coding, including multihour autonomous coding agents and large refactors.

In practice, that leaves four reasons to reach for Fable 5.1:

  • Your evals at Opus 5.5 xhigh or max still fail. This is Anthropic’s own trigger, and the only one backed by your data rather than a benchmark.
  • The task is research or document work as much as code. Anthropic positions Fable 5.1 for work that ends in a finished report or deck. Even here, Anthropic’s own research-report test favoured Opus 5.5, so check before assuming.
  • You need a stronger result at low effort. Fable 5.1’s low setting outscored Opus 5.5’s on Terminal-Bench, though Opus 5.5 at medium was both better and cheaper.
  • The work spans domains beyond code. Vals AI’s broad index ranks Fable 5.1 first of 60 models at 68.83% and Opus 5.5 fourth at 66.16%, and Vals found Opus 5.5 weaker than Opus 5 on medical coding, legal research and tax tasks.
One Anthropic page hasn’t caught up yet

Anthropic’s cost-optimisation guide still says “For most agent workloads, start with Claude Fable 5.1 at low effort,” and supports it with comparisons against Opus 5, not Opus 5.5. The model selection guide, updated for Opus 5.5, says most workloads should start on Opus 5.5. The Fable 5.1 model page, too, still points readers to Opus 5 as the starting model. We read the newer guidance as current.

05The mechanicsHow to escalate from Opus 5.5 to Fable 5.1

The two models behave differently in production, and one of the differences makes the escalation pattern work.

Sources: Anthropic model pages for Opus 5.5 and Fable 5.1, What’s new in Claude Opus 5.5, fast-mode docs. September 22, 2026.
AreaOpus 5.5Fable 5.1
Default effortmediumhigh
Anthropic’s latency labelModerateSlower
Fast modeYes, Claude API only, 2× priceNot offered
Data retentionZero data retention available30-day retention; no zero retention unless Anthropic authorises it
Reads the other model’s thinkingNo, not Fable’sYes, Opus 5.5’s (Claude API)

The last row is the useful one. On the Claude API, Fable 5.1 can read the thinking blocks Opus 5.5 produces, so a conversation can start on Opus 5.5 and move up to Fable 5.1 without losing the reasoning so far. The reverse doesn’t work: moving from Fable 5.1 down to Opus 5.5 drops Fable’s reasoning. So start on Opus 5.5 and switch up when a task stalls; don’t start on Fable 5.1 and try to save money later in the same conversation.

One route that isn’t available yet: Anthropic’s advisor tool, which lets a cheaper model consult a stronger one mid-task, does not list Opus 5.5 in its compatibility table as of September 22. Until it does, escalation means switching the conversation’s model, not pairing the two.

For regulated work, the retention row may decide the question before any benchmark does. According to its model page, Fable 5.1 carries 30-day data retention and isn’t available under zero data retention unless Anthropic expressly authorises it. Opus 5.5 is.

06ConclusionOpus 5.5 is the better default for agentic coding; Fable 5.1 is the escalation

Everyday agentic coding: features, fixes, refactors, reviews
Run Opus 5.5 at medium, and move to high for multi-step changes. On Artificial Analysis’s data this matches or beats Fable 5.1 at roughly a third of the cost per task.
Opus 5.5
A long-horizon task still fails on Opus 5.5 at xhigh or max
Switch the conversation up to Fable 5.1 on the Claude API; it keeps Opus 5.5’s reasoning. Record which tasks needed it, so you know how often you are paying the premium.
Escalate to Fable 5.1
Your contract requires zero data retention
Use Opus 5.5. Fable 5.1 needs Anthropic’s express authorisation for zero retention.
Opus 5.5
You run Fable 5.1 as your default today
Run your own coding tasks on Opus 5.5 at medium and high beside your Fable 5.1 baseline. If quality holds, you cut input and output prices by 60% and, on this evidence, the per-task bill by more.
Test the switch
What to do this week

Make Opus 5.5 the default for coding agents, and keep Fable 5.1 for the tasks that fail on it

Fable 5.1 is still a strong model, and for the hardest long-horizon work it may still be the right call. But on every coding benchmark Anthropic published, and on Artificial Analysis’s independent runs, Opus 5.5 does at least as well for well under half the cost. Pay for Fable 5.1 when your own evals show you need it, not by default. For the full launch detail see our Opus 5.5 launch post, and for how it compares outside Anthropic, our cost-per-task comparison.

Digital Applied

Route each task to the model it actually needs.

We measure your coding agents' tasks on each model and effort level, set the default, and build the escalation path, so you pay for the top model only when a task needs it.

Model evaluationEffort tuningEscalation routing
Your next project

A coding agent that pays for Fable only when it must

  • Your tasks, measured on both models
  • A default model and effort level
  • An escalation rule with a record
Questions and answers

The questions we get about Opus 5.5 and Fable 5.1 for coding

On published evidence, yes. It scores higher on Terminal-Bench 4.0, CursorBench 4.0 and FrontierCode in Anthropic's table, on Artificial Analysis's independent Terminal-Bench runs Opus 5.5 at high effort beat Fable 5.1's best score, and on Vals AI's ProgramBench it fully solved 37 of 200 tasks to Fable 5.1's 14. Anthropic itself says the real-world gap is narrower than the benchmarks suggest.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading