AI DevelopmentDecision Matrix6 min readPublished September 28, 2026

40 chart points · 4 benchmarks · 5 effort levels · the cheaper model is not always the cheaper answer

Sonnet 5.5 or Opus 5.5: Which Effort Level to Pay For

Anthropic's charts give 40 score and cost points for Sonnet 5.5 and Opus 5.5. Up to $2 a task Sonnet wins three of four; at $5, Opus wins all four.

DA
Digital Applied Team
Research and practical guidance
DataAnthropic, September 28, 2026
Points40 (plus 40 comparison points)

Claude Sonnet 5.5 has half Opus 5.5’s input and output token price, but that does not make it half the price per finished task. Anthropic’s launch page for Sonnet 5.5, published on September 28, 2026, plots each model’s score against its cost per task at five effort levels on four benchmarks. We extracted the exact points from the page’s chart data.

They give a clear answer. Up to about $2 a task, Sonnet 5.5 buys the higher score on three of the four benchmarks. At $5 a task, Opus 5.5 at medium or high buys the higher score on all four. Sonnet 5.5 at max is poor value on three of the four charts, where a cheaper Opus 5.5 setting scores higher. The exception is the single highest Terminal-Bench score.

Key takeaways
  1. 01
    Under about $2 a task, Sonnet 5.5 at medium or high is the better buy.It leads that budget on Terminal-Bench 4.0, CursorBench 4.0 and AA-Briefcase. On FrontierCode, Opus 5.5 at medium wins at every budget from $0.80 up.
  2. 02
    At $5 a task, Opus 5.5 at medium or high wins on all four charts.On CursorBench, Opus 5.5 at high scores 56.0% for $3.97; Sonnet 5.5 needs max, at $9.67, to reach 55.5%.
  3. 03
    At the Claude API defaults, Opus 5.5 scores higher on every chart.The API defaults Sonnet 5.5 to high and Opus 5.5 to medium. Opus costs 11% to 90% more per task at those settings.
  4. 04
    Every point is vendor-run, and two charts carry caveats.Sonnet 5.5 scores lower at max than at xhigh on FrontierCode, and AA-Briefcase ran on a pre-release deployment with a since-fixed bug.

01 — The dataWhat the charts measure

An effort level tells a Claude model how long to think before and between actions. Higher levels usually score better and cost more, because the model spends more output tokens. Anthropic’s announcement charts plot one point per model per level, from low to max. The x-axis is cost per task in US dollars, or cost per attempt on Terminal-Bench. The y-axis is the score, or an Elo rating on AA-Briefcase.

Anthropic’s own reading is that Sonnet 5.5 is most useful beside Opus 5.5 at the low end, where it is cheaper per task, and roughly level on score and cost at the high end. The points support the first half. The second half is generous to Sonnet 5.5: on three of the four charts, Opus 5.5 at a lower level matches or beats one of Sonnet 5.5’s two top settings for less money. That happens at xhigh on Terminal-Bench, at xhigh and max on FrontierCode, and at max on CursorBench.

02 — The findingThe best buy at each budget

For each benchmark, the table shows the highest-scoring configuration that costs no more than $1, $2 or $5 per task. We included Sonnet 5, GPT-6 Sol and GPT-5.6 Sol where Anthropic charted them; in this table, only AA-Briefcase’s $1 column goes to a non-Claude model.

Our selection from Anthropic’s chart data, September 28, 2026. Terminal-Bench costs are per attempt.
BenchmarkUp to $1Up to $2Up to $5
Terminal-Bench 4.0Sonnet medium · 28.8%Sonnet high · 43.0%Opus high · 64.2%
FrontierCode 1.1Opus medium · 54.6%Opus medium · 54.6%Opus medium · 54.6%
CursorBench 4.0Sonnet medium · 39.2%Sonnet high · 47.8%Opus high · 56.0%
AA-Briefcase 1.1 (Elo)GPT-6 Sol high · 1289Sonnet medium · 1461Opus medium · 1642

The pattern holds because the two models’ curves cross. Sonnet 5.5 is cheap at low and medium and climbs quickly to high. Opus 5.5 starts higher and flattens out, so its medium or high point often matches Sonnet 5.5’s xhigh or max at a lower price. On FrontierCode and CursorBench, spending past $5 buys at most 1.8 points. On Terminal-Bench and AA-Briefcase it keeps paying: Opus 5.5 at max rates 180 Elo points above the best AA-Briefcase result under $5, for $21.05 a task.

The one exception

Terminal-Bench 4.0 is the only chart where the top score belongs to Sonnet 5.5: 70.6% at max, for $12.54 per attempt. Opus 5.5 peaks at 66.4% at xhigh for $7.35. If a long-running terminal agent has to finish as many tasks as possible and cost is secondary, Sonnet 5.5 at max is the choice the data supports.

03 — CodingThe three coding benchmarks, point by point

Terminal-Bench 4.0 measures multi-step work in a command line. Opus 5.5 at high scores 64.2% for $3.88, above Sonnet 5.5 at xhigh (61.5% for $5.30) at 27% less per attempt.

Score and cost per attempt. Source: Anthropic, Sonnet 5.5 announcement chart data, September 28, 2026.
EffortSonnet 5.5Opus 5.5
Low20.0% · $0.7638.5% · $1.29
Medium28.8% · $0.8357.6% · $2.94
High43.0% · $1.9464.2% · $3.88
Xhigh61.5% · $5.3066.4% · $7.35
Max70.6% · $12.5464.8% · $11.24

FrontierCode 1.1 asks whether a code change could be merged without human edits. Opus 5.5 at medium has the best score on the chart, 54.6%, for $0.80. Sonnet 5.5 at high is the budget pick: 49.4% for $0.42, about what Opus 5.5 at low costs. Sonnet 5.5 at max costs $20.78 and scores below its own high result; Anthropic’s footnote attributes this to the model running a multi-agent code review at max, which in two examined cases led to a timeout or to edits outside the task.

Score and cost per task. Source: Anthropic, Sonnet 5.5 announcement chart data, September 28, 2026.
EffortSonnet 5.5Opus 5.5
Low29.3% · $0.1947.3% · $0.40
Medium36.5% · $0.2454.6% · $0.80
High49.4% · $0.4254.0% · $1.09
Xhigh52.1% · $1.5951.4% · $2.25
Max46.2% · $20.7854.4% · $6.19

CursorBench 4.0 uses tasks from real Cursor coding sessions. The two models’ curves cross near $3 to $4: Sonnet 5.5 at xhigh scores 53.1% for $3.88, while Opus 5.5 at high scores 56.0% for $3.97. Sonnet 5.5 at max (55.5%, $9.67) costs 2.4 times as much as Opus 5.5 at high for a slightly lower score.

Score and cost per task. Source: Anthropic, Sonnet 5.5 announcement chart data, September 28, 2026.
EffortSonnet 5.5Opus 5.5
Low35.8% · $0.5043.7% · $1.17
Medium39.2% · $0.7052.5% · $2.91
High47.8% · $1.6756.0% · $3.97
Xhigh53.1% · $3.8856.0% · $6.98
Max55.5% · $9.6757.8% · $13.43

04 — Knowledge workKnowledge work, where Sonnet 5.5 stays close for longer

AA-Briefcase 1.1 is Artificial Analysis’ test of long-running knowledge work, scored as an Elo rating. Here the two Claude models stay close through high: Sonnet 5.5 at high rates 1634 for $3.95, and Opus 5.5 at medium rates 1642 for $4.40. At the top, the order flips on cost. Opus 5.5 at max rates 1822 for $21.05; Sonnet 5.5 at max rates 1811 for $29.19.

Elo rating and cost per task. Source: Anthropic, Sonnet 5.5 announcement chart data, September 28, 2026; run by Artificial Analysis on a pre-release deployment.
EffortSonnet 5.5Opus 5.5
Low1264 · $0.871285 · $1.15
Medium1461 · $1.641642 · $4.40
High1634 · $3.951705 · $6.27
Xhigh1746 · $9.631780 · $12.27
Max1811 · $29.191822 · $21.05

GPT-6 Sol is far cheaper on this chart. At max it rates 1483 for $2.67, and at high it rates 1289 for $0.63, which is the best result under $1. Sonnet 5.5 at medium comes within 22 points of GPT-6 Sol’s best rating for $1.64, and passes it at high. Two caveats apply to this chart. Anthropic says the Sonnet 5.5 run used a pre-release deployment with a structured-outputs bug, since fixed, which it expects to understate Sonnet 5.5 slightly. And OpenAI recently fixed an image-understanding bug in GPT-6 Sol that published scores may not yet reflect, though Artificial Analysis does not expect a major impact on AA-Briefcase.

05 — DefaultsAt default settings, the two models are not like for like

The Claude API does not give both models the same default. According to Anthropic’s models overview, Sonnet 5.5 defaults to high and Opus 5.5 to medium. In Claude Code and the Claude apps, Sonnet 5.5 defaults to medium. A team that swaps models without setting effort is therefore comparing Sonnet 5.5 at high with Opus 5.5 at medium:

  • Terminal-Bench 4.0: Sonnet 43.0% for $1.94; Opus 57.6% for $2.94.
  • FrontierCode: Sonnet 49.4% for $0.42; Opus 54.6% for $0.80.
  • CursorBench: Sonnet 47.8% for $1.67; Opus 52.5% for $2.91.
  • AA-Briefcase: Sonnet 1634 for $3.95; Opus 1642 for $4.40.

Opus 5.5 scores higher on all four, and costs between 11% and 90% more per task. Whether that is worth paying depends on how much a failed task costs you. For a coding agent whose failures a person has to fix, the higher Opus score at medium may be the cheaper option overall. For high-volume work that can be checked automatically, Sonnet 5.5 at high with a retry on failure is the natural starting point. Our Sonnet 5.5 launch post covers the price and the headline table, and the Opus 5.5 launch post covers Opus 5.5’s own effort curves.

06 — MethodHow the points were collected

Methodology

Every value is Anthropic’s, as published on its Sonnet 5.5 launch page. The selections and ratios are ours.

What was collected
All 80 points on Anthropic’s four score-against-cost charts: Sonnet 5.5, Opus 5.5, Sonnet 5 and one OpenAI model, each at five effort levels. The tables print the 40 Sonnet 5.5 and Opus 5.5 points.
Sources
Anthropic, “Introducing Claude Sonnet 5.5,” September 28, 2026: the accessible labels on each chart point, which give the score and cost to the cent; and Anthropic’s models overview for default effort. Nothing was estimated from a chart image.
As-of date
Read on September 28, 2026, the day of the launch.
Units
US dollars per task, or per attempt on Terminal-Bench 4.0, as Anthropic’s axes label them. Scores are percentages, except AA-Briefcase, which is an Elo rating. The OpenAI model is GPT-6 Sol on FrontierCode and AA-Briefcase, and GPT-5.6 Sol on Terminal-Bench and CursorBench, where no public GPT-6 Sol result exists.
Known limitations
All runs are vendor-reported and none has been replicated. Anthropic does not state the prices behind each cost, the number of tasks per point or the error bars. The budget table ignores differences smaller than the likely noise.
Refresh
Updated in place if Anthropic revises the charts or publishes Haiku 5.5 on the same axes.

07 — ConclusionSonnet 5.5 wins the cheap end, Opus 5.5 the middle

High-volume work checked automatically
Sonnet 5.5 at medium or high, with a retry at a higher level only on failure.
Under $2 a task
Coding agents whose failures a person fixes
Opus 5.5 at medium or high. On FrontierCode and CursorBench it reaches Sonnet 5.5's best scores for less.
About $3 to $5 a task
Long terminal tasks where completion matters most
Sonnet 5.5 at max holds the top Terminal-Bench score, at the highest cost per attempt of either model.
Above $10 an attempt
You swap models without setting effort
Set effort explicitly. The API defaults compare Sonnet 5.5 at high with Opus 5.5 at medium.
Before any comparison
What to do this week

Pick a cost per task you can afford, then choose the model and effort level that scores best at that budget

Anthropic’s charts are vendor-run, but they are detailed enough to reason from. Stop treating Sonnet 5.5 as the cheap default and Opus 5.5 as the premium option. Run your own tasks at two levels on each model, around the budget you can carry, and compare cost per completed task. The cross-vendor view of effort levels is in our reasoning-effort guide.

Digital Applied

Spend on the effort level that pays.

We run effort sweeps on your own tasks and build the routing that sends each job to the cheapest model and level that passes.

Effort sweepsModel routingCost per task tracking
Your next project

A model budget built on data

  • →Your tasks at every effort level
  • →Cost per completed task
  • →Routing rules you can audit
Questions and answers

The questions we get about Sonnet 5.5 and Opus 5.5 effort

At the same effort level, yes. For the same score, not always. On Anthropic's charts, Opus 5.5 at medium or high reaches Sonnet 5.5's best score for less on FrontierCode and CursorBench.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading