Claude Sonnet 5.5 has half Opus 5.5’s input and output token price, but that does not make it half the price per finished task. Anthropic’s launch page for Sonnet 5.5, published on September 28, 2026, plots each model’s score against its cost per task at five effort levels on four benchmarks. We extracted the exact points from the page’s chart data.
They give a clear answer. Up to about $2 a task, Sonnet 5.5 buys the higher score on three of the four benchmarks. At $5 a task, Opus 5.5 at medium or high buys the higher score on all four. Sonnet 5.5 at max is poor value on three of the four charts, where a cheaper Opus 5.5 setting scores higher. The exception is the single highest Terminal-Bench score.
- 01Under about $2 a task, Sonnet 5.5 at medium or high is the better buy.It leads that budget on Terminal-Bench 4.0, CursorBench 4.0 and AA-Briefcase. On FrontierCode, Opus 5.5 at medium wins at every budget from $0.80 up.
- 02At $5 a task, Opus 5.5 at medium or high wins on all four charts.On CursorBench, Opus 5.5 at high scores 56.0% for $3.97; Sonnet 5.5 needs max, at $9.67, to reach 55.5%.
- 03At the Claude API defaults, Opus 5.5 scores higher on every chart.The API defaults Sonnet 5.5 to high and Opus 5.5 to medium. Opus costs 11% to 90% more per task at those settings.
- 04Every point is vendor-run, and two charts carry caveats.Sonnet 5.5 scores lower at max than at xhigh on FrontierCode, and AA-Briefcase ran on a pre-release deployment with a since-fixed bug.
01 — The dataWhat the charts measure
An effort level tells a Claude model how long to think before and between actions. Higher levels usually score better and cost more, because the model spends more output tokens. Anthropic’s announcement charts plot one point per model per level, from low to max. The x-axis is cost per task in US dollars, or cost per attempt on Terminal-Bench. The y-axis is the score, or an Elo rating on AA-Briefcase.
Anthropic’s own reading is that Sonnet 5.5 is most useful beside Opus 5.5 at the low end, where it is cheaper per task, and roughly level on score and cost at the high end. The points support the first half. The second half is generous to Sonnet 5.5: on three of the four charts, Opus 5.5 at a lower level matches or beats one of Sonnet 5.5’s two top settings for less money. That happens at xhigh on Terminal-Bench, at xhigh and max on FrontierCode, and at max on CursorBench.
02 — The findingThe best buy at each budget
For each benchmark, the table shows the highest-scoring configuration that costs no more than $1, $2 or $5 per task. We included Sonnet 5, GPT-6 Sol and GPT-5.6 Sol where Anthropic charted them; in this table, only AA-Briefcase’s $1 column goes to a non-Claude model.
| Benchmark | Up to $1 | Up to $2 | Up to $5 |
|---|---|---|---|
| Terminal-Bench 4.0 | Sonnet medium · 28.8% | Sonnet high · 43.0% | Opus high · 64.2% |
| FrontierCode 1.1 | Opus medium · 54.6% | Opus medium · 54.6% | Opus medium · 54.6% |
| CursorBench 4.0 | Sonnet medium · 39.2% | Sonnet high · 47.8% | Opus high · 56.0% |
| AA-Briefcase 1.1 (Elo) | GPT-6 Sol high · 1289 | Sonnet medium · 1461 | Opus medium · 1642 |
The pattern holds because the two models’ curves cross. Sonnet 5.5 is cheap at low and medium and climbs quickly to high. Opus 5.5 starts higher and flattens out, so its medium or high point often matches Sonnet 5.5’s xhigh or max at a lower price. On FrontierCode and CursorBench, spending past $5 buys at most 1.8 points. On Terminal-Bench and AA-Briefcase it keeps paying: Opus 5.5 at max rates 180 Elo points above the best AA-Briefcase result under $5, for $21.05 a task.
Terminal-Bench 4.0 is the only chart where the top score belongs to Sonnet 5.5: 70.6% at max, for $12.54 per attempt. Opus 5.5 peaks at 66.4% at xhigh for $7.35. If a long-running terminal agent has to finish as many tasks as possible and cost is secondary, Sonnet 5.5 at max is the choice the data supports.
03 — CodingThe three coding benchmarks, point by point
Terminal-Bench 4.0 measures multi-step work in a command line. Opus 5.5 at high scores 64.2% for $3.88, above Sonnet 5.5 at xhigh (61.5% for $5.30) at 27% less per attempt.
| Effort | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Low | 20.0% · $0.76 | 38.5% · $1.29 |
| Medium | 28.8% · $0.83 | 57.6% · $2.94 |
| High | 43.0% · $1.94 | 64.2% · $3.88 |
| Xhigh | 61.5% · $5.30 | 66.4% · $7.35 |
| Max | 70.6% · $12.54 | 64.8% · $11.24 |
FrontierCode 1.1 asks whether a code change could be merged without human edits. Opus 5.5 at medium has the best score on the chart, 54.6%, for $0.80. Sonnet 5.5 at high is the budget pick: 49.4% for $0.42, about what Opus 5.5 at low costs. Sonnet 5.5 at max costs $20.78 and scores below its own high result; Anthropic’s footnote attributes this to the model running a multi-agent code review at max, which in two examined cases led to a timeout or to edits outside the task.
| Effort | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Low | 29.3% · $0.19 | 47.3% · $0.40 |
| Medium | 36.5% · $0.24 | 54.6% · $0.80 |
| High | 49.4% · $0.42 | 54.0% · $1.09 |
| Xhigh | 52.1% · $1.59 | 51.4% · $2.25 |
| Max | 46.2% · $20.78 | 54.4% · $6.19 |
CursorBench 4.0 uses tasks from real Cursor coding sessions. The two models’ curves cross near $3 to $4: Sonnet 5.5 at xhigh scores 53.1% for $3.88, while Opus 5.5 at high scores 56.0% for $3.97. Sonnet 5.5 at max (55.5%, $9.67) costs 2.4 times as much as Opus 5.5 at high for a slightly lower score.
| Effort | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Low | 35.8% · $0.50 | 43.7% · $1.17 |
| Medium | 39.2% · $0.70 | 52.5% · $2.91 |
| High | 47.8% · $1.67 | 56.0% · $3.97 |
| Xhigh | 53.1% · $3.88 | 56.0% · $6.98 |
| Max | 55.5% · $9.67 | 57.8% · $13.43 |
04 — Knowledge workKnowledge work, where Sonnet 5.5 stays close for longer
AA-Briefcase 1.1 is Artificial Analysis’ test of long-running knowledge work, scored as an Elo rating. Here the two Claude models stay close through high: Sonnet 5.5 at high rates 1634 for $3.95, and Opus 5.5 at medium rates 1642 for $4.40. At the top, the order flips on cost. Opus 5.5 at max rates 1822 for $21.05; Sonnet 5.5 at max rates 1811 for $29.19.
| Effort | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Low | 1264 · $0.87 | 1285 · $1.15 |
| Medium | 1461 · $1.64 | 1642 · $4.40 |
| High | 1634 · $3.95 | 1705 · $6.27 |
| Xhigh | 1746 · $9.63 | 1780 · $12.27 |
| Max | 1811 · $29.19 | 1822 · $21.05 |
GPT-6 Sol is far cheaper on this chart. At max it rates 1483 for $2.67, and at high it rates 1289 for $0.63, which is the best result under $1. Sonnet 5.5 at medium comes within 22 points of GPT-6 Sol’s best rating for $1.64, and passes it at high. Two caveats apply to this chart. Anthropic says the Sonnet 5.5 run used a pre-release deployment with a structured-outputs bug, since fixed, which it expects to understate Sonnet 5.5 slightly. And OpenAI recently fixed an image-understanding bug in GPT-6 Sol that published scores may not yet reflect, though Artificial Analysis does not expect a major impact on AA-Briefcase.
05 — DefaultsAt default settings, the two models are not like for like
The Claude API does not give both models the same default. According to Anthropic’s models overview, Sonnet 5.5 defaults to high and Opus 5.5 to medium. In Claude Code and the Claude apps, Sonnet 5.5 defaults to medium. A team that swaps models without setting effort is therefore comparing Sonnet 5.5 at high with Opus 5.5 at medium:
- Terminal-Bench 4.0: Sonnet 43.0% for $1.94; Opus 57.6% for $2.94.
- FrontierCode: Sonnet 49.4% for $0.42; Opus 54.6% for $0.80.
- CursorBench: Sonnet 47.8% for $1.67; Opus 52.5% for $2.91.
- AA-Briefcase: Sonnet 1634 for $3.95; Opus 1642 for $4.40.
Opus 5.5 scores higher on all four, and costs between 11% and 90% more per task. Whether that is worth paying depends on how much a failed task costs you. For a coding agent whose failures a person has to fix, the higher Opus score at medium may be the cheaper option overall. For high-volume work that can be checked automatically, Sonnet 5.5 at high with a retry on failure is the natural starting point. Our Sonnet 5.5 launch post covers the price and the headline table, and the Opus 5.5 launch post covers Opus 5.5’s own effort curves.
06 — MethodHow the points were collected
Every value is Anthropic’s, as published on its Sonnet 5.5 launch page. The selections and ratios are ours.
- What was collected
- All 80 points on Anthropic’s four score-against-cost charts: Sonnet 5.5, Opus 5.5, Sonnet 5 and one OpenAI model, each at five effort levels. The tables print the 40 Sonnet 5.5 and Opus 5.5 points.
- Sources
- Anthropic, “Introducing Claude Sonnet 5.5,” September 28, 2026: the accessible labels on each chart point, which give the score and cost to the cent; and Anthropic’s models overview for default effort. Nothing was estimated from a chart image.
- As-of date
- Read on September 28, 2026, the day of the launch.
- Units
- US dollars per task, or per attempt on Terminal-Bench 4.0, as Anthropic’s axes label them. Scores are percentages, except AA-Briefcase, which is an Elo rating. The OpenAI model is GPT-6 Sol on FrontierCode and AA-Briefcase, and GPT-5.6 Sol on Terminal-Bench and CursorBench, where no public GPT-6 Sol result exists.
- Known limitations
- All runs are vendor-reported and none has been replicated. Anthropic does not state the prices behind each cost, the number of tasks per point or the error bars. The budget table ignores differences smaller than the likely noise.
- Refresh
- Updated in place if Anthropic revises the charts or publishes Haiku 5.5 on the same axes.
07 — ConclusionSonnet 5.5 wins the cheap end, Opus 5.5 the middle
Pick a cost per task you can afford, then choose the model and effort level that scores best at that budget
Anthropic’s charts are vendor-run, but they are detailed enough to reason from. Stop treating Sonnet 5.5 as the cheap default and Opus 5.5 as the premium option. Run your own tasks at two levels on each model, around the budget you can carry, and compare cost per completed task. The cross-vendor view of effort levels is in our reasoning-effort guide.