Anthropic released Claude Sonnet 5.5 on September 28, 2026, six days after Opus 5.5. It is the second model in the Claude 5.5 family and it costs exactly what Sonnet 5 costs: $2 per million input tokens, $10 per million output tokens and $0.20 per million for cache reads. Anthropic says it costs up to 30% less per task anyway, because it needs fewer tokens to finish the same work.
The benchmark jump is large. On Terminal-Bench 4.0, a test of multi-step work in a command line, Sonnet 5.5 scores 70.6% against Sonnet 5’s 10.3%, and above Opus 5.5’s 66.4%. That figure comes from the highest effort setting, where one attempt costs more than six times what it does at the API default. This post separates the price from the per-task claim, shows which setting each score comes from, and lists what changes for code that runs on Sonnet 5 today.
Every score here is Anthropic’s own or a partner’s as quoted by Anthropic. We have not re-run any of them.
- 01The per-token price did not change; the up-to-30% saving is a per-task claim.Sonnet 5.5 lists at $2 / $10 with $0.20 cache reads, the same as Sonnet 5. The saving depends on it using fewer tokens on your work.
- 02The headline scores are max-effort results, not what the API default delivers.On Anthropic's own chart, Terminal-Bench 4.0 falls from 70.6% at max to 43.0% at high, the Claude Platform default.
- 03Opus 5.5 stays ahead on seven of the eight rows, mostly by narrow margins.Sonnet 5.5 leads only on Terminal-Bench 4.0. Anthropic says Opus 5.5 is still clearly stronger at open-ended work that needs sustained judgment.
- 04It is the first Sonnet with cyber fallbacks and reasoning-extraction classifiers.Higher-risk security requests fall back to Sonnet 5, and thinking blocks are now tied to the account that created them.
01 — The releaseWhat shipped
- API IDAmazon Bedrock uses anthropic.claude-sonnet-5-5. No date suffix.
- claude-sonnet-5-5
- Context / max outputUp to 300K output on the Batch API with the output-300k-2026-03-24 beta header.
- 1M / 128K
- Knowledge cutoffReliable knowledge and training data, per Anthropic's models overview.
- June 2026
- Default effortHigh on the Claude Platform; medium in Claude Code and the Claude apps.
- high / medium
- RetirementEarliest date Anthropic commits to on the models overview.
- ≥ Sep 28, 2027
Anthropic positions the two 5.5 models differently. In the announcement, Anthropic pitches Opus 5.5 at complex work where judgment matters and Sonnet 5.5 at clearly defined daily tasks: fixing bugs and producing documents, slide decks and spreadsheets. Claude Haiku 5.5 is due “in the coming weeks.” Anthropic also says Sonnet 5.5 generates output more than 30% faster than Sonnet 5, which makes it the fastest Sonnet so far.
The model is available on the Claude Platform, Amazon Web Services, Google Cloud and Microsoft Azure, with zero data retention available as it is for Opus 5.5 and Sonnet 5. On OpenRouter, the route anthropic/claude-sonnet-5.5 was listed at 18:04 UTC on launch day, with a :batch route at exactly half price. At our check at 21:28 UTC, OpenRouter’s ~anthropic/claude-sonnet-latest alias pointed at Sonnet 5.5, so anyone calling that alias changed model without changing code. Our alias and retirement ledger tracks moves like that one.
02 — The invoiceThe price is unchanged, so the saving has to come from tokens
| Line | Sonnet 5.5 | Sonnet 5 | Opus 5.5 |
|---|---|---|---|
| Input | $2 | $2 | $4 |
| Output | $10 | $10 | $20 |
| Cache write, 5 minutes | $2.50 | $2.50 | $5 |
| Cache write, 1 hour | $4 | $4 | $8 |
| Cache read | $0.20 | $0.20 | $0.20 |
| Batch API, input / output | $1 / $5 | $1 / $5 | $2 / $10 |
Every line matches Sonnet 5. Anthropic’s pricing page also notes that Sonnet 5’s $2 / $10 rate, first sold as introductory pricing, became the standard price, and the planned rise to $3 / $15 on September 1 did not happen. Opus 5.5 costs twice as much on input, output and cache writes, and the same $0.20 on cache reads.
One small cost change sits outside the price table. The minimum prompt that can be cached drops to 512 tokens, from 1,024 on Sonnet 5, so shorter system prompts can now be cached.
The “up to 30% less per task” figure is therefore about token use, not price. Anthropic says the model “typically needs far fewer tokens to do the same work.” Three launch partners gave numbers for their own workloads, each quoted by Anthropic:
- Joe Poirier of Balyasny Asset Management says that on 2,441 finance tasks Sonnet 5.5 used about 121,000 tokens per answer, where Sonnet 5 used 497,000, and scored higher.
- Curtis Allen of Slack says it beat Sonnet 5 on almost all of Slackbot’s offline evaluations with about 14% fewer output tokens.
- Yashodha Bhavnani of Box reports 12% fewer total tokens and 2.4 times the speed of the model Box used before.
The spread between a 12% and a 75% reduction is the point: the saving depends on the workload. A team that runs Sonnet 5 at scale should re-baseline cost per completed task rather than assume Anthropic’s figure. Current prices for every major model are in our frontier model price index.
03 — The numbersThe benchmarks, and the notes that qualify them
Anthropic compared Sonnet 5.5 with Sonnet 5, Opus 5.5 and GPT-6 Sol. The table keeps the three Claude models; the GPT-6 Sol figures follow below it.
| Benchmark | Sonnet 5.5 | Sonnet 5 | Opus 5.5 |
|---|---|---|---|
| Terminal-Bench 4.0 (agentic coding) | 70.6% | 10.3% | 66.4% |
| FrontierCode 1.1, main set (agentic coding) | 46.2% max · 52.1% xhigh | 42.4% | 54.4% |
| CursorBench 4.0 (agentic coding) | 55.5% | 34.1% | 57.8% |
| GDPval-AA v2.1 (knowledge work, Elo) | 1844 | 1449 | 1846 |
| AA-Briefcase v1.1 (long knowledge work, Elo) | 1811 | 1359 | 1822 |
| Humanity's Last Exam, with tools | 64.5% | 54.9% | 67.7% |
| OSWorld 2.1, partial (computer use) | 80.1% | 57.0% | 81.8% |
| Chartography, no tools (chart reading) | 61.6% | 15.6% | 64.4% |
Sonnet 5.5 leads Opus 5.5 only on Terminal-Bench 4.0. On every other row Opus 5.5 is ahead. On the percentage-scored tests the margin runs from 1.7 points on OSWorld to 8.2 points on FrontierCode (2.3 against Sonnet 5.5’s better xhigh score). On the two Elo-scored tests it is 2 and 11 points. Anthropic’s own reading is that Sonnet 5.5 at max effort “even performs comparably” to Opus 5.5 on several tests, and that Opus 5.5 is still clearly stronger on complex, open-ended work.
Four notes on Anthropic’s page change how the table reads:
- FrontierCode is lower at max than at xhigh. The test asks whether a code change could be merged without human edits, so it penalises changes outside the task. At max effort, Sonnet 5.5 more often ran Claude Code’s code-review skill across many subagents, which in two cases led to a timeout or to extra edits. That is why 46.2% sits below 52.1%.
- The knowledge-work scores came from a buggy deployment. Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release deployment that had a bug affecting structured outputs. Anthropic expects the effect to be small and to understate Sonnet 5.5, and says the bug is fixed.
- Two coding rows compare against GPT-5.6 Sol. Terminal-Bench and CursorBench have no public GPT-6 Sol result, so Anthropic’s charts use GPT-5.6 Sol on those tests.
- GPT-6 Sol’s image scores may be stale. OpenAI recently fixed a bug that degraded GPT-6 Sol’s image understanding, and the published third-party scores may predate the fix. Artificial Analysis does not expect a large effect on its two tests.
Where Anthropic does print GPT-6 Sol, it scores 49.3% on FrontierCode, 1487 on GDPval-AA, 1483 on AA-Briefcase and 53.6% on Chartography. Sonnet 5.5 leads on three of them, by more than 300 Elo points on the two Artificial Analysis knowledge-work tests. On FrontierCode its max score of 46.2% trails GPT-6 Sol; its xhigh score of 52.1% leads.
04 — The settingsThe effort level behind the scores
Anthropic also published charts that plot each model’s score against its cost per task at five effort levels, from low to max. The exact points are in the chart data on the announcement page. On the four tests charted, the Sonnet 5.5 figure in the headline table is the max point. The bars below show what the same model scores on Terminal-Bench 4.0 at each setting.
Sonnet 5.5 on Terminal-Bench 4.0, by effort level
Anthropic, Sonnet 5.5 announcement chart data, September 28, 2026. Cost is per attempt, as Anthropic's chart labels it.At the API default of high, Sonnet 5.5 scores 43.0% on this test, not 70.6%, and Opus 5.5 at high scores 64.2%. The gap between the two models is much wider at the settings most integrations actually use. Sonnet 5.5 still beats Sonnet 5’s best result on this test at its cheapest settings. At medium it scores 28.8% for $0.83 an attempt, against Sonnet 5’s 10.3% at max for $11.62. Comparisons like that one are behind Anthropic’s claim that Sonnet 5.5 at low or medium beats Sonnet 5’s best score “for about a tenth of the cost per task.”
The pattern differs by test, and on some of them Sonnet 5.5 at a high setting costs more than Opus 5.5 does for the same score. Our effort-level comparison works through all 40 points. For a launch-day decision, the short version is to benchmark Sonnet 5.5 at the effort level you will actually run, not at the level in the headline table.
Sonnet 5's tendency to reach for web search too often and its high token use are both gone in this new model.David Loker, VP of AI, CodeRabbit, quoted in Anthropic's Sonnet 5.5 announcement, September 28, 2026
05 — The safeguardsSafeguards: the first Sonnet with cyber fallbacks
Anthropic says Sonnet 5.5’s cybersecurity capability is comparable to Opus 5’s, which is why it is the first Sonnet model to launch with the kind of cyber safeguards used on the most capable Claude models. The system card adds that Sonnet 5.5 is far better than Sonnet 5 at building sophisticated exploits, though it remains behind Opus 5.5 and Mythos 5.1 on cyber evaluations.
Cybersecurity
Finding and fixing bugs in your own code stays on Sonnet 5.5. Higher-risk security tasks fall back to Sonnet 5, and Anthropic says the fallback is visible. Cyberdefenders will be able to apply to an expanded Cyber Verification Program for tiered access.
Biology
The same biology safeguards as Sonnet 5. Anthropic warns that some microbiology and virology requests may be flagged in error; organisations can apply to the Life Sciences Verification Program.
Distillation
The first Sonnet with classifiers that block requests to reproduce its reasoning. Its thinking blocks also work only in the account that produced them, or a linked account.
The account binding is the one most teams could notice. If a conversation moves between accounts, including switching accounts mid-session in Claude Code, the API drops Sonnet 5.5’s earlier thinking blocks and the model continues without that reasoning. The request still succeeds.
On risk thresholds, the system card says Sonnet 5.5 does not cross any new ones under Anthropic’s Responsible Scaling Policy. Anthropic treats it as meeting the first-level thresholds for chemical and biological risk (CB-1) and autonomy (Autonomy-1) and applies the matching mitigations. Its automated behavioural audit covered roughly 1,850 scenarios, and Anthropic reports that Sonnet 5.5 matched or improved on Sonnet 5 on most measures while still trailing Opus 5.5 overall.
06 — The migrationFive settings that now return a 400 error
Code that runs on Sonnet 5 will not always run on Sonnet 5.5 after a model-ID swap. Anthropic’s migration guide names five request settings that the new model rejects: thinking budgets, sampling parameters such as temperature, assistant prefill, forced tool choice, and thinking: {"type": "disabled"}. Teams coming from Sonnet 5 will mostly hit the last two, because Sonnet 5 already rejected the first three.
Replace "disabled" with thinking: {"type": "between_tools"}, which keeps up-front thinking off. It works only at low, medium and high effort; at xhigh or max it returns a 400, and with it effort cannot change mid-conversation. Progress updates the model writes between tool calls still come back as thinking blocks.
Two more changes affect agents. Computer use on the Claude API and Google Cloud now needs the computer_toolset_20260801 toolset, and the advisor tool no longer accepts Sonnet 5 or Opus 4.8 as the advisor. Our Opus 5.5 launch post covers the same family-wide changes from the Opus side, and the Sonnet 5 launch post is the baseline most teams are moving from. Every error string and fix, by starting model, is in our Sonnet 5.5 migration guide.
07 — ConclusionSonnet 5.5 is cheaper than Opus 5.5 only at the lower effort levels
Re-run your evaluations at the effort level you will actually use, and compare cost per completed task, not price per token
The price is settled: it is Sonnet 5’s. What is not settled is how many tokens Sonnet 5.5 spends on your work, and launch partners report reductions from 12% to about 75%. Measure that at medium and high, and treat the max-effort headline scores as a ceiling rather than a forecast. Batch work gets the same 50% discount as before; the rates for every provider are in our batch pricing landscape.