AI DevelopmentNew Release14 min readPublished July 25, 2026

Frontier-class results at $5/$25 per Mtok · half of Fable 5, same as Opus 4.8

Claude Opus 5: Frontier Intelligence at Half the Price

Anthropic shipped Claude Opus 5 on July 24, 2026 at $5 per million input tokens and $25 per million output — identical to Opus 4.8, and exactly half of Fable 5. The company calls it state of the art on coding and knowledge work while positioning it as the model you run every day rather than the one you save for hard problems. On Anthropic’s own table it more than doubles Opus 4.8 on agentic terminal coding and posts the lowest misalignment score the company has measured.

DA
Digital Applied Team
Senior strategists · Published July 25, 2026
PublishedJuly 25, 2026
Read time14 min
SourcesAnthropic primary + system card
Price / M tokens
$5/25
unchanged from Opus 4.8
Half of Fable 5
Frontier-Bench v0.1
43.3%
Opus 4.8 scored 21.1%
+22.2 pts
ARC-AGI-3
30.2%
GPT-5.6 Sol scored 7.8%
Anthropic: 3× next-best
Misalignment score
2.3
automated behavioural audit
Lowest Anthropic has measured

Claude Opus 5 launched on July 24, 2026, and the framing Anthropic chose is unusual for a flagship release. The company did not lead with a new capability ceiling. It led with economics: Opus 5, in Anthropic’s words, “comes close to the frontier intelligence of Claude Fable 5 at half the price,” and it is “designed to be used every day.” At $5 per million input tokens and $25 per million output, it costs exactly what Opus 4.8 cost, and exactly half what Fable 5 costs.

That positioning is the story. For most of 2026 the frontier tier has been something teams rationed — a model you route to for the hard 5% of tasks while a cheaper workhorse handles volume. Opus 5 is Anthropic arguing that the rationing layer can collapse, because the model that wins the benchmark is now also the model that is cheap enough to leave switched on. It is the new default on Claude Max and the strongest model available on Claude Pro.

This analysis covers what actually shipped, the full benchmark table against Fable 5, Opus 4.8, and GPT-5.6 Sol, the price arithmetic that makes the “half” claim load-bearing, the four evaluations where Opus 5 does not come first, the behavioural shift Anthropic documents around self-verification, the safeguards and alignment results, and a routing take for teams deciding what to change on Monday. For the wider Claude line, our when-to-use-which guide across Sonnet 5, Opus 4.8, and Fable 5 is the companion piece this release reshuffles.

Key takeaways
  1. 01
    Same price as Opus 4.8, half the price of Fable 5.$5 per million input tokens and $25 per million output, unchanged from its predecessor. Fable 5 sits at $10/$50. Fast mode runs at roughly 2.5× the default speed for twice the base price.
  2. 02
    It more than doubles Opus 4.8 on agentic terminal coding.Frontier-Bench v0.1: 43.3% for Opus 5 against 21.1% for Opus 4.8, 33.7% for Fable 5, and 34.4% for GPT-5.6 Sol. On GDPval-AA v2 knowledge work it scores 1861 against Fable 5's 1747.
  3. 03
    The ARC-AGI-3 gap is the most striking number.Opus 5 scores 30.2% on novel problem-solving where Opus 4.8 managed 1.5% and GPT-5.6 Sol 7.8%. Anthropic describes the result as three times the next-best model's score.
  4. 04
    It is not a clean sweep — four evaluations go elsewhere.GPT-5.6 Sol leads DeepSWE v1.1 (72.7% vs 68.8%), Fable 5 edges Humanity's Last Exam without tools, FrontierCode v1.1, and the held-out Legal Agent Benchmark, and Mythos 5 leads HealthBench Professional.
  5. 05
    Anthropic calls it the most aligned model it has released.Its automated behavioural audit scores Opus 5 at 2.3 on overall misaligned behaviour, the lowest of any recent Anthropic model, with the lowest rates of deceptive behaviour and the least susceptibility to being tricked into misuse.

01What ShippedA flagship priced like a workhorse.

Opus 5 is available today across all platforms, with the API model ID claude-opus-5. Pricing is $5 per million input tokens and $25 per million output — the same rates Opus 4.8 has carried since May 28, 2026. Fast mode returns on the same terms as its predecessor: roughly 2.5× the default speed at twice Opus 5’s base price, available on the Claude Platform and through usage credits in Claude Code.

Two plan-level details matter more than they might appear. Opus 5 is the new default model on Claude Max, and the strongest model on Claude Pro — which means a large population of paying users gets moved onto it without changing a setting. And, consistent with prior Opus models, it carries no data retention requirements for general access. That last point is a real differentiator: Mythos-class models including Fable 5 come with a mandatory 30-day retention policy that zero-data-retention agreements do not override, which has been a live blocker for regulated buyers since June.

Positioning
The everyday frontier model
$5 in · $25 out / 1M tokens

Anthropic's framing is explicit: Opus 5 is 'designed to be used every day' and 'works more efficiently than other models.' The comparison target is Fable 5's intelligence at Opus 4.8's price, not a new ceiling above Fable 5.

claude-opus-5 · live on the Claude API
Access
Default on Max, strongest on Pro
Available today across all platforms

It becomes the default model on Claude Max and the top model on Claude Pro, so a large share of subscriber traffic moves onto it without anyone changing a setting. Fast mode runs at roughly 2.5× the default speed for twice the base price.

Fast mode · 2.5× speed at 2× price
Data handling
No retention requirement
General access · no mandatory retention

Unlike Mythos-class models, Opus 5 has no data retention requirement for general access — the same posture as previous Opus releases, and a meaningful gate-opener for regulated workloads that could not adopt Fable 5.

Fable 5 carries a 30-day retention policy

The effort setting is central to reading every chart Anthropic published. Rather than a single score per model, the launch post presents performance as a curve against effort — the control that lets you trade intelligence against token spend. Opus 5’s claims are frequently of the form “better performance at a given cost,” which is a statement about the whole curve, not a single peak. If your team has not yet operationalised effort as a routing variable, this release makes that overdue; we covered the mechanics when the control landed alongside Opus 4.8 and dynamic workflows.

02The NumbersThe full table, including the rows Opus 5 loses.

Anthropic published a head-to-head table covering Opus 5, Fable 5, Opus 4.8, and GPT-5.6 Sol across twelve evaluation families. Every figure below is vendor-reported. Reproduced in full, including the four rows where another model wins — a completeness most launch coverage skips.

Anthropic-reported benchmark comparison of Claude Opus 5, Claude Fable 5, Claude Opus 4.8, and GPT-5.6 Sol across twelve evaluation families, from the Claude Opus 5 launch announcement of July 24, 2026. Health and biology rows compare against Mythos 5 rather than Fable 5, as published.
EvaluationOpus 5Fable 5Opus 4.8GPT-5.6 Sol
Agentic terminal codingFrontier-Bench v0.143.3%33.7%21.1%34.4%
Knowledge workGDPval-AA v21861174715931736
Novel problem-solvingARC-AGI-330.2%1.5%7.8%
Agentic searchBrowseComp90.8%87.4%84.3%90.4%
Multidisciplinary reasoningHumanity’s Last Exam · no tools / with tools56.3% / 64.7%56.5% / 63.9%49.8% / 57.9%— / —
Computer useOSWorld 2.070.6%66.1%55.7%62.6%
Agentic codingDeepSWE v1.168.8%69.7%59.0%72.7%
Agentic codingFrontierCode v1.1, Main53.4%53.5%46.5%47.5%
Business workflowsAutomationBench26.0%17.4%17.0%18.1%
LegalLegal Agent Benchmark, held-out11.7%13.3%10.4%2.5%
HealthHealthBench Professional59.8%66.0%Mythos 557.4%60.5%
BiologyBioMysteryBench · hard / human-solved49.4% / 90.1%46.5% / 89.0%Mythos 5 on human-solved42.4% / 88.5%— / —

Read down the Opus 4.8 column and the generational story is unambiguous. Frontier-Bench doubles. ARC-AGI-3 goes from 1.5% to 30.2%. OSWorld 2.0 adds nearly fifteen points. GDPval-AA v2 adds 268 points and clears Fable 5 by 114. Anthropic’s claim that Opus 5 “more than doubles Opus 4.8’s performance at a lower cost per task” on Frontier-Bench is borne out by its own table.

Read across the Fable 5 column and the story is subtler. Opus 5 wins most rows, but the margins on the coding benchmarks are inside noise — 53.4% against 53.5% on FrontierCode is a tie in everything but typography. Anthropic is candid about this on CursorBench 3.2, where it says Opus 5 at max effort performs within 0.5% of Fable 5’s peak score at half the cost per task. The pitch was never that Opus 5 beats Fable 5 outright. It is that the gap has narrowed to the point where paying double stops making sense for most work.

Generational gains · Opus 5 vs Opus 4.8

Source: Anthropic, Introducing Claude Opus 5, July 24, 2026 — vendor-reported
Frontier-Bench v0.1 · Opus 4.8Agentic terminal coding
21.1%
Frontier-Bench v0.1 · Opus 5More than double its predecessor
43.3%
ARC-AGI-3 · Opus 4.8Novel problem-solving
1.5%
ARC-AGI-3 · Opus 5Anthropic: three times the next-best model
30.2%
OSWorld 2.0 · Opus 4.8Computer use
55.7%
OSWorld 2.0 · Opus 5Beats Fable 5's 66.1% and GPT-5.6 Sol's 62.6%
70.6%
AutomationBench · Opus 4.8End-to-end business workflows
17.0%
AutomationBench · Opus 5Highest of the four models by 7.9 points
26.0%
Read the footnote before quoting the coding number
Anthropic discloses that its Frontier-Bench v0.1 figures come from an internal run on the mini-SWE-agent harness with a GKE backend, taking mean reward over five attempts per task — and that Opus 4.8 served as the fallback on safety-classifier refusals for both Opus 5 and Fable 5. That fallback detail means the headline coding scores are not pure single-model results. It is a well-disclosed methodology, but it is a vendor harness, and the 43.3% should be cited as such.

03The EconomicsWhy “half the price” is the entire argument.

The arithmetic is simple enough to state in one line: Fable 5 costs $10 per million input tokens and $50 per million output. Opus 5 costs $5 and $25. For a workload that spends $4,000 a month on Fable 5 tokens, the equivalent Opus 5 spend is $2,000 — before any efficiency gain, purely on list rates.

But per-token price is the least interesting half of the story, and Anthropic knows it. The launch post repeatedly frames results as cost per task rather than cost per token, because a model that finishes in fewer turns with fewer tool calls costs less even at identical rates. Customer reports in the announcement put numbers on that: one financial-modelling team reported averaging nine percentage points higher accuracy with a third fewer turns and tool calls and 60% less time; a legal team reported similar performance while generating 26% fewer tokens on average than Opus 4.8 at max reasoning; a trading firm reported roughly a seventh of the reasoning tokens and under half the latency of Opus 4.8.

Those are self-reported figures from named early-access customers, not audited benchmarks, and they should be read as directional. But the direction is consistent, and it compounds with the rate cut. If you are moving from Fable 5, the halved rate and the token reduction multiply. This is exactly the calculation our per-task and per-user cost framework exists to make repeatable — sticker price alone will mislead you on every release this year.

vs Fable 5
Lower list rate, comparable results
50%

$5/$25 against $10/$50. On CursorBench 3.2 at max effort, Anthropic puts Opus 5 within 0.5% of Fable 5's peak score at half the cost per task.

Exactly half
vs Opus 4.8
No price increase for the generation
$0

Identical rates to the model it replaces, with more than double the Frontier-Bench score. The generational upgrade is free at the token level.

Same $5/$25
Fast mode
Speed at twice the base price
2.5×

Available on the Claude Platform and through usage credits in Claude Code — the same structure Opus 4.8 used, applied to Opus 5's base rate.

2× price

04The GapsFour rows where Opus 5 comes second.

A launch table with no losses is a marketing document. This one has four, and they are the rows worth knowing if you are routing production traffic.

DeepSWE v1.1 goes to GPT-5.6 Sol at 72.7% against Opus 5’s 68.8% — a four-point gap on an agentic coding benchmark, and the only row where an OpenAI model leads outright. Notably Fable 5 also edges Opus 5 here at 69.7%, making this the clearest case where the “close to Fable 5” framing costs you something real. FrontierCode v1.1 is a statistical tie at 53.4% against 53.5%. Humanity’s Last Exam without tools goes to Fable 5 by two-tenths of a point, though Opus 5 retakes it with tools enabled at 64.7% against 63.9% — which tells you the gap is in raw recall rather than in tool-augmented reasoning.

The held-out Legal Agent Benchmark is the widest loss: 11.7% against Fable 5’s 13.3%. Absolute scores are low across the board here — GPT-5.6 Sol manages 2.5% — so this is a benchmark where every model is failing most tasks, and a 1.6-point spread on a hard held-out set is thin ground for a routing decision either way. And on HealthBench Professional, the comparison column is Mythos 5 at 66.0% rather than Fable 5, with Opus 5 at 59.8% behind even GPT-5.6 Sol’s 60.5%.

The pattern across all four: specialist and domain-heavy evaluations. Opus 5 dominates general agentic capability — terminal coding, computer use, novel problem-solving, business workflows — and gives ground on narrow professional domains. If your workload is legal-agent or clinical, the frontier tier above still buys you something. If it is engineering or operations, it increasingly does not.

05What ChangedThe real upgrade is self-verification.

The most useful part of the announcement is not a benchmark. It is the section describing what Opus 5 does differently, which Anthropic summarises as being “much stronger at verifying its work and iterating carefully until it succeeds.” Three documented examples make the behaviour concrete.

On one Frontier-Bench task, the model was given a drawing of a machine part and asked to write code rebuilding it as a 3D FreeCAD model — but was deliberately given no way to view the drawing. Rather than failing or guessing, Opus 5 wrote its own computer vision pipeline to extract the geometry from raw pixels, then reconstructed the part. Anthropic reports it did so repeatedly, and that no competing model with the same setup solved it in five attempts. In a second case, given a real bug in a popular open-source package manager, it found the root cause and fixed an edge case the community’s own patch had missed, while a competing model fixed the surface symptom and declared the bug resolved. In a third, an engineer at a trading firm used it to build a market data feed for a new exchange in one session; finding no live feed to validate against, the model built its own test harness to confirm it was parsing the exchange’s data correctly.

The common thread is a model that constructs its own verification apparatus when none is provided. That is the difference between an assistant you supervise turn by turn and one you can hand a multi-hour task. It is also the capability that makes longer autonomous runs economically sensible, because the failure mode of long agentic sessions has always been confident wrongness compounding unchecked.

“Claude Opus 5 is much stronger at verifying its work and iterating carefully until it succeeds.”— Anthropic, Introducing Claude Opus 5, July 24, 2026

Two customer observations in the announcement sharpen the point further. One reviewer described handing off a pull request and finding the model verified the branches, checked the template, and reasoned about test implications rather than rushing to publish — where older models “jumped ahead and got caught on our checks.” Another described the model pushing back on a proposed design and not folding when pressed: it narrowed its objection to a single design question and proposed a compromise. For teams running agents with reduced oversight, that kind of calibrated disagreement is worth more than a benchmark point.

Anthropic also flags stronger visual outputs — sophisticated visualisations and interactive illustrations — and a broad improvement in scientific research, with every life-sciences evaluation improving over Opus 4.8. The standouts are organic chemistry, where inferring molecular structures from spectroscopy data scores 10.2 percentage points higher than Opus 4.8 on Anthropic’s internal benchmark, and protein sequence-variation prediction at 7.7 points higher.

06Alignment & SafeguardsThe most aligned model Anthropic has measured.

On Anthropic’s automated behavioural audit, Opus 5 scores 2.3 on overall misaligned behaviour — the lowest of its recent models. The company states it adheres to Claude’s Constitution better than Opus 4.8, Sonnet 5, or Fable 5, exhibits the lowest rates of deceptive behaviour, is the least susceptible to being tricked into misuse, and is its safest model yet at avoiding reckless actions with hard-to-reverse side effects. That last property is the one that matters operationally: it is the difference between an agent you let touch production and one you do not.

On dangerous capabilities, Anthropic’s position is that Opus 5 does not advance the frontier. Evaluated alongside private-sector and government partners, it remains behind Mythos 5 in both biology research and offensive cybersecurity. The nuance is worth stating precisely, because it is where the interesting engineering sits: Anthropic deliberately avoided training Opus 5 on cyber tasks, yet it improved substantially anyway as a by-product of general capability, and now comes close to Mythos 5 at finding vulnerabilities. It remains substantially behind on exploiting them. On the OSS-Fuzz evaluation, the two models identify vulnerabilities with similar success while Opus 5’s exploit-development score sits far behind.

What the cyber safeguards actually block
Opus 5’s cyber classifiers are proportionally less restrictive than Fable 5’s — Anthropic expects them to intervene around 85% less often. They permit finding vulnerabilities in source code, but block binary-based vulnerability scanning, penetration testing, and exploit generation. In Claude.ai, Claude Code, and Claude Cowork, flagged requests fall back to Opus 4.8 by default, and that fallback can also be enabled on the API. Enterprises and researchers already in Anthropic’s Cyber Verification Program get immediate access to a version with fewer restrictions.

On biology the change is a loosening. Because Opus 5 carries a similar safeguard suite to Opus 4.8 rather than Fable 5’s stricter one, it is now Anthropic’s most capable generally available model for scientific research — and biology-related requests blocked on Fable 5 now route to Opus 5 rather than Opus 4.8. Anthropic is explicit that limitations remain on long-running autonomous research tasks, where it considers the substantial biology risk to sit, and that Mythos 5 remains stronger for that work. If you have been living with classifier friction on Fable 5, this release is a material quality-of-life change; we wrote up the production patterns for handling it in our guide to fallback models and production resilience.

07PlatformTwo betas that quietly change agent architecture.

Alongside the model, Anthropic shipped two beta platform features that will matter more to builders than to end users.

Mid-conversation tool changes. Developers can now change which tools Claude can use within a conversation without invalidating the prompt cache. That sounds like plumbing; it is not. Until now, an agent that needed a different toolset for a later phase of a task either carried every tool from the start — paying context and confusion costs throughout — or paid a full cache invalidation to swap. Phase-based agents with narrow, stage-specific toolsets now become cheap to build.

Automatic fallbacks on the API. Requests flagged by safety classifiers on Opus 5, or on Fable 5, can now be routed automatically to another model rather than blocked. With fallbacks on, API requests always route to the best available model by default instead of failing. For anyone who has built retry logic around classifier refusals — a recurring tax on Fable 5 since launch — this removes a category of production error handling.

Anthropic also published a dedicated prompting guide for Opus 5. Given how much of this release’s value is expressed through the effort setting rather than raw capability, that guide is worth reading before you port prompts across from Opus 4.8 wholesale.

08ImplicationsWhat to actually change on Monday.

The routing consequences differ sharply depending on which model you are on today.

Already on Opus 4.8
General agentic and coding workloads

The clearest call in the release. Identical pricing, more than double the Frontier-Bench score, +15 points on OSWorld 2.0, and a large jump on novel problem-solving. Validate on your own evals, then move the default — there is no cost argument for staying.

Migrate to Opus 5
Paying for Fable 5
Engineering and operations work

Test the swap seriously. Opus 5 wins most rows and ties the coding benchmarks at half the rate, plus it drops the Mythos-class 30-day retention requirement. Keep Fable 5 for the narrow domains where it still leads — and note DeepSWE is one of them.

Test the downgrade
Legal, clinical, specialist domains
Narrow professional benchmarks

This is where the four losing rows live. Fable 5 leads the held-out Legal Agent Benchmark, Mythos 5 leads HealthBench Professional, and absolute scores on both are low enough that human review is non-negotiable regardless of which you pick.

Stay and verify
Long autonomous runs
Multi-hour agentic sessions

The self-verification behaviour is the reason to revisit tasks you previously judged too long to delegate. Build the harness that lets the model check its own work, and re-test the ceiling on session length before assuming last quarter's limits still hold.

Re-test your ceiling

The broader signal is one we have been tracking across every frontier release this year: the competitive axis has moved from headline intelligence to cost per completed task. Anthropic did not claim Opus 5 is smarter than Fable 5. It claimed Opus 5 gets close enough for half the money, and built its entire chart set around performance-per-dollar curves rather than peak scores. When the vendor with the strongest model on the market argues its cheaper model is the right default, the era of rationing the frontier tier is ending. If your team wants help benchmarking this on your real workloads and rebuilding routing around per-task economics, that evaluation is where our AI transformation engagements begin.

09ConclusionThe frontier tier stops being a luxury.

Claude Opus 5, July 2026

Anthropic priced the frontier for everyday use — and mostly earned the claim.

Opus 5 is the rare flagship whose headline is a price rather than a capability. At $5/$25 it costs what Opus 4.8 cost while more than doubling its agentic coding score, clearing Fable 5 on knowledge work, computer use, agentic search, and business workflows, and posting a 30.2% ARC-AGI-3 result against its predecessor’s 1.5%. For engineering and operations workloads, the case for paying double has genuinely weakened.

The asterisks are real and worth carrying. Every figure is vendor-reported, the flagship coding number comes from an internal harness with Opus 4.8 as a refusal fallback, and the customer efficiency claims are self-reported. Four evaluations still go elsewhere, clustered in exactly the specialist domains where a wrong answer costs most. And the safeguard changes cut both ways: looser cyber classifiers and biology routing are a usability win, but they are a policy change to read carefully rather than a pure upgrade.

The practical move is the same one every release this year has called for, and it has not stopped being right: run your own evals on your own prompts, price the work per completed task rather than per million tokens, and move defaults only where your numbers agree with the vendor’s. On this release, for most engineering teams, they probably will.

Route on economics, not headlines

The best model is the one your workload can afford to leave switched on.

Our team helps businesses benchmark frontier models on real workloads, rebuild routing around per-task economics, and ship agentic systems on the model tier the numbers actually support.

Free consultationExpert guidanceTailored solutions
What we work on

Model-economics engagements

  • Per-task cost benchmarking — Claude / GPT / Gemini
  • Opus 4.8 and Fable 5 migration audits
  • Effort-level tuning and routing policy
  • Agentic harness design for long autonomous runs
  • Token-spend observability and cost governance
FAQ · Claude Opus 5

The questions we get every week.

Claude Opus 5 is Anthropic's newest Opus-tier model, released July 24, 2026 and available across all platforms from launch under the API model ID claude-opus-5. Anthropic describes it as a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price, and as state of the art on coding and knowledge work evaluations. It is now the default model on Claude Max and the strongest model available on Claude Pro. Unlike Fable 5, it is positioned for everyday use rather than as a model reserved for the hardest tasks.
Related dispatches

Continue exploring frontier releases.