AI DevelopmentDecision Matrix7 min readPublished September 22, 2026

$4 / $20 vs $10 / $50 · 7 shared benchmarks · the vendors disagree by up to 5.4 points on the same model

Claude Opus 5.5 vs GPT-6 Astra: Benchmarks, Price and Fit

Claude Opus 5.5 lists at 40% of GPT-6 Astra's price with no long-context surcharge. Where each leads, and where the vendors' benchmarks disagree.

DA
Digital Applied Team
Research and practical guidance
Opus 5.5 releasedSeptember 22, 2026
Astra releasedSeptember 3, 2026

Claude Opus 5.5, released on September 22, 2026, lists at $4 per million input tokens and $20 per million output tokens. GPT-6 Astra, released on September 3, lists at $10 and $50, and bills the whole request at $20 and $75 once input passes 272,000 tokens. On price, Opus 5.5 is the cheaper model by a wide margin at every request size.

The benchmarks are closer, and harder to read. Only seven tests appear in both vendors’ published tables. Opus 5.5 leads on four, GPT-6 Astra leads on two, and one isn’t comparable because the two vendors used different task sets. On the tests where both vendors ran the same Claude model, their scores differed by up to 5.4 points, which is more than the gap between the two models on some rows.

Every figure below is a vendor’s own published number, from Anthropic’s announcement and OpenAI’s launch post, and none has been independently replicated. OpenAI’s launch post predates Opus 5.5, so it has no Opus 5.5 figures.

Key takeaways
  1. 01
    Opus 5.5 costs 40% of Astra’s list price on input and output, and 20% on cached input.Above 272,000 input tokens Astra’s whole request moves to $20 / $75. Anthropic bills Opus 5.5 at the same rate across its full 1M-token window.
  2. 02
    Opus 5.5 leads on terminal coding, knowledge work and Humanity’s Last Exam.Terminal-Bench 4.0 66.4% vs 57.9%, GDPval-AA 1846 vs 1542 Elo, HLE with tools 67.7% vs 57.2%. FrontierCode is within noise.
  3. 03
    Astra leads on business automation and scientific terminal work.AutomationBench 41.4% vs 40.0% and Terminal-Bench-Science 64.6% vs 58.7%. Astra also publishes long-context and ARC-AGI-3 results that Opus 5.5 has no figure for.
  4. 04
    Vendor harnesses disagree enough to flip close rows.Anthropic scored Opus 5 at 48.0% on FrontierCode Main; OpenAI scored the same model at 53.4%. Treat any gap under about five points as unresolved until you test.

01The invoiceThe price gap, including the 272K line

According to OpenAI’s model page, prompts with more than 272,000 input tokens are priced at 2× the input and cache rates and 1.5× the output rate “for the full request,” not just the tokens above the line. Anthropic’s pricing page says Claude models from 4.6 onward include the full 1M-token window at standard pricing, so a 900,000-token request costs the same per token as a 9,000-token one.

Per million tokens. Sources: Anthropic pricing and fast-mode docs; OpenAI GPT-6 Astra model page and launch post. Checked September 22, 2026.
LineOpus 5.5Astra ≤272KAstra >272K
Input$4$10$20
Cached input (cache read)$0.20$1$2
Cache write$5 (5 min) / $8 (1 hr)$12.50$25
Output$20$50$75
Batch50% of standard50% (Batch and Flex)50% (Batch and Flex)
Fast mode2× price, up to 2.5× speed2× price, up to 2× speed2× the applicable rate

The table below prices one illustrative agent task at the same token counts on both models. The counts are invented for the example; the rates are the published ones. In practice the counts won’t be identical, because the two models use different tokenizers and each vendor says its model uses fewer tokens per task. Treat this as the list-price gap, not a prediction of your bill.

Illustrative token counts at published rates, September 22, 2026. The last column assumes every request in the task is over 272K input tokens.
Illustrative taskOpus 5.5Astra ≤272KAstra >272K
Cache reads · 8,000,000 tokens$1.60$8.00$16.00
Uncached input · 400,000 tokens$1.60$4.00$8.00
Cache writes · 600,000 tokens$3.00$7.50$15.00
Output incl. reasoning · 300,000 tokens$6.00$15.00$22.50
Total$12.20$34.50$61.50

At the same token counts, Astra costs about 2.8× as much as Opus 5.5 on this task below the line, and about 5× as much when every request crosses it. The 272,000-token line matters most for coding agents, because an agent that loads a large repository into context can cross it early in a session and stay above it. Our long-context pricing reference lists where each vendor draws its line.

Each vendor’s cost claim has a different target

Anthropic says that at default effort Opus 5.5 beats Astra on FrontierCode at “roughly 20% of the cost per task” and matches it on Terminal-Bench 4.0 for “about 40% of the cost.” OpenAI’s cost comparisons were published before Opus 5.5 existed and are made against Claude Fable 5.1, which lists at the same $10 / $50 as Astra, and GPT-5.6 Sol. Neither vendor’s claim has been checked independently, and the two aren’t measured against the same model.

02The numbersThe seven benchmarks both vendors published

Both vendors report each model at its best setting. Anthropic’s footnote says Opus 5.5 was run at max effort except on Terminal-Bench 4.0 (xhigh), and that the Astra figures on Terminal-Bench come from OpenAI. OpenAI’s table says its scores are “the maximum at any effort.” So this is a comparison of each model’s best case, not of the settings you are likely to run.

Sources: Anthropic, Introducing Claude Opus 5.5 (September 22, 2026); OpenAI, GPT-6 Astra (September 3, 2026).
BenchmarkOpus 5.5AstraHow to read it
Terminal-Bench 4.066.4%57.9%Opus 5.5 at xhigh, Astra at high; each model’s best score. Astra figure as reported by OpenAI
GDPval-AA v2.1 (Elo)18461542Both figures from Anthropic’s table; OpenAI does not report GDPval-AA
Humanity’s Last Exam, with tools67.7%57.2%Astra figure matches OpenAI’s own table
FrontierCode v1.1 Main54.4%53.3%Within noise. OpenAI ran Astra with a Codex-style developer message
AutomationBench40.0%41.4%Run by Zapier; Opus 5.5 was run without fallback models, so safeguard interventions counted as failures
Terminal-Bench-Science 0.158.7%64.6%Astra figure as reported by OpenAI; standard error ±3.5–5 points per model
OSWorld 2.0, partial score81.8%72.6%Not comparable: OpenAI’s figure is on the offline task set

Three of Opus 5.5’s leads are large enough to survive the caveats below: Terminal-Bench 4.0 (8.5 points, against a stated standard error of ±2.6), GDPval-AA (304 Elo) and Humanity’s Last Exam with tools (10.5 points). Astra’s lead on Terminal-Bench-Science is 5.9 points, a little above the ±3.5–5 point standard error Anthropic gives for that test. AutomationBench (1.4 points) and FrontierCode (1.1 points) are too close to call from vendor numbers.

OSWorld is the row to be careful with. OpenAI reports Astra at 72.6% on the offline subset of OSWorld 2.0, with Claude models run on the official settings. Anthropic reports Opus 5.5 at 81.8% and leaves Astra’s cell blank. OpenAI’s own footnote says the Fable 5.1 system card used modified tasks and grading. Until someone runs both models on the same set, there is no computer-use comparison to draw.

At these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences.Anthropic, Introducing Claude Opus 5.5, September 22, 2026

03The harnessesWhen both vendors score the same model

OpenAI’s appendix scores several Claude models, and Anthropic’s table scores the same models, which lets us check how far the two harnesses agree. On most of these benchmarks they agree to within about a point. On FrontierCode Main they are 5.4 points apart on Opus 5, and on OSWorld the task sets differ.

Same Claude model, each vendor’s published score. Gap in percentage points, our arithmetic. Sources as above.
Model · benchmarkAnthropicOpenAIGap
Opus 5 · Terminal-Bench 4.052.3%52.6%0.3
Opus 5 · FrontierCode v1.1 Main48.0%53.4%5.4
Fable 5.1 · FrontierCode v1.1 Main50.3%50.9%0.6
Opus 5 · Terminal-Bench-Science 0.129.0%30.0%1.0
Fable 5.1 · Humanity’s Last Exam, with tools65.6%65.0%0.6
Opus 5 · OSWorld 2.0, partial74.0%70.2% (offline set)n/a

The FrontierCode gap is the one that matters here. Opus 5.5 leads Astra on that benchmark by 1.1 points, and the two vendors’ scores for Opus 5 differ by nearly five times that. OpenAI’s footnote explains one difference: Astra was run with a developer message modelled on Codex that asks for clean, mergeable code. A developer message changes the result, so the benchmark alone can’t tell you which model writes better code in your repository.

The practical rule is simple. Where the gap between two models is larger than the gap between two vendors’ runs of the same model, it probably reflects a real difference. Where it is smaller, it could reflect the harness as much as the model, and only your own tasks can settle it. Our guide to testing a model on your own traffic covers how to set that up.

04The gapsBenchmarks only one side ran

Most of what each vendor published has no counterpart from the other. These are the results a buyer is most likely to ask about. Where OpenAI scored Opus 5, its figure is shown as the nearest reference. Opus 5.5 may score differently.

Sources: OpenAI GPT-6 Astra launch appendix; Anthropic Claude Opus 5.5 announcement.
Benchmark (model, reported by)ScoreWhat exists for the other side
ARC-AGI-3 (Astra, OpenAI)99.9%Run with a Responses API harness that changes two settings. OpenAI lists Opus 5 at 30.2%
MRCR v2 8-needle, 512K–1M (Astra, OpenAI)96.3%OpenAI’s long-context retrieval test; no Claude model listed
DeepSWE v1.1 (Astra, OpenAI)74.1%OpenAI lists Opus 5 at 73.7%
BrowseComp (Astra, OpenAI)91.5%OpenAI lists Opus 5 at 90.8%
CursorBench 4.0 (Opus 5.5, Anthropic)57.8%No Astra figure in Anthropic’s table
Chartography, with tools (Opus 5.5, Anthropic)89.0%No Astra figure in Anthropic’s table

Two of these matter for model choice. The long-context retrieval score (96.3% at 512K–1M tokens) is evidence that Astra reads a very long prompt accurately, though it also means using the part of Astra’s context window that costs the most. The DeepSWE and BrowseComp rows show Astra only narrowly ahead of Opus 5 in OpenAI’s own runs. That makes those benchmarks worth running on Opus 5.5 before assuming Astra keeps the lead. Our Astra launch guide covers the rest of OpenAI’s appendix.

05The buildIntegration differences and what happens when safeguards trigger

Both models use the same five effort names, from low to max, and both can change effort mid-conversation without breaking the prompt cache. The practical differences are in integration and safeguards. Astra’s details come from OpenAI’s model guide; Opus 5.5’s come from Anthropic’s documentation.

Sources: Anthropic Opus 5.5 docs and announcement; OpenAI GPT-6 Astra model page, model guide and launch post. September 22, 2026.
AreaOpus 5.5GPT-6 Astra
Effort levelslow to max; default medium; thinking always onlow to max; the none level is not supported
Changing effort mid-conversationPer-message effort (beta) keeps the cacheconfiguration_update items keep the cache
Integration constraintsNo forced tool_choice (any or tool); thinking blocks bound to the conversationTool calling requires the Responses API; temperature, top_p and top_logprobs removed
Context · output · knowledge cutoff1M · 128K · June 20261.05M (922K max input) · 128K · April 30, 2026
Where you can buy itClaude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud, Microsoft FoundryOpenAI API, Microsoft Azure, Amazon Bedrock
When a safeguard triggersThe request falls back to another model; most cybersecurity tasks go to Opus 4.8In the API “the task will stop”; advanced cybersecurity tasks are refused
Fast mode limitsClaude API onlyNot with EU data residency; no latency SLA

The safeguard row matters most for unattended agents. When an Opus 5.5 safeguard triggers, the request is answered by another Claude model and the run continues. When Astra’s misalignment monitoring intervenes in the API, OpenAI says the task stops. Neither approach is simply better. A fallback keeps work moving but changes which model did it. A stop is easier to audit but interrupts the run. Plan your agent loop for whichever behaviour you choose.

On safety evidence, the two vendors again used different tests. OpenAI rates Astra “Critical” for cybersecurity under its Preparedness Framework. On its computer-use stress test, Astra produced misaligned outcomes 2.4% of the time against 11.5% for Opus 5; OpenAI hasn’t tested Opus 5.5. Anthropic reports that Opus 5.5 tried to get around containment boundaries about 85% less often than Opus 5. No test covers both models. Both vendors offer zero data retention to eligible API customers.

06ConclusionOpus 5.5 wins on price everywhere and on most shared benchmarks; Astra keeps two

Your requests regularly exceed 272,000 input tokens
Start with Opus 5.5. Astra bills those whole requests at $20 / $75, while Opus 5.5 stays at $4 / $20 across its 1M window.
Opus 5.5
Terminal coding agents and long knowledge-work tasks
Opus 5.5 leads on Terminal-Bench 4.0 and GDPval-AA by margins larger than the harness noise, at 40% of Astra’s list price. Confirm on your own repository.
Opus 5.5 first
Scientific terminal workflows and multi-app business automation
Astra leads on Terminal-Bench-Science and narrowly on AutomationBench. Run both on your tasks; weigh Astra’s higher price against the score gap.
Test Astra first
You are already on OpenAI’s Responses API or Azure OpenAI
Price the switch, not just the model. Opus 5.5 needs a different SDK, different tool-choice handling and a thinking-block policy. Google Cloud buyers can only choose Opus 5.5.
Your platform
What to do this week

Run both models at two effort levels on twenty of your own tasks, and compare cost per completed task

The price gap is certain and large: Opus 5.5 costs 40% of Astra’s list price, and much less once Astra’s requests pass 272,000 tokens. The benchmark picture is mixed. Opus 5.5 leads clearly on terminal coding and knowledge work, Astra leads on scientific terminal work, and the rest is too close to call from vendor numbers. A short test on your own work settles the close rows. For each model’s full launch detail, see our Opus 5.5 launch post and the frontier price index.

Digital Applied

Pick the model your own tasks favour.

We run side-by-side evaluations on your real work, price the result per completed task, and set up routing with a fallback, so a model switch is a measured decision.

Side-by-side evalsCost per taskModel routing
Your next project

A model choice backed by your own data

  • Twenty tasks that reflect your work
  • Two effort levels per model
  • A routing plan with a fallback
Questions and answers

The questions we get about Opus 5.5 and GPT-6 Astra

Yes, at list price. Opus 5.5 is $4 input and $20 output per million tokens against Astra's $10 and $50, and $0.20 against $1 for cached input. Above 272,000 input tokens Astra bills the whole request at $20 and $75, while Opus 5.5 keeps the same rates up to 1M tokens. Actual bills also depend on how many tokens each model uses for your tasks.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading