AI DevelopmentComparison6 min readPublished September 22, 2026

$2 / $10 against $4 / $20 · one independent harness · Sol is cheaper, until you need a higher score

GPT-6 Sol vs Claude Opus 5.5: Cost per Task and Benchmarks

Artificial Analysis ran GPT-6 Sol and Claude Opus 5.5 on the same tests. Sol is cheaper up to a point; Opus 5.5 at its default outscores Sol at max.

DA
Digital Applied Team
Research and practical guidance
Both releasedSeptember 22, 2026
Data readSeptember 22, 2026

OpenAI released GPT-6 Sol and Anthropic released Claude Opus 5.5 on the same day, September 22, 2026. Sol lists at $2 per million input tokens and $10 per million output tokens, half of Opus 5.5’s $4 and $20. Neither vendor compared the two: OpenAI’s charts show Opus 5, and Anthropic’s table shows GPT-5.6 Sol.

The independent evaluator Artificial Analysis has now run both models on the same ten tests at every effort level. On its index, Sol is the cheaper way to reach any score up to about 44. Above that, Sol runs out of headroom. Its best result, 47.5 at max for $1.06 a task, is below Opus 5.5 at its default medium setting, 51.2 for $1.34.

Every Artificial Analysis figure below was read from its GPT-6 Sol and Claude Opus 5.5 model pages on September 22, Intelligence Index version 4.3.2. It measures cost per task at list prices, including cache reads and writes.

Key takeaways
  1. 01
    Up to an index score of about 44, GPT-6 Sol reaches it for less.Sol at high scores 42.8 for $0.37 a task. Opus 5.5 at low scores 42.3 for $0.55, about 50% more.
  2. 02
    Opus 5.5 at its default outscores Sol at its maximum.Opus 5.5 at medium: 51.2 for $1.34. Sol at max: 47.5 for $1.06. That is 3.7 more points for 26% more money.
  3. 03
    The biggest gaps are coding and knowledge work; business workflows are level.On Terminal-Bench 4.0, Opus 5.5 at medium scores 52.5% against Sol's 43.9% at max. On AutomationBench-AA, Sol at xhigh matches Opus 5.5 at medium for 40% of the cost.
  4. 04
    Half the list price does not mean half the bill.Cache reads cost $0.20 on both. On a cache-heavy agent task Sol costs about 57% of Opus 5.5, and requests over 272K input tokens erase the gap.

01The ladderThe same tests, at every effort level

The Intelligence Index averages ten evaluations: agentic coding, business workflows, knowledge-work documents, science, long-context reasoning and factual knowledge. Both models default to medium, and both offer low through max. Each cell shows the index score and the average cost per task.

Index score · cost per task at list prices. Source: Artificial Analysis Intelligence Index v4.3.2, read September 22, 2026.
EffortClaude Opus 5.5GPT-6 Sol
noneNot offered28.1 · $0.33
low42.3 · $0.5533.9 · $0.13
medium (default)51.2 · $1.3439.8 · $0.25
high53.6 · $1.8242.8 · $0.37
xhigh56.0 · $3.4644.1 · $0.53
max57.6 · $5.9847.5 · $1.06

Intelligence Index score, sorted by cost per task

Artificial Analysis Intelligence Index v4.3.2, read September 22, 2026
Sol · low$0.13 a task
33.9
Sol · medium$0.25 a task
39.8
Sol · high$0.37 a task
42.8
Sol · xhigh$0.53 a task
44.1
Opus 5.5 · low$0.55 a task
42.3
Sol · max$1.06 a task
47.5
Opus 5.5 · medium$1.34 a task
51.2
Opus 5.5 · high$1.82 a task
53.6
Opus 5.5 · xhigh$3.46 a task
56.0
Opus 5.5 · max$5.98 a task
57.6

Sorted by cost, the two models barely overlap. Every Sol setting below max costs less than the cheapest Opus 5.5 setting, and Sol at xhigh (44.1 for $0.53) already beats Opus 5.5 at low (42.3 for $0.55). Everything Opus 5.5 does from medium upwards scores higher than anything Sol can reach. The choice is less “which model is better” than “which price band does this task belong in.”

One result goes against intuition. Sol with reasoning switched off scored 28.1 for $0.33 a task, below Sol at low (33.9) and more than twice as expensive. Almost all of the extra cost is input. One likely reason is that a model without reasoning takes more steps to finish, and each step re-sends the conversation. Turning reasoning off is not a reliable way to save money on agent work.

02The detailTest by test: where the gap opens and where it closes

The fairest pairing is each model at the setting a team is most likely to use for demanding work: Opus 5.5 at its default medium, against Sol at max, which costs about the same. Sol at xhigh is added as the cost-conscious option.

Source: Artificial Analysis model pages, Intelligence Index v4.3.2, read September 22, 2026. Costs are averages across the whole index.
EvaluationOpus 5.5 mediumSol maxSol xhigh
Cost per index task$1.34$1.06$0.53
Output tokens per task25.7k31.2k16.0k
Intelligence Index v4.3.251.247.544.1
Terminal-Bench 4.0 (agentic coding)52.5%43.9%30.3%
AutomationBench-AA (business workflows, partial credit)61.2%61.6%61.7%
GDPval-AA (knowledge work, Elo)157614871437
AA-Briefcase (work documents, Elo)164214831364
Humanity’s Last Exam54.7%47.9%46.3%
SciCode (scientific coding)59.3%57.6%55.1%
CritPt (physics research)27.7%30.9%28.0%
AA-LCR (long-context reasoning)84.3%83.7%81.3%
AA-Omniscience accuracy64.5%54.5%53.8%
AA-Omniscience hallucination rate (lower is better)68.4%60.1%58.9%

Coding is the widest gap. On Terminal-Bench 4.0, Opus 5.5 at medium scores 52.5% against Sol’s 43.9% at max, and Opus 5.5 reaches 59.6% at xhigh. Sol falls away fast at lower settings: 30.3% at xhigh and 26.3% at high. For coding agents that run in a terminal, Sol’s lower price buys noticeably weaker results.

Knowledge work also favours Opus 5.5. On GDPval-AA and AA-Briefcase, which rate work documents in head-to-head comparisons scored as Elo, Opus 5.5 at medium leads Sol at max by about 90 and 160 Elo points. At max, Opus 5.5 reaches 1846 on GDPval-AA.

Business workflows are level. On AutomationBench-AA, Artificial Analysis’s partial-credit run of business tasks across apps, Sol at xhigh (61.7%) matches Opus 5.5 at medium (61.2%) at 40% of the average cost per task. Long-context reasoning is also level, and at these settings Sol leads on CritPt, a physics research test.

Knowledge versus guessing

AA-Omniscience rewards right answers, penalises wrong ones and does not penalise declining. Opus 5.5 answers more questions correctly (64.5% against 54.5%). When it doesn’t know, it gives a wrong answer 68.4% of the time, against 60.1% for Sol at max and 50.7% for Sol at low. Sol knows less but guesses less. Neither rate is low enough to skip verification on factual work.

03The launchesThe vendors’ own numbers, and why they don’t line up

Each vendor published results for the benchmarks both models share. They point the same way as the independent data, but they are not a clean comparison.

  • AutomationBench (Zapier). Anthropic reports Opus 5.5 at 40.0% at max, from Zapier’s early-access run. OpenAI’s chart puts Sol’s best at 33.2%, at xhigh. These are Zapier’s own scores, which run far lower than Artificial Analysis’s partial-credit version in the table above.
  • FrontierCode 1.1 Main (Cognition). Anthropic reports Opus 5.5 at 54.4%; OpenAI’s chart puts Sol at 49.3%, both at max. But the two vendors disagree about the same model: Anthropic lists Opus 5 at 48.0%, OpenAI’s chart at 53.4%. A five-point gap between vendors’ runs is larger than the gap between these two models, so treat this row as inconclusive.
  • OSWorld 2.0. Anthropic reports 81.8% on the partial-credit score; OpenAI reports 64.4% for Sol on the offline set. These are different task sets and cannot be compared.

Each launch also claims cost wins, but against other models: OpenAI against Opus 5, Anthropic against GPT-6 Astra and GPT-5.6 Sol. Our GPT-6 Sol and Luna launch analysis and Opus 5.5 launch analysis read those claims against each vendor’s own chart data.

At these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences.Anthropic, Introducing Claude Opus 5.5, September 22, 2026

04The invoiceThe price sheets: half the list price is not half the bill

Per million tokens. Sources: Anthropic pricing and OpenAI pricing, September 22, 2026. Ratios are our arithmetic.
LineOpus 5.5GPT-6 SolSol as % of Opus
Input$4$250%
Cache read$0.20$0.20100%
Cache write (Anthropic: 5-minute)$5$2.5050%
Output$20$1050%
Request over 272K input, input / output$4 / $20$4 / $15100% / 75%
Batch, input / output$2 / $10$1 / $550%
Fast mode, input / output$8 / $40$4 / $2050%

Two lines break the “half price” rule. Cache reads cost $0.20 per million tokens on both models, and for agents, cache reads are usually the largest share of the bill. And Anthropic charges the full 1M-token window at standard rates, while OpenAI bills any request over 272K input tokens at twice the input and cache rates and 1.5 times the output rate.

The table below reuses the illustrative agent task from our Opus 5.5 launch post. The token counts are invented for the example and kept the same for both models; the rates are the published ones.

Illustrative token counts at list prices, September 22, 2026. The last column assumes every request carries more than 272K input tokens.
Illustrative taskOpus 5.5SolSol, long prompts
Cache reads · 8,000,000$1.60$1.60$3.20
Uncached input · 400,000$1.60$0.80$1.60
Cache writes · 600,000$3.00$1.50$3.00
Output incl. reasoning · 300,000$6.00$3.00$4.50
Total$12.20$6.90$12.30

At equal token counts, Sol costs 57% of Opus 5.5 on this cache-heavy task, not 50%. If every request goes past 272K input tokens, Sol costs slightly more than Opus 5.5. Real token counts differ between the models, and Artificial Analysis’s measured figures already include that difference. At their defaults, Sol costs $0.25 a task against $1.34 for Opus 5.5, about a fifth, but for an index score of 39.8 against 51.2.

05The plumbingIntegration differences that decide a migration

Sources: Anthropic’s Opus 5.5 model and migration pages; OpenAI’s GPT-6 Sol model page and announcement, September 22, 2026.
ItemClaude Opus 5.5GPT-6 Sol
Reasoning offNot possible; low is the minimumeffort: none
Changing effort mid-conversationInvalidates the prompt cache; a per-message effort beta avoids itKeeps the prompt cache, per OpenAI
Tool callingtool_choice any or tool returns a 400Responses API; Chat Completions only with effort none
Context / max output1M at standard rates / 128K1.05M, surcharge above 272K / 128K
Knowledge cutoffJune 2026April 20, 2026
Where to run itClaude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud, Microsoft FoundryOpenAI API
SubscriptionsClaude apps and Claude CodeChatGPT Work and Codex on Plus, Pro, Business, Enterprise and Edu

The cache behaviour matters most for routers. A system that raises effort on hard turns and lowers it on easy ones keeps its cache on Sol. On Opus 5.5 it needs the per-message effort beta, or it pays to rebuild the cache after every change. Opus 5.5 also routes most cybersecurity tasks to Opus 4.8 behind the scenes, so security teams should test that workload directly before choosing.

06ConclusionSol wins the low price band; Opus 5.5 owns everything above it

High-volume agents with a tight per-task budget
GPT-6 Sol at medium to xhigh. Below about $0.55 a task, it scores higher than any Opus 5.5 setting on Artificial Analysis's index.
GPT-6 Sol
Coding agents and terminal work
Claude Opus 5.5 at medium, stepping up to xhigh for failures. On Terminal-Bench 4.0 it scores 8.6 points above Sol at max, for about a quarter more per index task.
Claude Opus 5.5
Reports, analysis and knowledge work
Claude Opus 5.5, with higher effort where quality matters. From its default medium setting upwards, it scores above Sol's best on GDPval-AA and AA-Briefcase.
Claude Opus 5.5
Business workflows across apps
Test Sol at xhigh first. It matches Opus 5.5 at medium on AutomationBench-AA for about 40% of the cost.
GPT-6 Sol, then compare
What to do this week

Route by price band: send budget work to Sol and demanding work to Opus 5.5, then check both on your own tasks

One independent harness is a strong start, but not a verdict on your workload. Take ten to twenty real tasks, run Sol at xhigh and Opus 5.5 at medium, and compare the cost of each completed task, not the cost per token. If you also use Anthropic’s top model, our Opus 5.5 and GPT-6 Astra comparison covers the next tier up, and the frontier model price index tracks list prices across vendors.

Digital Applied

Route each task to the model that finishes it for less.

We build the eval sets, effort sweeps and routing rules that decide which model handles which task, based on cost per completed task.

Model evaluationEffort tuningModel routing
Your next project

Two vendors, one routing plan

  • A task set drawn from your work
  • Cost per completed task for each model
  • Routing rules with a fallback
Questions and answers

The questions we get about GPT-6 Sol vs Claude Opus 5.5

Not at the top end. On Artificial Analysis's index, Sol's best score (47.5 at max) is below Opus 5.5 at its default medium setting (51.2). Sol is the cheaper way to reach scores up to about 44, and it matches Opus 5.5 at medium on business workflows.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading