AI DevelopmentDecision Matrix6 min readPublished September 22, 2026

$4 / $20 vs $2 / $6 vs $1.25 / $4.25 · the cheapest list price is not the cheapest task

Opus 5.5 vs Grok 4.7 vs Muse Spark 1.3: Real Cost per Task

Grok 4.7 lists at under a third of Opus 5.5's output price, yet costs twice as much per task in Artificial Analysis tests. Where Muse Spark 1.3 fits.

DA
Digital Applied Team
Research and practical guidance
IndexArtificial Analysis v4.3.2
Data readSeptember 22, 2026

Grok 4.7 lists at $2 per million input tokens and $6 per million output tokens. Claude Opus 5.5 lists at $4 and $20. On price per token, Grok 4.7 looks like the cheaper model. On Artificial Analysis’s Intelligence Index, it costs about twice as much per task at its default effort ($2.73 against $1.34 for Opus 5.5 at its default) and scores lower (46.3 against 51.2).

The difference is token use. To run the index at its default effort, Grok 4.7 generated about 65,900 output tokens per task. Opus 5.5 at its default generated about 25,700. Muse Spark 1.3, the cheapest of the three at list price, sits between them: it scores above Grok 4.7 and costs less per task, but not less than Opus 5.5 at its default.

Every figure in this post comes from Artificial Analysis, an independent benchmarking firm that runs every model through the same ten evaluations and prices each run at the vendor’s list rates. We read the data on September 22, 2026, from index version 4.3.2. Opus 5.5 was added the day it launched, so its numbers may be revised as more runs complete.

Key takeaways
  1. 01
    At each model’s default effort, Grok 4.7 cost $2.73 per task and Opus 5.5 cost $1.34.Grok 4.7 generated about 2.6× as many output tokens per task (65.9k against 25.7k), and 85% of its cost per task was input re-read across turns. It also scored lower (46.3 against 51.2).
  2. 02
    Grok 4.7 used 1.8× to 2.1× the output tokens per task of Grok 4.6 for about 2 index points.Cost per task rose 47% at high effort and 61% at xhigh; time per task rose from 619 to 1,068 seconds at high.
  3. 03
    Muse Spark 1.3 beats Grok 4.7 on both score and cost, but not Opus 5.5 at its default.Muse Spark 1.3 at max scored 48.1 for $1.60 per task, and Artificial Analysis measured it at about 213 output tokens per second.
  4. 04
    Opus 5.5 is token-efficient at medium and high, not at max.At max it used 119.2k output tokens per task, the most of any run here, for a top score of 57.6 at $5.98 per task.

01The rate cardsList prices, per million tokens

Sources: Anthropic’s pricing page, xAI’s release notes and Meta’s Model API page, all read on September 22, 2026. Meta also sells a Muse Spark 1.3 Contributor tier at $0.10 / $0.20 whose prompts and outputs Meta may use to improve its products, so we leave it out of this comparison.

Published list prices per million tokens, September 22, 2026. Grok 4.7’s higher band applies to the whole request once the prompt passes 200,000 tokens.
LineOpus 5.5Grok 4.7Muse Spark 1.3
Input$4$2$1.25
Cached input$0.20$0.50$0.15
Output$20$6$4.25
Long promptsSame rate to 1M$4 / $1 / $12 above 200KOne rate listed
Context window1M500K1M

One line in this table reverses the usual order. Opus 5.5’s cached input costs $0.20 per million tokens, less than Grok 4.7’s $0.50, even though Opus 5.5’s uncached input costs twice as much. Agent loops re-send their context on every turn, so cached input is often the largest part of an agent’s bill. That is one reason list prices predict agent costs so poorly.

02The measurementWhat each model cost to run the same ten evaluations

Artificial Analysis runs each model at several effort levels and records the tokens it used and the cost at list prices. The table shows every run it has published for these three models. Cost per task is Artificial Analysis’s own figure. It tested Muse Spark 1.3 only at xhigh and max.

Source: Artificial Analysis Intelligence Index v4.3.2 model pages for Opus 5.5, Grok 4.7 and Muse Spark 1.3, read September 22, 2026. Output tokens are Artificial Analysis’s per-task average and include reasoning tokens.
Model · effortIndexOutput tokens / taskCost per task
Opus 5.5 · low42.310.2k$0.55
Opus 5.5 · medium (default)51.225.7k$1.34
Opus 5.5 · high53.635.6k$1.82
Opus 5.5 · xhigh56.065.7k$3.46
Opus 5.5 · max57.6119.2k$5.98
Grok 4.7 · high (default)46.365.9k$2.73
Grok 4.7 · xhigh46.480.6k$3.74
Muse Spark 1.3 · xhigh45.154.5k$1.37
Muse Spark 1.3 · max48.160.2k$1.60

Output tokens per Intelligence Index task (thousands)

Artificial Analysis Intelligence Index v4.3.2, read September 22, 2026
Opus 5.5 · mediumIndex 51.2
25.7k
Opus 5.5 · highIndex 53.6
35.6k
Grok 4.6 · highIndex 44.3
35.8k
Muse Spark 1.3 · maxIndex 48.1
60.2k
Grok 4.7 · highIndex 46.3
65.9k
Grok 4.7 · xhighIndex 46.4
80.6k
Opus 5.5 · maxIndex 57.6
119.2k

The cost split explains most of the gap, and it isn’t where the list prices suggest. Per task, 85% of Grok 4.7’s cost at high effort ($2.33 of $2.73) was input: context sent to the model again on each turn. Its output cost per task, $0.40, was lower than Opus 5.5’s $0.51 at medium. Opus 5.5 spent 61% of its $1.34 on input. Longer reasoning and more turns both mean the same context is paid for again, and cheap output tokens don’t offset that. Artificial Analysis doesn’t publish turn counts, so the mechanism is our reading of its cost split.

OpenRouter listed Grok 4.7 at $1.60 / $4.80 on September 22, 20% below xAI’s list price on every line. Applied to Artificial Analysis’s figure, that would bring Grok 4.7 at high effort to about $2.18 per task (our arithmetic), still above Opus 5.5’s $1.34.

Opus 5.5 is not token-efficient at every setting. At max it generated about 119,200 output tokens per task, more than Grok 4.7 at xhigh (80,600), and cost $5.98 per task. It also posted the highest score here, 57.6. The step from xhigh to max added 1.6 index points and cost 73% more per task. That matches Anthropic’s own advice to reserve the top settings for work where you’ve measured a gain.

03The regressionGrok 4.7 against Grok 4.6: small gains, much larger bills

xAI kept Grok 4.7 at Grok 4.6’s list price, which our Grok 4.7 launch post covered. The Artificial Analysis runs show what the same price buys per task.

Source: Artificial Analysis Intelligence Index v4.3.2, read September 22, 2026. Changes are our arithmetic.
Metric · effortGrok 4.6Grok 4.7Change
Intelligence Index · high44.346.3+2.0 points
Output tokens per task · high35.8k65.9k1.8×
Cost per task · high$1.86$2.73+47%
Time per task · high619 s1,068 s+73%
Intelligence Index · xhigh44.246.4+2.2 points
Output tokens per task · xhigh37.6k80.6k2.1×
Cost per task · xhigh$2.32$3.74+61%

For about two index points, Grok 4.7 roughly doubles its output and takes 65–73% longer per task. The gains are real on some evaluations: Terminal-Bench 4.0 rose from 21.2% to 24.7% at high effort and GDPval-AA from 1605 to 1694 Elo. For a team already on Grok 4.6, the question is whether those specific gains matter enough to pay 47–61% more per task for them.

04The scorecardWhere each model leads

The index is an average, and the averages hide real differences. The table shows the evaluations inside it, with each model at its default effort. Muse Spark 1.3 is shown at max, its highest-scoring tested setting.

Opus 5.5 at medium, Grok 4.7 at high, Muse Spark 1.3 at max. Source: Artificial Analysis v4.3.2, read September 22, 2026.
EvaluationOpus 5.5Grok 4.7Muse 1.3
Terminal-Bench 4.052.5%24.7%33.3%
Humanity’s Last Exam54.7%42.3%48.7%
CritPt (physics research)27.7%18.0%24.9%
AA-LCR (long-context reasoning)84.3%77.0%83.0%
GDPval-AA v2.1 (Elo)157616941674
AutomationBench-AA (partial score)61.2%63.5%57.9%
SciCode59.3%57.8%58.8%
AA-Omniscience index40.330.925.0

At the default settings, Grok 4.7 leads on two evaluations: GDPval-AA, which scores real professional tasks, and AutomationBench. Opus 5.5 at medium leads everywhere else, by about 28 points on Terminal-Bench 4.0 and 12 points on Humanity’s Last Exam. The Grok leads are not fixed. Opus 5.5 at high, which still costs less per task than Grok 4.7 at high ($1.82 against $2.73), scores 1692 on GDPval-AA and 63.2% on AutomationBench, level with Grok 4.7 on both.

Muse Spark 1.3 is the middle model on most rows. Its advantage is speed. Artificial Analysis measured it at about 213 output tokens per second at max, and it averaged 278 seconds per index task, against 1,068 seconds for Grok 4.7 at high. Artificial Analysis had not yet published a time per task for Opus 5.5. Our Muse Spark 1.3 launch post covers Meta’s own token claims.

The one row where Opus 5.5 looks worse

Artificial Analysis’s knowledge test, AA-Omniscience, rewards right answers and penalises wrong ones, and doesn’t penalise declining to answer. Opus 5.5 answered the most questions correctly (64.5% at medium, against 47.8% for Grok 4.7 and 43.6% for Muse Spark 1.3). But when it didn’t know, it gave a wrong answer 68.4% of the time, against about 32–33% for the other two, which more often declined. For work where a confident wrong answer is costly, that matters as much as the score.

05The caveatsHow to read these numbers

Artificial Analysis’s numbers are independent, which is why this post uses them, but they measure one fixed workload. Four things to keep in mind:

  • Its harness scores lower than the vendors’. On Terminal-Bench 4.0 at xhigh, Artificial Analysis measured Opus 5.5 at 59.6% where Anthropic reports 66.4%, and Grok 4.7 at 25.8% where xAI reports 38.0%. The drop is larger for Grok 4.7.
  • Opus 5.5 was run with fallback on. Artificial Analysis labels its Opus 5.5 runs “Default Fallback,” meaning requests that tripped a safeguard were answered by another Claude model, as they would be in production with that option enabled.
  • Scores change between index versions. Version 4.3.2 includes ten evaluations, among them AA-Briefcase, GDPval-AA, AutomationBench-AA and Terminal-Bench 4.0. Earlier versions used a different set of evaluations. Our September 2 Muse Spark 1.3 post quoted an older version, and its figures can’t be compared with the ones here.
  • Cost depends on your mix. Cost per task reflects this set of evaluations at list prices. A workload with shorter contexts, more caching or fewer turns will cost differently, and the ranking can change.

For more on each model, see our Opus 5.5 launch post and Opus 5.5 vs GPT-6 Astra. The raw Artificial Analysis pages for Opus 5.5, Grok 4.7 and Muse Spark 1.3 will carry any updates.

06ConclusionToken use, not list price, decides which model is cheapest per task

You want the best score per dollar at default settings
Opus 5.5 at medium scored 51.2 for $1.34 per task, above both other models at any setting tested, and cheaper per task than every other run here except Opus 5.5’s own low setting.
Opus 5.5 · medium
Latency matters more than the last few points
Muse Spark 1.3 generated about 213 tokens per second at max and averaged 278 seconds per index task. Check whether it clears your quality bar; if so, it is the fastest measured option here.
Muse Spark 1.3
You run Grok 4.6 today
Measure before you upgrade. On this workload Grok 4.7 cost 47–61% more per task for about two index points. If your tasks resemble GDPval or AutomationBench, the gain may be worth it.
Test Grok 4.7
A confident wrong answer is expensive
Weigh the AA-Omniscience result. Grok 4.7 and Muse Spark 1.3 declined more often when unsure; Opus 5.5 answered more often, correctly and incorrectly. Add verification to Opus 5.5 outputs in that kind of work.
Add a check
What to do this week

Price your own tasks per completed job, not per million tokens, before you pick the cheaper-looking model

On Artificial Analysis’s workload, the model with the highest list price cost the least per task at its default setting, because it used far fewer tokens to finish. That won’t hold for every workload, and it didn’t hold for Opus 5.5 at max. Run a sample of your own tasks on each model at its default and one higher effort level, record total cost per completed task, and compare those numbers. Our price index keeps the list prices current.

Digital Applied

Know what a model costs per finished task.

We measure token use, cost per completed task and quality on your own work across the models you are considering, then set up routing that uses the cheapest one that meets the bar.

Cost per taskToken auditsModel routing
Your next project

A model bill you can predict

  • Token use measured on your tasks
  • Cost per completed job, not per token
  • Routing to the cheapest model that passes
Questions and answers

The questions we get about Opus 5.5, Grok 4.7 and Muse Spark 1.3

Per token, yes: $2 / $6 against $4 / $20 per million input and output tokens. Per task, not on Artificial Analysis's Intelligence Index. At each model's default effort, Grok 4.7 cost $2.73 per task and Opus 5.5 cost $1.34, because Grok 4.7 generated about 2.6× as many output tokens per task and paid for its context again on more turns.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading