AI DevelopmentNew Release8 min readPublished September 2, 2026

Fewer tokens in Meta’s loop, more tokens on a fixed task

Muse Spark 1.3 Asks Before It Acts. Fewer Tokens, Really?

Meta released Muse Spark 1.3 on September 2 with a behaviour change and an efficiency claim: the model now asks before it acts, and Meta’s engineers saw about 25% fewer tokens than 1.2. The same day, Artificial Analysis measured about 57% more input tokens per task. Neither is wrong. They count different things, and the difference is what a buyer needs to understand.

DA
Digital Applied Team
Senior strategists · Published Sep 2, 2026
PublishedSep 2, 2026
Read time8 min
EventSep 2, 2026
Meta’s efficiency claim vs 1.2
−25% tokens
and ~20% fewer tool calls, in comparisons by Meta engineers (vendor-run)
Artificial Analysis, input tokens per task
+57%
on agentic evaluations vs 1.2; output tokens up ~8% (independent)
Intelligence Index, xhigh reasoning
61
ties GPT-5.6 Sol (max) and Grok 4.6 (high); 1.2 scored 57
Cost per index task
$0.55
cheapest of any model at 59 or above, per Artificial Analysis

On September 2, 2026 Meta released Muse Spark 1.3 into Muse Code and the Meta Model API, its fourth Muse Spark release in five months by Artificial Analysis’s count. The launch post leads with behaviour rather than benchmarks: the model “asks clarifying questions when prompts are ambiguous, invokes help from the user when stuck, and confirms before taking consequential actions.” On coding work Meta’s engineers found it “significantly faster and more efficient, using ~20% fewer tool calls and ~25% fewer tokens” than Muse Spark 1.2.

The same day Artificial Analysis published its independent run. Muse Spark 1.3 scores 61 on the Intelligence Index at the xhigh reasoning level that is available today, up four points from 1.2, at $0.55 per index task, which the firm says no model scoring 59 or above beats. It also found the model consumed about 57% more input tokens per task than 1.2 on agentic evaluations. A reader who sees “25% fewer tokens” in one tab and “57% more tokens” in the other is entitled to ask which is true. Both are, and the reason is worth a post because it applies to every efficiency claim you will read this year.

Key takeaways
  1. 01
    Meta measured its own loop. Artificial Analysis measured fixed tasks.Meta’s ~25% comes from engineer comparisons on coding work where the model takes fewer turns. Artificial Analysis runs the same task set for every model and counts every token, and there the input side grew about 57%. Different denominators, both reported honestly.
  2. 02
    Cheapest model at its intelligence level, at an unchanged price.$1.25 input and $4.25 output per million tokens, cache hits $0.15, the same as 1.2. At 61 on the index the peers cost $0.95 (GPT-5.6 Sol, max) and $0.94 (Grok 4.6, high) per task against $0.55 here.
  3. 03
    The gains are agentic, and they cost tokens.τ³-Bench Banking rose 12 points to 47%, Terminal-Bench 2.1 from 80% to 85%, GDPval-AA v2 from 1,615 to 1,709 Elo. Cost per task rose from $0.40 to $0.55 in the process. Two evaluations regressed slightly, partly because the model abstains more.
  4. 04
    Asking first is a harness requirement, not just a feature.A model that stops to ask needs something on the other end that can answer. Unattended pipelines built for 1.2 should be tested for stalls before switching; the contributor tier and max reasoning both carry their own caveats.

01The releaseWhat Meta shipped.

Muse Spark 1.3 is a point release positioned around long-horizon agentic work. Meta says it was trained “across a diverse set of harnesses” to generalise, that it keeps track of what it has learned across a long thread, and that it maps incoming prompts to the right task more accurately when a user interrupts or steers. On coding, Meta says it was trained on more long-horizon tasks and “takes fewer turns where not needed and is less verbose.” The model is available today at the previously offered reasoning levels; max reasoning is “coming shortly after we finish additional safety testing,” and Artificial Analysis describes the max variant as a limited preview for Meta’s partners.

The safety paragraph is unusually specific for a launch post: Meta claims stronger resistance to prompt injection and “better calibration on what constitutes irreversible actions.” That pairs with the confirmation behaviour and is the part of the release most relevant to anyone who gave an agent write access. Our Muse Spark 1.2 launch guide covers the Muse Code tooling this ships into; nothing there changes with 1.3 beyond the model.

Trained to more actively collaborate with the user, Muse Spark 1.3 asks clarifying questions when prompts are ambiguous, invokes help from the user when stuck, and confirms before taking consequential actions.Meta, “Introducing Muse Spark 1.3,” September 2, 2026

02The citable partTwo token numbers that both hold.

Meta’s figure is a comparison by its own engineers on coding workflows, where the model’s new habit of taking fewer turns directly reduces the tool calls and tokens a session consumes. If a task that took eight tool calls on 1.2 takes six on 1.3, and each call carries the whole context back in, the token saving follows from the turn saving. Artificial Analysis does something different: it runs a fixed set of tasks against every model under the same harness and counts every token the model reads and writes. On that basis 1.3 read about 57% more input tokens per task than 1.2 on the agentic evaluations, wrote about 8% more, and consequently cost $0.55 per task where 1.2 cost $0.40. The firm attributes the higher agentic scores of the max variant to “more turns and total reasoning tokens,” 62% more reasoning on GDPval-AA v2 than xhigh.

So the model saves tokens when it can shorten a loop and spends tokens when a fixed task rewards more reading. Both are what a more capable agent looks like from the outside. The mistake would be to take either number as a forecast of your bill. The table sets the two measurements side by side with what each one counts.

The two efficiency measurements for Muse Spark 1.3 against 1.2, both published September 2, 2026. Meta’s figures are vendor-run; Artificial Analysis’s are independent and cover its Intelligence Index task set at xhigh reasoning.
MeasurementMeta (vendor-run)Artificial Analysis (independent)
What was comparedCoding workflows, “in comparisons by Meta engineers”Fixed Intelligence Index tasks, same harness for every model
Tool calls~20% fewermore turns on agentic evaluations
Tokens~25% fewerinput +57%, output +8%
Cost per tasknot stated$0.40 → $0.55
Why the number movesFewer turns where not needed; less verbose outputMore reading and more turns on tasks that reward them
What it predicts for youInteractive coding sessions may get cheaperLong autonomous tasks may get dearer and better

03EvidenceThe independent scorecard.

Artificial Analysis places Muse Spark 1.3 at xhigh on 61, tied with GPT-5.6 Sol at max, Grok 4.6 at high and Claude Opus 5 at high, and behind Claude Fable 5.1 at max on 66, Opus 5 at max on 63 and Fable 5 at max on 62. The max variant, in limited preview, lands on 62. The firm’s summary is that the gains “are concentrated in agentic evaluations”: τ³-Bench Banking from 35% to 47% at xhigh and 52% at max, which it calls the top score on that test; Terminal-Bench 2.1 from 80% to 85%; GDPval-AA v2 from 1,615 to 1,709 Elo. Scientific reasoning rose too, led by an eight-point gain on CritPt to 26% and GPQA Diamond to 94%.

Two evaluations moved the other way. AA-LCR, a long-context test, fell four points to 79% for both variants, and AA-Omniscience accuracy fell three points at xhigh, which Artificial Analysis attributes to “a higher abstention rate,” the model declining to answer when unsure, which also lowered its hallucination rate. That is the same trait as “asks when stuck,” showing up in a benchmark. Meta’s own scorecard image compares 1.3 against 1.2, GPT-5.6 Sol at max and Opus 5 at max across agent, coding, instruction-following and long-context tests; we do not reproduce its numbers because they are vendor-run and the image carries no methodology beyond a linked report.

04CostPrice per token, price per task.

The list price did not move: $1.25 per million input tokens, $4.25 per million output tokens, cache hits at $0.15, per Artificial Analysis and the OpenRouter listing that appeared at 19:45 UTC on September 2. What moved is the value of that price. At 61 on the index, the nearest models by score cost $0.94 and $0.95 per task; below it, Gemini 3.8 Flash at 59 costs $0.58 and GPT-5.6 Sol at xhigh costs $0.63. Muse Spark 1.3 at $0.55 sits on the intelligence-versus-cost frontier, and its neighbour on that frontier, Gemini 3.8 Flash, shipped the same morning with a price rise already scheduled.

A second listing followed at 20:38 UTC: a contributor tier at $0.10 input and $0.20 output per million, with cached input at $0.002. The tier is not new; it is the arrangement where Meta trains on your traffic in exchange for the discount, and our August note on the contributor tier’s economics explains why it is unsuitable for client work regardless of the price. Both rows are on the frontier model API price index.

Tier
Muse Spark 1.3 (xhigh)
$1.25 / $4.25 per M · cache $0.15

The generally available model. Index 61, $0.55 per index task, 1M-token context, text, image and video in. Available in the Meta Model API and Muse Code from September 2.

Available now
Tier
Muse Spark 1.3 (max)
Price not announced

Index 62 with 62% more reasoning tokens on GDPval-AA v2 than xhigh. Limited preview for Meta’s partners; Meta says it arrives “shortly after we finish additional safety testing.” Excluded from cost comparisons because there is no price.

Limited preview
Tier
Contributor tier
$0.10 / $0.20 per M · cache $0.002

Listed on OpenRouter at 20:38 UTC on September 2. Meta trains on the traffic. Cached input is priced at 2% of the base input rate, and it is the wrong tier for anything that touches a client’s data.

Trains on traffic
Open weights

Meta’s post lists “the Muse Spark open weights release” in its roadmap with no date. That is a row on our withheld-weights ledger, not a claim in this post. Until a date exists, plan as if the model is API-only.

05DesignWhat “asks before it acts” changes.

A model that asks clarifying questions, asks for help when stuck and confirms before consequential actions is what most teams said they wanted after a summer of agents that did the opposite. It also changes the contract with the harness around it. A question is a turn that ends with the model waiting. In Muse Code, with a person at the keyboard, that is the point. In an unattended pipeline built for 1.2, it is a stall unless the orchestrator can answer, escalate or time out. Meta says the model adapts to user preferences between frequent updates and working silently, so the behaviour is steerable, but it has to be steered.

Interactive coding in Muse Code or a chat harness
Switch. This is the setting Meta’s efficiency claim was measured in, and the clarifying-question behaviour has a human to answer it. Expect fewer wasted tool calls on ambiguous asks.
1.3 · xhigh
Unattended agent pipelines that ran on 1.2
Test before switching. Give the orchestrator a way to answer or reject a clarifying question, set a timeout for a waiting turn, and measure input tokens per task on your own workload; Artificial Analysis’s +57% is on its task set, not yours.
Trial, then 1.3
Agents with write access to production systems
Welcome the confirmation step and wire it to a real approval, not an auto-yes. Meta’s “better calibration on what constitutes irreversible actions” is a vendor claim; your approval log is the evidence.
1.3 with approvals
Anyone routing by cost across vendors
At 61 on the index, $0.55 per task is the lowest available, and the price did not change. Route long agentic work here and keep Gemini 3.8 Flash for the 59-level tier until its January price rise.
Router

The general rule this release illustrates is simple to state and easy to forget: a token-efficiency claim is only meaningful with its denominator attached. When a vendor says fewer tokens, ask fewer tokens per what. When a benchmark says more tokens, ask on which tasks. Then measure your own. That instrumentation is what our AI transformation practice builds first on every agent engagement, because every model change after that becomes a comparison rather than a guess.

06ConclusionFewer per loop, more per task.

Muse Spark 1.3

Meta’s 25% fewer and Artificial Analysis’s 57% more are both true. One counts a loop, the other counts a task.

Muse Spark 1.3 is a real step: four index points, a top score on a banking tool-use test, an unchanged price that makes it the cheapest model at its level, and a behaviour change that most agent operators asked for. It reads more to get there, and on a fixed task that shows up as tokens.

Switch interactive work now. Trial unattended work with a harness that can answer a question. Skip the contributor tier for anything that is not yours. And attach a denominator to every efficiency number you repeat, including these.

Measure before you switch

Swap models on evidence, not press releases.

We instrument every agent route for tokens per task before a model swap, so that a vendor’s efficiency claim and a benchmark’s token count can both be checked against your own workload.

Free consultationExpert guidanceTailored solutions
What we work on

Agent model-migration engagements

  • Tokens-per-task instrumentation per route
  • Harness support for clarifying questions and approvals
  • Cross-vendor routing by measured cost per task
  • Contributor-tier and data-use policy reviews
  • Model swap trials with before-and-after evidence
FAQ · Muse Spark 1.3

The questions we get about Muse Spark 1.3.

In Meta’s own comparisons on coding work, yes: about 20% fewer tool calls and 25% fewer tokens, because the model takes fewer turns where they are not needed. On Artificial Analysis’s fixed task set it used about 57% more input tokens and 8% more output tokens per task, lifting cost per task from $0.40 to $0.55. Which applies to you depends on whether your workload looks like an interactive loop or a fixed long task.
Related dispatches

Continue exploring Meta’s models and agent pricing.