On September 2, 2026 Meta released Muse Spark 1.3 into Muse Code and the Meta Model API, its fourth Muse Spark release in five months by Artificial Analysis’s count. The launch post leads with behaviour rather than benchmarks: the model “asks clarifying questions when prompts are ambiguous, invokes help from the user when stuck, and confirms before taking consequential actions.” On coding work Meta’s engineers found it “significantly faster and more efficient, using ~20% fewer tool calls and ~25% fewer tokens” than Muse Spark 1.2.
The same day Artificial Analysis published its independent run. Muse Spark 1.3 scores 61 on the Intelligence Index at the xhigh reasoning level that is available today, up four points from 1.2, at $0.55 per index task, which the firm says no model scoring 59 or above beats. It also found the model consumed about 57% more input tokens per task than 1.2 on agentic evaluations. A reader who sees “25% fewer tokens” in one tab and “57% more tokens” in the other is entitled to ask which is true. Both are, and the reason is worth a post because it applies to every efficiency claim you will read this year.
- 01Meta measured its own loop. Artificial Analysis measured fixed tasks.Meta’s ~25% comes from engineer comparisons on coding work where the model takes fewer turns. Artificial Analysis runs the same task set for every model and counts every token, and there the input side grew about 57%. Different denominators, both reported honestly.
- 02Cheapest model at its intelligence level, at an unchanged price.$1.25 input and $4.25 output per million tokens, cache hits $0.15, the same as 1.2. At 61 on the index the peers cost $0.95 (GPT-5.6 Sol, max) and $0.94 (Grok 4.6, high) per task against $0.55 here.
- 03The gains are agentic, and they cost tokens.τ³-Bench Banking rose 12 points to 47%, Terminal-Bench 2.1 from 80% to 85%, GDPval-AA v2 from 1,615 to 1,709 Elo. Cost per task rose from $0.40 to $0.55 in the process. Two evaluations regressed slightly, partly because the model abstains more.
- 04Asking first is a harness requirement, not just a feature.A model that stops to ask needs something on the other end that can answer. Unattended pipelines built for 1.2 should be tested for stalls before switching; the contributor tier and max reasoning both carry their own caveats.
01 — The releaseWhat Meta shipped.
Muse Spark 1.3 is a point release positioned around long-horizon agentic work. Meta says it was trained “across a diverse set of harnesses” to generalise, that it keeps track of what it has learned across a long thread, and that it maps incoming prompts to the right task more accurately when a user interrupts or steers. On coding, Meta says it was trained on more long-horizon tasks and “takes fewer turns where not needed and is less verbose.” The model is available today at the previously offered reasoning levels; max reasoning is “coming shortly after we finish additional safety testing,” and Artificial Analysis describes the max variant as a limited preview for Meta’s partners.
The safety paragraph is unusually specific for a launch post: Meta claims stronger resistance to prompt injection and “better calibration on what constitutes irreversible actions.” That pairs with the confirmation behaviour and is the part of the release most relevant to anyone who gave an agent write access. Our Muse Spark 1.2 launch guide covers the Muse Code tooling this ships into; nothing there changes with 1.3 beyond the model.
Trained to more actively collaborate with the user, Muse Spark 1.3 asks clarifying questions when prompts are ambiguous, invokes help from the user when stuck, and confirms before taking consequential actions.Meta, “Introducing Muse Spark 1.3,” September 2, 2026
02 — The citable partTwo token numbers that both hold.
Meta’s figure is a comparison by its own engineers on coding workflows, where the model’s new habit of taking fewer turns directly reduces the tool calls and tokens a session consumes. If a task that took eight tool calls on 1.2 takes six on 1.3, and each call carries the whole context back in, the token saving follows from the turn saving. Artificial Analysis does something different: it runs a fixed set of tasks against every model under the same harness and counts every token the model reads and writes. On that basis 1.3 read about 57% more input tokens per task than 1.2 on the agentic evaluations, wrote about 8% more, and consequently cost $0.55 per task where 1.2 cost $0.40. The firm attributes the higher agentic scores of the max variant to “more turns and total reasoning tokens,” 62% more reasoning on GDPval-AA v2 than xhigh.
So the model saves tokens when it can shorten a loop and spends tokens when a fixed task rewards more reading. Both are what a more capable agent looks like from the outside. The mistake would be to take either number as a forecast of your bill. The table sets the two measurements side by side with what each one counts.
| Measurement | Meta (vendor-run) | Artificial Analysis (independent) |
|---|---|---|
| What was compared | Coding workflows, “in comparisons by Meta engineers” | Fixed Intelligence Index tasks, same harness for every model |
| Tool calls | ~20% fewer | more turns on agentic evaluations |
| Tokens | ~25% fewer | input +57%, output +8% |
| Cost per task | not stated | $0.40 → $0.55 |
| Why the number moves | Fewer turns where not needed; less verbose output | More reading and more turns on tasks that reward them |
| What it predicts for you | Interactive coding sessions may get cheaper | Long autonomous tasks may get dearer and better |
03 — EvidenceThe independent scorecard.
Artificial Analysis places Muse Spark 1.3 at xhigh on 61, tied with GPT-5.6 Sol at max, Grok 4.6 at high and Claude Opus 5 at high, and behind Claude Fable 5.1 at max on 66, Opus 5 at max on 63 and Fable 5 at max on 62. The max variant, in limited preview, lands on 62. The firm’s summary is that the gains “are concentrated in agentic evaluations”: τ³-Bench Banking from 35% to 47% at xhigh and 52% at max, which it calls the top score on that test; Terminal-Bench 2.1 from 80% to 85%; GDPval-AA v2 from 1,615 to 1,709 Elo. Scientific reasoning rose too, led by an eight-point gain on CritPt to 26% and GPQA Diamond to 94%.
Two evaluations moved the other way. AA-LCR, a long-context test, fell four points to 79% for both variants, and AA-Omniscience accuracy fell three points at xhigh, which Artificial Analysis attributes to “a higher abstention rate,” the model declining to answer when unsure, which also lowered its hallucination rate. That is the same trait as “asks when stuck,” showing up in a benchmark. Meta’s own scorecard image compares 1.3 against 1.2, GPT-5.6 Sol at max and Opus 5 at max across agent, coding, instruction-following and long-context tests; we do not reproduce its numbers because they are vendor-run and the image carries no methodology beyond a linked report.
04 — CostPrice per token, price per task.
The list price did not move: $1.25 per million input tokens, $4.25 per million output tokens, cache hits at $0.15, per Artificial Analysis and the OpenRouter listing that appeared at 19:45 UTC on September 2. What moved is the value of that price. At 61 on the index, the nearest models by score cost $0.94 and $0.95 per task; below it, Gemini 3.8 Flash at 59 costs $0.58 and GPT-5.6 Sol at xhigh costs $0.63. Muse Spark 1.3 at $0.55 sits on the intelligence-versus-cost frontier, and its neighbour on that frontier, Gemini 3.8 Flash, shipped the same morning with a price rise already scheduled.
A second listing followed at 20:38 UTC: a contributor tier at $0.10 input and $0.20 output per million, with cached input at $0.002. The tier is not new; it is the arrangement where Meta trains on your traffic in exchange for the discount, and our August note on the contributor tier’s economics explains why it is unsuitable for client work regardless of the price. Both rows are on the frontier model API price index.
Muse Spark 1.3 (xhigh)
The generally available model. Index 61, $0.55 per index task, 1M-token context, text, image and video in. Available in the Meta Model API and Muse Code from September 2.
Muse Spark 1.3 (max)
Index 62 with 62% more reasoning tokens on GDPval-AA v2 than xhigh. Limited preview for Meta’s partners; Meta says it arrives “shortly after we finish additional safety testing.” Excluded from cost comparisons because there is no price.
Contributor tier
Listed on OpenRouter at 20:38 UTC on September 2. Meta trains on the traffic. Cached input is priced at 2% of the base input rate, and it is the wrong tier for anything that touches a client’s data.
Meta’s post lists “the Muse Spark open weights release” in its roadmap with no date. That is a row on our withheld-weights ledger, not a claim in this post. Until a date exists, plan as if the model is API-only.
05 — DesignWhat “asks before it acts” changes.
A model that asks clarifying questions, asks for help when stuck and confirms before consequential actions is what most teams said they wanted after a summer of agents that did the opposite. It also changes the contract with the harness around it. A question is a turn that ends with the model waiting. In Muse Code, with a person at the keyboard, that is the point. In an unattended pipeline built for 1.2, it is a stall unless the orchestrator can answer, escalate or time out. Meta says the model adapts to user preferences between frequent updates and working silently, so the behaviour is steerable, but it has to be steered.
The general rule this release illustrates is simple to state and easy to forget: a token-efficiency claim is only meaningful with its denominator attached. When a vendor says fewer tokens, ask fewer tokens per what. When a benchmark says more tokens, ask on which tasks. Then measure your own. That instrumentation is what our AI transformation practice builds first on every agent engagement, because every model change after that becomes a comparison rather than a guess.
06 — ConclusionFewer per loop, more per task.
Meta’s 25% fewer and Artificial Analysis’s 57% more are both true. One counts a loop, the other counts a task.
Muse Spark 1.3 is a real step: four index points, a top score on a banking tool-use test, an unchanged price that makes it the cheapest model at its level, and a behaviour change that most agent operators asked for. It reads more to get there, and on a fixed task that shows up as tokens.
Switch interactive work now. Trial unattended work with a harness that can answer a question. Skip the contributor tier for anything that is not yours. And attach a denominator to every efficiency number you repeat, including these.