Meta released Muse Spark 1.3 on 2 September, and the clearest numbers in the announcement are about cost rather than capability: roughly 20% fewer tool calls and 25% fewer tokens than its predecessor on the same coding work.

Efficiency as the headline

Those two figures are unusual for a model launch because they are the ones a buyer can act on. A model that reaches the same result through fewer calls and fewer tokens is cheaper per task whatever it scores on a leaderboard, and the saving compounds in agentic workloads where a single job may run for hundreds of steps.

An abstract pattern of connected nodes on a dark background.
Fewer tool calls and fewer tokens lower the cost of a task regardless of leaderboard position. Ann H · pexels · Pexels License

What Meta says it improved

Beyond efficiency, the company describes longer-horizon execution across multi-step tasks, better collaboration — the model asking clarifying questions and asking for help when it needs it — improved multitasking, and better awareness of its own capabilities and limits.

That last one is the hardest to evaluate and the most consequential. A model that knows when it is out of its depth fails differently from one that does not.

Meta also reports improved safety training targeting adversarial robustness and prompt injection resistance, and says the model was trained across diverse harnesses for agentic environments.

The scorecard without scores

The announcement carries a benchmark scorecard comparing Muse Spark 1.3 against its own 1.2, GPT 5.6 Sol at max, and Opus 5 at max, across agent, coding, instruction-following and long-context evaluations.

The specific numbers are not in the post. A comparison chart naming two named competitors, without the figures written out, asks readers to accept a visual rather than a result.

A team collaborating around a laptop in an office.
The model is rolling out in Muse Code and the Meta Model API; open weights are on the roadmap only. Andrea Piacquadio · pexels · Pexels License

Availability

The model is rolling out in Muse Code and the Meta Model API. A max reasoning mode is described as coming shortly. Open weights appear on the roadmap but are not available now, and no licence terms have been stated.

What to watch

Whether the efficiency claim survives contact with other people’s workloads. A 25% token reduction measured on Meta’s own agentic harnesses is a claim about those harnesses until someone else measures it on theirs.