
Share
Meta is touting "frontier performance almost too cheap to meter" for Muse Spark 1.3. The catch: its strongest scores come from a max reasoning configuration still in safety testing, not the version shipping to developers today.
Meta shipped Muse Spark 1.3 this week, and the headline claim from CEO Mark Zuckerberg is bold: "frontier performance almost too cheap to meter," and the "biggest jump" yet for Meta in coding and agentic work. There's real substance behind that. But the version you can actually call through the API right now isn't the one generating Meta's most impressive numbers.
Here's the split. Meta's strongest Muse Spark 1.3 benchmark results come from a "max" reasoning configuration. That version is still completing safety testing and will arrive "shortly," according to Meta. Third-party benchmarking firm Artificial Analysis says it evaluated max only in a limited partner preview and currently lists zero API providers offering it. The model actually rolling out this week through Muse Code and the Meta Model API uses Meta's previously available reasoning settings, including a mode called xhigh.
Meta isn't hiding this distinction. It discloses results for both configurations in its evaluation report. But the launch materials lean heavily on max, and several of the eye-catching scores belong to that unshippable variant.
The gaps aren't trivial. Meta reports GDPval-AA v2 scores of 1,754 Elo for max versus 1,709 for xhigh. On OSWorld 2.0, a benchmark for autonomous computer-use tasks, it's 66.9 versus 57.2. On JobBench, 64.9 versus 61.2. Some evaluations barely move: DeepSearchQA ties at 89.4, and xhigh actually edges max on Terminal-Bench 2.1, 89.2 to 88.8.
Artificial Analysis puts Muse Spark 1.3 max at 62 on its Intelligence Index, with the shipping xhigh version at 61. That 61 ties GPT-5.6 Sol max, Grok 4.6 high, and Claude Opus 5 high. Anthropic still sits at the top of the leaderboard, though: Claude Fable 5.1 hits 66 at max and 65 at xhigh, while Claude Opus 5 reaches 63 at both settings.
So xhigh is legitimately in the frontier cluster. It just isn't setting the frontier.
That said, the jump from Muse Spark 1.2 is real. VentureBeat's coverage of last month's 1.2 launch found Meta fielding a credible coding contender that still generally trailed Anthropic's best model, scoring 82.9% on Terminal-Bench 2.1 against Opus 5's 86.7%. With 1.3, Meta is trading wins with OpenAI and Anthropic on several coding and agentic evaluations instead of just showing up to the contest.
Meta says the model itself has gotten easier to run in production, not just smarter. Muse Spark 1.3 is trained to hold multiple workflows in a single long thread, gather context using tools, flag gaps in its own plans, ask clarifying questions when needed, and confirm before taking consequential actions. Internally, Meta engineers measured roughly 20% fewer tool calls and 25% fewer tokens than 1.2 during coding work.

For any enterprise running thousands or millions of agent loops a day, that kind of behavioral tightening can matter more than a leaderboard point. Fewer wasted tool calls and tokens compounds fast at scale.
Pricing didn't move, though. Meta kept Standard API pricing for 1.3 identical to 1.2: $1.25 per million input tokens, $4.25 per million output tokens, $0.15 per million cached input tokens. Against the current price-performance field, that's mid-pack, cheaper than Claude Opus 5's $30 blended rate but pricier than Gemini 3.7/3.8 Flash's promotional $4.50, or DeepSeek-V4-Flash off-peak at $0.88.
Artificial Analysis measures Muse Spark 1.3 xhigh at 235.2 output tokens per second and an estimated $0.55 per Intelligence Index task. At a score of 61, that's currently the lowest cost-per-task of any model at that intelligence tier, which is presumably what Zuckerberg means by "almost too cheap to meter." It's not about the per-token rate. It's about throughput of useful work per dollar.
Here's the wrinkle: Muse Spark 1.2 cost only $0.40 per task at a score of 57. So even with flat token pricing, the average cost per completed task went up generation over generation. Artificial Analysis attributes that to heavier input-token consumption on agentic evaluations. That doesn't contradict Meta's 25%-fewer-tokens claim, since Meta is measuring internal coding workflows while Artificial Analysis is measuring a broader mix of reasoning and agentic tasks. But it's a good reminder that "cheap" gets slippery fast once a model is chaining tool calls and retries across long agent loops. Token rate is just one input into real cost.
Meta also kept its aggressive Contributor tier alive: $0.10 per million input tokens, $0.20 per million output tokens, in exchange for letting Meta use your prompts and completions for training. Fine for prototyping. A much harder sell for enterprises running proprietary code or sensitive internal data through it.
Meta chief AI officer Alexandr Wang was less measured than Zuckerberg. After Artificial Analysis posted its numbers, Wang reposted them on X with: "i really hate to say it, but... gemini who?" The timing was pointed. Google released Gemini 3.8 Flash the same day, pitched at nearly the identical workload: long-horizon software engineering, autonomous agents, multi-step professional reasoning. Google calls it its best reasoning and coding Flash model yet, and its third Flash release in six weeks.
The independent numbers back Wang up, mildly. Artificial Analysis scores Muse Spark 1.3 xhigh at 61 on the Intelligence Index for $0.55 per task, versus Gemini 3.8 Flash high at 59 for $0.58. Meta edges Google on both intelligence and cost at those specific settings. Google wins on speed by a wide margin, though: roughly 305 output tokens per second versus Meta's 235, about 30% faster. Google's introductory pricing is also lower for now, $0.75 input and $3.75 output per million tokens versus Meta's $1.25 and $4.25.
Muse Spark 1.3 is a genuine step forward for Meta, closing much of the gap with Anthropic and OpenAI on coding and agentic benchmarks while trimming tool calls and token use in real workflows. But the "frontier" framing rests heavily on a max configuration that isn't broadly available yet, still in safety testing with no listed API provider. The model developers can actually use today, xhigh, is solidly in the frontier cluster but not leading it. Enterprises evaluating Muse Spark 1.3 should benchmark against the shipping xhigh configuration, not the marketing materials built around max, and factor in that per-task cost, not per-token pricing, is the number that actually predicts your bill.
Tags
Original Sources
Meta says Muse Spark 1.3 has frontier performance — but its best results come from a model developers can’t broadly use yet
↗ https://venturebeat.com/technology/meta-says-muse-spark-1-3-has-frontier-performance-but-its-best-results-come-from-a-model-developers-cant-broadly-use-yet
About the author
Kai built ML infrastructure at a Bay Area startup before developing an obsession with transformer architectures and inference optimisation that eventually pulled him out of product work entirely. A stint at a compute research lab sharpened his instinct for what actually matters in a model release versus what is marketing. He writes from the inside — from the perspective of someone who has debugged the systems he is describing at three in the morning. He is allergic to hype and instinctively drawn to the unglamorous plumbing questions that everyone else skips over.
More from The Engineer →This Week's Edition
8 September 2026
41 articles
Related Articles
Related Articles
More Stories
© 2026 Cedar & Bloom. All rights reserved.