Meta has officially pulled the curtain back on its latest artificial intelligence advancement, Muse Spark 1.3. Unveiled yesterday, the new model represents a significant evolution in the company’s AI strategy, offering faster processing speeds and superior performance on third-party benchmarks compared to its predecessor. However, the release comes with a distinct nuance regarding the availability of its most powerful configurations and the broader strategic direction of Meta’s model development.
"Muse Spark 1.3 is rolling out today with frontier performance almost too cheap to meter," Meta co-founder and CEO Mark Zuckerberg announced on X. In his statement, he characterized the update as Meta’s "biggest jump" yet in the fields of coding and agentic work—the latter referring to AI systems capable of executing multi-step tasks autonomously.
There is significant substance backing these claims. Muse Spark 1.3 delivers marked improvements over the 1.2 version released just last month, particularly in the realm of long-running agentic tasks. For developers and enterprises currently navigating the crowded landscape of LLM APIs, the version available today represents one of the most compelling price-performance offerings near the pinnacle of independent industry rankings.
The Max Configuration: A Future Promise
The most impressive benchmark results cited by Meta originate from its "max" reasoning configuration. According to the company, this version is currently undergoing final safety testing and is expected to be released "shortly." Independent benchmarking firm Artificial Analysis, which evaluated the max configuration during a limited partner preview, notes that it is not currently available through any public API provider.
For the time being, the version rolling out to the public via the Muse Code harness and the Meta Model API utilizes the company’s established reasoning settings, including the "xhigh" configuration. This reality shifts the conversation for enterprise leaders: the question is no longer whether Muse Spark 1.3 can reach the upper echelons of "frontier" performance, but rather how close the currently deployable version gets to that peak and at what tangible cost.
Benchmarking the Shipping Model
While Meta has been transparent in its technical documentation, disclosing results for both the "max" and "xhigh" configurations, its marketing materials have leaned heavily into the performance metrics of the unreleased max variant. This is a common practice in the industry, but it underscores the need for nuance when evaluating the model’s current capabilities.
In Meta’s internal evaluation report, the max configuration recorded a GDPval-AA v2 score of 1,754 Elo, compared to 1,709 for the xhigh version. Similarly, on the OSWorld 2.0 evaluation, the max variant achieved 66.9 against the xhigh’s 57.2. However, the gap is not uniform across all domains. On the DeepSearchQA benchmark, the two versions are tied at 89.4, and in a slight reversal, the xhigh configuration actually outperformed the max variant on Terminal-Bench 2.1, scoring 89.2 to 88.8.
Artificial Analysis currently places the Muse Spark 1.3 max configuration at 62 on its Intelligence Index, while the shipping xhigh version sits at 61. At this level, the xhigh model is competitive with other industry leaders, tying with GPT-5.6 Sol max, Grok 4.6 high, and Claude Opus 5 high. Despite these gains, Anthropic remains at the top of the leaderboard; Claude Fable 5.1 reaches a score of 66 at max and 65 at xhigh, while Claude Opus 5 maintains a score of 63 in both configurations.
These figures illustrate that while the shipping version of Muse Spark 1.3 is firmly within the "frontier cluster" of high-performing models, it is not currently the undisputed leader. Nevertheless, this represents a substantial leap from the 1.2 release. In previous evaluations, Meta’s offerings were viewed as credible challengers that generally trailed the top-tier models from Anthropic. With the 1.3 update, Meta is no longer just competing; it is trading wins with the industry’s most prominent players.
Beyond raw scores, Meta has focused on usability. The underlying model has been refined to better maintain complex workflows within long-running threads. It is now more adept at gathering context via tools, identifying gaps in its own logic, seeking user clarification, and verifying actions before committing to high-consequence tasks. In internal tests conducted by Meta engineers, the 1.3 model utilized approximately 20% fewer tool calls and 25% fewer tokens than the 1.2 version during coding tasks. For enterprises managing millions of agentic loops, these behavioral efficiencies are likely to provide more long-term value than minor fluctuations in leaderboard standings.
The Economics of "Cheap to Meter"
The phrase "almost too cheap to meter" used by Zuckerberg has sparked debate regarding the actual cost of deployment. Notably, Meta has not implemented an API price cut for Muse Spark 1.3. The standard pricing remains fixed at $1.25 per million input tokens, $4.25 per million output tokens, and $0.15 per million cached input tokens.

This indicates that Zuckerberg’s assertion is less about a reduction in per-token costs and more about the efficiency of the model’s output. By requiring fewer tokens and fewer tool calls to accomplish the same work, the effective cost per completed task can be lower, even if the price per token remains unchanged.
However, data from Artificial Analysis introduces a layer of complexity. The firm measured the Muse Spark 1.3 xhigh version at 235.2 output tokens per second, with an estimated cost of $0.55 per Intelligence Index task. At a score of 61, this model provides the lowest cost per task among currently measured models at that specific level of intelligence. Interestingly, despite the efficiency improvements, the cost of completing an average task actually rose compared to Muse Spark 1.2, which cost roughly $0.40 per task. Artificial Analysis attributes this increase to heavier input-token consumption during complex agentic evaluations.
This highlights the difficulty of predicting AI costs in an agentic world. Token rates, the duration of reasoning efforts, the number of turns, and the frequency of tool usage all coalesce to determine the final invoice. Additionally, Meta continues to offer a "Contributor" tier—priced at $0.10 per million input tokens and $0.20 per million output tokens—which is exceptionally low. However, this tier requires developers to grant Meta permission to use their prompts and completions for future model training. While potentially attractive for early-stage prototyping, this creates a significant data-governance challenge for enterprises handling sensitive intellectual property or proprietary code.
The Competitive Landscape
The launch of Muse Spark 1.3 was met with characteristic bravado from Meta’s chief AI officer, Alexandr Wang. After the benchmark results were published, Wang posted on X, "i really hate to say it, but… gemini who?"
The comment was pointed, as Google simultaneously released its Gemini 3.8 Flash, a model positioned for similar workloads: long-horizon software engineering, autonomous agents, and multi-step professional reasoning. While Wang’s remark was dismissive, the independent data suggests a very close race.
According to Artificial Analysis, the Muse Spark 1.3 xhigh version and Gemini 3.8 Flash high-reasoning mode are nearly neck-and-neck. Meta holds a slight edge in intelligence scores (61 vs. 59) and cost per task ($0.55 vs. $0.58). However, Google dominates in terms of raw throughput, with Gemini 3.8 Flash performing at approximately 305 output tokens per second, compared to 235 for Muse Spark—a 30% speed advantage. Furthermore, Google is currently offering promotional pricing that undercuts Meta’s standard rates until the end of 2026.
This snapshot highlights the narrow margins at the frontier of AI development. Meta is currently winning on independent intelligence measures, while Google is providing higher throughput and more aggressive short-term pricing. For architects designing enterprise AI systems, the choice between these models will depend on whether their priority is raw speed or the qualitative nuances of high-effort reasoning.
The Future of Open Weights
Perhaps the most significant ambiguity surrounding this release is the status of Meta’s commitment to open-weight models. When the company launched Muse Code and Muse Spark 1.2 in August, it signaled a departure from the open-weights strategy that had made the Llama series a foundational element of the open-source AI community. These were served exclusively as proprietary API products.
Meta briefly pivoted back to openness shortly after, releasing the Muse Glimmer model under an Apache 2.0 license and promising that open weights for Muse Spark 1.2 would arrive in "the coming weeks." However, with the release of Muse Spark 1.3 as a proprietary model, the timeline and the nature of those releases have become less certain.
The company’s latest communications no longer specify a version or a release date for the open weights, stating only that a "Muse Spark open weights release" is in the roadmap. Zuckerberg similarly echoed this on social media, using the plural "releases" to describe upcoming efforts without providing specific parameters.
For the developer community, which has long relied on Meta to provide downloadable weights for self-hosting and customization, this ambiguity is a point of concern. While Muse Spark 1.3 demonstrates that Meta has successfully accelerated its ability to iterate and deploy proprietary frontier models, the company now faces a different challenge: proving that it can provide a transparent and reliable roadmap for developers. The ultimate test for Meta will be whether it can bridge the gap between its proprietary, high-performance API offerings and the open-weights ecosystem it helped build, ensuring that enterprises can plan their long-term infrastructure with confidence.

