Meta closes in on the top with Muse Spark 1.3, and undercuts rivals on price
Key Points
Meta has released Muse Spark 1.3. The xhigh tier is available now, while the more powerful max version runs as a limited preview for now.
The model improves most on agentic tasks but still trails top performers like Claude Fable 5.1 across most benchmarks.
At $0.55 per task, it's currently the cheapest model in its performance class. Meta also says an open-weights version is coming.
Meta has released Muse Spark 1.3, its fourth model in five months. Independent testing shows solid gains on agentic tasks, but a gap to the top stays. What sells the model is the price.
Meta has released Muse Spark 1.3 through Muse Code and the Meta Model API. The series launched in April, with version 1.1 following in July and 1.2 in August. The xhigh tier is available now, while the more compute-heavy max tier arrives only after further safety testing and currently runs as a limited partner preview, according to Artificial Analysis.
At an unchanged $1.25 and $4.25 per million input and output tokens, one index task costs $0.55. No model scoring 59 points or higher is cheaper, and rivals at the same index level run between $0.94 and $1.23. Muse Spark 1.3 does cost more than version 1.2, which ran $0.40. Ad
The gains cluster where the index pays off
On the Intelligence Index , max scores 62 points and xhigh 61, up from 57 in August and 53 in July. The jump comes down to how the index weights its tests. GDPval-AA v2 counts for 20 percent, Terminal-Bench 2.1 for 16 percent, and τ³-Bench Banking for 14 percent, and Meta's biggest gains land in exactly those three tests. Ad
On τ³-Banking, where agents operate tools in a simulated banking scenario, max hits 52 percent. That's number one right now, according to Artificial Analysis, and it's the only outright lead the model holds. The available xhigh tier reaches 47 percent, tying Claude Fable 5.1 (max) and GLM-5.3-Flash rather than leading. The predecessor 1.2 sat at 35 percent.
Terminal-Bench 2.1, which tests coding in the terminal, climbs from 80 to 85 percent on xhigh and 86 on max, but Claude Fable 5.1 still holds the top spot at 91.4 percent in its max tier, 91.0 at xhigh, and 89.9 at high. On the index's highest-weighted test, GDPval-AA v2, Meta improves from 1,615 to 1,709 and 1,754 on a scale calibrated to human expert performance at 1,000 across 220 real-world professional tasks. Claude Fable 5.1 (max) sits at 1,853. Meta buys the max variant's edge with compute, burning 62 percent more reasoning tokens than xhigh. Ad
Outside the agentic tests, the model stays mid-pack
On GPQA Diamond, which poses expert-level science questions, Muse Spark rises from 90 to 94 percent. That's the top group, but still below Gemini 3.8 Flash (high) at 95.3 and Grok 4.6 (high) at 94.9. CritPt, covering research physics, jumps from 18 to 26 percent, well behind GPT-5.6 Sol (max) at 32.3 and Claude Fable 5.1 (xhigh) at 31.1 percent.
Two scores actually drop against 1.2. AA-LCR falls from 83 to 79 percent. Factual accuracy in AA-Omniscience slips by up to three points, because the model more often declines to answer when it's unsure. Ad
Neither Meta nor Artificial Analysis has named a price for the max variant yet. Larger models and an open-weights release are on the way. Ad
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.