Benchmarks

Muse Spark 1.3 sat above GPT-6 for two days. Here's why

Muse Spark 1.3 scored 62 on Artificial Analysis, above GPT-6 Astra, then fell to ninth in a rescore. I traced where the score came from and what it costs.

On this page
  1. What happened between 2 and 4 September 2026
  2. What Meta shipped, and what it claimed
  3. Why did the index put Muse Spark next to Fable and Astra? Four effects
  4. What changed in version 4.2, and what it did not change
  5. The number that survived: price and speed
  6. What developers say after using it
  7. Where I would use it

Meta released Muse Spark 1.3 on 2 September 2026, and Artificial Analysis scored its top setting at 62 on the Intelligence Index, third behind Claude Fable 5.1 and Opus 5 and one point above the GPT-6 Astra that OpenAI shipped the next day. [1] [4] On 4 September the index moved to version 4.2 and Muse Spark dropped to 53 and ninth place. [2] [3] The score was real, and so was the drop. Four things explain both, and none of them is cheating.

What happened between 2 and 4 September 2026

Muse Spark 1.3 held a top-three position on the most quoted independent index for about 48 hours, then the index itself changed. The sequence matters more than any single number, so here it is in order.

  1. Muse Spark 1.0

    First model from Meta Superintelligence Labs, closed weights.

  2. Muse Spark 1.2 and Muse Code

    Meta's own coding agent ships with the model.

  3. Muse Spark 1.3

    Artificial Analysis scores the max setting 62, third of all models.

  4. GPT-6 Astra

    OpenAI's flagship enters the same index at 61.

  5. Index v4.2

    Rescore: Muse Spark 1.3 at 53, ninth. The max setting goes public the same day.

Figure 1. The Muse Spark releases and the two days that produced the confusion.

Figure 1 shows why so many people saw contradictory numbers in the same week. On 2 September, Artificial Analysis wrote that Muse Spark 1.3 at max “scores 62 on the Artificial Analysis Intelligence Index, behind only Claude Fable 5.1 and Claude Opus 5”, with the generally available xhigh setting at 61. [1] When GPT-6 Astra arrived on 3 September, the same index gave it 61, and screenshots of Meta above OpenAI went around within hours. [4] On 4 September Artificial Analysis published version 4.2 of the index. On that version Muse Spark 1.3 at max scores 53 and ranks ninth of 202 models, Fable 5.1 scores 57, Astra 55, Opus 5 54, and GPT-5.6 Sol 51. [2] [3] Both sets of numbers are correct for the version they belong to, and quoting one without the version is how the argument started.

What Meta shipped, and what it claimed

Muse Spark 1.3 is a closed-weights model from Meta Superintelligence Labs, sold through Meta’s own API and its Muse Code agent at $1.25 per million input tokens and $4.25 per million output tokens, with a contributor tier at $0.10 and $0.20 for developers who let Meta train on their prompts. [5] [7] Meta’s claim was frontier performance at a fraction of the price. Its own comparison table supports part of that and quietly contradicts the rest.

The background matters for how people read the numbers. Meta introduced Muse Spark in April 2026 as the first model in a new series from Meta Superintelligence Labs, the lab it formed in mid 2025 around Alexandr Wang after the Llama 4 generation disappointed. [8] [27] Llama has not been formally discontinued, but there is no migration path from it, because Muse Spark is proprietary and served only by Meta. [9] Mark Zuckerberg has promised open weights for Muse Spark 1.2, and in August Meta released a separate 30 billion parameter model, Muse Glimmer, under an open license, but as of 5 September no Muse Spark weights are public. [10] [11]

The launch claims were large. Zuckerberg said Muse Spark 1.3 was “rolling out today with frontier performance almost too cheap to meter”, and Meta’s table compares it against GPT-5.6 Sol and Claude Opus 5, both at their maximum reasoning setting. [6] [11] On that table Muse Spark wins DeepSWE at 75.4% against 73.0% for Sol and 74.0% for Opus, ties Sol on Terminal-Bench 2.1 at 88.8%, and posts 98.1% on a long-context recall test between 512K and 1M tokens where Sol manages 73.8%. Opus 5 wins most of the agent rows, including GDPval-AA at 1824 Elo against 1754, OSWorld computer use, and AutomationBench. [6] So Meta’s own numbers describe a model that leads on coding and long context and trails Anthropic on agent work, which is a narrower claim than the launch language.

Two choices in Meta’s methodology deserve a mention because they shape every row. Meta reports its own model at max but its previous version, Muse Spark 1.2, at xhigh, so part of the version-over-version gain is a settings change. And for competitors it reports “the highest comparable primary-metric value available” from its own runs, the official leaderboard, or the vendor’s self-reported figure, whichever is best. [12] The table also compares against no model newer than late July: neither Claude Fable 5.1, released 1 September, nor GPT-6 Astra appears in it. [6]

Why did the index put Muse Spark next to Fable and Astra? Four effects

The 62 came from a real evaluation, run by a third party, on a model that is genuinely good at some things. It also came from a specific index version, a specific effort setting, a specific harness, and a set of comparisons Meta chose. Each of those pushed the number up, and each is documented.

Measure Muse Spark 1.3Fable 5.1Opus 5GPT-6 AstraGPT-5.6 Sol
Intelligence Index v4.1.1, 2 Sep 62 66 (Best value in this row) 63 61 61
Intelligence Index v4.2, 5 Sep 53 57 (Best value in this row) 54 55 51
Cost per index task, $ 0.95 (Best value in this row) 6.12 4.21 2.57 1.25
Time to first token, s 18.7 (Best value in this row) 266.5 91.2 463.7 163.0
Output speed, tokens/s 190.1 (Best value in this row) 68.7 59.1 87.5 85.4
Figure 2. The same five models on both index versions, with the cost and speed figures that did not change. Artificial Analysis, 2 to 5 September 2026.

Figure 2 is the map for the rest of this section. The two index rows show the drop, and the three rows under them show what stayed true through it. All five models are at their maximum effort setting, and the Muse Spark column is the max variant. [1] [3] [4]

The index rewards agent tasks, and that is where Meta improved. Artificial Analysis builds the Intelligence Index from four categories, with agents at 30% and general reasoning at 30%, coding and science at 20% each, and it says outright that “the weighting emphasizes agentic tasks.” [13] Muse Spark 1.3’s gain over 1.2 sits almost entirely in that agent slice: its score on a banking customer-service benchmark went from 35% to 52%, which made it first among all models, and its rating on the GDPval professional-work test rose by about 139 Elo. Meanwhile it lost four points on long-context reasoning and one to three on the knowledge test. [1] Artificial Analysis also explains how the max setting earns its extra points: by “using more turns and total reasoning tokens, reasoning 62% more on GDPval-AA v2 and 28% more on Tau3-Bench Banking” than xhigh. [1] A composite can rise on two rows while the rows a developer cares about fall.

The 62 was measured on a setting nobody could buy. The max variant was in limited preview for Meta’s partners at launch, and VentureBeat reported that Meta said it was “still completing additional safety testing” and would arrive shortly. It went public on 4 September. [11] The setting developers could actually call, xhigh, scored 61 on the old index and 52 on the new one. [1] [3] That is not unusual, since Anthropic and OpenAI also publish their best-of-any-effort numbers, but it means that “Muse Spark beats Astra” described a comparison between a preview build and a shipping product.

The coding scores come from Meta’s own harness. Artificial Analysis runs coding-agent tests inside each vendor’s agent where one exists, so Muse Spark runs inside Muse Code, and Meta says the model was trained across its own harnesses. [5] [18] Vals AI runs every model on the same Terminal-Bench 2.1 harness, and there Muse Spark 1.3 scores 72.28% and ranks sixteenth of 62, while Astra, Sol, Fable 5.1 and Opus 5 all land between 85% and 87%. [14] Meta’s own table reports 88.8% on the same benchmark. [6] The gap is a harness effect, and Cline demonstrated it in August: the team copied the instructions from Meta’s Muse Code system prompt into its own agent and ran the same Muse Spark 1.2 model on the same task. Tokens fell from 19.7 million to 7.2 million, the run finished in 24 minutes instead of 49, and the cost fell from $7.69 to $3.25 with nothing but the prompting changed. [15]

Meta’s comparisons were chosen to be favorable. The max versus xhigh mismatch and the best-available rule are described above. One row shows what that does in practice. Meta’s table gives Muse Spark 59.4 on Scale’s codebase-question benchmark, ahead of Opus 5 at 52.7 and Sol at 53.5. Scale’s own leaderboard, measured on a different metric, has Opus 5 at 63.17 and Sol at 46.00, and does not list Muse Spark 1.3 at all. [6] [16] Nobody falsified anything, but the table compares Meta’s own run against figures taken from other vendors’ documents, and it presents the result as a win.

What changed in version 4.2, and what it did not change

The 4 September rescore doubled the weight of private test sets to 40% “to prevent gaming”, removed GPQA Diamond because it “has now been saturated”, and added two new evaluations. [2] Muse Spark 1.3 had scored 94% on GPQA. Under the new weighting it fell from third to ninth, and every model’s score compressed, so the gap to Opus 5 narrowed to one point even as Astra moved ahead. [2] [3]

Artificial Analysis names no lab in that announcement, and I am not going to name one for it. Part of the rank change is simply that GPT-6 Astra entered the board on 3 September. But the timing is what it is: a model that leaned on a saturated public benchmark and on agent rows that were partly public lost ground the moment the index leaned on private data. The new private benchmark, a business-document test called AA-Briefcase, places Muse Spark fourth, behind Fable 5.1, Opus 5 and Astra, so it did not collapse on held-out work. [3] It just stopped being third.

A second evaluator says roughly the same thing without any version change. Vals AI’s own index, which weights tasks by economic value and runs everything on one harness, puts Muse Spark 1.3 eighth of 54 at $2.90 per test, with Fable 5.1 first at $28.92. [17] Two independent methods now agree on about ninth place. My read is that ninth is where the model was all along, and the index version 4.1.1 result was a two-day artifact of weighting and effort settings rather than a two-day period in which Meta led OpenAI.

The number that survived: price and speed

Every rescore left one claim standing. Muse Spark 1.3 is the cheapest model near the top of the board by a wide margin, and it is the fastest to answer. That is the honest version of “frontier performance almost too cheap to meter.”

per Intelligence Index task
$0.95
Fable 5.1 costs $6.12, Astra $2.57
to first token
18.7 s
Astra waits 464 s, Fable 5.1 267 s
output speed
190 tok/s
about twice Astra, three times Opus 5
Figure 3. Cost and speed at maximum effort. Artificial Analysis model pages, 5 September 2026.

Figure 3 puts the price advantage in the same units as the intelligence score. On the current index Muse Spark 1.3 scores four points below Fable 5.1 at about one sixth of the cost per task, and two points below Astra at little more than a third. [3] The speed numbers are the ones users notice first: a first token in under 19 seconds where Astra at maximum effort takes almost eight minutes. [3] [4] Artificial Analysis said on launch day that no model scoring 59 or above cost less per task, and the version 4.2 announcement still lists Meta among the four labs on its cost-per-task frontier. [1] [2]

The coding comparison follows the same pattern. On the Coding Agent Index, Muse Spark 1.3 inside Muse Code scores 64 at $1.62 per task and about 13 minutes, against Fable 5.1 inside Claude Code at 70 for $9.18 and 24 minutes, and Astra inside Codex at 67 for $4.72. [18] Six points behind the leader for less than a fifth of the money is a real position, even with the harness caveat above attached.

The contributor tier is the part I would read twice. At $0.10 in and $0.20 out it is roughly twelve times cheaper than the standard tier, and the condition is that Meta may use your prompts and completions to train future models. [7] OpenCode made the contributor model free for a period after launch, and the r/opencode thread announcing it called it the most capable free model available. [22] For a hobby project that is a bargain. For a repository that contains a client’s code it is a data-handling decision, not a pricing one.

What developers say after using it

The people who actually ran Muse Spark 1.3 describe a fast, cheap and obedient workhorse that does not feel like a top model, and that description fits the index better than either the launch claims or the dismissals. The dismissals were louder, though.

The most upvoted verdict came in r/codex on 3 September, in a thread pointing out that Artificial Analysis had Astra level with Sol and below Muse Spark. The reply, at 204 points, was blunt: “This score doesn’t mean much anymore. If you have used Muse Spark 1.3, you’d know it’s not even close to being a top 10 model.” [20] The best hands-on account came from r/singularity, where one user posted four times across an evening. Before testing, they were “having a hard time believing this.” An hour in: “feels mechanical (excellent for work, idk for strategy) and very capable. Seems very fast.” Their final read was that it “raises the bar on Grok 4.6; really smart but doesn’t feel deep. Very very capable of work.” [21] On Hacker News, a developer who had been using the 1.2 contributor tier for development wrote that it was “not a frontier model by any means” but “felt like it knew its weaknesses and didn’t try to impose its opinions on me”. Another summarized the launch better than most press did: “You shouldn’t think ‘wow, Facebook has caught up’. You should think, ‘wow, Facebook is less than 6 months behind the frontier’.” [19]

The accusations of cheating are a different matter, because I found no evidence behind them. The top comment in one Reddit thread was a link titled “Pretraining on the Test Set Is All You Need”, and “benchmaxxed” appears in nearly every thread, but nobody produced a contamination finding against any Muse model, and Artificial Analysis added reward-hacking detection to its coding index without flagging one. [29] The one specific methodology complaint is narrower. In July, a Hacker News commenter who described themselves as a former Meta employee pointed out that Meta’s Muse Spark 1.1 evaluation report ran Terminal-Bench tasks with 6 CPU cores and 8 GB of RAM, above the caps most of those tasks specify, and argued that this explained why the model was absent from the official leaderboard. Meta’s report does say exactly that, and the complaint was contested in the same thread by someone who said the limits are recommendations that do not shift results. [25] [26] I read it as a real question about one benchmark, not as proof of anything wider.

The distrust has a source, and it is not 2026. In April 2025 Meta submitted an experimental chat-tuned version of Llama 4 Maverick to the LMArena leaderboard, where it ranked second, while the version it shipped to developers later placed thirty-second. LMArena said Meta’s interpretation of its policy “did not match what we expect from model providers”, and a later paper counted 27 private Meta variants tested on the arena before the Llama 4 release. [23] [24] Nothing like that has surfaced for Muse Spark, and Meta now publishes a methodology document with each release. [12] But when a lab with that history posts a number above OpenAI’s new flagship, the internet reaches for the old explanation first. My read is that the 2025 episode is why the 2026 skepticism was instant, and the four effects above are why it was partly right for the wrong reasons.

Where I would use it

I would use Muse Spark 1.3 for high-volume agent work where speed and price matter more than the last few points of judgment, and I would keep Fable 5.1, Opus 5 or Astra for anything long-horizon or anything I could not afford to redo. The scores support that split better than they support either “Meta caught up” or “Meta gamed it.”

TaskFitWhy
Bulk agent runs: classification, extraction, routine ticketsGoodAbout $0.95 per index task and a first token in 19 seconds
Iterating quickly inside Muse CodeGoodThe harness it was trained with, where its coding scores are measured
Long-horizon coding across a large repositoryWeakSixteenth on a neutral terminal harness, behind on the private agent test
Anything with client code or sensitive dataStandard tier onlyThe contributor discount is paid for with training rights
Replacing a frontier model because of the indexNot yetNinth on two independent indexes once effort and version are held constant

Two practical notes. Ask which effort setting a number was measured at before you trust it, because max and xhigh are anywhere from one to ten points apart on this model depending on the test. [3] [11] And treat every index score as carrying a version number, because the one you read in a screenshot may already have been replaced. Zvi Mowshowitz’s launch-week comment was that Grok 4.6 also scored 61 once and was mostly never heard from again, and that the sensible default was to keep an eye on Muse Spark and assume it was not good enough yet. [28] After three days of reading the numbers behind the numbers, that is roughly where I landed too, with one addition: at this price, “not good enough yet” is still worth a slot in the toolbox.

Sources

  1. Muse Spark 1.3: Meta reaches the frontierArtificial Analysis · 2026-09-02
  2. Announcing Artificial Analysis Intelligence Index v4.2Artificial Analysis · 2026-09-04
  3. Muse Spark 1.3 (max) model analysisArtificial Analysis · 2026-09-05
  4. Benchmarking GPT-6 AstraArtificial Analysis · 2026-09-03
  5. Introducing Muse Spark 1.3Meta · 2026-09-02
  6. Muse Spark model pageMeta Developers · 2026-09-05
  7. Pricing and rate limitsMeta Developers · 2026-09-05
  8. Introducing Muse Spark: Scaling towards personal superintelligenceMeta · 2026-04-08
  9. Meta abandons open-source Llama for proprietary Muse SparkThe New Stack · 2026-04-30
  10. Meta to open source its most powerful AI model as it takes swipe at OpenAI, AnthropicCNBC · 2026-08-10
  11. Meta says Muse Spark 1.3 has frontier performance, but its best results come from a model developers can't broadly use yetVentureBeat · 2026-09-03
  12. Muse Spark 1.3 evaluation methodologyMeta · 2026-09-02
  13. Intelligence benchmarking methodologyArtificial Analysis · 2026-09-04
  14. Terminal-Bench 2.1 leaderboardVals AI · 2026-09-04
  15. On running Muse Spark 1.2 with Muse Code prompts inside ClineX, Cline · 2026-08-06
  16. SWE Atlas: Codebase QnAScale Labs · 2026-09-05
  17. Vals IndexVals AI · 2026-09-04
  18. Coding Agent IndexArtificial Analysis · 2026-09-05
  19. Muse Spark 1.3 discussionHacker News · 2026-09-02
  20. Astra is good, but maybe we should calm down with the hypeReddit, r/codex · 2026-09-03
  21. Meta's muse spark 1.3 surpassed fable 5 and GPT 5.6 solReddit, r/singularity · 2026-09-02
  22. Meta Muse Spark 1.3 is FREE on OpenCode ZenReddit, r/opencode · 2026-09-03
  23. Meta's benchmarks for its new AI models are a bit misleadingTechCrunch · 2025-04-06
  24. The Leaderboard IllusionarXiv, Cohere Labs et al. · 2025-04-29
  25. On the Terminal-Bench resource caps in the Muse Spark 1.1 evaluation reportHacker News · 2026-07-09
  26. Muse Spark 1.1 evaluation reportMeta Superintelligence Labs · 2026-07-09
  27. The future of Meta Superintelligence: a 1 year progress updateSemiAnalysis · 2026-07-09
  28. AI #184: Post Post MortemZvi Mowshowitz · 2026-09-03
  29. Coding agents benchmarking methodologyArtificial Analysis · 2026-09-05