Muse Spark 1.3 sat above GPT-6 for two days. Here's why
Muse Spark 1.3 scored 62 on Artificial Analysis, above GPT-6 Astra, then fell to ninth in a rescore. I traced where the score came from and what it costs.
benchmarks
New benchmark numbers read against hands-on experience: where they hold up and where they do not.
Muse Spark 1.3 scored 62 on Artificial Analysis, above GPT-6 Astra, then fell to ninth in a rescore. I traced where the score came from and what it costs.
DeepSWE puts Luna Max 2.2 points behind Sol High at roughly one-sixth the attempt cost. I explain why the models can still feel far apart in repository work.
I compared Sol, Terra, Opus 5 and Fable 5 across coding benchmarks. The winner changes with the task, effort setting, agent setup and budget.
Kimi K3 ties GPT-5.6 medium but takes 4.6 times as long. GLM-5.2 is cheap per token yet costly per task. I checked where both models still win.