Model results · Meta

Muse Spark 1.2

Report19
Muse Spark 1.2 across the effort knob
Meta’s flagship runs the steepest backward reasoning dial measured on v3: 87.0% at low, 52.2% at xhigh, −34.8 points. Higher effort converts wrong answers into wall-clock timeouts under the same budgets every model gets, and ten of the fifteen timed-out runs never wrote a line.
2026-08-25 · 23 tasks · 69 runs · v3 suite

How would Muse Spark 1.2 do on your stack?

Public results answer "which model is stronger on this suite." Whether Muse Spark 1.2 belongs in your routing table depends on your languages, your codebases, and your task mix: that gets measured, not guessed. On this suite the difference between its best and worst effort setting is 35 points, so the integration detail matters as much as the model choice. Grading rules are in the methodology.