Public results answer "which model is stronger on this suite." Whether Muse Spark 1.2 belongs in your routing table depends on your languages, your codebases, and your task mix: that gets measured, not guessed. On this suite the difference between its best and worst effort setting is 35 points, so the integration detail matters as much as the model choice. Grading rules are in the methodology.