Every measured run, newest first. Each links to its model card, findings, and report.
Looking for the ranking?
The combined leaderboard is offline while a new eval suite is built. Eval Suite 3 is frozen, and the reports below
carry the per-task detail, cost, speed, and effort curves for every model that ran on it. A new board goes up when
the new suite has enough measured columns to be worth ranking.