VulcanBench Frontier v4 · current suite

Leaderboard

Updated 2026-10-04 · 23 tasks · 11 models · 53 model×effort columns · 1209 runs · v4 suite

VulcanBench Frontier v4

11 models, 53 model×effort columns, 1,209 runs. Combined score is 50% functional correctness, 8.5% lint and complexity, 8.5% security and 33% Code quality, judged for a human reader by Muse Spark 1.3 and Grok 4.6 under one frozen protocol (v3.4 to v3.7, v3.15 to v3.18 and v3.20 apply the same rubric, controls and gates to each population; v3.20 seats GPT-6.1 Sol in place of Grok 4.6 to judge Grok 4.7, see ◊). The chart plots combined score against cost per task, $0 on the left, one line per model from Low to Max; the cost axis shows up to $5 per task and scrolls sideways for anything costlier. The table below carries every column. Completion tokens are the model's own output per task, reasoning included. $/task is API-equivalent at list rates from the solver receipts; every model here ran on a subscription. Grok 4.7 has no list price, so its $/task is unavailable and it appears on the Tokens and Minutes views only.

VulcanBench Frontier v4 board: every model and effort level, ranked by combined score
#Model / harnessEffortCombinedSECode qualityPassedMin/task$/task
1Grok 4.7 Cursorbestextra-high◊93.150.3284.3523/2328.5unavailable
2Grok 4.7 Cursorhigh◊92.710.4183.4322/2325.4unavailable
3Grok 4.7 Cursormedium◊92.300.8683.9021/2327.2unavailable
4Fable 5.1 Claude Codebestmax†91.840.4782.4323/2327.1$9.06
5Opus 5.5 Claude Codebesthigh91.110.5777.9922/2221.4$3.27
6Opus 5.5 Claude Codemedium90.860.6477.3223/2318.6$2.84
7Fable 5.1 Claude Codeextra-high†90.750.8281.9322/2338.8$14.49
8Opus 5.5 Claude Codeextra-high90.681.1079.9822/2320.7$4.11
9Opus 5.5 Claude Codemax90.221.0279.7721/2337.5$8.75
10Fable 5.1 Claude Codemedium†90.130.9480.1420/2331.7$9.19
11Fable 5.1 Claude Codehigh†90.070.7880.8620/2328.8$9.70
12Grok 4.7 Cursorlow◊89.421.5581.4118/2320.2unavailable
13GPT-6 Astra Codexbestmax89.300.3776.2623/2310.3$2.57
14GPT-6 Astra Codexextra-high89.160.3075.4723/238.1$2.30
15Fable 5.1 Claude Codelow†89.121.1578.3619/2326.9$7.96
16GPT-5.6 Terra Codexbestmax‡89.050.7172.5223/2319.7$1.88
17GPT-6.1 Sol Codexbestextra-high88.780.4074.9923/2315.4$0.43
18GPT-6.1 Sol Codexmax88.350.3373.8023/2314.5$0.46
19GPT-6.1 Sol Codexhigh88.230.4073.6723/2310.2$0.33
20GPT-6 Astra Codexhigh88.160.4273.4823/234.9$1.71
21GPT-6 Astra Codexmedium87.730.3271.1923/233.8$1.48
22GPT-6 Astra Codexlow87.430.5170.3122/234.2$1.72
23GPT-5.6 Sol Codexbestmax§87.180.6267.5821/2211.6$1.77
24GPT-6 Sol Codexbestmax86.820.9170.7819/2324.5$2.52
25GPT-5.6 Sol Codexextra-high86.470.8067.3121/2310.6$1.61
26GPT-6.1 Sol Codexmedium*86.440.7870.2522/2310.6$0.40
27Opus 5.5 Claude Codelow86.361.3072.5317/2313.7$1.70
28GPT-6.1 Sol Codexlow86.220.8270.7920/236.9$0.31
29GPT-6 Sol Codexextra-high85.941.3870.7718/2320.9$2.00
30GPT-5.6 Sol Codexhigh85.760.9867.6520/2312.0$2.05
31GPT-5.6 Luna Codexbestmax84.151.3565.3519/2344.1$0.45
32GPT-5.6 Terra Codexextra-high83.462.7273.1314/2319.2$2.12
33GPT-6 Sol Codexhigh82.832.4370.0915/2323.8$2.33
34GPT-5.6 Sol Codexmedium82.561.2568.1912/2311.1$1.72
35GPT-6 Luna Codexmax¶81.422.8871.5112/2360.5$0.10
36GPT-6 Sol Codexmedium‖80.602.5966.7913/2317.5$1.76
37GPT-5.6 Luna Codexextra-high79.113.0768.0513/2325.4$0.27
38GPT-6 Luna Codexbestextra-high¶78.793.0271.4911/2352.6$0.16
39GPT-5.5 Codexbestextra-high78.462.5167.6711/2320.8$4.56
40GPT-5.6 Terra Codexhigh76.683.4773.9011/2312.6$1.18
41GPT-5.6 Luna Codexhigh71.223.7165.799/2321.0$0.22
42GPT-5.5 Codexhigh69.923.8162.937/2318.2$4.20
43GPT-5.6 Sol Codexlow69.053.5064.986/239.8$1.36
44GPT-6 Sol Codexlow67.623.5962.534/2313.0$1.02
45GPT-5.6 Terra Codexmedium66.353.8270.935/237.9$0.69
46GPT-5.5 Codexmedium63.543.8464.653/2312.4$2.71
47GPT-5.6 Terra Codexlow59.703.4372.312/236.7$0.58
48GPT-6 Luna Codexhigh57.353.8066.402/2315.6$0.05
49GPT-5.6 Luna Codexmedium53.282.9465.581/237.4$0.07
50GPT-5.5 Codexlow49.802.9361.751/238.2$1.79
51GPT-6 Luna Codexmedium45.721.3269.330/233.6$0.01
52GPT-5.6 Luna Codexlow41.290.9671.730/232.7$0.02
53GPT-6 Luna Codexlow40.832.0465.840/235.6$0.03

† Fable 5.1 runs include 11 disclosed Opus 4.8 fallbacks across the sweep; they stay in the population. ‡ GPT-5.6 Terra at max includes paddockcore, run on September 17 on a second ChatGPT account after the first hit its quota window and judged under the v3.6.1 top-up with the same judges and calibration. § GPT-5.6 Sol at max is judged on 22 of 23 tasks: on codeccore, Grok 4.6's intent probe quoted an excerpt absent from the code on both attempts, so the v3.7 protocol publishes no Code quality score for that run; the run passed its tests and is priced. ¶ GPT-6 Luna at extra-high and max is judged on 21 and 19 of 23 tasks: the other 2 and 4 runs hit the flat 3-hour task bound while still working, have no finished code to judge, count as failed tasks and are unpriced ($/task covers the finished runs; Min/task covers all 23). Counting each timeout as a combined score of 0 over all 23 runs gives 71.94 at extra-high and 67.26 at max; its best tag and effort suggestions use these figures. Ran on Codex CLI 0.155.0, the first release that serves GPT-6 Luna on a ChatGPT plan; the extra-high pacecore run was retried after an 88-minute Codex client stall, and the retry counts. ‖ GPT-6 Sol at medium is judged on 22 of 23 tasks: on codeccore, Grok 4.6's intent probe gave no valid answer (one malformed attempt, one quoting code absent from the run), so the v3.17 protocol publishes no Code quality score for that run; the run failed its tests, counts as a failed task in Passed (over all 23) and is priced. Ran on Codex CLI 0.155.0, the first release that serves GPT-6 Sol on a ChatGPT plan; no run reached the 3-hour bound. * GPT-6.1 Sol at medium includes paddockcore with Code quality from Muse Spark 1.3 alone: Grok 4.6's review of that run had no valid response (both attempts quoted a changed line), so the v3.18 protocol scores it from the one valid review, as it does for any such run. Every GPT-6.1 Sol cell is judged on 23 runs. Ran on Codex CLI 0.159.0, the first release that serves GPT-6.1 Sol on a ChatGPT plan; no run reached the 3-hour bound. ◊ Grok 4.7 (xAI) ran in Cursor's agent CLI 2026.10.01, which offers Low to Extra High and no Max. It is judged by Muse Spark 1.3 and GPT-6.1 Sol under v3.20, because Grok 4.6, the second judge for every other column, is not neutral for an xAI model. GPT-6.1 Sol rates Grok 4.7's code about 6 points above Muse does, so its Code quality and combined score are not strictly comparable with the other columns. Rescored from Muse Spark 1.3 alone for every column, Grok 4.7 scores 88.64 to 92.27: still first at Medium, High and Extra-high, and second at Low behind Fable 5.1 (89.46). Medium is judged on 22 of 23 tasks: lodgecore hit the 3-hour bound, counts as a failed task in Passed (over all 23) and in Min/task, and has no score. $/task is unavailable, not $0: VulcanBench has no list price for Grok 4.7 and the sweep ran on the Cursor subscription, so it is not on the cost chart and its effort suggestions are picked on time alone. SE is one task standard error of the combined score. Astra’s $/task is the central estimate; its report carries a long-context upper bound. The best tag marks each model’s highest-scoring effort level. Per-run records, judge sub-scores and pricing are in each report’s evidence bundle: Astra vs. Fable 5.1, GPT-5.5 vs. Luna, Terra, Sol, Opus 5.5, GPT-6 Luna, GPT-6 Sol, GPT-6.1 Sol, Grok 4.7. Download the board as CSV.

VulcanBench Routine v1

Frontier v4 asks which model and effort level can do hard work. Routine v1 asks the everyday question: what is the cheapest effort level that is enough for an ordinary ticket. It is 12 private routine tickets on small Python packages (a targeted bug fix, a small feature behind a flag, input validation, an edge case in dates, money or text), each admitted because a frontier model at its lowest effort finds it easy. 8 models, 38 model×effort columns, 456 runs, one attempt per task and level, each model through its own CLI on a subscription. The tasks stay private so they cannot leak into training data, so only per-model, per-effort aggregates are published.

Suggested effort for routine work, by the same rule as the Frontier board: the cheapest level, then the fastest, within 3 combined-score points of the model’s best.

VulcanBench Routine v1: every model and effort level on twelve routine tickets
Model / harnessEffortPassedCombinedSECode qualitySec/task$/task
Fable 5.1 Claude Codesuggestedlow12/1294.240.7585.5040$0.22
Fable 5.1 Claude Codemedium12/1294.750.5186.9845$0.26
Fable 5.1 Claude Codehigh12/1295.720.5690.0253$0.32
Fable 5.1 Claude Codeextra-high12/1296.420.5492.19107$0.63
Fable 5.1 Claude Codemax12/1296.900.3293.84242$1.46
Opus 5.5 Claude Codesuggestedlow12/1294.220.7385.5032$0.072
Opus 5.5 Claude Codemedium12/1295.300.7788.7242$0.11
Opus 5.5 Claude Codehigh12/1295.900.6090.7151$0.14
Opus 5.5 Claude Codeextra-high12/1296.740.3493.3266$0.20
Opus 5.5 Claude Codemax12/1296.500.4592.53342$1.12
GPT-6 Astra Codexsuggestedlow12/1294.790.6687.7674$0.26
GPT-6 Astra Codexmedium12/1294.140.6885.8576$0.26
GPT-6 Astra Codexhigh12/1295.010.4488.2885$0.28
GPT-6 Astra Codexextra-high12/1295.180.5988.72121$0.42
GPT-6 Astra Codexmax12/1294.740.4287.33177$0.56
GPT-5.6 Terra Codexsuggestedlow12/1293.130.8482.47102$0.074
GPT-5.6 Terra Codexmedium12/1294.000.7485.16112$0.076
GPT-5.6 Terra Codexhigh12/1294.320.8586.72126$0.083
GPT-5.6 Terra Codexextra-high12/1294.390.7786.37161$0.10
GPT-5.6 Terra Codexmax12/1294.520.5686.72236$0.13
GPT-5.6 Sol Codexsuggestedlow12/1294.390.7986.37118$0.13
GPT-5.6 Sol Codexmedium12/1294.180.6785.59137$0.15
GPT-5.6 Sol Codexhigh12/1294.550.6086.72161$0.18
GPT-5.6 Sol Codexextra-high12/1294.130.7785.85216$0.18
GPT-5.6 Sol Codexmax12/1294.510.6686.81226$0.21
GPT-5.6 Luna Codexsuggestedlow12/1292.500.8680.8260$0.006
GPT-5.6 Luna Codexmedium12/1292.031.0379.6090$0.008
GPT-5.6 Luna Codexhigh12/1292.860.9482.03110$0.009
GPT-5.6 Luna Codexextra-high12/1292.670.8181.42146$0.010
GPT-5.6 Luna Codexmax12/1293.080.9082.38175$0.013
GPT-5.5 Codexsuggestedlow12/1294.860.7587.7676$0.21
GPT-5.5 Codexmedium12/1294.860.6287.8584$0.22
GPT-5.5 Codexhigh12/1295.080.6488.5490$0.23
GPT-5.5 Codexextra-high12/1295.240.6388.89110$0.27
Grok 4.7 Cursorsuggestedlow12/1297.290.5894.6266unavailable
Grok 4.7 Cursormedium12/1297.210.5594.27109unavailable
Grok 4.7 Cursorhigh12/1297.370.4294.62189unavailable
Grok 4.7 Cursorextra-high12/1297.140.5593.92241unavailable

Combined score uses the Frontier weights: 50% functional correctness from hidden tests, 8.5% lint and complexity, 8.5% security and 33% Code quality, judged by Muse Spark 1.3 and Grok 4.6 under Code quality protocol v3.8 (Opus 5.5 under v3.14, the same protocol on its own population, judged in a separate session) with the same rubric, controls, gates and calibration exam as Frontier v4. Grok 4.7 is judged under v3.22 by Muse Spark 1.3 alone: Grok 4.6 is not neutral for an xAI model, and GPT-6.1 Sol, seated in its place, failed that protocol’s calibration exam on two gates, so the single-panel rule publishes from Muse. Muse rates code above Grok 4.6, so rescoring every other column from Muse alone puts their lowest tested levels at 93.1 to 95.6, against Grok 4.7’s 97.3 at Low: still first. Routine and Frontier Code quality are not comparable. On Frontier v4 part of Code quality measures whether a reviewer can recover each task’s deliberate legacy quirks; routine tickets have no such quirks by design, so the protocol’s own pre-registered rule scores Routine Code quality from the reviewed panel alone. Compare levels and models within this table, never across the two boards. SE is one task standard error of the combined score. Sec/task is mean wall clock. $/task is API-equivalent at list rates from the solver receipts, not a subscription bill; Grok 4.7 has no list price and ran on the Cursor subscription, so its cost is unavailable, not $0, and its suggestion is made on time alone. Download this table as CSV.

VulcanBench-SWE v3 (retired in August 2026)

The final SWE v3 board, fifteen models and 42 model×effort columns on 23 tasks from real merged open-source PRs ranked by pass@1, is frozen and published as machine-readable JSON. The numbered reports on the Benchmarks page carry its per-task detail. SWE v3 scores are not comparable with Frontier v4.

VulcanBench Frontier v4: 23 hard tasks that ask a model to rebuild a retired program whose real behaviour drifted from its written spec. Each column aggregates every run of that model at that effort level in its own harness. Combined score is 50% functional correctness from hidden tests, 8.5% lint and complexity, 8.5% security and 33% Code quality judged for a human reader by Muse Spark 1.3 and Grok 4.6 under one frozen protocol. Cost is API-equivalent at list rates from the solver receipts, not a subscription charge; runtime is sandbox wall-clock. How the suite is built and gated is on the methodology page, and each report on the Benchmarks page carries the per-run evidence.