Open-source coding benchmark

Benchmarking models
on real engineering work.

Real engineering tasks, from rebuilding retired legacy binaries to merged open-source pull requests, graded by deterministic hidden tests.

vulcanbench: suite: v1-micro · sandbox: docker · network: off

The harness, live: an actual v1-micro run on Claude Sonnet 5 (low): 26 sandboxed tasks, graded by hidden tests, 25 pass at $0.82 total.

Benchmarking what matters to engineering teams: accuracy, token use, time, and cost, all on real engineering tasks, no puzzles, no math problems.

Hi I'm Morgan, and I run the open source benchmarking tool and lab that is VulcanBench. My motivation behind building this developed organically through my own need, as a Founder and CTO, to help my team leverage the best models for the work that we do.

The core problem I ran into with agentic coding benchmarks is that a good chunk of the tasks in the evals are things like puzzles and math problems, that don't represent the kind of work engineering teams do every day. At the same time, benchmarks that I like, don't show how models compared across token use and time, which is critical for engineering leaders to know.

So, I started testing models myself, on real engineering tasks, and looking at not just task completion scores, but also at token use, cost, and time to complete tasks. If I am comparing two models, and one is 5% better than the other, but that model also has an overthinking problem and uses 5x the number of tokens and takes much longer to solve problems, that would be important to know, right?

I also learned through my own testing that two models that have different scores on a popular benchmark, might actually produce the same accuracy on regular every day engineering tasks.

And what started as a bunch of Python scripts I was piecing together to test different models for my team, turned into fully open source benchmarking software, and open, transparent evals.

The harness and published task suites are open source. Each report links its evidence and explains what can be independently checked, what is withheld, and what would be needed to reproduce a run. Alongside scores, I report token use, estimated API cost and time to complete tasks.

My goal is simple: let's build better benchmarks that represent the real work engineering teams do, and that go beyond scores, showing things like token use, cost, and time to complete a task.

Live long and benchmark 🖖

What every comparison makes clear.

01

Suite-specific scoring

Functional outcomes come from deterministic tests. SWE v4 combined scores also include lint and complexity checks, a security scan and, from September 2026, 33% Code quality judged for a human reader by a calibrated panel from labs with no model on the board. Earlier reports retain their own scoring methods.

02

Disclosed evidence

Reports identify tasks, models, harnesses, fallbacks and audit limits. Post-cutoff dates and source hashes alone cannot prove training-data absence or rule out prohibited access.

03

Economics as a result

Cost, wall-clock time, agent steps, and tokens are recorded alongside accuracy. When correctness converges, efficiency is the signal.

Recent reports by suite.

Suitev4
GPT-5.6 Terra across every effort level
114 Codex runs at five effort levels with Code quality at 33% judged by Muse Spark 1.3 and Grok 4.6. The steepest effort curve on the SWE v4 board: 59.70 at Low to 89.46 at Max, where Terra passes every judged task, at $0.58 to $2.12 per task.
September 16, 2026 · 23 tasks · 114 runs · VulcanBench-SWE v4 · Code quality protocol v3.6
Suitev4
GPT-5.5 vs. GPT-5.6 Luna across every effort level
207 Codex runs at every effort level each API offers, with Code quality at 33% judged by Muse Spark 1.3 and Grok 4.6. GPT-5.5 leads at Low and Medium; Luna edges ahead from High up and its Max level is the top cell, at 17 to 75 times less per task.
September 15, 2026 · 23 tasks · 207 runs · VulcanBench-SWE v4 · Code quality protocol v3.5
Suitev4
GPT-6 Astra vs. Fable 5.1 under a neutral Code quality panel
230 runs from the September 2026 effort sweep, scored with Code quality at 33% and judged for a human reader by Muse Spark 1.3 and Grok 4.6. Fable leads combined score and Code quality at every effort; Astra is faster at every effort.
September 9, 2026 · 23 matched tasks · 230 runs · VulcanBench-SWE v4 · Code quality protocol v3.4
Report20
Muse Spark 1.2 in Pi vs. a bare-bones harness
Our first open-source-harness study: Pi beats the bare loop at every effort level, 4.3 to 17.4 points. And the integrity audit caught the model hunting the host for answer keys: 17 of 69 cells were replaced by kernel-confined, audited-clean reruns.
2026-08-28 · 23 tasks · 138 scored cells · v3 suite · Harness Study No. 04
Report19
Muse Spark 1.2 across the effort knob
Meta’s flagship runs the steepest backward reasoning dial measured on v3: 87.0% at low, 52.2% at xhigh, −34.8 points. Higher effort converts wrong answers into wall-clock timeouts under the same budgets every model gets.
2026-08-25 · 23 tasks · 69 runs · v3 suite
Report18
GLM 5.3 in ZCode vs. a bare-bones harness
Our first subscription-harness study: the same GLM 5.3 through Z.ai’s own ZCode harness and through a bare-bones API loop. The effort knob points opposite directions, and at max the harness is worth 21.8 points, 65.2% to 87.0%, with zero timed-out runs.
2026-08-24 · 23 tasks · 138 runs · v3 suite · Harness Study No. 03

Why this exists.

Model routing and eval suites built on your codebase.

A private eval suite built from your repositories and the pull requests your engineers actually ship, run across models and effort levels in the same harness as the public reports, and returned as a routing policy your team can deploy, with the measurements and cost model behind it.

Have your own model measured.

Have your own model, fine-tune, or agent scaffold measured on the same suite, in a private report graded to the same standard as everything published here.

VulcanBench is free. Sponsoring keeps it running.

Sponsorships run entirely through GitHub Sponsors, at whatever amount and cadence you choose. Sponsors are named on GitHub as a thank you.