Real engineering tasks, from rebuilding retired legacy binaries to merged open-source pull requests, graded by deterministic hidden tests.
The harness, live: an actual v1-micro run on Claude Sonnet 5 (low): 26 sandboxed tasks, graded by hidden tests, 25 pass at $0.82 total.
Hi I'm Morgan, and I run the open source benchmarking tool and lab that is VulcanBench. My motivation behind building this developed organically through my own need, as a Founder and CTO, to help my team leverage the best models for the work that we do.
The core problem I ran into with agentic coding benchmarks is that a good chunk of the tasks in the evals are things like puzzles and math problems, that don't represent the kind of work engineering teams do every day. At the same time, benchmarks that I like, don't show how models compared across token use and time, which is critical for engineering leaders to know.
So, I started testing models myself, on real engineering tasks, and looking at not just task completion scores, but also at token use, cost, and time to complete tasks. If I am comparing two models, and one is 5% better than the other, but that model also has an overthinking problem and uses 5x the number of tokens and takes much longer to solve problems, that would be important to know, right?
I also learned through my own testing that two models that have different scores on a popular benchmark, might actually produce the same accuracy on regular every day engineering tasks.
And what started as a bunch of Python scripts I was piecing together to test different models for my team, turned into fully open source benchmarking software, and open, transparent evals.
The harness and published task suites are open source. Each report links its evidence and explains what can be independently checked, what is withheld, and what would be needed to reproduce a run. Alongside scores, I report token use, estimated API cost and time to complete tasks.
My goal is simple: let's build better benchmarks that represent the real work engineering teams do, and that go beyond scores, showing things like token use, cost, and time to complete a task.
Live long and benchmark 🖖
Functional outcomes come from deterministic tests. SWE v4 combined scores also include lint and complexity checks, a security scan and, from September 2026, 33% Code quality judged for a human reader by a calibrated panel from labs with no model on the board. Earlier reports retain their own scoring methods.
Reports identify tasks, models, harnesses, fallbacks and audit limits. Post-cutoff dates and source hashes alone cannot prove training-data absence or rule out prohibited access.
Cost, wall-clock time, agent steps, and tokens are recorded alongside accuracy. When correctness converges, efficiency is the signal.
A private eval suite built from your repositories and the pull requests your engineers actually ship, run across models and effort levels in the same harness as the public reports, and returned as a routing policy your team can deploy, with the measurements and cost model behind it.
Have your own model, fine-tune, or agent scaffold measured on the same suite, in a private report graded to the same standard as everything published here.
Sponsorships run entirely through GitHub Sponsors, at whatever amount and cadence you choose. Sponsors are named on GitHub as a thank you.