I built VulcanBench to help engineering teams choose models based on the work they need to deliver. As Cofounder and CTO of Bold Metrics, I wanted to measure task correctness alongside code quality, security, time, token usage, and cost.
Each result identifies the model, coding harness, effort setting, and benchmark version, so teams can compare configurations they could actually use. Read why I started VulcanBench for the longer story.
VulcanBench uses two task formats. Tasks from merged pull requests start from a real open-source repository before a fix. Binary-parity tasks ask an agent to repair a replacement implementation until its behavior matches a compiled legacy program.
The agent receives a starting repository and an issue describing the required change. Reference solutions and hidden tests are held separately for grading. Historical suites remain available under their recorded versions; comparisons must identify the exact task set.
This suite focuses on behavioral reconstruction. Each task includes a compiled binary, a written specification that has drifted from its behavior, and a replacement implementation based on that specification. The agent can probe or disassemble the binary to recover the behavior needed to fix the replacement.
Difficulty comes from the reconstruction work: discovering interacting rules, testing hypotheses, and preserving existing behavior. Source-code repair can also be difficult; this format targets a different engineering challenge.
Candidates are evaluated with three attempts per reference model under their respective coding harnesses. Admission requires both conditions: GPT 5.6 Sol solves at most one of three attempts, and Claude Opus 5 either needs a median of at least ten minutes or fails at least one attempt. These are reference measurements at admission, not a guarantee of difficulty for later models.
Separately, the reference solution must score 1.0 and the unmodified base 0.0 in three repeated checks. Randomized input sweeps compare the reference solution with the binary. Task changes require a new version, preserving the earlier version for historical comparisons. Reference measurements and execution details are in the technical notes.
VulcanBench prepares the task, captures the resulting patch, and runs verification and scoring. Each report discloses the execution mode:
Results measure a model together with its harness, including its system prompt, tools, context management, and any product routing. Runtime comparisons also depend on hardware and concurrency. Host runs do not currently pin CPU and memory.
Effort settings map to the provider's supported controls. Reports record the requested setting and, where available, the setting reported by the CLI. Unsupported settings and unavailable confirmations are disclosed.
Effort labels are not calibrated across vendors. Within a sweep, the suite and configured task budgets stay fixed while effort varies. Higher effort can increase tokens and runtime; we measure both rather than assume they rise with every setting.
Frontier tasks have a ten-hour wall-clock limit. Time-based results such as pass@1 within ten or thirty minutes are calculated afterward from recorded durations, alongside the full-budget result. These reporting thresholds do not stop a run early.
New-behavior tests are validated to fail on the starting code and pass on the reference solution. Regression tests check that working behavior remains intact. A regression failure gates the functional score to zero.
Individual test groups can produce a partial functional score. A task counts as solved for pass@1 only when its functional score is 1.0. Subjective review cannot turn a failed task into a functional pass.
The default correctness grader is deterministic. Tasks may explicitly opt into an independent model grader for hidden acceptance criteria; those results must identify that exception and its validation. This is separate from the Code quality metric described below.
A partial suite must not be presented as a complete result. Retry counts, exclusions, and the final denominator belong beside the reported score.
Functional pass@1 measures how often a task is fully solved in one attempt. With one attempt per task, it is tasks solved divided by tasks scored. With repeated attempts, the harness averages each task's success rate across tasks, giving tasks equal weight.
The combined score measures functional correctness, lint and complexity, security, and Code quality. Passing every task does not imply a combined score of 100. Functional correctness retains half the weight; Code quality contributes a third.
Code quality assesses the code beyond test outcomes: whether a person who has never seen it could read it, learn the real contract from it, and change it safely. It supplements functional grading rather than changing which tasks count as solved.
Bars use a shared 0% to 100% scale. For each run, the combined score is
100 × (0.50F + 0.085Q + 0.085S + 0.33C), with each factor on a 0 to 1 scale.
F is the partial-credit functional score, not the binary solved indicator used by pass@1. Q is the lint and
complexity score from static tools (ruff and radon for Python) and S is the security scan.
C is Code quality, composed of three layers: a reviewed panel score (15 points), an intent recovery
probe scored against each task's known spec departures (6 points), and a measured maintenance test in
which a fixed agent applies a follow-up change and regenerated hidden tests check it (12 points). Until the
measured layer is built, the reviewed and intent layers carry 24 and 9 points, disclosed on each card.
Reports average attempts within each task, then average tasks equally. All factors must be present;
a missing factor makes this combined score unavailable rather than silently changing its weights.
Weight moved from the lint and complexity metric because its maintainability index rewards fewer lines and per-function complexity stays low when each dense line does something different, so compressed code scored higher than the same logic written for a reader. A third is a benchmark policy choice, locked before any submission was rescored. Reports published before September 2026 used a 20% Code quality weight with a 50/15/15/20 split and are not rewritten; new cards report both profiles side by side.
This is model-based code review, not a human rating. The rubric is identical for every task and frozen by hash before any review. It names the reader it scores for: an engineer who has never seen the code, reads it without running it, and must make a correct change in one sitting. It scores six dimensions from 0 to 4 with written anchors: naming, presentation, and intent form a human readability sub-score; structure, changeability, and verifiability form a maintainability sub-score. Every score must cite an exact excerpt and a concrete consequence for that reader. The host computes the sub-scores; the judge does no arithmetic.
The scored panel is drawn from labs with no model on the board being compared, so no judge grades its own relative. Models from the labs being compared may also review every submission, but only as disclosed sensitivity panels outside the combined score; their gap against the neutral panel on their own family's code is published as a self-preference estimate.
Before a judge scores a single submission it takes a calibration exam on ten held-out programs that all do the same thing, from clear to compressed to over-engineered to deceptively commented, five reviews each, against twenty gates fixed in advance. A judge that fails is published as failed. Reviewers receive the issue, the complete saved patch, and the reconstructed final source; solver model and effort labels are omitted; tools and external access are disabled; every raw prompt, response, and receipt is kept.
Review does not modify the submitted solution or its functional result. These ratings are reviewed for a human reader by a blinded model panel; they are not human validation. The report for each comparison names the exact protocol version, judges, calibration outcome, and every operator intervention. The first results under the September 2026 protocol are the Astra and Fable 5.1 rescoring of September 9, 2026; cards published earlier describe the prior three-persona review.
Independent reviewer validation uses a different model with the same rubric, full patches, and test evidence. Reviewers do not receive the solver identity, effort, or previous reviewer scores. Reports distinguish validation-only reviews from ratings included in the combined score.
Validation compares task-level ratings and rationales, not only overall averages. A different provider reduces reliance on self-review but does not eliminate shared biases. Agreement is supporting evidence, not proof of correctness; disagreements identify cases for closer inspection.
Reports identify the judge model, effort, and aggregation rule. Additional reviewer ratings are not automatically blended into the combined score. The scoring formula and judge protocol identify the exact calculation used for each result.
Quality uses lint, complexity, and maintainability tools; security uses static analysis. These are automated indicators and do not establish that code is production-ready or free of vulnerabilities. Efficiency is reported separately through tokens, wall-clock time, and cost, not blended into the combined score.
Comparisons require the same task set, scoring formula, and judge protocol. Each report identifies these settings. Functional pass@1 is reported separately from the combined score so readers can distinguish task completion from code quality.
We reduce training-data contamination risk through original task authorship, recorded provenance, publication dates, and comparisons with documented model cutoffs. These controls cannot prove that a task was absent from training data. A published knowledge cutoff is not an independently verified inventory of training material.
For tasks from public repositories, we record the upstream merge date and whether the task meets the benchmark's decontamination criteria. A later merge date alone cannot rule out earlier public discussion or code. Hand-authored tasks provide original scenarios, but publication can expose them to future training.
Canary strings and task identifiers can help detect exposure. Failure to recover a canary is not proof that exposure never occurred. Suite versions and task hashes identify the exact material used in each evaluation.
Training exposure and access during evaluation are separate risks. Agent workspaces are placed outside the benchmark checkout. Harness-specific controls restrict web access and access to reference solutions and hidden tests; cross-session memory is disabled for independent attempts.
Trace audits check web and filesystem activity. A denied request is not successful access. A clean audit means no prohibited access was detected in the available evidence, not proof that every possible exposure channel was eliminated. Historical incidents and audit details are documented in the technical notes.
Cost per solved task is total cost divided by fully solved tasks, including the cost of unsuccessful attempts. If no task is solved, this ratio is undefined. Reports identify the cost basis and distinguish total elapsed wall-clock time from summed or per-task execution time.
API usage: estimates use the provider's published rates and recorded usage, including cache discounts when the necessary breakdown is available. An estimate is labeled as such; it is not a billing receipt.
Subscription usage: included usage is not presented as zero economic cost. Reports distinguish observed overage or marginal cash, any allocated subscription fee, and estimated API-equivalent cost. These are separate views and are not added together.
Token usage, cache breakdowns, and quota consumption are reported when the provider exposes them. Missing fields remain unavailable. Prices, assumptions, and measurement limitations accompany cost estimates.
The repository contains task definitions and grading code. Run artifacts retain available traces, patches, test results, usage, and configuration. Reports identify the suite version, task hashes, model, harness version, effort, execution environment, and attempt counts. Unsupported provider fields are marked unknown.
Reproducibility means the evaluation can be inspected and repeated under a documented setup. Deterministic tests can reproduce a verdict for the same patch and environment; a new model run may produce a different patch or result. Optional model graders are also non-deterministic.
Reports should include task coverage, repeat counts, uncertainty where applicable, and any exclusions. Results from different suite versions, hardware, or harnesses require explicit qualification.
The probability of solving a task in one attempt, estimated from the observed attempts and averaged across tasks. A functional score of 1.0 is required for a solve.
The probability that at least one of k independent attempts solves the task. It measures success when retries and a way to identify a successful result are available.
The probability that all k independent attempts succeed, a measure of consistency across repeated attempts.
A provider-specific control over reasoning effort. Labels and their effects vary by model and harness.
Controls and provenance checks intended to reduce prior-exposure risk. They do not certify absence from a model's training data.
Matching a replacement program's outputs to a reference binary for the tested inputs and workflows.
A task's measured difficulty threshold under specified reference models, harnesses, and budgets.
Choosing a model and effort setting for a task class based on measured quality and resource use.
Evaluation tests withheld from the agent's workspace and used to assess its submitted patch.