Most engineering teams picked one model and one effort setting at some point and have used it for everything since: a config change, a null check, a new endpoint, a migration across three services, all routed the same way. The work isn’t uniform, and the public reports here show that models and effort levels aren’t interchangeable across task types either. But those reports run on public tasks. Whether the same patterns hold on a Go monorepo with a custom ORM, or a Rails app with fifteen years of history, is a separate question, and the only way to answer it is to measure on that codebase. This work started on my own team’s codebase, for exactly this reason: how a $300,000 model bill turned into a benchmark.
The evals come from your merged pull requests. A bug fix that shipped last month becomes a task: the repository at the parent commit, a description of the problem at the level of detail your engineers would give a new hire, and hidden tests that fail before the change and pass after it. A feature PR becomes a feature task; a test-only PR becomes a test-writing task; a dependency bump or a schema migration becomes a migration task. The suite ends up with the same distribution of work your team produces, in your languages, against your frameworks, at the scale of your repositories.
Some PRs don’t convert cleanly: they have no test coverage, or they touch infrastructure the sandbox can’t reproduce, or they’re entangled with three other changes. For those we build matched equivalents on the same stack and at the same scale, held to the same standard. The task classes are whatever falls out of your actual mix, not a generic taxonomy, and if a platform team on Go and a product team on TypeScript don’t share a codebase they don’t share a suite.
Which model, at which effort, for each class of work in your suite, with escalation rules for when the first route fails its tests. Delivered as a config file for whatever gateway, router, or agent harness you already run, with the reasoning written down next to it. See a sample row.
Tasks, hidden tests, and the harness to run them, built from your code and kept private. It’s yours to rerun when a new model ships, when a vendor changes pricing, or when your own codebase drifts far enough that last quarter’s answer stops applying.
Per task class and per model-and-effort cell: pass@1, pass^k reliability across repeats, tokens, cost, and wall-clock. Plus your task volume priced through the policy against what you pay today, as a spreadsheet you can change the assumptions in, and a written report in the same format as the public ones.
A conversation about your stack and how your engineers use models today: languages, frameworks, repository size and shape, which tools and harnesses are in use, what the default model and effort are, and roughly how the work breaks down between fixes, features, refactors, tests, and migrations. From that we propose the task classes and how many tasks each needs.
Tasks from your merged PRs where they convert, matched equivalents where they don’t. Every task has hidden tests validated to fail before the fix and pass after, the same bar as the public suite. The build happens against a copy of your repositories under whatever access arrangement your security team needs; nothing from it becomes public eval material.
Every model under consideration, at every effort level worth testing, with repeats, in the network-isolated Docker sandbox with deterministic grading and no LLM judge. If your engineers work through a specific harness (Claude Code, Cursor, an internal agent), the matrix can run through that harness rather than a reference loop, since the harness itself changes the results.
The matrix collapses into one route per task class plus escalation rules. We ship that as config, the measurements and cost model as files, and the report as a document, then walk your team through how each row was decided and where the results were close enough that the call could go either way.
| Task class | Route to | Escalate on miss | What the measurement showed |
|---|---|---|---|
| Routine bug fixes | Mid-tier · low | Frontier · low | Same pass rate as the frontier model on this class, at a fraction of the cost |
| Test writing | Small · low | Mid-tier · low | Highest-volume class; the small model passed reliably across repeats |
| Feature work | Mid-tier · medium | Frontier · medium | Effort improved results up to medium; high and max did not |
| Refactors & migrations | Frontier · low | Frontier · high | Model choice mattered, effort level did not |
| Novel, cross-cutting work | Frontier · high | Human review | The only cell where these tasks passed at all |
Illustrative. The task classes, the routes, and the escalation rules in your policy come out of your measurements, and they can land differently from this in every row.
The table ships as a policy file your router reads. One row of the sample, with the measured cell that justifies it:
routine_fix:
route: {provider: openai, model: gpt-5.6-terra, effort: low}
escalate:
- {provider: anthropic, model: claude-fable-5, effort: low}
budget: {max_wall_clock_s: 900, max_attempts: 2}
measured: {pass_at_1: 0.88, pass_pow_3: 0.79, cost_per_task_usd: 0.14, p50_minutes: 2.6}
why: "Accuracy parity with the frontier baseline (0.89) at a fraction of the cost."
The full sample has the classifier that assigns work to a class, the escalation conditions, the constraints the measurement respected (allowed providers, data residency, wall-clock ceiling), the cost model priced through your volume mix, and the rerun triggers. The README beside it covers wiring it into whatever gateway you run.
Tasks built from your repositories stay private: they are not published, not reused in other engagements, and never appear on the leaderboard. No model vendor pays for placement in a routing policy. Grading is by hidden tests, so the result owes nothing to me or to anyone whose model appears in it.
Scoping, the suite build, the full matrix, the routing policy, and a walkthrough with your team. A focused engagement (one codebase, a handful of task classes, a four-model matrix) starts in the mid four figures. Larger matrices and more task classes cost more; another team on a different stack adds a suite build, not a second engagement.
Optional, after the assessment. When a new frontier model ships or a vendor changes pricing, your suite is rerun against it and you get an updated policy with a short note on what changed and whether it’s worth switching. Priced per rerun or as a quarterly plan, at a fraction of the initial build, because the suite already exists.
The suite, the hidden tests, and the harness are yours after the assessment. If your team would rather rerun the matrix in house, nothing stops you, and you don’t owe anything for it. The rerun plan is there for teams that would rather not spend engineering time on it.
Compute is metered and itemized the same way it is in the public reports, so you can see what the matrix itself cost separately from the engineering time.
Private engagements are new, so there are no case studies on this page yet. Until there are, the public reports are the track record: each one publishes its method, its costs, and its raw data, and a private engagement is held to exactly the same standard. If you want to judge the work before talking, start with the methodology and any report on the benchmarks page.
The suite is built against a copy of your repositories under whatever access arrangement your security team needs. Tasks built from your code are never published, never reused in another engagement, and never appear on the leaderboard.
Usually yours, so the spend lands on your own bills and under the vendor agreements you already have. If that doesn’t work for you, the compute can be run separately and billed as itemized cost instead.
Yes, because the harness changes the results. The matrix can run through the tool your engineers actually use rather than a reference loop, so the policy reflects how the work really gets done.
It depends mostly on the size of the suite and the matrix. The reply to your first message includes a proposed schedule along with the suite shape and the price.
A few answers is enough to come back with a proposed suite shape, a matrix, and a price. Every field is optional. Nothing is sent from this page: the button opens an email in your own mail client with your answers filled in, and you can edit it before sending.
Or email hello@vulcanbench.com directly. If you’re building a model rather than choosing between them, see benchmark your model.