For engineering teams

Which model, at which effort, does your team actually need?

Most engineering teams picked one model and one effort setting at some point and have used it for everything since: a config change, a null check, a new endpoint, a migration across three services, all routed the same way. The work isn’t uniform, and the public reports here show that models and effort levels aren’t interchangeable across task types either. But those reports run on public tasks. Whether the same patterns hold on a Go monorepo with a custom ORM, or a Rails app with fifteen years of history, is a separate question, and the only way to answer it is to measure on that codebase. This work started on my own team’s codebase, for exactly this reason: how a $300,000 model bill turned into a benchmark.

Worth it for some teams, not for every team.

A good fit

  • Your model spend is large enough that moving part of it to a cheaper route would matter to your budget.
  • Your engineers default to one model and one effort level for everything, and nobody has measured whether that default is right for each kind of work.
  • You have a history of merged pull requests with tests in them, which is what the suite is built from.
  • You want the answer as something you can deploy and rerun, not a one-off opinion.

Probably not yet

  • Your model spend is small. The public leaderboard and reports will likely answer your question for free, and an engagement would cost more than it could save.
  • Your code has very little test coverage. A suite can still be built from matched equivalents, but the result says less about your own work.
  • You need the answer by next week. Building the suite properly takes time, and a rushed suite gives a confident answer to the wrong question.

Built from what your engineers actually do.

The evals come from your merged pull requests. A bug fix that shipped last month becomes a task: the repository at the parent commit, a description of the problem at the level of detail your engineers would give a new hire, and hidden tests that fail before the change and pass after it. A feature PR becomes a feature task; a test-only PR becomes a test-writing task; a dependency bump or a schema migration becomes a migration task. The suite ends up with the same distribution of work your team produces, in your languages, against your frameworks, at the scale of your repositories.

Some PRs don’t convert cleanly: they have no test coverage, or they touch infrastructure the sandbox can’t reproduce, or they’re entangled with three other changes. For those we build matched equivalents on the same stack and at the same scale, held to the same standard. The task classes are whatever falls out of your actual mix, not a generic taxonomy, and if a platform team on Go and a product team on TypeScript don’t share a codebase they don’t share a suite.

01

A routing policy

Which model, at which effort, for each class of work in your suite, with escalation rules for when the first route fails its tests. Delivered as a config file for whatever gateway, router, or agent harness you already run, with the reasoning written down next to it. See a sample row.

02

The eval suite itself

Tasks, hidden tests, and the harness to run them, built from your code and kept private. It’s yours to rerun when a new model ships, when a vendor changes pricing, or when your own codebase drifts far enough that last quarter’s answer stops applying.

03

The measurements

Per task class and per model-and-effort cell: pass@1, pass^k reliability across repeats, tokens, cost, and wall-clock. Plus your task volume priced through the policy against what you pay today, as a spreadsheet you can change the assumptions in, and a written report in the same format as the public ones.

1

Scoping

A conversation about your stack and how your engineers use models today: languages, frameworks, repository size and shape, which tools and harnesses are in use, what the default model and effort are, and roughly how the work breaks down between fixes, features, refactors, tests, and migrations. From that we propose the task classes and how many tasks each needs.

2

Suite build

Tasks from your merged PRs where they convert, matched equivalents where they don’t. Every task has hidden tests validated to fail before the fix and pass after, the same bar as the public suite. The build happens against a copy of your repositories under whatever access arrangement your security team needs; nothing from it becomes public eval material.

3

The matrix

Every model under consideration, at every effort level worth testing, with repeats, in the network-isolated Docker sandbox with deterministic grading and no LLM judge. If your engineers work through a specific harness (Claude Code, Cursor, an internal agent), the matrix can run through that harness rather than a reference loop, since the harness itself changes the results.

4

Policy and walkthrough

The matrix collapses into one route per task class plus escalation rules. We ship that as config, the measurements and cost model as files, and the report as a document, then walk your team through how each row was decided and where the results were close enough that the call could go either way.

One row per class of work.

One row per class of work
Task classRoute toEscalate on missWhat the measurement showed
Routine bug fixesMid-tier · lowFrontier · lowSame pass rate as the frontier model on this class, at a fraction of the cost
Test writingSmall · lowMid-tier · lowHighest-volume class; the small model passed reliably across repeats
Feature workMid-tier · mediumFrontier · mediumEffort improved results up to medium; high and max did not
Refactors & migrationsFrontier · lowFrontier · highModel choice mattered, effort level did not
Novel, cross-cutting workFrontier · highHuman reviewThe only cell where these tasks passed at all

Illustrative. The task classes, the routes, and the escalation rules in your policy come out of your measurements, and they can land differently from this in every row.

The table ships as a policy file your router reads. One row of the sample, with the measured cell that justifies it:

routine_fix:
  route:    {provider: openai,    model: gpt-5.6-terra, effort: low}
  escalate:
    - {provider: anthropic, model: claude-fable-5, effort: low}
  budget:   {max_wall_clock_s: 900, max_attempts: 2}
  measured: {pass_at_1: 0.88, pass_pow_3: 0.79, cost_per_task_usd: 0.14, p50_minutes: 2.6}
  why: "Accuracy parity with the frontier baseline (0.89) at a fraction of the cost."

The full sample has the classifier that assigns work to a class, the escalation conditions, the constraints the measurement respected (allowed providers, data residency, wall-clock ceiling), the cost model priced through your volume mix, and the rerun triggers. The README beside it covers wiring it into whatever gateway you run.

Private, and not sponsored by anyone in the table

Tasks built from your repositories stay private: they are not published, not reused in other engagements, and never appear on the leaderboard. No model vendor pays for placement in a routing policy. Grading is by hidden tests, so the result owes nothing to me or to anyone whose model appears in it.

Build it once, rerun it when the models change.

01

The assessment

Scoping, the suite build, the full matrix, the routing policy, and a walkthrough with your team. A focused engagement (one codebase, a handful of task classes, a four-model matrix) starts in the mid four figures. Larger matrices and more task classes cost more; another team on a different stack adds a suite build, not a second engagement.

02

Reruns when models ship

Optional, after the assessment. When a new frontier model ships or a vendor changes pricing, your suite is rerun against it and you get an updated policy with a short note on what changed and whether it’s worth switching. Priced per rerun or as a quarterly plan, at a fraction of the initial build, because the suite already exists.

03

Or rerun it yourself

The suite, the hidden tests, and the harness are yours after the assessment. If your team would rather rerun the matrix in house, nothing stops you, and you don’t owe anything for it. The rerun plan is there for teams that would rather not spend engineering time on it.

Compute is metered and itemized the same way it is in the public reports, so you can see what the matrix itself cost separately from the engineering time.

A new offering, held to the public standard

Private engagements are new, so there are no case studies on this page yet. Until there are, the public reports are the track record: each one publishes its method, its costs, and its raw data, and a private engagement is held to exactly the same standard. If you want to judge the work before talking, start with the methodology and any report on the benchmarks page.

Does our code have to leave our environment?

The suite is built against a copy of your repositories under whatever access arrangement your security team needs. Tasks built from your code are never published, never reused in another engagement, and never appear on the leaderboard.

Whose API keys does the matrix run on?

Usually yours, so the spend lands on your own bills and under the vendor agreements you already have. If that doesn’t work for you, the compute can be run separately and billed as itemized cost instead.

Our engineers work in Cursor, Claude Code, or an internal agent. Does that matter?

Yes, because the harness changes the results. The matrix can run through the tool your engineers actually use rather than a reference loop, so the policy reflects how the work really gets done.

How long does it take?

It depends mostly on the size of the suite and the matrix. The reply to your first message includes a proposed schedule along with the suite shape and the price.

Tell me about your stack.

A few answers is enough to come back with a proposed suite shape, a matrix, and a price. Every field is optional. Nothing is sent from this page: the button opens an email in your own mail client with your answers filled in, and you can edit it before sending.

Open a discussion on GitHub

Or email hello@vulcanbench.com directly. If you’re building a model rather than choosing between them, see benchmark your model.