All VulcanBench-SWE v4 reports

September 15, 2026 · 207 runs · Code quality protocol v3.5

GPT-5.5 vs. GPT-5.6 Luna across every effort level

207 runs through the Codex CLI on a ChatGPT subscription, 23 tasks at each effort level the API offers: four for GPT-5.5 and five for GPT-5.6 Luna. Scored with Code quality at 33% of the combined score and judged for a human reader by Muse Spark 1.3 and Grok 4.6 under the same frozen protocol as the Astra and Fable 5.1 report.

Across every effort level

Each model ran every task once at each effort level its API accepts, 23 tasks per cell: GPT-5.5 at Low, Medium, High and Extra-high, and GPT-5.6 Luna at those four plus Max. Both ran through the same Codex CLI on the same ChatGPT subscription, so this compares two models in one harness. Combined score weights are 50% functional correctness, 8.5% lint and complexity, 8.5% security and 33% Code quality, judged for a named human reader by Muse Spark 1.3 and Grok 4.6, two judges from labs with no model on this board, with a ground-truth intent-recovery probe.

Combined score

GPT-5.5 49.80 to 78.46 · Luna 41.29 to 84.15

The effort knob decides the winner. GPT-5.5 leads at Low and Medium by 8.51 and 10.26 points; Luna edges ahead at High and Extra-high by 1.30 and 0.65, inside one standard error. Luna's Max level, which GPT-5.5 does not offer, is the top cell.

Tasks passed

GPT-5.5 1 to 11 of 23 · Luna 0 to 19 of 23

Both models fail most tasks below High. Functional correctness climbs steeply with effort for both: Luna passes none at Low and 19 at Max; GPT-5.5 passes one at Low and 11 at Extra-high.

Code quality

GPT-5.5 61.75 to 67.67 · Luna 65.35 to 71.73

Luna is rated higher at every matched effort, by under three points from Medium up. The ten-point gap at Low is mostly a denominator effect: nine Luna Low runs passed no quirk family, so their Code quality is the reviewed score alone. GPT-5.5's Code quality rises about six points from Low to Extra-high; Luna's does not track effort.

Runtime and cost

GPT-5.5 8.2 to 20.8 min · Luna 2.7 to 44.1 min

Luna is faster at Low and Medium and slower from High up, reaching 44.1 minutes per task at Max. At list API rates Luna costs 16.6 to 74.8 times less per task at matched effort.

Combined score, Code quality, tasks passed and runtime at each effort level, n=23 per cell
MetricLowMediumHighExtra-highMax
GPT-5.5 combined score49.8063.5469.9278.46no level
Luna combined score41.2953.2871.2279.1184.15
Luna minus GPT-5.5, combined-8.51-10.26+1.30+0.65
GPT-5.5 tasks passed1/233/237/2311/23no level
Luna tasks passed0/231/239/2313/2319/23
GPT-5.5 Code quality61.7564.6562.9367.67no level
Luna Code quality71.7365.5865.7968.0565.35
Luna minus GPT-5.5, Code quality+9.98+0.93+2.86+0.38
GPT-5.5 minutes per task8.212.418.220.8no level
Luna minutes per task2.77.421.025.444.1

Under the prior 50/15/15/20 profile applied to the same Code quality scores, the same picture holds: GPT-5.5 ahead at Low and Medium, Luna ahead at High, Extra-high and Max. The card and the full per-effort table follow.

GPT-5.5 and GPT-5.6 Luna across effort levels under the v3.5 Code quality protocol. Combined scores range from 41.29 to 84.15. GPT-5.5 has the higher combined score at Low and Medium; Luna has the higher score at High, Extra-high and Max. Luna has the lower mean runtime at Low and Medium and the higher runtime from High up. A table shows Code quality components at each model's best effort. Exact values and limitations are provided in the adjacent accessible tables.
Combined score uses a focused scale; runtime starts at zero. Whiskers show ±1 task standard error, not statistical significance. Open full-size card.

GPT-5.5 has the higher combined score at Low and Medium effort. Luna has the higher combined score at High and Extra-high, by less than one standard error, and its Max level is the highest cell on the board. Luna is rated higher on Code quality at every matched effort. Both judges agree on every Code quality ranking. These runs compare model-and-harness combinations, not isolated base models.

Combined score and Code quality by effort

n=23 at every effort. Combined (33%) is the published score. Combined (20% profile) applies the prior 50/15/15/20 weights to the same Code quality scores for comparison. Code quality, human readability, maintainability and intent recovery are out of 100 and equal means of the two judges. Passed counts tasks with a perfect functional score. Runtime is solver wall-clock per task.

VulcanBench-SWE v4: effort results under Code quality protocol v3.5
Model / harnessEffortCombined (33%)SECombined (20% profile)Code qualityHuman readabilityMaintainabilityIntent recoveryPassedMin/task
GPT-5.5 / CodexLow49.802.9353.4461.7556.265.963.51/238.2
GPT-5.5 / CodexMedium63.543.8466.7864.6557.167.770.83/2312.4
GPT-5.5 / CodexHigh69.923.8173.3262.9355.265.669.87/2318.2
GPT-5.5 / CodexExtra-high78.462.5181.2567.6759.669.576.011/2320.8
GPT-5.6 Luna / CodexLow41.290.9643.7271.7364.973.380.8*0/232.7
GPT-5.6 Luna / CodexMedium53.282.9456.4665.5862.267.866.4*1/237.4
GPT-5.6 Luna / CodexHigh71.223.7174.3265.7958.667.872.69/2321.0
GPT-5.6 Luna / CodexExtra-high79.113.0781.8768.0561.069.076.213/2325.4
GPT-5.6 Luna / CodexMax84.151.3587.1865.3555.864.379.519/2344.1

*Intent recovery is scored over the specification departures a submission actually passed tests for. 9 of 23 Luna Low runs and 1 Luna Medium run passed none, so they have no intent-recovery score and their Code quality is the reviewed score alone; the starred means cover the remaining runs. No GPT-5.5 run was affected.

SE is one sample standard error of the combined score across 23 tasks. It describes task sampling, not judge uncertainty, repeated runs or a significance test. No solver fallback occurred in either sweep, and neither judge produced a reviewer fallback on any of the 207 submissions.

Cost and tokens across effort levels

The same 207 runs priced from their Codex receipts at list API rates, cache-aware, with judging excluded. These are API-equivalent estimates, not subscription charges: both models ran on a ChatGPT subscription, and what the plan actually bills is not observable from the receipts. GPT-5.5 lists at 25 times Luna's input and output rates, and its per-task cost is 16.6 to 74.8 times Luna's at matched effort. Luna uses more tokens than GPT-5.5 from High up and reaches 11.88M per task at Max, most of it cache reads.

GPT-5.5 and GPT-5.6 Luna API-equivalent cost per task and raw tokens per task at every effort level, with a table of cost, cost ratio and tokens at each effort and for the full sweeps. Exact values are in the adjacent accessible table.
Whiskers show ±1 task standard error. Open full-size card.
API-equivalent cost, raw tokens and runtime at each effort level, n=23 per cell
EffortGPT-5.5 $/taskLuna $/taskCost ratioGPT-5.5 tokens/taskLuna tokens/taskGPT-5.5 min/taskLuna min/task
Low$1.79$0.0274.8x1.72M0.42M8.22.7
Medium$2.71$0.0740.8x2.74M1.35M12.47.4
High$4.20$0.2219.4x4.52M5.89M18.221.0
Extra-high$4.56$0.2716.6x4.72M7.58M20.825.4
Maxno level$0.4511.88M44.1
Full sweeps, 92 and 115 runs$305.04$23.7012.9x315M624M22.9 h38.6 h

Pricing. Rates checked 2026-09-11 from OpenAI: GPT-5.5 at $5 input, $0.50 cached input and $30 output per million tokens; Luna at $0.20, $0.02 and $1.20. Codex receipts report input, cached input, output and reasoning tokens per run, so cached input is billed at the cache-read rate and the rest at the standard rate. Per-request context sizes are not exposed, so no long-context premium is applied. One GPT-5.5 Extra-high task was re-run after its first attempt overran the cap under a harness fault; the re-run is the priced and judged run.

Tokens are Codex's raw totals including cache reads. Per-run estimates, token usage breakdowns and the rate table are in economics.json and runs.csv.

What Code quality measures

Every judge receives the same rubric, frozen by hash before any review. It names the reader it scores for: an engineer who has never seen the code, reads it top to bottom without running it, and must make a correct change in one sitting. It tells the judge that its own ease at parsing dense code is not evidence of readability. Six dimensions are scored 0 to 4 in half steps against written anchors: naming, presentation and intent form the human readability sub-score; structure, changeability and verifiability form the maintainability sub-score. Every score must cite an exact excerpt from the code and a concrete consequence for that reader. The host computes the sub-scores; the judge does no arithmetic.

The 33 points have three designed layers. The reviewed panel carries 24 points. An intent-recovery probe carries 9: every task in this suite is a rewrite of a retired binary whose real behaviour departs from its written specification in documented ways, and the judge, given only the specification and the code, must list where the code departs from the specification. A separate call matches that list against a frozen answer key, over the departures the submission passed tests for. The third layer, measured maintenance, is designed for 12 points but is not yet built, so the pre-registered 24 plus 9 split is in force and is stated on the card.

Code quality components at each model's best effort, the setting on the card; every effort is in the table above
ComponentGPT-5.5, Extra-highLuna, MaxLuna minus GPT-5.5
Code quality67.6765.35-2.32
Human readability59.655.8-3.8
Maintainability69.564.3-5.2
Intent recovery76.079.5+3.5
Rated by Muse Spark 1.362.858.0-4.8
Rated by Grok 4.666.362.1-4.2
Standard error of Code quality1.491.33

At each model's best effort the reviewed scores are close and GPT-5.5 reads slightly better; Luna recovers more of the hidden contract. Across matched efforts Luna's Code quality and human readability are higher at every level, while maintainability and intent recovery trade places; both judges rank the reviewed score the same way in every cell. Both models sit well below the Astra and Fable 5.1 board under the same rubric, 62 to 72 here against 70 to 82 there: the judges find their code harder for a person to read. Per-run sub-scores from each judge are in runs.csv; the exact rubric, system text, probe and match instructions are in judge-protocols.json.

Judges and calibration

The scored panel is Muse Spark 1.3 (Meta) through the Muse CLI on its Standard tier, and Grok 4.6 (xAI) through the Cursor CLI at medium effort, with equal weight. Neither lab has a model on this board, so neither judge grades a relative. Every session is fresh, tools are disabled, the workspace is empty, model identity is checked per call from the CLI's own records, and solver labels are withheld. Protocol v3.5 changes nothing in the rubric, controls, gates, repeats, seed or judge settings from v3.4; it applies the protocol to this population, and both judges retook the calibration exam under it before any counted call.

Before scoring a single submission, each judge reviews ten held-out programs that implement the same ledger specification: clear, compressed, compressed then auto-formatted, verbose with duplicated policy, needlessly abstracted, misleadingly commented, narrated with a comment on every line, a documented legacy quirk, an embedded instruction to give full marks, and hidden module state. Each program is reviewed five times in a seeded order. Twenty gates fixed in advance check that the judge sees the construct; a judge may miss at most one gate by at most half a point.

Calibration verdicts
JudgeProtocolCallsResultAllowance usedFailing gates
Muse Spark 1.3code-quality-maintenance-v3.580passedyes, g11 by 0.02g11_repeatability
Grok 4.6code-quality-maintenance-v3.580passednonone

Muse Spark 1.3 used the pre-registered allowance: its five repeats on one control spread 0.02 of a point more than gate 11 permits, within the half-point allowance, and every other gate passed. Grok 4.6 passed every gate. Both judges separate the clear program from the compressed one by more than two points on naming and about two points on presentation, credit formatting on presentation but not on naming, penalise misleading comments on intent, and penalise hidden state on verifiability. Every gate value and control mean is in calibration.json.

Each judge made 718 calls: 80 in calibration, 207 primary reviews, 9 repeats, 8 pairwise checks, 207 intent probes and 207 answer-key matches. Muse needed a second attempt on five calls (two unsupported excerpts, two malformed JSON responses, one missing consequence in calibration) and Grok on six (four unsupported excerpts, one malformed response, and one transport fault when the Cursor CLI could not resolve its API host). The transport fault stopped the run for a person; the operator wrapper gained a rule that gives one fresh attempt after a judge CLI network fault, with the receipt retained, and the call was retried under it. Every second attempt is archived beside the first in the harness run directory.

Evidence and reproduction

Public files contain all 207 run measurements, both judges' six-dimension sub-scores and intent-recovery scores per run, raw tokens and API-equivalent cost estimates per run, nine aggregates, every calibration gate value, the exact judge protocol text and source hashes. Each run's Code quality is (0.24 × reviewed score + 0.09 × intent recovery) / 0.33, where both terms are the equal mean of the two judges, or the reviewed score alone where intent recovery has no denominator. The combined score is 100 × (0.50F + 0.085Q + 0.085S + 0.33C / 100). The reproduction guide gives the public arithmetic checks and links the protocol documents, controls, runner and operator wrapper.

No human rated anything; the scores are model judgment for a human reader, not human validation. Runs were not repeated, and effort labels are Codex's own. Hash checks bind the export to frozen files; they do not prove that judges were unbiased or that no training overlap exists.

Withheld: raw prompts and responses, submitted patches and reconstructed sources, the quirk answer keys (they describe hidden-test behaviour), and reviewer session identifiers and usage receipts. Summed solver time is 61.42 hours; judging is excluded.