All VulcanBench-SWE v4 reports

September 9, 2026 · 230 runs rescored · Code quality protocol v3.4

GPT-6 Astra vs. Fable 5.1 under a neutral Code quality panel

230 runs from the September 2026 effort sweep, 23 tasks at five effort levels for each model, scored with Code quality at 33% of the combined score and judged for a human reader by Muse Spark 1.3 and Grok 4.6.

Across five effort levels

Each model ran every task at Low, Medium, High, Extra-high and Max effort, 23 tasks per cell. No solver ran for this report: functional, lint and complexity, and security measurements come from that sweep. What is new is the Code quality factor, now 33% of the combined score, scored for a named human reader by Muse Spark 1.3 and Grok 4.6, two judges from labs with no model on this board, with a ground-truth intent-recovery probe. Combined score weights are 50% functional correctness, 8.5% lint and complexity, 8.5% security and 33% Code quality.

Combined score

Astra 87.43 to 89.30 · Fable 5.1 89.12 to 91.84

Fable leads at every effort, by 1.59 to 2.54 points. Both models gain from effort: Astra +1.87 and Fable +2.72 from Low to Max.

Code quality

Astra 70.31 to 76.26 · Fable 5.1 78.36 to 82.43

Fable leads at every effort, by 6.17 to 8.95 points. Astra gains more from effort (+5.95 Low to Max) than Fable (+4.07) but starts further back.

Human readability

Astra 61.0 to 69.1 · Fable 5.1 76.9 to 80.8

The widest gap at every effort: 11.7 to 17.1 points in Fable's favour. Intent recovery, the ground-truth layer, stays within about two points at every effort.

Runtime

Astra 3.8 to 10.3 min · Fable 5.1 26.9 to 38.8 min

Astra is faster at every effort. Astra's runtime climbs from Extra-high upward; Fable's does not track effort.

Combined score and Code quality at each effort level, n=23 per cell
MetricLowMediumHighExtra-highMax
Astra combined score87.4387.7388.1689.1689.30
Fable 5.1 combined score89.1290.1390.0790.7591.84
Fable minus Astra, combined+1.69+2.40+1.91+1.59+2.54
Astra Code quality70.3171.1973.4875.4776.26
Fable 5.1 Code quality78.3680.1480.8681.9382.43
Fable minus Astra, Code quality+8.05+8.95+7.38+6.46+6.17
Astra minutes per task4.23.84.98.110.3
Fable 5.1 minutes per task26.931.728.838.827.1

Applying the prior 50/15/15/20 profile to the same Code quality scores still puts Fable ahead at every effort, so the ranking does not depend on the weight change. The card and the full per-effort table follow.

Astra and Fable 5.1 across five effort levels under the v3.4 Code quality protocol. Combined scores range from 87.43 to 91.84. Fable has the higher combined score at every matched effort; Astra has the lower mean runtime at every matched effort. A table shows Code quality components at each model's best effort. Exact values and limitations are provided in the adjacent accessible tables.
Combined score uses a focused scale; runtime starts at zero. Whiskers show ±1 task standard error, not statistical significance. Open full-size card.

Fable 5.1 has the higher combined score and the higher Code quality at every matched effort. Astra has the lower mean runtime at every matched effort. Both judges agree on every ranking. These runs compare model-and-harness combinations, not isolated base models.

Combined score and Code quality by effort

n=23 at every effort. Combined (33%) is the published score. Combined (20% profile) applies the prior 50/15/15/20 weights to the same new Code quality scores for comparison. Code quality, human readability, maintainability and intent recovery are out of 100 and equal means of the two judges. Runtime is solver wall-clock per task.

VulcanBench-SWE v4: matched effort results under Code quality protocol v3.4
Model / harnessEffortCombined (33%)SECombined (20% profile)Code qualityHuman readabilityMaintainabilityIntent recoveryMin/task
Astra / CodexLow87.430.5189.2970.3161.072.979.34.2
Astra / CodexMedium87.730.3289.3671.1961.774.080.13.8
Astra / CodexHigh88.160.4289.2473.4863.977.281.24.9
Astra / CodexExtra-high89.160.3090.2575.4767.878.481.88.1
Astra / CodexMax89.300.3790.2076.2669.179.881.110.3
Fable 5.1* / Claude CodeLow89.121.1590.5678.3676.979.778.526.9
Fable 5.1* / Claude CodeMedium90.130.9491.1080.1478.881.580.131.7
Fable 5.1* / Claude CodeHigh90.070.7890.3880.8679.882.979.628.8
Fable 5.1* / Claude CodeExtra-high90.750.8290.9281.9380.983.980.738.8
Fable 5.1* / Claude CodeMax91.840.4792.3182.4380.884.581.827.1

*Fallbacks: 11 of 115 Fable solver runs include Opus 4.8 fallback usage. They remain in the population and are labeled in the run records. Neither judge produced a reviewer fallback on any of the 230 submissions.

SE is one sample standard error of the combined score across 23 tasks. It describes task sampling, not judge uncertainty, repeated runs or a significance test.

Cost and tokens across effort levels

The same 230 runs priced from their solver receipts at list API rates, cache-aware, with judging excluded. These are API-equivalent estimates, not subscription charges: both models ran on subscriptions, and what a plan actually bills is not observable from the receipts. Astra is cheaper at every effort, by 3.5 to 6.3 times per task, and its cost rises with effort while Fable's does not track it.

Astra and Fable 5.1 API-equivalent cost per task and raw tokens per task at five effort levels, with a table of cost, cost ratio and tokens at each effort and for the full sweep. Exact values are in the adjacent accessible table.
Ticks above Astra's bars mark its long-context upper bound; whiskers show ±1 task standard error. Open full-size card.
API-equivalent cost, raw tokens and runtime at each effort level, n=23 per cell
EffortAstra $/taskAstra upper boundFable $/taskCost ratioAstra tokens/taskFable tokens/taskAstra min/taskFable min/task
Low$1.72$3.28$7.964.6x0.89M4.38M4.226.9
Medium$1.48$2.79$9.196.2x0.68M6.97M3.831.7
High$1.71$3.22$9.705.7x0.77M5.49M4.928.8
Extra-high$2.30$4.25$14.496.3x0.93M11.44M8.138.8
Max$2.57$4.69$9.063.5x0.95M3.82M10.327.1
Full sweep, 115 runs each$225.04$419.42$1,159.105.2x97M738M12.0 h58.7 h

Pricing. Rates checked 2026-09-06 from OpenAI and Anthropic. Astra input includes cache reads and output includes reasoning; Claude pricing covers observed five-minute and one-hour cache writes, Fable, Opus 5, Opus 4.8 fallback and auxiliary Haiku usage. Astra's receipts do not record per-request sizes, so the central estimate uses standard rates and the upper bound prices every run that exceeded 272k cumulative input tokens at the long-context tier. Astra stays cheaper at every effort under that bound.

Tokens are the solver CLI's raw totals including cache reads. Fable's cached prefixes make its token count a poor cost proxy: its tokens are 4 to 12 times Astra's while its cost is 3.5 to 6.3 times. Per-run estimates, token usage breakdowns and the rate tables are in economics.json and runs.csv.

What Code quality measures now

Every judge receives the same rubric, frozen by hash before any review. It names the reader it scores for: an engineer who has never seen the code, reads it top to bottom without running it, and must make a correct change in one sitting. It tells the judge that its own ease at parsing dense code is not evidence of readability. Six dimensions are scored 0 to 4 in half steps against written anchors: naming, presentation and intent form the human readability sub-score; structure, changeability and verifiability form the maintainability sub-score. Every score must cite an exact excerpt from the code and a concrete consequence for that reader. The host computes the sub-scores; the judge does no arithmetic.

The 33 points have three designed layers. The reviewed panel above carries 24 points. An intent-recovery probe carries 9: every task in this suite is a rewrite of a retired binary whose real behaviour departs from its written specification in documented ways, and the judge, given only the specification and the code, must list where the code departs from the specification. A separate call matches that list against a frozen answer key. The third layer, measured maintenance, is designed for 12 points but is not yet built, so the pre-registered 24 plus 9 split is in force and is stated on the card.

Code quality components at Max effort, the setting on the card; every effort is in the table above
ComponentGPT-6 AstraFable 5.1Fable minus Astra
Code quality76.2682.43+6.17
Human readability69.180.8+11.7
Maintainability79.884.5+4.7
Intent recovery81.181.8+0.8
Rated by Muse Spark 1.373.683.4+9.9
Rated by Grok 4.675.481.9+6.5
Standard error of Code quality0.881.10

Both judges agree on every ranking and differ by a few points on levels. The largest gap between the models is human readability. Intent recovery, the only layer with ground truth, is close. Per-run sub-scores from each judge are in runs.csv; the exact rubric, system text, probe and match instructions are in judge-protocols.json.

Judges and calibration

The scored panel is Muse Spark 1.3 (Meta) through the Muse CLI on its Standard tier, and Grok 4.6 (xAI) through the Cursor CLI at medium effort, with equal weight. Neither lab has a model on this board, so neither judge grades a relative. Every session is fresh, tools are disabled, the workspace is empty, model identity is checked per call from the CLI's own records, and solver labels are withheld. Astra and Claude Opus 5 also passed the calibration exam under an earlier protocol version but were retired from the score by the benchmark owner, because Astra would otherwise have judged its own code while Fable never did.

Before scoring a single submission, each judge reviews ten held-out programs that implement the same ledger specification: clear, compressed, compressed then auto-formatted, verbose with duplicated policy, needlessly abstracted, misleadingly commented, narrated with a comment on every line, a documented legacy quirk, an embedded instruction to give full marks, and hidden module state. Each program is reviewed five times in a seeded order. Twenty gates fixed in advance check that the judge sees the construct; a judge may miss at most one gate by at most half a point.

Calibration verdicts
JudgeProtocolCallsResultAllowance usedFailing gates
Muse Spark 1.3code-quality-maintenance-v3.480passednonone
Grok 4.6code-quality-maintenance-v3.380passednonone
GLM 5.3code-quality-maintenance-v3.33failedn/ag01_validity

GLM 5.3 was tried first and failed on validity after fabricating an excerpt in both attempts on one control; Muse Spark 1.3 replaced it under amendment v3.4. Both scored judges separate the clear program from the compressed one by more than two points on naming and presentation, credit formatting on presentation but not on naming, penalise misleading comments on intent, and penalise hidden state on verifiability. Every gate value and control mean is in calibration.json.

Evidence and reproduction

Public files contain all 230 run measurements, both judges' six-dimension sub-scores and intent-recovery scores per run, raw tokens and API-equivalent cost estimates per run, ten aggregates, every calibration gate value, the exact judge protocol text and source hashes. Each run's Code quality is (0.24 × reviewed score + 0.09 × intent recovery) / 0.33, where both terms are the equal mean of the two judges. The combined score is 100 × (0.50F + 0.085Q + 0.085S + 0.33C / 100). The reproduction guide gives the public arithmetic checks and links the protocol documents, controls, runner and operator wrapper.

No human rated anything; the scores are model judgment for a human reader, not human validation. Runs were not repeated, CLI versions were not held fixed across weeks, and effort labels are harness-specific. Hash checks bind the export to frozen files; they do not prove that judges were unbiased or that no training overlap exists.

Withheld: raw prompts and responses, submitted patches and reconstructed sources, the quirk answer keys (they describe hidden-test behaviour), reviewer session identifiers and usage receipts, and the retired Astra and Opus 5 reviews. Summed solver time is unchanged at 70.75 hours; judging is excluded.