September 19, 2026 · 115 runs · Code quality protocol v3.7
GPT-5.6 Sol across every effort level
115 runs through the Codex CLI on a ChatGPT Pro subscription, 23 tasks at each of the five effort levels the API offers. Scored with Code quality at 33% of the combined score and judged for a human reader by Muse Spark 1.3 and Grok 4.6 under the same frozen protocol as the other Frontier v4 reports, so Sol takes its place on the Frontier v4 leaderboard beside Terra, Luna, GPT-5.5, Astra and Fable 5.1, completing the GPT-5.6 family.
Across every effort level
Sol ran every task once at each of its five effort levels, 23 tasks per cell, through the Codex CLI on a ChatGPT Pro subscription on September 17 to 18, 2026. Combined score weights are 50% functional correctness, 8.5% lint and complexity, 8.5% security and 33% Code quality, judged for a named human reader by Muse Spark 1.3 and Grok 4.6, two judges from labs with no model on this board, with a ground-truth intent-recovery probe. Judging ran September 18 to 19. One Max run, codeccore, has no published Code quality score: Grok 4.6's intent probe quoted an excerpt absent from the code on both attempts and the frozen protocol accepts no invented characters, so the Max cell's Code quality and combined score cover 22 of its 23 runs. The run passed its tests and is counted in every runtime, token and cost figure.
Combined score
69.05 at Low to 87.18 at Max
Most of the ladder is one step: 13.5 points from Low to Medium, then 3.2 to High, 0.7 to Extra-high and 0.7 to Max. Above High the levels sit within one standard error of each other.
Tasks passed
6 of 23 at Low to 21 of 22 judged at Max
Functional correctness drives the curve. Sol passes 6 of 23 at Low, 12 at Medium and 20 at High; Extra-high passes 21 of 23 and Max 21 of the 22 judged tasks (22 of 23 in the sweep itself).
Code quality
64.98 to 68.19
Flat within noise: the judges rate Sol's code within about three points at every effort, with standard errors of 1.2 to 1.6. The effort knob buys correctness, not readability.
Runtime and cost
9.8 to 12.0 min · $1.36 to $2.05 per task
Neither runtime nor cost tracks the ladder. High is the slowest and most expensive level (12.0 minutes, $2.05); Extra-high and Max come in at $1.61 and $1.77.
| Metric | Low | Medium | High | Extra-high | Max |
|---|---|---|---|---|---|
| Combined score | 69.05 | 82.56 | 85.76 | 86.47 | 87.18 |
| Standard error | 3.50 | 1.25 | 0.98 | 0.80 | 0.62 |
| Tasks passed | 6/23 | 12/23 | 20/23 | 21/23 | 21/22 |
| Code quality | 64.98 | 68.19 | 67.65 | 67.31 | 67.58 |
| Human readability | 57.4 | 61.2 | 58.8 | 59.2 | 59.8 |
| Minutes per task | 9.8 | 11.1 | 12.0 | 10.6 | 11.6 |
| API-equivalent $ per task | $1.36 | $1.72 | $2.05 | $1.61 | $1.77 |
Under the prior 50/15/15/20 profile applied to the same Code quality scores the ladder is the same shape, 72.24 to 89.90. The card and the full per-effort table follow.
Sol's combined score gains most of its range in the first step of the ladder and flattens from High upward, where the three top levels sit within 1.5 points of each other. Code quality does not move with effort, and neither runtime nor cost rises with it. These runs compare a model-and-harness combination, not an isolated base model; the Frontier v4 leaderboard places every Sol effort level beside the other models judged under the same protocol.
Combined score and Code quality by effort
n=23 at Low, Medium, High and Extra-high; n=22 at Max, where one run has no published Code quality score (see the note below). Combined (33%) is the published score. Combined (20% profile) applies the prior 50/15/15/20 weights to the same Code quality scores for comparison. Code quality, human readability, maintainability and intent recovery are out of 100 and equal means of the two judges. Passed counts judged tasks with a perfect functional score. Runtime is solver wall-clock per task over every run in the cell.
| Model / harness | Effort | Combined (33%) | SE | Combined (20% profile) | Code quality | Human readability | Maintainability | Intent recovery | Passed | Min/task |
|---|---|---|---|---|---|---|---|---|---|---|
| GPT-5.6 Sol / Codex | Low | 69.05 | 3.50 | 72.24 | 64.98 | 57.4 | 66.9 | 72.4 | 6/23 | 9.8 |
| GPT-5.6 Sol / Codex | Medium | 82.56 | 1.25 | 85.33 | 68.19 | 61.2 | 69.8 | 75.3 | 12/23 | 11.1 |
| GPT-5.6 Sol / Codex | High | 85.76 | 0.98 | 88.51 | 67.65 | 58.8 | 67.8 | 79.3 | 20/23 | 12.0 |
| GPT-5.6 Sol / Codex | Extra-high | 86.47 | 0.80 | 89.23 | 67.31 | 59.2 | 67.8 | 77.4 | 21/23 | 10.6 |
| GPT-5.6 Sol / Codex | Max | 87.18 | 0.62 | 89.90 | 67.58 | 59.8 | 68.3 | 77.1 | 21/22§ | 11.6 |
§Max is judged on 22 of 23 tasks. On codeccore, Grok 4.6's intent probe quoted memo = record[31:46].rstrip(".",") on both attempts where the source line is memo = record[31:46].rstrip(".,"). A character is inserted inside the string literal; it is neither a re-wrap, an omission nor an escape spelling, so no recovery rule accepts it and none was added. No answer-key match was made and the frozen summary, which requires a valid match from every passing panel, publishes no Code quality or combined score for that run. Both judges' reviews and the Muse probe are retained in the harness archive. The run passed its tests, so the sweep itself passed 22 of 23 at Max; its runtime, tokens and cost are in every economics figure. The Max cell's Code quality, combined score and standard error are means over the 22 judged runs.
Intent recovery is scored over the specification departures a submission actually passed tests for. Every judged Sol run passed at least one, so no run fell back to the reviewed score alone. SE is one sample standard error of the combined score across the cell's tasks. It describes task sampling, not judge uncertainty, repeated runs or a significance test. No solver fallback occurred, and neither judge produced a reviewer fallback on any of the 115 submissions.
Cost and tokens across effort levels
The same 115 runs priced from their Codex receipts at list API rates, cache-aware, with judging excluded. These are API-equivalent estimates, not subscription charges: Sol ran on a ChatGPT Pro subscription, and what the plan actually bills is not observable from the receipts. Cost per task does not follow the ladder: $1.36 at Low, $1.72 at Medium, $2.05 at High, the most expensive level, then $1.61 at Extra-high and $1.77 at Max. Runtime moves the same way, from 9.8 to 12.0 minutes with High the slowest. The whole sweep comes to $195.55 for 115 runs, every one of them priced, including the Max run without a published Code quality score.
| Effort | Runs | $/task | Level total | Tokens/task | Min/task |
|---|---|---|---|---|---|
| Low | 23 | $1.36 | $31.26 | 1.79M | 9.8 |
| Medium | 23 | $1.72 | $39.45 | 2.26M | 11.1 |
| High | 23 | $2.05 | $47.04 | 2.83M | 12.0 |
| Extra-high | 23 | $1.61 | $37.01 | 1.91M | 10.6 |
| Max | 23 | $1.77 | $40.79 | 2.11M | 11.6 |
| Full sweep | 115 | $1.70 | $195.55 | 251M | 21.1 h |
Pricing. Rates checked 2026-09-11 from OpenAI: Sol at $4 input, $0.40 cached input and $20 output per million tokens. Codex receipts report input, cached input, output and reasoning tokens per run, so cached input is billed at the cache-read rate and the rest at the standard rate. Per-request context sizes are not exposed, so no long-context premium is applied. The sweep stamped each run at these rates at run time; the export recomputed every run from its receipt and matched every stamp.
Tokens are Codex's raw totals including cache reads. Per-run estimates, token usage breakdowns and the rate table are in economics.json and runs.csv.
What Code quality measures
Every judge receives the same rubric, frozen by hash before any review. It names the reader it scores for: an engineer who has never seen the code, reads it top to bottom without running it, and must make a correct change in one sitting. It tells the judge that its own ease at parsing dense code is not evidence of readability. Six dimensions are scored 0 to 4 in half steps against written anchors: naming, presentation and intent form the human readability sub-score; structure, changeability and verifiability form the maintainability sub-score. Every score must cite an exact excerpt from the code and a concrete consequence for that reader. The host computes the sub-scores; the judge does no arithmetic.
The 33 points have three designed layers. The reviewed panel carries 24 points. An intent-recovery probe carries 9: every task in this suite is a rewrite of a retired binary whose real behaviour departs from its written specification in documented ways, and the judge, given only the specification and the code, must list where the code departs from the specification. A separate call matches that list against a frozen answer key, over the departures the submission passed tests for. The third layer, measured maintenance, is designed for 12 points but is not yet built, so the pre-registered 24 plus 9 split is in force and is stated on the card.
| Component | Low | Medium | High | Extra-high | Max |
|---|---|---|---|---|---|
| Code quality | 64.98 | 68.19 | 67.65 | 67.31 | 67.58 |
| Human readability | 57.4 | 61.2 | 58.8 | 59.2 | 59.8 |
| Maintainability | 66.9 | 69.8 | 67.8 | 67.8 | 68.3 |
| Intent recovery | 72.4 | 75.3 | 79.3 | 77.4 | 77.1 |
| Rated by Muse Spark 1.3 | 60.2 | 64.3 | 62.0 | 61.8 | 63.0 |
| Rated by Grok 4.6 | 64.1 | 66.8 | 64.6 | 65.3 | 65.1 |
| Standard error of Code quality | 1.56 | 1.30 | 1.17 | 1.39 | 1.39 |
Code quality sits between 64.98 and 68.19 at every effort, with the two judges within four points of each other in every cell and Grok 4.6 the higher rater at each level. Intent recovery, the ground-truth layer, rises from Low to High as more of the hidden contract is implemented and holds there; human readability and maintainability do not move with the knob. Per-run sub-scores from each judge are in runs.csv; the exact rubric, system text, probe and match instructions are in judge-protocols.json.
Judges and calibration
The scored panel is Muse Spark 1.3 (Meta) through the Muse CLI on its Standard tier, and Grok 4.6 (xAI) through the Cursor CLI at medium effort, with equal weight. Neither lab has a model on this board, so neither judge grades a relative. Every session is fresh, tools are disabled, the workspace is empty, model identity is checked per call from the CLI's own records, and solver labels are withheld. Protocol v3.7 changes nothing in the rubric, controls, quirk keys, gates, repeats, seed or judge settings from v3.6; it applies the protocol to this population, and both judges retook the calibration exam under it before any counted call.
Before scoring a single submission, each judge reviews ten held-out programs that implement the same ledger specification: clear, compressed, compressed then auto-formatted, verbose with duplicated policy, needlessly abstracted, misleadingly commented, narrated with a comment on every line, a documented legacy quirk, an embedded instruction to give full marks, and hidden module state. Each program is reviewed five times in a seeded order. Twenty gates fixed in advance check that the judge sees the construct; a judge may miss at most one gate by at most half a point.
| Judge | Protocol | Calls | Result | Allowance used | Failing gates |
|---|---|---|---|---|---|
| Muse Spark 1.3 | code-quality-maintenance-v3.7 | 80 | passed | no | none |
| Grok 4.6 | code-quality-maintenance-v3.7 | 80 | passed | no | none |
Both judges passed every gate with no allowance used; under v3.5 and v3.6 Muse Spark 1.3 had needed the pre-registered allowance on the repeatability gate, and under v3.7 it did not. Both judges separate the clear program from the compressed one by more than two points on naming and by 2.0 to 2.1 points on presentation, credit formatting mainly on presentation, penalise misleading comments on intent, and penalise hidden state on verifiability. Every gate value and control mean is in calibration.json.
Under v3.7 Muse Spark 1.3 made 440 calls: 80 in calibration, 115 primary reviews, 5 repeats, 10 pairwise checks, 115 intent probes and 115 answer-key matches. Grok 4.6 made 439, with 114 matches, since no match was made for the invalid codeccore probe. Muse needed a second attempt on nine calls (six primary reviews and three probes): four unsupported excerpts and five malformed JSON responses. Grok needed a second attempt on nine: five unsupported excerpts (two primary reviews, one calibration control and two probes, one of them the codeccore probe, which was unsupported on both attempts and is recorded as invalid), one malformed JSON probe, two transport faults when the Cursor CLI could not resolve its API host (a primary review and a calibration pair), each retried under the network-fault rule, and one calibration probe the provider blocked before any model output, retried once under a new wrapper rule recorded in the operations log. Neither judge produced a reviewer fallback. Every attempt is archived beside its replacement in the harness run directory.
Evidence and reproduction
Public files contain all 115 run measurements, both judges' six-dimension sub-scores and intent-recovery scores per published run, raw tokens and API-equivalent cost estimates per run, five aggregates, every calibration gate value, the exact judge protocol text, the operator's finding on the unpublished row and source hashes. Each run's Code quality is (0.24 × reviewed score + 0.09 × intent recovery) / 0.33, where both terms are the equal mean of the two judges, or the reviewed score alone where intent recovery has no denominator. The combined score is 100 × (0.50F + 0.085Q + 0.085S + 0.33C / 100). The reproduction guide gives the public arithmetic checks and links the protocol documents, controls, runner and operator wrapper.
No human rated anything; the scores are model judgment for a human reader, not human validation. Runs were not repeated, and effort labels are Codex's own. The Max cell's Code quality and combined score are means over 22 of 23 runs, for the reason given under the effort table; the frozen summary's publication flag is false only because of that strict count, and the export asserts the shortfall is exactly that row. Hash checks bind the export to frozen files, including the operator's invalid-probe marker; they do not prove that judges were unbiased or that no training overlap exists.
Withheld: raw prompts and responses, submitted patches and reconstructed sources, the quirk answer keys (they describe hidden-test behaviour), and reviewer session identifiers and usage receipts. Summed solver time is 21.11 hours; judging is excluded.