September 30, 2026 · 115 runs · Code quality protocol v3.17
GPT-6 Sol across every effort level
115 runs of GPT-6 Sol through the Codex CLI on a ChatGPT Pro subscription, 23 tasks at each of five effort levels, Low to Max. GPT-6 Sol is a new model, not the GPT-5.6 Sol already on the board; with GPT-6 Luna it is the second GPT-6 model on Frontier v4. Code quality carries 33% of the combined score and is judged for a human reader by Muse Spark 1.3 and Grok 4.6 under the same frozen protocol as the other Frontier v4 reports, so GPT-6 Sol takes its place on the Frontier v4 leaderboard. Every run finished inside the 3-hour task bound.
How the test works
A retired program to rebuild
Each of the 23 tasks is a legacy binary with a written specification and an issue. Its real behaviour departs from the specification in documented ways, and hidden tests check the rewrite against the binary.
One attempt per task and level
GPT-6 Sol works in Codex with the binary, the specification and the code, once per task at each effort level, inside a flat 3-hour bound.
Tests and two judges score it
Hidden tests give functional correctness. Two calibrated judges from labs with no model on the board rate the code for a human reader and try to recover the documented departures.
Scores, time and cost per level
Each level gets a combined score, tasks passed, minutes and API-equivalent cost. A run the judges could not score still counts for tasks passed, time and cost.
Combined score = 50% functional correctness + 8.5% lint and complexity + 8.5% security + 33% Code quality, averaged over the judged runs in each level, as on every other board column. Tasks passed, minutes and cost cover all 23 runs.
Across every effort level
GPT-6 Sol ran every task once at each of its five effort levels, 23 tasks per cell, through Codex CLI 0.155.0 on a ChatGPT Pro subscription from September 28 to 29, 2026. Combined score weights are 50% functional correctness, 8.5% lint and complexity, 8.5% security and 33% Code quality, judged for a named human reader by Muse Spark 1.3 and Grok 4.6 with a ground-truth intent-recovery probe. Judging ran September 29 to 30. No run reached the flat 3-hour task bound; the longest, Max paddockcore, took 129 minutes. One Medium run, codeccore, has no published Code quality score because Grok 4.6's intent probe gave no valid answer, so the Medium cell's Code quality and combined score cover 22 of its 23 runs. That run failed its tests and counts in every pass count, runtime and cost figure.
Combined score
67.62 at Low to 86.82 at Max
The score rises at every step: 13.0 points from Low to Medium, then 2.2 to High, 3.1 to Extra-high and 0.9 to Max. Extra-high and Max sit within one standard error of each other.
Tasks passed
4 of 23 at Low to 19 of 23 at Max
Functional correctness drives the curve: 4 tasks pass at Low, 13 at Medium, 15 at High, 18 at Extra-high and 19 at Max.
Code quality
62.53 to 70.78
The judges rate the code higher as effort rises, from 62.53 at Low to about 70 from High up, with standard errors of 1.3 to 1.6.
Runtime and cost
13.0 to 24.5 min · $1.02 to $2.52 per task
Low is the fastest and cheapest level. High costs more than Extra-high ($2.33 against $2.00), and Max is the slowest and most expensive. The whole sweep prices at $221.38, $1.93 per task.
| Metric | Low | Medium | High | Extra-high | Max |
|---|---|---|---|---|---|
| Combined score | 67.62 | 80.60 | 82.83 | 85.94 | 86.82 |
| Standard error | 3.59 | 2.59 | 2.43 | 1.38 | 0.91 |
| Judged runs | 23 | 22 | 23 | 23 | 23 |
| Tasks passed | 4/23 | 13/23 | 15/23 | 18/23 | 19/23 |
| Code quality | 62.53 | 66.79 | 70.09 | 70.77 | 70.78 |
| Human readability | 53.8 | 60.0 | 62.1 | 61.7 | 63.4 |
| Minutes per task | 13.0 | 17.5 | 23.8 | 20.9 | 24.5 |
| API-equivalent $ per task | $1.02 | $1.76 | $2.33 | $2.00 | $2.52 |
Minutes and cost per task cover all 23 runs in each cell, including the unscored Medium run. Under the prior 50/15/15/20 profile applied to the same Code quality scores the ladder has the same shape, 71.23 to 89.19. The card and the full per-effort table follow.
GPT-6 Sol gains most of its range in the first step of the ladder, from Low to Medium, and keeps climbing more slowly to Max. Both functional correctness and Code quality rise with effort. These runs compare a model-and-harness combination, not an isolated base model; the Frontier v4 leaderboard places every GPT-6 Sol effort level beside the other models judged under the same protocol.
Combined score and Code quality by effort
n=23 at Low, High, Extra-high and Max; n=22 at Medium, where one run has no published Code quality score (see the note below). Combined (33%) is the published score. Combined (20% profile) applies the prior 50/15/15/20 weights to the same Code quality scores for comparison. Code quality, human readability, maintainability and intent recovery are out of 100 and equal means of the two judges over judged runs. Passed counts all 23 runs. Runtime is solver wall-clock per task over every run in the cell.
| Model / harness | Effort | Combined (33%) | SE | Combined (20% profile) | Code quality | Human readability | Maintainability | Intent recovery | Judged | Passed | Min/task |
|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-6 Sol / Codex | Low | 67.62 | 3.59 | 71.23 | 62.53 | 53.8 | 65.5 | 70.2 | 23/23 | 4/23 | 13.0 |
| GPT-6 Sol / Codex | Medium | 80.60 | 2.59 | 83.77 | 66.79 | 60.0 | 68.5 | 73.6 | 22/23‖ | 13/23 | 17.5 |
| GPT-6 Sol / Codex | High | 82.83 | 2.43 | 85.46 | 70.09 | 62.1 | 72.5 | 77.5 | 23/23 | 15/23 | 23.8 |
| GPT-6 Sol / Codex | Extra-high | 85.94 | 1.38 | 88.38 | 70.77 | 61.7 | 72.7 | 80.3 | 23/23 | 18/23 | 20.9 |
| GPT-6 Sol / Codex | Max | 86.82 | 0.91 | 89.19 | 70.78 | 63.4 | 73.6 | 76.9 | 23/23 | 19/23 | 24.5 |
‖Medium is judged on 22 of 23 tasks. On codeccore, Grok 4.6's intent probe produced no valid answer after the protocol's single retry: the first attempt was not valid JSON, and the second quoted a line that is not in the code, memo = record[31:46].rstrip(". ,".replace(" ", "")) if False else record[31:46].rstrip(".,"). No recovery rule accepts an invented excerpt, so no answer-key match was made, and the frozen summary, which requires a valid match from every passing panel, publishes no Code quality or combined score for that run. It is the same codeccore line Grok 4.6 misquoted for GPT-5.6 Sol at Max under v3.7, with the same outcome. Both judges' reviews and the Muse probe are retained in the harness archive. The run failed its tests, so it counts as a failed task, and its runtime, tokens and cost are in every economics figure. The Medium cell's Code quality, combined score and standard error are means over the 22 judged runs.
Intent recovery is scored over the specification departures a submission actually passed tests for. Every judged GPT-6 Sol run passed at least one, so no run fell back to the reviewed score alone. SE is one sample standard error of the combined score across the cell's tasks. It describes task sampling, not judge uncertainty, repeated runs or a significance test. No solver fallback occurred, and neither judge produced a reviewer fallback on any of the 115 submissions.
Cost and tokens across effort levels
All 115 runs priced from their Codex receipts at list API rates, cache-aware, with judging excluded. These are API-equivalent estimates, not subscription charges: GPT-6 Sol ran on a ChatGPT Pro subscription, and what the plan actually bills is not observable from the receipts. Cost per task is $1.02 at Low, $1.76 at Medium, $2.33 at High, $2.00 at Extra-high and $2.52 at Max, following token use: High uses more tokens per task than Extra-high. The whole sweep comes to $221.38 for 115 runs, every one of them priced, including the Medium run without a published Code quality score.
| Effort | Runs | $/task | Level total | Tokens/task | Min/task |
|---|---|---|---|---|---|
| Low | 23 | $1.02 | $23.45 | 3.26M | 13.0 |
| Medium | 23 | $1.76 | $40.44 | 6.16M | 17.5 |
| High | 23 | $2.33 | $53.60 | 8.28M | 23.8 |
| Extra-high | 23 | $2.00 | $45.91 | 6.68M | 20.9 |
| Max | 23 | $2.52 | $57.97 | 8.32M | 24.5 |
| Full sweep | 115 | $1.93 | $221.38 | 752M | 38.2 h |
Pricing. Rates checked 2026-09-25 from OpenAI: GPT-6 Sol at $2.00 input, $0.20 cached input and $10.00 output per million tokens. Codex receipts report input, cached input, output and reasoning tokens per run, so cached input is billed at the cache-read rate and the rest at the standard rate. OpenAI bills prompts over 272K input tokens at a higher rate, but Codex keeps each request within its 272K context window, so no long-context premium applies. The sweep stamped each run at these rates at run time; the export recomputed every run from its receipt and matched every stamp.
Tokens are Codex's raw totals including cache reads. Runtime covers all 115 runs, 38.2 hours in all. Per-run estimates, token usage breakdowns and the rate table are in economics.json and runs.csv.
Beside GPT-5.6 Sol
GPT-5.6 Sol ran the same 23 tasks through Codex in September and was judged by the same two judges under v3.7, with Max judged on 22 of 23 after an invalid Grok probe on the same codeccore line. GPT-6 Sol's combined score is slightly behind GPT-5.6 Sol's at every level, by 1.43 points at Low, 1.96 at Medium, 2.93 at High, 0.53 at Extra-high and 0.36 at Max; the gaps at Extra-high and Max are within one standard error of either model. GPT-6 Sol passes one more task at Medium (13 against 12) and fewer at every other level. Its Code quality is higher from High up (70.09 to 70.78 against 67.31 to 67.65).
| Metric | Low | Medium | High | Extra-high | Max |
|---|---|---|---|---|---|
| Combined, GPT-6 Sol | 67.62 | 80.60 | 82.83 | 85.94 | 86.82 |
| Combined, GPT-5.6 Sol | 69.05 | 82.56 | 85.76 | 86.47 | 87.18 |
| Standard error, GPT-6 Sol | 3.59 | 2.59 | 2.43 | 1.38 | 0.91 |
| Standard error, GPT-5.6 Sol | 3.50 | 1.25 | 0.98 | 0.80 | 0.62 |
| Tasks passed, GPT-6 Sol | 4/23 | 13/23 | 15/23 | 18/23 | 19/23 |
| Tasks passed, GPT-5.6 Sol | 6/23 | 12/23 | 20/23 | 21/23 | 22/23 |
| Code quality, GPT-6 Sol | 62.53 | 66.79 | 70.09 | 70.77 | 70.78 |
| Code quality, GPT-5.6 Sol | 64.98 | 68.19 | 67.65 | 67.31 | 67.58 |
| Minutes per task, GPT-6 Sol | 13.0 | 17.5 | 23.8 | 20.9 | 24.5 |
| Minutes per task, GPT-5.6 Sol | 9.8 | 11.1 | 12.0 | 10.6 | 11.6 |
| $ per task, GPT-6 Sol | $1.02 | $1.76 | $2.33 | $2.00 | $2.52 |
| $ per task, GPT-5.6 Sol | $1.36 | $1.72 | $2.05 | $1.61 | $1.77 |
| Tokens per task, GPT-6 Sol | 3.26M | 6.16M | 8.28M | 6.68M | 8.32M |
| Tokens per task, GPT-5.6 Sol | 1.79M | 2.26M | 2.83M | 1.91M | 2.11M |
GPT-6 Sol is slower at every level, 1.3 times GPT-5.6 Sol's mean minutes at Low and 2.1 times at Max. Its list price is half of GPT-5.6 Sol's ($2.00 against $4.00 input and $10.00 against $20.00 output per million tokens), but it uses 1.8 to 3.9 times the raw tokens per task, so it costs less per task only at Low. At Medium the two are close ($1.76 against $1.72); from High up GPT-6 Sol costs 14% to 42% more per task. Two cautions on the speed figures: GPT-6 Sol's sweep overlapped GPT-6 Luna's judging on the same machine (see How it was run), and the two sweeps ran on different Codex CLI versions (0.155.0 and 0.153.4). Each model was judged in its own protocol run of the same frozen rubric.
The GPT-6 family beside GPT-5.6
With GPT-6 Sol judged, both GPT-6 models on Frontier v4 can be set beside the GPT-5.6 models they follow. On judged runs, neither GPT-6 model scores above its GPT-5.6 predecessor at any effort level. The Sol gap is small, 0.4 to 2.9 points. The Luna gap is 0.3 to 0.5 points at Low and Extra-high, but 7.6 points at Medium, 13.9 at High and 2.7 at Max, before counting GPT-6 Luna's six 3-hour timeouts, which lower its Extra-high and Max figures to 71.94 and 67.26 when counted as 0. On price the two lines move apart: GPT-6 Luna costs less per priced task than GPT-5.6 Luna from Medium up, while GPT-6 Sol costs more than GPT-5.6 Sol from Medium up. The GPT-6 Luna report has its own comparison with GPT-5.6 Luna.
What Code quality measures
Every judge receives the same rubric, frozen by hash before any review. It names the reader it scores for: an engineer who has never seen the code, reads it top to bottom without running it, and must make a correct change in one sitting. It tells the judge that its own ease at parsing dense code is not evidence of readability. Six dimensions are scored 0 to 4 in half steps against written anchors: naming, presentation and intent form the human readability sub-score; structure, changeability and verifiability form the maintainability sub-score. Every score must cite an exact excerpt from the code and a concrete consequence for that reader. The host computes the sub-scores; the judge does no arithmetic.
The 33 points have three designed layers. The reviewed panel carries 24 points. An intent-recovery probe carries 9: the judge, given only the specification and the code, lists where the code departs from the specification, and a separate call matches that list against a frozen answer key, over the departures the submission passed tests for. The third layer, measured maintenance, is designed for 12 points but is not yet built, so the pre-registered 24 plus 9 split is in force.
| Component | Low | Medium | High | Extra-high | Max |
|---|---|---|---|---|---|
| Code quality | 62.53 | 66.79 | 70.09 | 70.77 | 70.78 |
| Human readability | 53.8 | 60.0 | 62.1 | 61.7 | 63.4 |
| Maintainability | 65.5 | 68.5 | 72.5 | 72.7 | 73.6 |
| Intent recovery | 70.2 | 73.6 | 77.5 | 80.3 | 76.9 |
| Rated by Muse Spark 1.3 | 57.9 | 63.1 | 66.6 | 66.7 | 68.2 |
| Rated by Grok 4.6 | 61.4 | 65.4 | 68.0 | 67.8 | 68.8 |
| Standard error of Code quality | 1.61 | 1.49 | 1.57 | 1.34 | 1.40 |
Code quality rises from 62.53 at Low to 70.09 at High and holds at about 70.8 at Extra-high and Max, with Grok 4.6 the higher rater at every level and the two judges within 3.6 points of each other in every cell. Unlike GPT-5.6 Sol, whose Code quality stayed within about three points at every effort, GPT-6 Sol's human readability and maintainability both rise with effort. Per-run sub-scores from each judge are in runs.csv; the exact rubric, system text, probe and match instructions are in judge-protocols.json.
Judges and calibration
The scored panel is Muse Spark 1.3 (Meta) through the Muse CLI on its Standard tier, and Grok 4.6 (xAI) through the Cursor CLI at medium effort, with equal weight, the same pinned binaries as under v3.7. Neither lab has a model on this board, so neither judge grades a relative. Every session is fresh, tools are disabled, the workspace is empty, model identity is checked per call from the CLI's own records, and solver labels are withheld. Protocol v3.17 changes nothing in the rubric, controls, quirk keys, gates, repeats, seed, weights or judges from v3.7; it applies the protocol to this population, and both judges retook the calibration exam under it before any counted call.
Before scoring a single submission, each judge reviews ten held-out programs that implement the same ledger specification: clear, compressed, compressed then auto-formatted, verbose with duplicated policy, needlessly abstracted, misleadingly commented, narrated with a comment on every line, a documented legacy quirk, an embedded instruction to give full marks, and hidden module state. Each program is reviewed five times in a seeded order. Twenty gates fixed in advance check that the judge sees the construct; a judge may miss at most one gate by at most half a point.
| Judge | Protocol | Calls | Result | Allowance used | Failing gates |
|---|---|---|---|---|---|
| Muse Spark 1.3 | code-quality-maintenance-v3.17 | 80 | passed | no | none |
| Grok 4.6 | code-quality-maintenance-v3.17 | 80 | passed | no | none |
Both judges passed every gate with no allowance used; under v3.16 each had needed the pre-registered one-gate allowance. Every gate value and control mean is in calibration.json.
Under v3.17 Muse Spark 1.3 made 440 counted calls: 80 in calibration, 115 primary reviews, 5 repeats, 10 pairwise checks, 115 intent probes and 115 answer-key matches. Grok 4.6 made 439, with 114 matches, since no match was made for the invalid codeccore probe. Muse needed a second attempt on six calls, all unsupported evidence excerpts: three primary reviews and three probes. Grok needed a second attempt on five: three primary reviews whose first reply was a complete JSON review followed by a stray closing brace, and two probes (one unsupported excerpt, answered validly on the retry, and the codeccore probe described above). Every Grok call reported the display label "Grok 4.6 Medium" for the pinned model id, which the operator wrapper's display-rename rule accepts. Neither judge produced a reviewer fallback. Every attempt is archived beside its replacement in the harness run directory.
How it was run
- Codex CLI 0.155.0. The earlier Codex columns on the board ran on 0.153.4, which refuses GPT-6 models on a ChatGPT account before any work. 0.155.0 is the lowest release that serves them, so it was installed beside the global CLI for the GPT-6 Luna and Sol sweeps only; incumbent Codex columns are not re-baselined.
- A formatting-only recovery on one Grok review. On High payrollcore both of Grok 4.6's attempts at the primary review returned a complete, valid JSON review followed by one stray closing brace, so parsing failed before any protocol check. A new operator rule, recorded in the operations log, decodes the first JSON object, drops a remainder that is only whitespace and closing braces, and then applies every normal check (session, subscription guard, no tool use, usage, display label, schema and excerpts). The first attempt passed and was selected, with a reviewed score of 79.17. No field was edited, and the dropped text is kept in the receipt.
- One invalid Grok probe. The Grok 4.6 intent probe on Medium codeccore had no valid answer, one malformed attempt and one quoting code absent from the run, so Medium is judged on 22 of 23, as described under the results table.
- The sweep overlapped another model's judging. The GPT-6 Sol solver sweep ran from September 28, 00:31 PDT, to September 29, 14:53 PDT. At the owner's request it ran alongside GPT-6 Luna's Code quality judging, September 28, 00:32 to 14:32 PDT, on the same machine. That matters for the wall-clock speed figures, especially at Low, Medium and early High, which ran inside that window.
- One infrastructure retry. The first Max depotcore attempt ended after eight minutes when the API answered that the model was at capacity. The harness records that as an infrastructure error and re-queues the task; the retry is the counted run, and the failed attempt is not part of any cell.
- Judging window. Judging ran from September 29, 14:54 PDT, to September 30, 06:51 PDT, resumed after the two Grok findings above. GPT-6.1 Sol's solver sweep ran during that window at the owner's request; the judges share no quota with it, and the GPT-6 Sol runs had finished.
- One attempt per task and level. Runs were not repeated, and effort labels are Codex's own.
Evidence and reproduction
Public files contain all 115 run measurements, both judges' six-dimension sub-scores and intent-recovery scores per published run, raw tokens and API-equivalent cost estimates per run, five aggregates, every calibration gate value, the exact judge protocol text, the operator's finding on the unpublished row, the one trailing-brace recovery and source hashes. Each run's Code quality is (0.24 × reviewed score + 0.09 × intent recovery) / 0.33, where both terms are the equal mean of the two judges, or the reviewed score alone where intent recovery has no denominator. The combined score is 100 × (0.50F + 0.085Q + 0.085S + 0.33C / 100). The reproduction guide gives the public arithmetic checks and links the protocol documents, controls, runner, operator wrapper and decision records.
No human rated anything; the scores are model judgment for a human reader, not human validation. The Medium cell's Code quality and combined score are means over 22 of 23 runs, for the reason given under the results table; the frozen summary's publication flag is false only because of that strict count, and the export asserts the shortfall is exactly that row. Hash checks bind the export to frozen files, including the operator's invalid-probe marker and the recovered review; they do not prove that judges were unbiased or that no training overlap exists.
Withheld: raw prompts and responses, submitted patches and reconstructed sources, the quirk answer keys (they describe hidden-test behaviour), and reviewer session identifiers and usage receipts. Summed solver time is 38.22 hours; judging is excluded.