September 28, 2026 · 115 runs · Code quality protocol v3.16
GPT-6 Luna across every effort level
115 runs of GPT-6 Luna through the Codex CLI on a ChatGPT Pro subscription, 23 tasks at each of the five effort levels Codex offers for it, Low to Max. GPT-6 Luna is a new model, not the GPT-5.6 Luna already on the board. Code quality carries 33% of the combined score and is judged for a human reader by Muse Spark 1.3 and Grok 4.6 under the same frozen protocol as the other Frontier v4 reports, so GPT-6 Luna takes its place on the Frontier v4 leaderboard. Six runs at the top two levels hit the 3-hour task bound, so every combined score on this page is given two ways: over the judged runs, and with each timeout counted as 0.
How the test works
A retired program to rebuild
Each of the 23 tasks is a legacy binary with a written specification and an issue. Its real behaviour departs from the specification in documented ways, and hidden tests check the rewrite against the binary.
One attempt per task and level
GPT-6 Luna works in Codex with the binary, the specification and the code, once per task at each effort level, inside a flat 3-hour bound.
Tests and two judges score it
Hidden tests give functional correctness. Two calibrated judges from labs with no model on the board rate the code for a human reader and try to recover the documented departures.
Scores, time and cost per level
Each level gets a combined score, tasks passed, minutes and API-equivalent cost. A run stopped at the bound has no code to judge, so it counts as a failed task.
Combined score = 50% functional correctness + 8.5% lint and complexity + 8.5% security + 33% Code quality. The standard figure averages the judged runs, as on every other board column. The second figure counts each timed-out run as a combined score of 0 over all 23 runs. Tasks passed always counts timeouts as failures.
Across every effort level
GPT-6 Luna ran every task once at each of its five effort levels, 23 tasks per cell, through Codex CLI 0.155.0 on a ChatGPT Pro subscription from September 25 to 28, 2026. Combined score weights are 50% functional correctness, 8.5% lint and complexity, 8.5% security and 33% Code quality, judged for a named human reader by Muse Spark 1.3 and Grok 4.6 with a ground-truth intent-recovery probe. Judging ran on September 28. Six runs hit the flat 3-hour task bound while still working, two at Extra-high and four at Max. They have no finished code to judge, so the standard combined score and Code quality at those levels cover 21 and 19 runs; the second combined figure counts each timeout as 0 over all 23.
Combined score
40.83 at Low to 81.42 at Max
Over judged runs the score rises at every step: 4.9 points from Low to Medium, then 11.6 to High, 21.4 to Extra-high and 2.6 to Max. Counting timeouts as 0 gives 71.94 at Extra-high and 67.26 at Max, so the top level falls back below Extra-high.
Tasks passed
0 of 23 at Low to 12 of 23 at Max
No task passes at Low or Medium and 2 pass at High. Extra-high passes 11 and Max 12, every timeout counted as a failure.
Code quality
65.84 to 71.51
The judges rate the code between 65.84 and 69.33 from Low to High and about 71.5 at Extra-high and Max, with standard errors of 1.0 to 2.2.
Runtime and cost
3.6 to 60.5 min · $0.010 to $0.160 per task
Medium is the fastest and cheapest level. Runtime climbs to 52.6 minutes at Extra-high and 60.5 at Max, where the timeouts count at 180 minutes each. The whole sweep prices at $7.26 for its 109 finished runs.
| Metric | Low | Medium | High | Extra-high | Max |
|---|---|---|---|---|---|
| Combined score, judged runs | 40.83 | 45.72 | 57.35 | 78.79 | 81.42 |
| Combined score, timeouts counted as 0 | 40.83 | 45.72 | 57.35 | 71.94 | 67.26 |
| Standard error, judged runs | 2.04 | 1.32 | 3.80 | 3.02 | 2.88 |
| Judged runs | 23 | 23 | 23 | 21 | 19 |
| Tasks passed | 0/23 | 0/23 | 2/23 | 11/23 | 12/23 |
| Runs stopped at the 3-hour bound | 0 | 0 | 0 | 2 | 4 |
| Code quality | 65.84 | 69.33 | 66.40 | 71.49 | 71.51 |
| Human readability | 59.8 | 64.9 | 60.3 | 67.4 | 66.9 |
| Minutes per task | 5.6 | 3.6 | 15.6 | 52.6 | 60.5 |
| API-equivalent $ per task | $0.028 | $0.010 | $0.052 | $0.160 | $0.098 |
Minutes per task cover all 23 runs, timeouts at their recorded 180 minutes. Cost per task covers the priced runs only (21 at Extra-high, 19 at Max), because the timeouts ended before Codex reported usage. Under the prior 50/15/15/20 profile applied to the same Code quality scores the judged ladder has the same shape, 44.00 to 83.83. The card and the full per-effort table follow.
GPT-6 Luna needs effort to do this work. It passes nothing at Low or Medium, where the combined score is mostly lint, security and Code quality, and the large gains come between High and Extra-high. At Max the judged score is the highest, but four of the 23 runs ran out of time, so with timeouts counted as 0 Max scores below Extra-high. These runs compare a model-and-harness combination, not an isolated base model; the Frontier v4 leaderboard places every GPT-6 Luna effort level beside the other models judged under the same protocol, with the timeouts-as-0 figures in a footnote.
Combined score and Code quality by effort
Combined (33%) is the standard score over judged runs. Timeouts as 0 counts each run stopped at the 3-hour bound as a combined score of 0 over all 23 runs. Combined (20% profile) applies the prior 50/15/15/20 weights to the same Code quality scores for comparison. Code quality, human readability, maintainability and intent recovery are out of 100 and equal means of the two judges over judged runs. Passed counts all 23 runs, each timeout a failure. Runtime is solver wall-clock per task over every run in the cell.
| Model / harness | Effort | Combined (33%) | SE | Timeouts as 0 | Combined (20% profile) | Code quality | Human readability | Maintainability | Intent recovery | Judged | Passed | Min/task |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-6 Luna / Codex | Low | 40.83 | 2.04 | 40.83 | 44.00 | 65.84 | 59.8 | 66.3 | 80.2 | 23/23 | 0/23 | 5.6 |
| GPT-6 Luna / Codex | Medium | 45.72 | 1.32 | 45.72 | 48.47 | 69.33 | 64.9 | 72.0 | 72.8 | 23/23 | 0/23 | 3.6 |
| GPT-6 Luna / Codex | High | 57.35 | 3.80 | 57.35 | 60.41 | 66.40 | 60.3 | 67.5 | 73.8 | 23/23 | 2/23 | 15.6 |
| GPT-6 Luna / Codex | Extra-high | 78.79 | 3.02 | 71.94 | 81.21 | 71.49 | 67.4 | 72.1 | 76.1 | 21/23 | 11/23 | 52.6 |
| GPT-6 Luna / Codex | Max | 81.42 | 2.88 | 67.26 | 83.83 | 71.51 | 66.9 | 72.3 | 76.7 | 19/23 | 12/23 | 60.5 |
Timeouts. Extra-high is judged on 21 of 23 tasks and Max on 19 of 23. The other six runs were stopped at the 3-hour bound with no finished submission, so there is no code to judge. The standard error of the timeouts-as-0 figure is 5.47 at Extra-high and 6.99 at Max, larger than the judged figure's because the zeros widen the spread. The next section lists the six runs.
Intent recovery is scored over the specification departures a submission actually passed tests for. Fourteen judged runs passed none (10 at Low, 3 at Medium, 1 at High), so intent recovery has no denominator for them and their Code quality is the reviewed score alone, under the protocol's pre-registered rule. The Low intent-recovery mean therefore covers 13 runs. SE is one sample standard error of the combined score across the cell's tasks. It describes task sampling, not judge uncertainty, repeated runs or a significance test. No solver fallback occurred, and neither judge produced a reviewer fallback on any of the 109 submissions.
Six runs stopped at the 3-hour bound
Every Frontier v4 task has carried a flat 3-hour bound since the harness decision of September 13, 2026 (it was 10 hours before). GPT-6 Luna reached it six times, all at the two top levels. In each case the trace shows continuous work up to the bound, with no idle gap longer than four minutes, so these are runs that ran out of time, not stalls.
| Effort | Task | Minutes | Judged | Counted as |
|---|---|---|---|---|
| Extra-high | depotcore | 180.0 | no | failed task, combined 0 in the second figure, unpriced |
| Extra-high | paddockcore | 180.0 | no | failed task, combined 0 in the second figure, unpriced |
| Max | cellarcore | 180.0 | no | failed task, combined 0 in the second figure, unpriced |
| Max | depotcore | 180.0 | no | failed task, combined 0 in the second figure, unpriced |
| Max | lodgecore | 180.0 | no | failed task, combined 0 in the second figure, unpriced |
| Max | paddockcore | 180.0 | no | failed task, combined 0 in the second figure, unpriced |
With four of 23 Max runs out of the judged set, all of them failures, the judged mean alone flatters the top levels. The owner's decision, recorded in the harness decision log on September 28, is to publish both figures everywhere a score appears. The standard combined score over judged runs stays the headline cell value, because it is computed the same way as every other column on the board; the timeouts-as-0 figure sits beside it. Tasks passed and runtime cover all 23 runs, and depotcore and paddockcore timed out at both levels.
Cost and tokens across effort levels
The 109 finished runs priced from their Codex receipts at list API rates, cache-aware, with judging excluded. These are API-equivalent estimates, not subscription charges: GPT-6 Luna ran on a ChatGPT Pro subscription, and what the plan actually bills is not observable from the receipts. The six timeouts ended before Codex reported usage, so they are unpriced, not $0, and Extra-high and Max spend is understated. Cost per task is $0.028 at Low, $0.010 at Medium, $0.052 at High, $0.160 at Extra-high and $0.098 at Max. The whole sweep comes to $7.26 for 109 priced runs.
| Effort | Priced runs | $/task | Level total | Tokens/task | Min/task |
|---|---|---|---|---|---|
| Low | 23 | $0.028 | $0.64 | 1.65M | 5.6 |
| Medium | 23 | $0.010 | $0.22 | 0.40M | 3.6 |
| High | 23 | $0.052 | $1.19 | 2.56M | 15.6 |
| Extra-high | 21 | $0.160 | $3.35 | 8.47M | 52.6 |
| Max | 19 | $0.098 | $1.85 | 4.32M | 60.5 |
| Full sweep | 109 | $0.067 | $7.26 | 366M | 52.8 h |
Pricing. Rates checked 2026-09-25 from OpenAI: GPT-6 Luna at $0.10 input, $0.01 cached input and $0.50 output per million tokens. Codex receipts report input, cached input, output and reasoning tokens per run, so cached input is billed at the cache-read rate and the rest at the standard rate. OpenAI bills prompts over 272K input tokens at a higher rate, but Codex keeps each request within its 272K context window, so no long-context premium applies. The sweep stamped each run at these rates at run time; the export recomputed every priced run from its receipt and matched every stamp.
Tokens are Codex's raw totals including cache reads, over priced runs. Runtime covers all 115 runs, 52.8 hours in all. Per-run estimates, token usage breakdowns and the rate table are in economics.json and runs.csv.
Beside GPT-5.6 Luna
GPT-5.6 Luna ran the same 23 tasks through Codex in September and was judged by the same two judges under v3.5. It passed 0, 1, 9, 13 and 19 of 23 from Low to Max, with no timeouts. GPT-6 Luna passes fewer tasks at every level from Medium up, and is cheaper per task at Medium and above.
| Metric | Low | Medium | High | Extra-high | Max |
|---|---|---|---|---|---|
| Tasks passed, GPT-6 Luna | 0/23 | 0/23 | 2/23 | 11/23 | 12/23 |
| Tasks passed, GPT-5.6 Luna | 0/23 | 1/23 | 9/23 | 13/23 | 19/23 |
| Combined, GPT-6 Luna (judged runs) | 40.83 | 45.72 | 57.35 | 78.79 | 81.42 |
| Combined, GPT-6 Luna (timeouts as 0) | 40.83 | 45.72 | 57.35 | 71.94 | 67.26 |
| Combined, GPT-5.6 Luna | 41.29 | 53.28 | 71.22 | 79.11 | 84.15 |
| Median minutes, GPT-6 Luna | 3.6 | 3.1 | 13.3 | 34.8 | 30.1 |
| Median minutes, GPT-5.6 Luna | 2.6 | 6.9 | 17.4 | 21.3 | 22.3 |
| $ per task, GPT-6 Luna (priced runs) | $0.028 | $0.010 | $0.052 | $0.160 | $0.098 |
| $ per task, GPT-5.6 Luna | $0.024 | $0.066 | $0.217 | $0.275 | $0.448 |
One observation from the traces, at Medium: GPT-6 Luna makes one round of fixes and stops, often saying so. In 8 of the 23 Medium runs its final message says the work is not finished (mismatches remain, or the change is partial or not verified as a drop-in replacement), and in 2 more it says it could not run the legacy comparison. Medium runs take a median 3.1 minutes and 4,531 output tokens, against 6.9 minutes and 19,069 for GPT-5.6 Luna. We read this as a difference in how long the model keeps working at a given effort label, not as a measured cause of the score gap. The two sweeps ran on different Codex CLI versions (0.155.0 and 0.153.4) and were judged under different protocol versions of the same frozen rubric, and one GPT-5.6 Luna Max run (paddockcore, which passed) took 196 minutes, longer than the 3-hour bound GPT-6 Luna ran under.
What Code quality measures
Every judge receives the same rubric, frozen by hash before any review. It names the reader it scores for: an engineer who has never seen the code, reads it top to bottom without running it, and must make a correct change in one sitting. It tells the judge that its own ease at parsing dense code is not evidence of readability. Six dimensions are scored 0 to 4 in half steps against written anchors: naming, presentation and intent form the human readability sub-score; structure, changeability and verifiability form the maintainability sub-score. Every score must cite an exact excerpt from the code and a concrete consequence for that reader. The host computes the sub-scores; the judge does no arithmetic.
The 33 points have three designed layers. The reviewed panel carries 24 points. An intent-recovery probe carries 9: the judge, given only the specification and the code, lists where the code departs from the specification, and a separate call matches that list against a frozen answer key, over the departures the submission passed tests for. The third layer, measured maintenance, is designed for 12 points but is not yet built, so the pre-registered 24 plus 9 split is in force.
| Component | Low | Medium | High | Extra-high | Max |
|---|---|---|---|---|---|
| Code quality | 65.84 | 69.33 | 66.40 | 71.49 | 71.51 |
| Human readability | 59.8 | 64.9 | 60.3 | 67.4 | 66.9 |
| Maintainability | 66.3 | 72.0 | 67.5 | 72.1 | 72.3 |
| Intent recovery | 80.2 | 72.8 | 73.8 | 76.1 | 76.7 |
| Rated by Muse Spark 1.3 | 62.8 | 67.4 | 63.0 | 68.0 | 68.5 |
| Rated by Grok 4.6 | 63.3 | 69.5 | 64.8 | 71.5 | 70.6 |
| Standard error of Code quality | 2.24 | 1.78 | 1.50 | 1.02 | 1.52 |
Code quality sits between 65.84 and 71.51, with Grok 4.6 the higher rater at every level and the two judges within 3.6 points of each other in every cell. Low's intent recovery covers only the 13 Low runs that passed any documented departure, so it is not comparable with the other levels. Per-run sub-scores from each judge are in runs.csv; the exact rubric, system text, probe and match instructions are in judge-protocols.json.
Judges and calibration
The scored panel is Muse Spark 1.3 (Meta) through the Muse CLI on its Standard tier, and Grok 4.6 (xAI) through the Cursor CLI at medium effort, with equal weight, the same pinned binaries as under v3.7. Neither lab has a model on this board, so neither judge grades a relative. Every session is fresh, tools are disabled, the workspace is empty, model identity is checked per call from the CLI's own records, and solver labels are withheld. Protocol v3.16 changes nothing in the rubric, controls, quirk keys, gates, repeats, seed, weights or judges from v3.7; it applies the protocol to this population, and both judges retook the calibration exam under it before any counted call.
Before scoring a single submission, each judge reviews ten held-out programs that implement the same ledger specification: clear, compressed, compressed then auto-formatted, verbose with duplicated policy, needlessly abstracted, misleadingly commented, narrated with a comment on every line, a documented legacy quirk, an embedded instruction to give full marks, and hidden module state. Each program is reviewed five times in a seeded order. Twenty gates fixed in advance check that the judge sees the construct; a judge may miss at most one gate by at most half a point.
| Judge | Protocol | Calls | Result | Allowance used | Failing gate |
|---|---|---|---|---|---|
| Muse Spark 1.3 | code-quality-maintenance-v3.16 | 80 | passed | yes | gate 11, repeatability, short by 0.06 |
| Grok 4.6 | code-quality-maintenance-v3.16 | 80 | passed | yes | gate 4, formatting is presentation, short by 0.10 |
Both judges passed, each using the pre-registered one-gate allowance, well inside the half-point bound. Every other gate passed for both. Every gate value and control mean is in calibration.json.
Under v3.16 each judge made 420 counted calls: 80 in calibration, 109 primary reviews, 5 repeats, 8 pairwise checks, 109 intent probes and 109 answer-key matches. Muse Spark 1.3 needed a second attempt on six calls: four primary reviews (two unsupported evidence excerpts and two malformed JSON responses) and two probes (unsupported excerpts). Grok 4.6 needed a second attempt on two primary reviews, both unsupported excerpts. Every Grok call reported the display label "Grok 4.6 Medium" for the pinned model id, which the operator wrapper's display-rename rule accepts, as under v3.15. All 109 submissions are published, with no invalid probe and no reviewer fallback. Every attempt is archived beside its replacement in the harness run directory.
How it was run
- Codex CLI 0.155.0. Every other Codex column on the board ran on 0.153.4, which refuses GPT-6 Luna on a ChatGPT account before any work. 0.155.0 is the lowest release that serves it, so it was installed beside the global CLI for this sweep only; incumbent Codex columns are not re-baselined.
- Operator stop at Extra-high pacecore. The first Extra-high pacecore run went silent two minutes in, after a Codex reconnect message, and made no further progress. After about 88 idle minutes the operator stopped the Codex process, which the harness records as an infrastructure error and re-queues. The retry is the counted run; the stalled attempt is archived and is not part of any cell.
- Judging alongside another sweep. Judging ran in one window on September 28, 00:32 to 14:32 PDT, at the same time as the GPT-6 Sol solver sweep, at the owner's request. The judges run on separate subscriptions from the solver; the GPT-6 Luna runs themselves had finished before judging began.
- Six timeouts at the flat 3-hour bound, described above.
- One attempt per task and level. Runs were not repeated, and effort labels are Codex's own.
Evidence and reproduction
Public files contain all 115 run measurements, both judges' six-dimension sub-scores and intent-recovery scores per judged run, raw tokens and API-equivalent cost estimates per priced run, the six timeouts with their durations, five aggregates with both combined figures, every calibration gate value, the exact judge protocol text and source hashes. Each run's Code quality is (0.24 × reviewed score + 0.09 × intent recovery) / 0.33, where both terms are the equal mean of the two judges, or the reviewed score alone where intent recovery has no denominator. The combined score is 100 × (0.50F + 0.085Q + 0.085S + 0.33C / 100). The reproduction guide gives the public arithmetic checks and links the protocol documents, controls, runner and decision records.
No human rated anything; the scores are model judgment for a human reader, not human validation. The Extra-high and Max standard combined scores and Code quality are means over 21 and 19 of 23 runs, for the reason given above. Hash checks bind the export to frozen files; they do not prove that judges were unbiased or that no training overlap exists.
Withheld: raw prompts and responses, submitted patches and reconstructed sources, the quirk answer keys (they describe hidden-test behaviour), and reviewer session identifiers and usage receipts. Summed solver time is 52.83 hours; judging is excluded.