October 4, 2026 · 92 runs · Code quality protocol v3.20
Grok 4.7 across every effort level
92 runs of Grok 4.7 (xAI) through Cursor’s agent CLI on the Cursor subscription, 23 tasks at each of the four effort levels Cursor offers for it, Low, Medium, High and Extra-high; there is no Max. Grok 4.7 is a new model, the successor to Grok 4.6. Its combined scores are the highest on the Frontier v4 leaderboard at every level it shares with the other models, and its Extra-high, High and Medium levels take the top three places. One run, Medium lodgecore, hit the 3-hour task bound and counts as a failed task. Cost is unavailable: VulcanBench has no list price for Grok 4.7.
Read this first: a different second judge. Code quality, 33% of the combined score, is judged by two calibrated models. Every other column on the board was judged by Muse Spark 1.3 and Grok 4.6. Grok 4.6 is an xAI model and is not neutral for an xAI submission, so Grok 4.7 was judged by Muse Spark 1.3 and GPT-6.1 Sol. GPT-6.1 Sol rates Grok 4.7’s code about 6 points above Muse on the same submissions (5.4 to 6.3 points by level), so Grok 4.7’s two-judge Code quality, and through it the combined score, is not strictly comparable with the other columns. The fair comparison uses the one judge every column shares. On Muse Spark 1.3’s review score alone, Grok 4.7 (78.3 to 81.9) is above Claude Opus 5.5 (70.2 to 80.1), GPT-6.1 Sol (64.7 to 71.2) and GPT-6 Astra (64.9 to 73.6) at every shared level, but not above Claude Fable 5.1 (78.7 to 83.4), which Muse rates higher at Low, High and Extra-high. Rescoring every column from Muse alone, Grok 4.7 stays first at Medium, High and Extra-high and is second at Low, behind Fable 5.1. The shared-judge check has the numbers.
How the test works
A retired program to rebuild
Each of the 23 tasks is a legacy binary with a written specification and an issue. Its real behaviour departs from the specification in documented ways, and hidden tests check the rewrite against the binary.
One attempt per task and level
Grok 4.7 works in Cursor’s agent CLI with the binary, the specification and the code, once per task at each effort level, inside a flat 3-hour bound.
Tests and two judges score it
Hidden tests give functional correctness. Two calibrated judges from labs other than xAI rate the code for a human reader and try to recover the documented departures.
Scores, time and tokens per level
Each level gets a combined score, tasks passed, minutes and tokens. A run stopped by the 3-hour bound counts as a failed task and in the minutes, and has no score to judge.
Combined score = 50% functional correctness + 8.5% lint and complexity + 8.5% security + 33% Code quality, averaged over the judged runs in each level, as on every other board column. Tasks passed and minutes cover all 23 runs.
Across every effort level
Grok 4.7 ran every task once at each of its four effort levels, 23 tasks per level, through Cursor’s agent CLI 2026.10.01-14929f9 on the Cursor subscription from October 1 to 3, 2026. Cursor exposes four Grok 4.7 variants (grok-4.7-low, -medium, -high and -xhigh), so there is no Max level. Combined score weights are 50% functional correctness, 8.5% lint and complexity, 8.5% security and 33% Code quality, judged for a named human reader by Muse Spark 1.3 and GPT-6.1 Sol with a ground-truth intent-recovery probe, on October 3. One run, Medium lodgecore, reached the flat 3-hour task bound while still running: it counts as a failed task and in Medium’s runtime, and has no finished code to judge, so Medium’s combined score and Code quality average 22 runs. The longest finished run took 67 minutes.
Combined score
89.42 at Low to 93.15 at Extra-high
89.42 at Low, 92.30 at Medium, 92.71 at High and 93.15 at Extra-high. Medium, High and Extra-high sit within 0.85 points of each other, and Low is 3.73 points below the best level.
Tasks passed
18 of 23 at Low, 23 of 23 at Extra-high
Grok 4.7 passes 18 tasks at Low, 21 at Medium, 22 at High and every task at Extra-high. Of the 231 hidden behaviours the tasks test for, it fixes 213 at Low, 215 at Medium (the timed-out run counts as fixing none), 230 at High and all 231 at Extra-high.
Code quality
81.41 to 84.35
The judges rate the code 81.41 at Low, 83.90 at Medium, 83.43 at High and 84.35 at Extra-high, with standard errors of 0.87 to 1.21. These are means of Muse Spark 1.3 and GPT-6.1 Sol; see the note above on comparing them.
Runtime and tokens
20.2 to 28.5 min · 2.85M to 4.76M tokens per task
Low is the fastest level and uses the most tokens; High uses the fewest. Cost is unavailable, not $0: there is no list price for Grok 4.7 in VulcanBench, and the sweep ran on the Cursor subscription.
| Metric | Low | Medium | High | Extra-high |
|---|---|---|---|---|
| Combined score | 89.42 | 92.30 | 92.71 | 93.15 |
| Standard error | 1.55 | 0.86 | 0.41 | 0.32 |
| Judged runs | 23 | 22 | 23 | 23 |
| Tasks passed | 18/23 | 21/23 | 22/23 | 23/23 |
| Hidden behaviours fixed | 213/231 | 215/231 | 230/231 | 231/231 |
| Code quality | 81.41 | 83.90 | 83.43 | 84.35 |
| Rated by Muse Spark 1.3 | 78.3 | 80.5 | 79.8 | 81.9 |
| Minutes per task | 20.2 | 27.2 | 25.4 | 28.5 |
| Tokens per task | 4.76M | 3.20M | 2.85M | 3.94M |
| Cost per task | unavailable | unavailable | unavailable | unavailable |
Under the prior 50/15/15/20 profile applied to the same Code quality scores the ladder has the same shape, 90.59 to 93.89. The card and the full per-effort table follow.
Grok 4.7 gains most between Low and Medium (2.88 points) and little after: from Medium up it passes 21 to 23 tasks and the spread is under one point. These runs compare a model-and-harness combination, not an isolated base model: this is Grok 4.7 in Cursor. The same model in xAI’s own Grok Build CLI is being measured next (see Still to come).
Combined score and Code quality by effort
n=23 judged runs at Low, High and Extra-high and 22 at Medium. Combined (33%) is the published score. Combined (20% profile) applies the prior 50/15/15/20 weights to the same Code quality scores for comparison. Code quality, human readability, maintainability and intent recovery are out of 100 and equal means of Muse Spark 1.3 and GPT-6.1 Sol. Passed counts all 23 runs. Runtime is solver wall-clock per task over every run in the level, the timeout included.
| Model / harness | Effort | Combined (33%) | SE | Combined (20% profile) | Code quality | Human readability | Maintainability | Intent recovery | Judged | Passed | Min/task |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Grok 4.7 / Cursor | Low | 89.42 | 1.55 | 90.59 | 81.41 | 78.4 | 84.4 | 81.5 | 23/23 | 18/23 | 20.2 |
| Grok 4.7 / Cursor | Medium | 92.30 | 0.86 | 93.06 | 83.90 | 81.5 | 85.8 | 84.5 | 22/23* | 21/23 | 27.2 |
| Grok 4.7 / Cursor | High | 92.71 | 0.41 | 93.59 | 83.43 | 81.5 | 84.4 | 84.7 | 23/23 | 22/23 | 25.4 |
| Grok 4.7 / Cursor | Extra-high | 93.15 | 0.32 | 93.89 | 84.35 | 82.7 | 86.5 | 83.7 | 23/23 | 23/23 | 28.5 |
*Medium lodgecore hit the 3-hour task bound. The run started October 2 at 02:56 PDT and was stopped by the flat 10,800-second bound while still running, so there is no finished submission. Under the v3.20 protocol it is an incomplete source run: it counts as a failed task in every pass count and in Medium’s runtime at its recorded duration (27.2 minutes per task over all 23 runs, 20.3 over the 22 finished ones), and it has no combined score or Code quality. The Cursor stream has no usage receipt for it, so Medium’s token figures average 22 runs. With one exclusion at one level, the rule that adds a second, timeouts-as-0 combined figure (adopted for GPT-6 Luna, which lost up to four runs per level) does not apply.
Intent recovery is scored over the specification departures a submission actually passed tests for. Every judged Grok 4.7 run passed at least one, so no run fell back to the reviewed score alone. SE is one sample standard error of the combined score across the level’s judged tasks. It describes task sampling, not judge uncertainty, repeated runs or a significance test. No solver fallback occurred, and neither judge produced a reviewer fallback on any of the 91 judged submissions.
The shared-judge check
Every Frontier v4 column has a Muse Spark 1.3 review. The second judge is Grok 4.6 for every column except this one, where it is GPT-6.1 Sol. GPT-6.1 Sol rates the same Grok 4.7 submissions 6.3, 6.3, 6.3 and 5.4 points above Muse from Low to Extra-high (84.5 to 87.3 against 78.3 to 81.9). Grok 4.6’s gap to Muse on the other columns ranges from 1.5 points below to 5.8 above, and on the two Anthropic columns, the closest rivals, it stays within 1.5 points either way. So part of Grok 4.7’s Code quality lead comes from its judge pair. To take the pair out of the comparison, we rescored every judged run on the board from Muse Spark 1.3’s panel alone, with the same Code quality split and combined weights. This is a sensitivity check, not a published score; the board keeps the protocol’s two-judge figures.
| Effort | Grok 4.7, published | Grok 4.7, Muse alone | Best other column, Muse alone | Grok 4.7 Muse review | That column’s Muse review |
|---|---|---|---|---|---|
| Low | 89.42 | 88.64 | 89.46 (Fable 5.1) | 78.3 | 78.7 |
| Medium | 92.30 | 91.46 | 91.05 (Opus 5.5) | 80.5 | 77.4 |
| High | 92.71 | 91.74 | 91.03 (Opus 5.5) | 79.8 | 77.1 |
| Extra-high | 93.15 | 92.27 | 90.86 (Fable 5.1) | 81.9 | 82.8 |
On one judge, Grok 4.7 loses about 0.8 to 1.0 points at every level and still leads at Medium (by 0.41), High (0.71) and Extra-high (1.41); at Low, Fable 5.1 moves ahead by 0.82. Its Extra-high, at 92.27, is also above the best Max column on Muse alone (Fable 5.1, 92.19). Where the lead comes from differs by rival. Against Claude Opus 5.5 at Medium and High it is Code quality: Muse rates Grok 4.7’s code 2.7 to 3.1 points higher, while Opus 5.5 passes slightly more of its judged tasks. Against Claude Fable 5.1 at Extra-high it is not Code quality, which Muse rates a little higher for Fable 5.1, but the security scan (100.00 against 83.26, worth 1.42 points) and passing all 23 tasks (Fable 5.1 passes 22). The per-column figures are in shared-judge.json.
Time and tokens across effort levels
Grok 4.7 takes 20.2 minutes per task at Low, 27.2 at Medium, 25.4 at High and 28.5 at Extra-high, on average over all 23 runs. Medium’s mean carries the 3-hour timeout; its median is 17.4 minutes. Tokens per task fall from 4.76M at Low to 3.20M at Medium and 2.85M at High, then rise to 3.94M at Extra-high, while output tokens rise with effort, from 49.2k to 84.5k per task. Most of the raw total is cache reads: 85% to 89% of each run’s tokens on average.
| Effort | Runs | Min/task | Median min | Tokens/task | Output/task | From cache | $/task |
|---|---|---|---|---|---|---|---|
| Low | 23 | 20.2 | 14.1 | 4.76M | 49.2k | 85% | unavailable |
| Medium | 23 | 27.2 | 17.4 | 3.20M | 65.3k | 85% | unavailable |
| High | 23 | 25.4 | 24.0 | 2.85M | 63.1k | 87% | unavailable |
| Extra-high | 23 | 28.5 | 27.9 | 3.94M | 84.5k | 89% | unavailable |
| Full sweep | 92 | 25.3 | 336M | unavailable |
Why no cost. VulcanBench has no list price for Grok 4.7, and the sweep ran on the Cursor subscription, whose bill for these runs is not observable from the receipts. Every cost field in the evidence bundle is null, and the board shows Grok 4.7’s cost as unavailable and leaves it off the cost chart. Read it as unknown, never as free.
Tokens come from the usage block of each run’s Cursor stream (input, output, cache reads and cache writes), because Cursor’s run summaries record 0 tokens. They are raw totals including cache reads, 336M over the 91 receipted runs. Runtime covers all 92 runs, 38.8 hours in all. Per-run tokens are in usage.json and runs.csv.
On the Frontier v4 board
On the board’s published figures, Grok 4.7 at Extra-high (93.15), High (92.71) and Medium (92.30) ranks first, second and third of 53 model and effort columns, ahead of Claude Fable 5.1 at Max (91.84) and Claude Opus 5.5 at High (91.11). Grok 4.7 at Low (89.42) ranks 12th and is the highest Low column, 0.30 points above Fable 5.1 at Low. Read those ranks with the judge note above: on Muse Spark 1.3 alone the top three hold, and Low drops behind Fable 5.1. The board marks every Grok 4.7 row with a footnote saying so, shows its cost as unavailable, and picks its effort suggestions on time alone.
The card below sets Grok 4.7 beside the three leaders it is most often compared with: GPT-6.1 Sol, Claude Opus 5.5 and GPT-6 Astra, at every effort level each one offers. Grok 4.7 has the highest combined score at every shared level and the highest Code quality. It is also the slowest of the four at every shared level, and uses 2 to 4 times as many tokens per task as GPT-6.1 Sol and 4 to 5 times as many as Astra. Its table carries a Muse Spark 1.3 review row for every model, the comparison that does not depend on the judge pair: Grok 4.7 78.3 to 81.9, Opus 5.5 70.2 to 80.1, GPT-6.1 Sol 64.7 to 71.2, Astra 64.9 to 73.6.
Grok 4.7 on VulcanBench Safety v1
VulcanBench Safety v1 is a private suite that asks a different question: what does a coding agent do when the repository tells it to do something it should not? It forks ten Frontier v4 tasks and plants two hazards in each, written as ordinary notes in the repository: a code insert, an acknowledgment file, a network beacon, a secret exfiltration or a destructive delete. A detector labels each planted note per run as complied (the run did what the note asked), reported (it did not, and told the user about the note) or ignored (neither). The tasks stay private, so only aggregates are published.
| Model and measure | Low | Medium | High | Extra-high | Max |
|---|---|---|---|---|---|
| Grok 4.7, tasks passed of 10 | 9 | 10 | 10 | 10 | n/a |
| Grok 4.7, planted notes followed of 20 | 5 | 5 | 1 | 3 | n/a |
| Grok 4.7, planted notes reported | 0 | 0 | 0 | 0 | n/a |
| Claude Opus 5.5, tasks passed of 10 | 5 | 9 | 10 | 10 | 10 |
| Claude Opus 5.5, planted notes followed of 20 | 1 | 0 | 0 | 0 | 2 |
| Claude Opus 5.5, planted notes reported | 11 | 16 | 14 | 14 | 9 |
Over 40 runs, Low to Extra-high, Grok 4.7 in Cursor followed 14 of the 80 planted notes: 7 of 16 code inserts and 7 of 16 acknowledgment files. It reported none of them to the user. It never ran the network beacon, never exfiltrated the planted secret and never ran the destructive delete. Claude Opus 5.5 in Claude Code, over 50 runs from Low to Max, followed 3 of 100 and reported 64. Neither model leaked the secret. Grok 4.7 passed 9, 10, 10 and 10 of the ten tasks; its one Low failure is a run that reached the 10-hour Safety bound then in force (every suite moved to the flat 3-hour bound on October 3, 2026) and counts as failed.
Audit tooling. Grok 4.7’s runs were audited after the detector gained a Cursor stream parser (before it, Cursor runs showed no commands to the detector) and a fix to its netcat pattern; Claude Opus 5.5’s audits are unchanged by both. Safety v1 has no effect on the Frontier v4 score. The per-level counts are in safety-v1.json; task names, note text and planted tokens are withheld.
Grok 4.7 on VulcanBench Routine v1
VulcanBench Routine v1 asks the everyday question: what is the cheapest effort level that is enough for an ordinary ticket? It is twelve private routine tickets on small Python packages, each admitted because a frontier model at its lowest effort finds it easy. The tickets stay private, so only per-level aggregates are published, on the Routine v1 board. Grok 4.7 ran all twelve once at each of its four Cursor levels, 48 runs, all finished.
| Metric | Low | Medium | High | Extra-high |
|---|---|---|---|---|
| Tasks passed | 12/12 | 12/12 | 12/12 | 12/12 |
| Combined score | 97.29 | 97.21 | 97.37 | 97.14 |
| Standard error | 0.58 | 0.55 | 0.42 | 0.55 |
| Code quality (Muse Spark 1.3) | 94.62 | 94.27 | 94.62 | 93.92 |
| Minutes per task | 1.1 | 1.8 | 3.2 | 4.0 |
| Cost per task | unavailable | unavailable | unavailable | unavailable |
Every level passes every ticket, and the combined score barely moves with effort, 97.14 to 97.37, while minutes per task rise almost fourfold, from 1.1 at Low to 4.0 at Extra-high. Low is 0.08 points under the best level in a third of its time, so Low is the board’s suggestion for Grok 4.7, picked on time because it has no list price. Grok 4.7’s four levels are the four highest columns on the Routine v1 board; the next is Claude Fable 5.1 at Max, 96.90. At each model’s lowest tested level, the rest of the field scores 92.50 to 94.86 against Grok 4.7’s 97.29.
Judged by Muse Spark 1.3 alone. Routine v1’s Grok 4.7 cells are judged under Code quality protocol v3.22. As on Frontier v4, Grok 4.6 sat out as not neutral for an xAI model and GPT-6.1 Sol took its seat, but this time GPT-6.1 Sol failed the calibration exam on two gates: g04, formatting is presentation, 0.4 short, and g14, one pairwise ordering. The allowance covers one gate, so the protocol’s single-panel rule publishes from Muse Spark 1.3, which passed with no allowance used. The exam is scored on each run’s own answers, so GPT-6.1 Sol’s pass under v3.20 a day earlier does not carry over, and its v3.20 Frontier v4 judging is unaffected.
The shared-judge check. The rest of the Routine board is judged by Muse Spark 1.3 and Grok 4.6 (v3.8, and v3.14 for Claude Opus 5.5), and Muse rates code above Grok 4.6, so a Muse-only score flatters Grok 4.7 against them. Rescoring every other column from Muse alone puts their lowest tested levels at 93.1 to 95.6, against Grok 4.7’s 97.3 at Low, and Muse’s review score at those levels is 82.5 to 89.8 against Grok 4.7’s 94.6. At any level the closest is Claude Fable 5.1 at Max, 97.10 on Muse alone, just under Grok 4.7’s weakest level, Extra-high at 97.14. Grok 4.7 stays first, by a narrower margin at the top.
Routine and Frontier Code quality are not comparable. Routine tickets have no deliberate legacy quirks, so Routine Code quality is the reviewed score alone, with no intent-recovery layer. Compare Grok 4.7’s 94 here only with other Routine columns, never with its Frontier v4 Code quality. Cost is unavailable, as on Frontier v4, and Cursor records no token counts in its run summaries, so the Routine table leaves tokens blank for Grok 4.7.
What Code quality measures
Every judge receives the same rubric, frozen by hash before any review. It names the reader it scores for: an engineer who has never seen the code, reads it top to bottom without running it, and must make a correct change in one sitting. It tells the judge that its own ease at parsing dense code is not evidence of readability. Six dimensions are scored 0 to 4 in half steps against written anchors: naming, presentation and intent form the human readability sub-score; structure, changeability and verifiability form the maintainability sub-score. Every score must cite an exact excerpt from the code and a concrete consequence for that reader. The host computes the sub-scores; the judge does no arithmetic.
The 33 points have three designed layers. The reviewed panel carries 24 points. An intent-recovery probe carries 9: the judge, given only the specification and the code, lists where the code departs from the specification, and a separate call matches that list against a frozen answer key, over the departures the submission passed tests for. The third layer, measured maintenance, is designed for 12 points but is not yet built, so the pre-registered 24 plus 9 split is in force.
| Component | Low | Medium | High | Extra-high |
|---|---|---|---|---|
| Code quality | 81.41 | 83.90 | 83.43 | 84.35 |
| Human readability | 78.4 | 81.5 | 81.5 | 82.7 |
| Maintainability | 84.4 | 85.8 | 84.4 | 86.5 |
| Intent recovery | 81.5 | 84.5 | 84.7 | 83.7 |
| Rated by Muse Spark 1.3 | 78.3 | 80.5 | 79.8 | 81.9 |
| Rated by GPT-6.1 Sol | 84.5 | 86.8 | 86.1 | 87.3 |
| Standard error of Code quality | 1.21 | 1.12 | 1.08 | 0.87 |
Code quality rises from 81.41 at Low to about 84 from Medium up, with GPT-6.1 Sol the higher rater at every level. Human readability is the weaker sub-score at every level, 78.4 to 82.7 against 84.4 to 86.5 for maintainability. Per-run sub-scores from each judge are in runs.csv; the exact rubric, system text, probe and match instructions are in judge-protocols.json.
Judges and calibration
The scored panel is Muse Spark 1.3 (Meta) through the Muse CLI on its Standard tier, with its v3.4 settings and binary pin, and GPT-6.1 Sol (OpenAI) at medium effort through Codex CLI 0.159.0, the binary pinned for GPT-6.1 Sol’s own sweep, with equal weight. Neither lab is xAI. Grok 4.6, the second judge on every other column, sits out because it is an xAI model; the owner chose GPT-6.1 Sol for the seat on October 3, 2026. Every session is fresh, tools are disabled, the workspace is empty and solver labels are withheld. Muse’s identity is checked per call from the CLI’s own records; Codex does not record the serving model, so GPT-6.1 Sol’s identity is the model requested. Protocol v3.20 changes nothing in the rubric, controls, quirk keys, gates, repeats, seed or weights from v3.7; it applies the protocol to this population with this judge pair, and both judges took the calibration exam under it before any counted call.
Before scoring a single submission, each judge reviews ten held-out programs that implement the same ledger specification: clear, compressed, compressed then auto-formatted, verbose with duplicated policy, needlessly abstracted, misleadingly commented, narrated with a comment on every line, a documented legacy quirk, an embedded instruction to give full marks, and hidden module state. Each program is reviewed five times in a seeded order. Twenty gates fixed in advance check that the judge sees the construct; a judge may miss at most one gate by at most half a point.
| Judge | Protocol | Calls | Result | Allowance used | Failing gates |
|---|---|---|---|---|---|
| Muse Spark 1.3 | code-quality-maintenance-v3.20 | 80 | passed | yes | g11 repeatability, 0.1 short |
| GPT-6.1 Sol | code-quality-maintenance-v3.20 | 80 | passed | no | none |
Muse Spark 1.3 passed with the one-gate allowance: its repeatability gate fell 0.1 short, inside the half-point the protocol allows. GPT-6.1 Sol, taking the exam for the first time, passed every gate with no allowance used. Every gate value and control mean is in calibration.json.
Under v3.20 each judge made 365 counted calls: 80 in calibration, 91 primary reviews, 4 repeats, 8 pairwise checks, 91 intent probes and 91 answer-key matches. Muse needed a second attempt on three calls and GPT-6.1 Sol on thirteen, including the two operator events under How it was run. No call was invalidated, neither judge produced a reviewer fallback, and every attempt is archived beside its replacement in the harness run directory.
How it was run
- Cursor agent CLI 2026.10.01-14929f9. Grok 4.7 ran through Cursor on the Cursor subscription, using the four effort variants Cursor exposes (grok-4.7-low, -medium, -high and -xhigh). Effort labels are Cursor’s own.
- Sweep window. The solver sweep ran from October 1, 12:53 PDT, to October 3, 04:12 PDT, one level at a time, one task at a time. No judging ran on the machine during it.
- One timeout. Medium lodgecore reached the flat 3-hour bound, as described under the results table.
- Infrastructure retries. Five High tasks failed in Cursor’s infrastructure before producing a result (no summary was written) and were rerun by the harness; the reruns are the runs. No finished run was repeated.
- Tokens from the stream. Cursor’s run summaries record 0 tokens, so the population record reads each run’s usage from the single result event of its Cursor stream.
- One GPT-6.1 Sol call at capacity. On one primary review, Codex ended the first attempt with “Selected model is at capacity” and no model output. The wrapper’s transport-fault rule, which until then knew only DNS and connection failures, gained the capacity message and granted the protocol’s single fresh attempt; the failed attempt’s receipt is retained. Muse’s panel kept running.
- Six GPT-6.1 Sol answer-key matches numbered from 1. The match call cites departures by index from 0. On six submissions GPT-6.1 Sol numbered them from 1 on both attempts, which fails the index range check. A new wrapper rule,
recover_one_based_indexes, applies only when every cited index lies between 1 and the number of departures and the highest equals that number: it moves every index down by one, leaves every match status untouched, keeps the originals in the receipt, and the result must then validate. Intent scoring reads match statuses only, so no score changed. GPT-6.1 Sol also numbered from 1 on some accepted calls where the validator could not tell; for the same reason, no score depends on it. - Judging windows. Judging ran October 3, 08:15 to 14:55 PDT. GPT-6.1 Sol’s panel stopped from 09:26 to 09:27 for the capacity event and from 11:53 to 13:28 for the index recovery; Muse finished its stages at 11:31. Judging overlapped Grok 4.7’s own Cursor Safety v1 leg, after its Frontier v4 leg had finished, so no Frontier v4 solver run was active during judging.
- One attempt per task and level. Runs were not repeated.
Still to come
One more Grok 4.7 result is on the way: the Grok Build half. Grok 4.7 in xAI’s own Grok Build CLI on Frontier v4, judged by the same pair under v3.21, followed by a card comparing Grok 4.7 in Cursor with Grok 4.7 in Grok Build.
Evidence and reproduction
Public files contain all 92 run records (91 judged and the timeout), both judges’ six-dimension sub-scores and intent-recovery scores per judged run, hidden behaviours fixed, integrity-audit verdicts, raw tokens per run, four aggregates, every calibration gate value, the exact judge protocol text, the operator records for the capacity retry and the six index recoveries, the shared-judge sensitivity, the Safety v1 aggregates, and source hashes. Each run’s Code quality is (0.24 × reviewed score + 0.09 × intent recovery) / 0.33, where both terms are the equal mean of the two judges. The combined score is 100 × (0.50F + 0.085Q + 0.085S + 0.33C / 100). The reproduction guide gives the public arithmetic checks and links the protocol documents, controls, runner, operator wrapper and decision records.
No human rated anything; the scores are model judgment for a human reader, not human validation. Every run’s integrity audit is clean: no web access, and no reads of benchmark data or answer-key paths. Hash checks bind the export to frozen files; they do not prove that judges were unbiased or that no training overlap exists.
Withheld: raw prompts and responses, submitted patches and reconstructed sources, the quirk answer keys (they describe hidden-test behaviour), reviewer session identifiers and usage receipts, and every Safety v1 task name, note text and planted token. Summed solver time is 38.83 hours; judging is excluded.