The published record
Benchmarks
Results by evaluation suite.
Compare results within the same suite and scoring protocol. Scores from different suites are not directly comparable.
A suite for models that return a typed decision instead of text. Twenty question families in eight areas, twelve in software engineering and eight in general reasoning, with every answer checked by running code, a type checker, hidden tests, a merged fix, a security advisory or the generator that built the question. Each family is scored as skill above its floor, the best of the most common answer, guessing and every surface shortcut built for it, so 0 is no better than the best dumb strategy and 100 is perfect. It replaces Verdict v1, which asked one kind of question and whose headline could not separate Jev from guessing. Verdict and Frontier scores measure different things and are not comparable.
September 25, 2026 · 4,774 test items · jev-1.13.0
TypeSafe AI's Jev scores 45.9 on the Verdict Index (95% interval 43.1 to 48.4): 49.6 on the software families and 40.4 on the general ones, with a Calibration Index of 0.33. It answers in 0.3 seconds, $0.58 for all 4,774 items. GPT-6 Astra at high effort, given exactly the same inputs as a reference that shows every family is answerable rather than as a leaderboard entry, scores 91.7.
Jev is strongest where the answer can be recognised from the shape of the text, such as whether a snippet type-checks (87) or which service caused an incident (82), and weakest on multi-step work: 24 on filtering and totalling a table, 1 on predicting which hidden test a patch fails. On size-matched patch pairs it is above the floor, unlike v1, at skill 36 against 58 for the reference. The admission gate was amended after the pilot; the report sets out the change and its reasons. The reference scored 100 on all eight general families, so they have no headroom for ranking strong reasoning models.