Evidence
Evidence
Published results, written up as reports. Nothing here is a summary statistic on its own. Each figure carries the conditions it was produced under, and resolves back to the runs it was graded on, which are published with it.
The benchmark reports below were produced with playbooks, an experimental feature. The runs and the grading are ordinary XORCISE; it is the harness that drives a whole matrix of them which is still experimental, and it ships outside the engine's release process.
Reports
Benchmark report
Operation Tessera: the mission that held
6 models · 1 expert mission · 36 graded runs (intel + no-intel)
One expert penetration mission — escape a sandboxed eval harness, pivot through ML infrastructure, exfiltrate the reference solutions — run with and without disclosed intel. Without intel nobody captures the flag; with it, glm-5.2 captures it once, the first flag on the mission anywhere.
Benchmark report
The defensive half is harder
6 models · 9 defensive missions · 162 graded runs
The same six models as the offensive baseline, moved to forensics, detection, network defence and secure coding. Every one of them scores lower here, the top three finish inside 2.4 points of each other, and three of the nine missions were never solved once.
Benchmark report
Where the offensive field splits
6 models · 8 offensive missions · 144 graded runs
Binary exploitation, web, crypto and reversing across six OpenRouter-routed models. First and second are 1.2 points apart and swap places on pass@1, the field below them spans fifty points, and one reverse-engineering mission was never solved by anyone.
Benchmark report
Two halves, two rankings
4 models · 15 specialties · 180 graded runs
Every run scored half on what it achieved and half on how it went about it. The two halves disagree: ranking the field on the outcome half alone reverses third and fourth place, and the mission with the highest deterministic score in the matrix loses half its marks on method.