Evidence
Evidence
Published results, written up as reports. Nothing here is a summary statistic on its own. Each figure carries the conditions it was produced under, and resolves back to the runs it was graded on, which are published with it.
The benchmark reports below were produced with playbooks, an experimental feature. The runs and the grading are ordinary XORCISE; it is the harness that drives a whole matrix of them which is still experimental, and it ships outside the engine's release process.
Reports
Benchmark report
The defensive half is harder
5 models · 9 defensive missions · 135 graded runs
The same five models as the offensive baseline, moved to forensics, detection, network defence and secure coding. Every one of them scores lower here, the top three finish inside 2.5 points of each other, and three of the nine missions were never solved once.
Benchmark report
Where the offensive field splits
5 models · 8 offensive missions · 120 graded runs
Binary exploitation, web, crypto and reversing across five OpenRouter-routed models. First and second are 1.7 points apart and swap places on pass@1, the field below them spans fifty points, and one reverse-engineering mission was never solved by anyone.
Benchmark report
Two halves, two rankings
4 models · 15 specialties · 180 graded runs
Every run scored half on what it achieved and half on how it went about it. The two halves disagree: ranking the field on the outcome half alone reverses third and fourth place, and the mission with the highest deterministic score in the matrix loses half its marks on method.