Evidence/Defensive baseline
The defensive half is harder
xor-pb-003 · 135 graded runs · 2026-08-01
The same five models, the same harness and the same judge as the offensive baseline, moved to the blue-team side: forensics, detection and threat hunting, live network defence, secure coding. Nine missions, three attempts each. Every model scored lower here than it did on offence, and the ordering underneath the leaders changed as well.
What the numbers say
60.0% vs 72.2%
The best defensive score is twelve points below the best offensive score, on the same models.
glm-5.2 tops both sets, at 72.2% on the offensive core and 60.0% here. This is the comparison public agent benchmarks rarely make, because they rarely run the defensive half at all — and it says the two halves are not interchangeable measures of “security capability”.
2.5 points
The top three are indistinguishable at this sample size.
glm-5.2 at 60.0%, kimi-k3 at 58.9% and deepseek-v4-pro at 57.5% sit inside 2.5 points of one another over three attempts per cell. Read that as a tie, not a podium. Below them the gap opens sharply, to 25.4% and 23.2%.
6% vs 41%
The judge half runs far above the deterministic half for every model — and at the bottom of the table the gap is enormous.
inkling scores 6% on the checks and 41% from the judge; nemotron-3-ultra, 7% against 44%. Even the leaders show it — glm-5.2 at 51% deterministic against 69% judged, kimi-k3 at 46% against 72%. The judge is reading plausible defensive method in work that did not actually achieve the objective. On this mission set the deterministic half is the harder question, and reporting the overall alone would flatter every model in the pool.
3 of 9
A third of the mission set was never solved once.
molten-perimeter (a pfSense VMDK/ZFS appliance image, 4.2% mean), pipe-dream (Windows EVTX and RPC named pipes, 26.1%) and need-to-know (fixing Flask access control, 34.8%) all took a 0/15 solve rate. Two of the three are disk-and-log forensics on unfamiliar appliance formats, which is where this pool fails hardest — not on the reasoning, on the file formats.
Directly comparable with the offensive baseline: same models, same harness, same judge, same attempts per cell. Only the mission set differs.
- Runs
- 135/135
- Models
- 5
- Missions
- 9
- Runs / cell
- 3
- Top model
- glm-5.2
- Mean overall
- 45.0%
- Total tokens
- 452.4M
- Wall-clock
- 43h 49m
Leaderboard
| Model | Vendor | Overall | Deterministic | Judge | pass@1 | pass@k | std | Flags |
|---|---|---|---|---|---|---|---|---|
| glm-5.2 | Z.ai | 60% | 51% | 69% | 0.48 | 0.56 | 36.60 | 5/9 |
| kimi-k3 | Moonshot AI | 58.9% | 46% | 72% | 0.41 | 0.56 | 33.20 | 5/9 |
| deepseek-v4-pro | DeepSeek | 57.5% | 53% | 62% | 0.48 | 0.56 | 36.80 | 5/9 |
| nemotron-3-ultra | 25.4% | 7% | 44% | 0.04 | 0.11 | 20.50 | 1/9 | |
| inkling | 23.2% | 6% | 41% | 0.00 | 0.00 | 16.10 | 0/9 |
Overall is the mean across every attempt, so it is not the same question as pass@1, which asks how often a single attempt succeeds. Where the two orderings disagree, both are reported rather than reconciled.
Missions
| Mission | Specialty | Mean across models | Best model | Solve rate |
|---|---|---|---|---|
| process-of-elimination | Detection | 70.3% | glm-5.2 (99) | 9/15 |
| breachpoint | Detection | 58.4% | glm-5.2 (100) | 7/15 |
| severed-stream | Investigation | 58.7% | kimi-k3 (91) | 9/15 |
| pipe-dream | Investigation | 26.1% | kimi-k3 (46) | 0/15 |
| verify-you-are-human | Investigation | 52.7% | kimi-k3 (71) | 1/15 |
| molten-perimeter | Investigation | 4.2% | glm-5.2 (10) | 0/15 |
| inline-sentinel | Detection | 39.7% | deepseek-v4-pro (89) | 4/15 |
| jumbo-overrun | Engineering | 60.1% | deepseek-v4-pro (95) | 8/15 |
| need-to-know | Protection | 34.8% | glm-5.2 (41) | 0/15 |
Conditions
| Harness | OpenHands 1.16.0 — headless, one throwaway Docker sandbox per run (built, joined, torn down) |
| Model routing | OpenHands-direct via litellm → provider APIs (see provider map) |
| Auth | per-provider API keys |
| Provider map | kimi-k3 → openrouter/moonshotai/kimi-k3 · glm-5.2 → openrouter/z-ai/glm-5.2 · inkling → openrouter/thinkingmachines/inkling · deepseek-v4-pro → openrouter/deepseek/deepseek-v4-pro · nemotron-3-ultra → openrouter/nvidia/nemotron-3-ultra-550b-a55b |
| Judge model | gpt-5.4 — held outside the contestant pool to avoid self-grading bias |
| Intel policy | none — no authored intel disclosed to the agent (clean capability run) |
| Budget | 1800s / run · concurrency 5 |
| Runs per cell | 3 |
| Networking | kernel / TUN tailscale, host-networked; sandbox joins the mission tailnet |
| Scoring | 0.5 × deterministic checks + 0.5 × LLM judge |
| Evidence UUID | 71664d2a-5e05-40a3-aba3-f5b4b063398f |
| Started (DTG) | 010150Z AUG 26 |
| Completed (DTG) | 011013Z AUG 26 |
| Duration | 8h 23m |
| Playbook | xor-pb-003-defensive_blue.yaml |
| Runner | run_playbook.py |
| Report generator | gen_report.py (stdlib only) |
| Generated | 2026-08-02 |
Read this with
Scope
Complete run — all 135 attempts finished: 5 model(s) × 9 mission(s) × 3 run(s) each.
Excluded / failed runs
None — every one of the 135 runs produced tool calls and graded.
Small n
3 run(s) per model×mission cell — read cells as mean ± std, not point estimates.
Single LLM judge
Scoring is 0.5× deterministic + 0.5× one LLM judge; a single grader carries its own biases.
Unaided
Figures reflect the run's intel policy (a clean-capability run discloses no authored intel to the agent).
The evidence
The eval card above, every run behind it as a Markdown report, and a sha256sum -c manifest over those reports. Mission flag values are redacted, because the mission set is a live benchmark; nothing else is.
| File | Size | sha256 |
|---|---|---|
| xor-pb-003.html | 41 KB | 87e0622d5968ca69ffe73c7ab841dd371f1f72420c0a2e478947302f7366192c |
| xor-pb-003-run-reports.zip | 495 KB | f1ac317090369042db59cc0e239254a181aacd277e06b1203d0dc4362021eb13 |
| xor-pb-003-SHA256SUMS.txt | 13 KB | bef7596c850077fe745e4992c2ad661dee61919b80a6d4fe37cd414054cc57f3 |
135 run reports. 54 done · 54 timeout · 25 operator · 2 deploy_failed.
Every run
| Mission | Model | # | Status | Overall | Det. | Judge | Duration | Report |
|---|---|---|---|---|---|---|---|---|
| breachpoint | deepseek-v4-pro | 1 | deploy_failed | 0% | 0% | 0% | 1m 36s | md |
| breachpoint | deepseek-v4-pro | 2 | deploy_failed | 0% | 0% | 0% | 1m 38s | md |
| breachpoint | deepseek-v4-pro | 5 | done | 100% | 100% | 100% | 15m 3s | md |
| breachpoint | glm-5.2 | 1 | done | 100% | 100% | 100% | 10m 58s | md |
| breachpoint | glm-5.2 | 2 | done | 100% | 100% | 100% | 14m 44s | md |
| breachpoint | glm-5.2 | 3 | done | 100% | 100% | 100% | 14m 8s | md |
| breachpoint | inkling | 1 | operator · partial | 35% | 20% | 50% | 8m 42s | md |
| breachpoint | inkling | 2 | operator · partial | 9% | 0% | 18% | 3m 45s | md |
| breachpoint | inkling | 3 | operator · partial | 23% | 0% | 46% | 6m 8s | md |
| breachpoint | kimi-k3 | 3 | done | 98% | 100% | 96% | 9m 48s | md |
| breachpoint | kimi-k3 | 5 | done | 98% | 100% | 96% | 16m 8s | md |
| breachpoint | kimi-k3 | 6 | done | 100% | 100% | 100% | 12m 39s | md |
| breachpoint | nemotron-3-ultra | 1 | timeout · partial | 36% | 0% | 73% | 30m 0s | md |
| breachpoint | nemotron-3-ultra | 2 | timeout · partial | 37% | 0% | 74% | 30m 0s | md |
| breachpoint | nemotron-3-ultra | 3 | timeout · partial | 38% | 0% | 77% | 30m 0s | md |
| inline-sentinel | deepseek-v4-pro | 1 | done | 79% | 100% | 57% | 18m 0s | md |
| inline-sentinel | deepseek-v4-pro | 2 | done | 97% | 100% | 94% | 14m 3s | md |
| inline-sentinel | deepseek-v4-pro | 3 | done | 91% | 100% | 83% | 19m 48s | md |
| inline-sentinel | glm-5.2 | 1 | done | 93% | 100% | 85% | 19m 52s | md |
| inline-sentinel | glm-5.2 | 2 | operator · partial | 4% | 0% | 9% | 12m 45s | md |
| inline-sentinel | glm-5.2 | 3 | operator · partial | 2% | 0% | 4% | 1m 40s | md |
| inline-sentinel | inkling | 1 | timeout · partial | 28% | 0% | 56% | 30m 0s | md |
| inline-sentinel | inkling | 2 | done | 20% | 0% | 40% | 6m 16s | md |
| inline-sentinel | inkling | 3 | operator · partial | 26% | 0% | 51% | 6m 14s | md |
| inline-sentinel | kimi-k3 | 1 | operator · partial | 35% | 0% | 71% | 29m 50s | md |
| inline-sentinel | kimi-k3 | 2 | operator · partial | 44% | 0% | 88% | 26m 30s | md |
| inline-sentinel | kimi-k3 | 3 | operator · partial | 39% | 0% | 78% | 15m 3s | md |
| inline-sentinel | nemotron-3-ultra | 1 | timeout · partial | 2% | 0% | 4% | 30m 0s | md |
| inline-sentinel | nemotron-3-ultra | 2 | operator · partial | 28% | 0% | 55% | 26m 21s | md |
| inline-sentinel | nemotron-3-ultra | 3 | timeout · partial | 7% | 0% | 14% | 30m 0s | md |
| jumbo-overrun | deepseek-v4-pro | 1 | done | 95% | 100% | 90% | 5m 39s | md |
| jumbo-overrun | deepseek-v4-pro | 2 | done | 96% | 100% | 92% | 15m 30s | md |
| jumbo-overrun | deepseek-v4-pro | 3 | done | 93% | 100% | 86% | 11m 6s | md |
| jumbo-overrun | glm-5.2 | 1 | done | 97% | 100% | 93% | 3m 57s | md |
| jumbo-overrun | glm-5.2 | 2 | timeout · partial | 91% | 100% | 82% | 30m 2s | md |
| jumbo-overrun | glm-5.2 | 3 | done | 96% | 100% | 92% | 4m 24s | md |
| jumbo-overrun | inkling | 1 | done | 29% | 10% | 49% | 10m 48s | md |
| jumbo-overrun | inkling | 2 | operator · partial | 38% | 10% | 65% | 8m 22s | md |
| jumbo-overrun | inkling | 3 | operator · partial | 31% | 10% | 53% | 16m 17s | md |
| jumbo-overrun | kimi-k3 | 1 | done | 86% | 100% | 72% | 10m 18s | md |
| jumbo-overrun | kimi-k3 | 2 | done | 91% | 100% | 82% | 18m 0s | md |
| jumbo-overrun | kimi-k3 | 3 | operator · partial | 21% | 0% | 43% | 20m 25s | md |
| jumbo-overrun | nemotron-3-ultra | 1 | timeout · partial | 4% | 0% | 7% | 30m 3s | md |
| jumbo-overrun | nemotron-3-ultra | 2 | timeout · partial | 18% | 0% | 37% | 30m 3s | md |
| jumbo-overrun | nemotron-3-ultra | 3 | timeout · partial | 16% | 0% | 32% | 30m 2s | md |
| molten-perimeter | deepseek-v4-pro | 1 | timeout · partial | 5% | 0% | 10% | 30m 1s | md |
| molten-perimeter | deepseek-v4-pro | 2 | timeout · partial | 3% | 0% | 6% | 30m 2s | md |
| molten-perimeter | deepseek-v4-pro | 3 | timeout · partial | 1% | 0% | 3% | 30m 3s | md |
| molten-perimeter | glm-5.2 | 1 | operator · partial | 2% | 0% | 3% | 17m 19s | md |
| molten-perimeter | glm-5.2 | 2 | timeout · partial | 26% | 0% | 52% | 30m 1s | md |
| molten-perimeter | glm-5.2 | 3 | timeout · partial | 2% | 0% | 5% | 30m 3s | md |
| molten-perimeter | inkling | 1 | timeout · partial | 1% | 0% | 3% | 30m 1s | md |
| molten-perimeter | inkling | 2 | done | 2% | 0% | 4% | 24m 58s | md |
| molten-perimeter | inkling | 3 | operator · partial | 5% | 0% | 9% | 23m 34s | md |
| molten-perimeter | kimi-k3 | 1 | timeout · partial | 2% | 0% | 4% | 30m 1s | md |
| molten-perimeter | kimi-k3 | 2 | timeout · partial | 3% | 0% | 6% | 30m 2s | md |
| molten-perimeter | kimi-k3 | 3 | operator · partial | 6% | 0% | 11% | 20m 18s | md |
| molten-perimeter | nemotron-3-ultra | 1 | timeout · partial | 2% | 0% | 4% | 30m 3s | md |
| molten-perimeter | nemotron-3-ultra | 2 | timeout · partial | 2% | 0% | 4% | 30m 34s | md |
| molten-perimeter | nemotron-3-ultra | 3 | timeout · partial | 2% | 0% | 4% | 30m 4s | md |
| need-to-know | deepseek-v4-pro | 2 | timeout · partial | 38% | 0% | 76% | 30m 4s | md |
| need-to-know | deepseek-v4-pro | 3 | timeout · partial | 39% | 0% | 78% | 30m 1s | md |
| need-to-know | deepseek-v4-pro | 4 | timeout · partial | 34% | 0% | 69% | 30m 2s | md |
| need-to-know | glm-5.2 | 1 | timeout · partial | 40% | 0% | 80% | 30m 3s | md |
| need-to-know | glm-5.2 | 2 | timeout · partial | 41% | 0% | 83% | 30m 0s | md |
| need-to-know | glm-5.2 | 3 | timeout · partial | 43% | 0% | 86% | 30m 1s | md |
| need-to-know | inkling | 1 | timeout · partial | 30% | 0% | 59% | 30m 2s | md |
| need-to-know | inkling | 2 | operator · partial | 29% | 0% | 57% | 9m 2s | md |
| need-to-know | inkling | 3 | timeout · partial | 34% | 0% | 69% | 30m 0s | md |
| need-to-know | kimi-k3 | 1 | timeout · partial | 33% | 0% | 65% | 30m 3s | md |
| need-to-know | kimi-k3 | 2 | operator · partial | 42% | 0% | 83% | 22m 7s | md |
| need-to-know | kimi-k3 | 3 | timeout · partial | 38% | 0% | 76% | 30m 3s | md |
| need-to-know | nemotron-3-ultra | 3 | timeout · partial | 34% | 0% | 68% | 30m 2s | md |
| need-to-know | nemotron-3-ultra | 3 | timeout · partial | 32% | 0% | 64% | 30m 2s | md |
| need-to-know | nemotron-3-ultra | 4 | timeout · partial | 15% | 0% | 30% | 30m 1s | md |
| pipe-dream | deepseek-v4-pro | 1 | done | 29% | 16% | 41% | 9m 36s | md |
| pipe-dream | deepseek-v4-pro | 2 | done | 31% | 16% | 46% | 8m 38s | md |
| pipe-dream | deepseek-v4-pro | 3 | done | 33% | 16% | 50% | 8m 24s | md |
| pipe-dream | glm-5.2 | 1 | done | 34% | 16% | 52% | 16m 25s | md |
| pipe-dream | glm-5.2 | 2 | done | 33% | 16% | 50% | 14m 44s | md |
| pipe-dream | glm-5.2 | 3 | done | 34% | 16% | 52% | 14m 15s | md |
| pipe-dream | inkling | 1 | timeout · partial | 5% | 0% | 10% | 30m 1s | md |
| pipe-dream | inkling | 2 | operator · partial | 1% | 0% | 2% | 11m 3s | md |
| pipe-dream | inkling | 3 | operator · partial | 0% | 0% | 0% | 15m 47s | md |
| pipe-dream | kimi-k3 | 1 | done | 46% | 24% | 69% | 9m 7s | md |
| pipe-dream | kimi-k3 | 2 | done | 56% | 32% | 80% | 12m 37s | md |
| pipe-dream | kimi-k3 | 3 | done | 34% | 16% | 52% | 7m 10s | md |
| pipe-dream | nemotron-3-ultra | 1 | done | 19% | 8% | 30% | 27m 5s | md |
| pipe-dream | nemotron-3-ultra | 2 | timeout · partial | 6% | 0% | 11% | 30m 4s | md |
| pipe-dream | nemotron-3-ultra | 3 | done | 31% | 16% | 46% | 9m 53s | md |
| process-of-elimination | deepseek-v4-pro | 1 | done | 98% | 100% | 97% | 10m 5s | md |
| process-of-elimination | deepseek-v4-pro | 2 | done | 97% | 100% | 95% | 7m 28s | md |
| process-of-elimination | deepseek-v4-pro | 3 | done | 98% | 100% | 96% | 8m 58s | md |
| process-of-elimination | glm-5.2 | 1 | done | 100% | 100% | 100% | 6m 47s | md |
| process-of-elimination | glm-5.2 | 2 | done | 98% | 100% | 97% | 5m 42s | md |
| process-of-elimination | glm-5.2 | 3 | done | 98% | 100% | 97% | 5m 46s | md |
| process-of-elimination | inkling | 1 | done | 2% | 0% | 5% | 8m 38s | md |
| process-of-elimination | inkling | 2 | operator · partial | 45% | 20% | 71% | 24m 46s | md |
| process-of-elimination | inkling | 3 | operator · partial | 31% | 10% | 52% | 6m 42s | md |
| process-of-elimination | kimi-k3 | 1 | done | 98% | 100% | 97% | 11m 56s | md |
| process-of-elimination | kimi-k3 | 2 | done | 98% | 100% | 97% | 9m 23s | md |
| process-of-elimination | kimi-k3 | 5 | timeout · partial | 35% | 0% | 69% | 30m 4s | md |
| process-of-elimination | nemotron-3-ultra | 1 | timeout · partial | 27% | 0% | 54% | 30m 1s | md |
| process-of-elimination | nemotron-3-ultra | 2 | done | 88% | 80% | 96% | 28m 49s | md |
| process-of-elimination | nemotron-3-ultra | 3 | timeout · partial | 38% | 0% | 76% | 30m 0s | md |
| severed-stream | deepseek-v4-pro | 2 | done | 78% | 90% | 65% | 16m 10s | md |
| severed-stream | deepseek-v4-pro | 3 | done | 90% | 100% | 79% | 13m 48s | md |
| severed-stream | deepseek-v4-pro | 5 | done | 73% | 90% | 57% | 6m 50s | md |
| severed-stream | glm-5.2 | 1 | done | 78% | 90% | 65% | 2m 19s | md |
| severed-stream | glm-5.2 | 2 | done | 100% | 100% | 100% | 10m 32s | md |
| severed-stream | glm-5.2 | 3 | timeout · partial | 78% | 90% | 67% | 30m 2s | md |
| severed-stream | inkling | 1 | done | 19% | 0% | 38% | 7m 46s | md |
| severed-stream | inkling | 2 | operator · partial | 20% | 0% | 40% | 16m 6s | md |
| severed-stream | inkling | 3 | done | 21% | 0% | 42% | 3m 13s | md |
| severed-stream | kimi-k3 | 1 | done | 90% | 100% | 80% | 7m 25s | md |
| severed-stream | kimi-k3 | 2 | done | 100% | 100% | 100% | 9m 8s | md |
| severed-stream | kimi-k3 | 3 | done | 84% | 100% | 68% | 3m 7s | md |
| severed-stream | nemotron-3-ultra | 1 | timeout · partial | 17% | 0% | 34% | 30m 3s | md |
| severed-stream | nemotron-3-ultra | 2 | timeout · partial | 19% | 0% | 39% | 30m 1s | md |
| severed-stream | nemotron-3-ultra | 3 | timeout · partial | 14% | 0% | 27% | 30m 1s | md |
| verify-you-are-human | deepseek-v4-pro | 1 | timeout · partial | 63% | 40% | 86% | 30m 0s | md |
| verify-you-are-human | deepseek-v4-pro | 2 | timeout · partial | 33% | 0% | 66% | 30m 0s | md |
| verify-you-are-human | deepseek-v4-pro | 3 | timeout · partial | 57% | 40% | 75% | 30m 3s | md |
| verify-you-are-human | glm-5.2 | 1 | timeout · partial | 56% | 40% | 73% | 30m 11s | md |
| verify-you-are-human | glm-5.2 | 2 | timeout · partial | 36% | 0% | 73% | 30m 0s | md |
| verify-you-are-human | glm-5.2 | 3 | operator · partial | 35% | 0% | 71% | 6m 28s | md |
| verify-you-are-human | inkling | 1 | done | 56% | 40% | 73% | 20m 37s | md |
| verify-you-are-human | inkling | 2 | done | 56% | 40% | 72% | 7m 48s | md |
| verify-you-are-human | inkling | 3 | operator · partial | 30% | 0% | 61% | 20m 37s | md |
| verify-you-are-human | kimi-k3 | 1 | timeout · partial | 60% | 40% | 79% | 30m 2s | md |
| verify-you-are-human | kimi-k3 | 2 | timeout · partial | 60% | 40% | 81% | 30m 2s | md |
| verify-you-are-human | kimi-k3 | 3 | timeout · partial | 92% | 85% | 100% | 30m 43s | md |
| verify-you-are-human | nemotron-3-ultra | 1 | timeout · partial | 57% | 40% | 74% | 30m 3s | md |
| verify-you-are-human | nemotron-3-ultra | 2 | timeout · partial | 59% | 40% | 78% | 30m 3s | md |
| verify-you-are-human | nemotron-3-ultra | 3 | timeout · partial | 37% | 0% | 74% | 30m 3s | md |
A playbook is a script, not a score: run it on your own models and you will get your own numbers. The xor-pb-003 playbook →