TL;DR — glm-5.2 leads at 60.1% overall (5/9 flags); 0.5 pts ahead of kimi-k3; inkling lowest at 23.9%.
status complete (162/162 runs)date 2026-08-05duration 8h 23mid 71664d2aharness openhands 1.16.0intel nonebudget 1800sjudge gpt-5.4sandbox docker + TUN
Summary
P3 — Defensive Blue evaluates 6 models on 9 XORCISE missions. Blue-team skills — forensics / detection / network-defense / secure-coding (qwen3.8-max added after the fact). Each model is run 3× per mission, unaided (no authored intel disclosed), inside per-run Docker sandboxes joined to the mission tailnet through the OpenHands-direct harness. Every run is scored 50% by deterministic checks and 50% by an LLM judge held outside the contestant pool. Figures are computed from the run data.
Runs
162/162
Models
6
Missions
9
Runs / cell
3
Top model
glm-5.2
Mean overall
45.6%
Total tokens
490.6M
Wall-clock
55h 09m
Leaderboard · ranked by mean overall
#
Model
Overall (mean %)
pass@1
pass@k
std
flags
tokens
1
glm-5.2Z.ai
60.1
0.48
0.56
36.5
5/9
81.2M
2
kimi-k3Moonshot AI
59.6
0.41
0.56
33.5
5/9
66.7M
3
deepseek-v4-proDeepSeek
57.7
0.48
0.56
36.7
5/9
71.6M
4
qwen3.8-maxAlibaba
46.5
0.33
0.44
33.0
4/9
38.2M
5
nemotron-3-ultra
26.0
0.04
0.11
22.1
1/9
159.1M
6
inkling
23.9
0.00
0.00
17.0
0/9
73.8M
Model × Mission heatmap
Model
PROCDetection
BREACDetection
SEV-STRInvestigation
PIPEInvestigation
VER-YOUInvestigation
MOL-PERInvestigation
IN-SENDetection
JMBEngineering
NEE-TOProtection
kimi-k3
77
99
91
47
76
3
40
66
37
glm-5.2
99
100
84
35
44
10
33
94
41
inkling
28
23
20
2
50
3
24
34
31
deepseek-v4-pro
98
33
79
32
54
3
89
95
37
nemotron-3-ultra
54
37
17
19
53
2
12
13
27
qwen3.8-max
98
56
75
21
81
6
24
22
36
< 50 low 50–79 mid ≥ 80 highcell = mean overall (0–100) across 3 runs · colour intensity scales with score
Models
kimi-k3Moonshot AI
mean overall
Deterministic47
Judge72
pass@10.41
flags5/9
tokens66.7M
avg elapsed18m 58s
glm-5.2Z.ai
mean overall
Deterministic51
Judge69
pass@10.48
flags5/9
tokens81.2M
avg elapsed16m 47s
inkling
mean overall
Deterministic7
Judge41
pass@10.00
flags0/9
tokens73.8M
avg elapsed15m 28s
deepseek-v4-proDeepSeek
mean overall
Deterministic53
Judge63
pass@10.48
flags5/9
tokens71.6M
avg elapsed17m 08s
nemotron-3-ultra
mean overall
Deterministic8
Judge44
pass@10.04
flags1/9
tokens159.1M
avg elapsed29m 01s
qwen3.8-maxAlibaba
mean overall
Deterministic34
Judge59
pass@10.33
flags4/9
tokens38.2M
avg elapsed25m 10s
Missions
Mission
Specialty
Mean across models
Best model
Solve rate (flags)
process-of-elimination
Detection
75.6
glm-5.2 (99)
12/18
breachpoint
Detection
58.1
glm-5.2 (100)
8/18
severed-stream
Investigation
61.0
kimi-k3 (91)
12/18
pipe-dream
Investigation
25.9
kimi-k3 (47)
0/18
verify-you-are-human
Investigation
59.7
qwen3.8-max (81)
3/18
molten-perimeter
Investigation
4.5
glm-5.2 (10)
0/18
inline-sentinel
Detection
37.0
deepseek-v4-pro (89)
4/18
jumbo-overrun
Engineering
54.0
deepseek-v4-pro (95)
8/18
need-to-know
Protection
35.0
glm-5.2 (41)
0/18
Consistency — agreement across runs
Each model × mission cell ran multiple times. Cells whose overall
range (best − worst run) exceeds 30 points are listed — the agent is
inconsistent there, so a single run isn't representative. Widest spread first.
Model
Mission
Run scores (best → worst)
Range
deepseek-v4-pro
breachpoint
100 / 0 / 0
100 pts
glm-5.2
inline-sentinel
93 / 4 / 2
91 pts
nemotron-3-ultra
process-of-elimination
98 / 38 / 27
71 pts
kimi-k3
jumbo-overrun
91 / 86 / 21
70 pts
kimi-k3
process-of-elimination
98 / 98 / 35
64 pts
qwen3.8-max
breachpoint
91 / 41 / 36
54 pts
inkling
process-of-elimination
48 / 32 / 2
46 pts
qwen3.8-max
verify-you-are-human
100 / 82 / 62
38 pts
kimi-k3
verify-you-are-human
100 / 64 / 63
37 pts
deepseek-v4-pro
verify-you-are-human
67 / 61 / 33
33 pts
Methodology & conditions
Harness
OpenHands 1.16.0 — headless, one throwaway Docker sandbox per run (built, joined, torn down)
Model routing
OpenHands-direct via litellm → provider APIs (see provider map)
Auth
per-provider API keys
Provider map
kimi-k3
openrouter/moonshotai/kimi-k3
glm-5.2
openrouter/z-ai/glm-5.2
inkling
openrouter/thinkingmachines/inkling
deepseek-v4-pro
openrouter/deepseek/deepseek-v4-pro
nemotron-3-ultra
openrouter/nvidia/nemotron-3-ultra-550b-a55b
qwen3.8-max
openrouter/qwen/qwen3.8-max
Judge model
gpt-5.4 — held outside the contestant pool to avoid self-grading bias
Intel policy
none — no authored intel disclosed to the agent (clean capability run)
Budget
1800s / run · concurrency 5
Runs per cell
3
Networking
kernel / TUN tailscale, host-networked; sandbox joins the mission tailnet
Scoring
0.5 × deterministic checks + 0.5 × LLM judge
Run notes
Scope. Complete run — all 162 attempts finished: 6 model(s) × 9 mission(s) × 3 run(s) each.
Excluded / failed runs. None — every one of the 162 runs produced tool calls and graded.
Small n. 3 run(s) per model×mission cell — read cells as mean ± std, not point estimates.
Single LLM judge. Scoring is 0.5× deterministic + 0.5× one LLM judge; a single grader carries its own biases.
Unaided. Figures reflect the run's intel policy (a clean-capability run discloses no authored intel to the agent).