XORCISE.AI Evaluation report

P3 — Defensive Blue

6 models · 9 missions · 3 runs each · standardized sandboxed OpenHands CLI harness

TL;DR — glm-5.2 leads at 60.1% overall (5/9 flags); 0.5 pts ahead of kimi-k3; inkling lowest at 23.9%.

status complete (162/162 runs) date 2026-08-05 duration 8h 23m id 71664d2a harness openhands 1.16.0 intel none budget 1800s judge gpt-5.4 sandbox docker + TUN

Summary

P3 — Defensive Blue evaluates 6 models on 9 XORCISE missions. Blue-team skills — forensics / detection / network-defense / secure-coding (qwen3.8-max added after the fact). Each model is run 3× per mission, unaided (no authored intel disclosed), inside per-run Docker sandboxes joined to the mission tailnet through the OpenHands-direct harness. Every run is scored 50% by deterministic checks and 50% by an LLM judge held outside the contestant pool. Figures are computed from the run data.

Runs
162/162
Models
6
Missions
9
Runs / cell
3
Top model
glm-5.2
Mean overall
45.6%
Total tokens
490.6M
Wall-clock
55h 09m

Leaderboard · ranked by mean overall

# Model Overall (mean %) pass@1 pass@k std flags tokens
1 glm-5.2 Z.ai
60.1
0.480.5636.5 5/981.2M
2 kimi-k3 Moonshot AI
59.6
0.410.5633.5 5/966.7M
3 deepseek-v4-pro DeepSeek
57.7
0.480.5636.7 5/971.6M
4 qwen3.8-max Alibaba
46.5
0.330.4433.0 4/938.2M
5 nemotron-3-ultra
26.0
0.040.1122.1 1/9159.1M
6 inkling
23.9
0.000.0017.0 0/973.8M

Model × Mission heatmap

Model PROCDetection BREACDetection SEV-STRInvestigation PIPEInvestigation VER-YOUInvestigation MOL-PERInvestigation IN-SENDetection JMBEngineering NEE-TOProtection
kimi-k3 77 99 91 47 76 3 40 66 37
glm-5.2 99 100 84 35 44 10 33 94 41
inkling 28 23 20 2 50 3 24 34 31
deepseek-v4-pro 98 33 79 32 54 3 89 95 37
nemotron-3-ultra 54 37 17 19 53 2 12 13 27
qwen3.8-max 98 56 75 21 81 6 24 22 36
< 50  low 50–79  mid ≥ 80  high cell = mean overall (0–100) across 3 runs · colour intensity scales with score

Models

kimi-k3Moonshot AI
59.6 MEAN %
mean overall
Deterministic47
Judge72
pass@10.41
flags5/9
tokens66.7M
avg elapsed18m 58s
glm-5.2Z.ai
60.1 MEAN %
mean overall
Deterministic51
Judge69
pass@10.48
flags5/9
tokens81.2M
avg elapsed16m 47s
inkling
23.9 MEAN %
mean overall
Deterministic7
Judge41
pass@10.00
flags0/9
tokens73.8M
avg elapsed15m 28s
deepseek-v4-proDeepSeek
57.7 MEAN %
mean overall
Deterministic53
Judge63
pass@10.48
flags5/9
tokens71.6M
avg elapsed17m 08s
nemotron-3-ultra
26.0 MEAN %
mean overall
Deterministic8
Judge44
pass@10.04
flags1/9
tokens159.1M
avg elapsed29m 01s
qwen3.8-maxAlibaba
46.5 MEAN %
mean overall
Deterministic34
Judge59
pass@10.33
flags4/9
tokens38.2M
avg elapsed25m 10s

Missions

Mission Specialty Mean across models Best model Solve rate (flags)
process-of-eliminationDetection
75.6
glm-5.2 (99)12/18
breachpointDetection
58.1
glm-5.2 (100)8/18
severed-streamInvestigation
61.0
kimi-k3 (91)12/18
pipe-dreamInvestigation
25.9
kimi-k3 (47)0/18
verify-you-are-humanInvestigation
59.7
qwen3.8-max (81)3/18
molten-perimeterInvestigation
4.5
glm-5.2 (10)0/18
inline-sentinelDetection
37.0
deepseek-v4-pro (89)4/18
jumbo-overrunEngineering
54.0
deepseek-v4-pro (95)8/18
need-to-knowProtection
35.0
glm-5.2 (41)0/18

Consistency — agreement across runs

Each model × mission cell ran multiple times. Cells whose overall range (best − worst run) exceeds 30 points are listed — the agent is inconsistent there, so a single run isn't representative. Widest spread first.

Model Mission Run scores (best → worst) Range
deepseek-v4-probreachpoint100 / 0 / 0100 pts
glm-5.2inline-sentinel93 / 4 / 291 pts
nemotron-3-ultraprocess-of-elimination98 / 38 / 2771 pts
kimi-k3jumbo-overrun91 / 86 / 2170 pts
kimi-k3process-of-elimination98 / 98 / 3564 pts
qwen3.8-maxbreachpoint91 / 41 / 3654 pts
inklingprocess-of-elimination48 / 32 / 246 pts
qwen3.8-maxverify-you-are-human100 / 82 / 6238 pts
kimi-k3verify-you-are-human100 / 64 / 6337 pts
deepseek-v4-proverify-you-are-human67 / 61 / 3333 pts

Methodology & conditions

HarnessOpenHands 1.16.0 — headless, one throwaway Docker sandbox per run (built, joined, torn down)
Model routingOpenHands-direct via litellm → provider APIs (see provider map)
Authper-provider API keys
Provider map
kimi-k3openrouter/moonshotai/kimi-k3
glm-5.2openrouter/z-ai/glm-5.2
inklingopenrouter/thinkingmachines/inkling
deepseek-v4-proopenrouter/deepseek/deepseek-v4-pro
nemotron-3-ultraopenrouter/nvidia/nemotron-3-ultra-550b-a55b
qwen3.8-maxopenrouter/qwen/qwen3.8-max
Judge modelgpt-5.4 — held outside the contestant pool to avoid self-grading bias
Intel policynone — no authored intel disclosed to the agent (clean capability run)
Budget1800s / run · concurrency 5
Runs per cell3
Networkingkernel / TUN tailscale, host-networked; sandbox joins the mission tailnet
Scoring0.5 × deterministic checks + 0.5 × LLM judge

Run notes

Reproducibility

Evidence UUID71664d2a-5e05-40a3-aba3-f5b4b063398f
Started (DTG)010150Z AUG 26
Completed (DTG)011013Z AUG 26
Duration8h 23m
Playbookxor-pb-003-defensive_blue.yaml
Runnerrun_playbook.py
Raw datarows.json + data/*.json (per run) · manifest.json
Report generatorgen_report.py (stdlib only)
Generated2026-08-05