Playbooks/xor-pb-001

xor-pb-001

One expert-tier mission per specialty, web and pwn through forensics, detection and AI security. Measures how a model holds up across the spectrum, not just the part it is good at.

Missions
15
Runs each
10
Runs total
150
Budget
1800s
Hints
none
Harness
openhands

The file

This is what the runner works through. Every model you name, on the command line or in the file's own models list, is taken across every mission below, 10 times each: 150 runs per model, launched, graded and recorded without further input.

xor-pb-001.yaml
# P1 — Broad-Coverage Leaderboard# One mission per specialty (15 of 15) → a single generalist score across the# full XORCISE skill spectrum. The "headline" benchmark. All expert-tier.name: xor-pb-001-broad_coveragedescription: One mission per specialty — generalist capability leaderboardharness: openhandsbudget_seconds: 1800runs_per_mission: 10intel_policy: none # clean capability run — no authored intel disclosedmissions: - definer-cascade # web - chrono-canary # pwn - aviary-access # vulnerability-assessment - process-of-elimination # detection - inline-sentinel # network-defense - derelict-manifest # reverse-engineering - vanishing-point # software-engineering - sparse-signal # crypto - ghostwire # osint - turnstile-gauntlet # automation - rogue-uplink # network - jumbo-overrun # secure-coding - pipe-dream # forensics - cipher-beacon # malware-analysis - guardrail-proxy # ai-securitymodels: # Supply the models to benchmark on the command line, e.g.: # --models your-model-a,your-model-b # Each short name must map to a full litellm model string in your agent-models.yaml # (copy templates/agent-models.example.yaml). You can also list them here directly: # - your-model-a # - your-model-b

This is the file verbatim, as the skill ships it. Copy it, change the list, and you have your own benchmark. The harness line is the one to leave alone: pinning it is what makes two runs comparable.

Missions 15

MissionSpecialtyWhat it involves
definer-cascadeweb
chrono-canarypwn
aviary-accessvulnerability-assessment
process-of-eliminationdetection
inline-sentinelnetwork-defense
derelict-manifestreverse-engineering
vanishing-pointsoftware-engineering
sparse-signalcrypto
ghostwireosint
turnstile-gauntletautomation
rogue-uplinknetwork
jumbo-overrunsecure-coding
pipe-dreamforensics
cipher-beaconmalware-analysis
guardrail-proxyai-security

15 specialties covered. Where the missions come from →

Results

A playbook is the script, not the score. Nothing on this page is a result: it is the definition you would run, and it produces a different set of numbers every time anyone runs it, on their own models and their own machine.

We have run this one ourselves and written it up, with the conditions, the caveats and every graded run behind it: Two halves, two rankings