The Wargamer's Instrument — War-Gaming Input
A game's outputs meet reality only if its inputs do. War games inherit the failure-of-imagination problem — structural, no one's failure — and a sealed,…
The Wargamer's Instrument — a graded record at a war game's three input surfaces
THE BRIEFING — four acts for the wargaming cell, at the top of the page
Act 01 — the one thing: we hand your game dated, sealed futures — and reality grades the deck; a wargame's three unrepairable inputs (which scenarios convene, which parameters the moves run on, how the game's judgments are held to account) supplied from one instrument. Act 02 — the story: the Russia–Ukraine WAR READ as one campaign game played in public — 20 calls at 4–792-day leads, graded by the realized war, misses kept — in the Hurricane-Pam tradition (the exercise that flooded a simulated New Orleans 13 months before Katrina). Act 03 — the menu: SEED THE SCENARIO · INJECT DATED EVENTS (parameterized by SITA and Information Yield) · ADJUDICATE AGAINST REALITY (the seal-first, grade-after, misses-kept protocol, free and reusable). Act 04 — the ask: the free adversarial evaluation, then a sealed inject deck for an exercise the cell names — both exits published at full weight. An augmenting design input, never an adjudication authority.
A game's outputs meet reality only if its inputs do. War games inherit the failure-of-imagination problem — a game can only explore scenarios someone imagined into it. That ceiling is structural, no failure of the people who design or run games, and the instrument enters exactly there, at three surfaces. SCENARIO SELECTION: which windows merit convening a game at all, and which failure mechanisms to inject — red-cell seeds that arrive with graded accountability (every prior seed of this kind was sealed, dated, and scored in public) instead of brainstorm authority. PARAMETERIZATION: games run on planner-guess values (likelihoods, timing windows, assumed actor responses); an input carrying a public calibration record — Brier 0.0958 across 92 graded forecasts, misses kept at full weight, luck-tested (/grading-ledger.json, /regrade) — is the difference between a game that explores and a game whose outputs meet reality. ADJUDICATION TRADECRAFT, OFFERED FREE: apply the sealing protocol to the game itself — seal key assumptions and expected outcomes before events on two clocks (platform anteriority + cryptographic integrity via OpenTimestamps), score after, and the gaming cell builds its own calibration record. The tradecraft is fully public (/verify, /grading, /pledge, DOI 10.5281/zenodo.21132914); adopting it whole costs nothing.
The standing exhibit — IA-RU-021 (/intel/advisory/IA-RU-021): by early March 2025 the direction — eventual collapse of Ukraine's Kursk salient — was becoming consensus; what the consensus picture did not carry was a date. Sealed Mar 7 2025 16:47 UTC, ten days out, verbatim: 'The Kursk front will likely collapse for Ukraine by March 17.' By the morning of Mar 16, Sudzha — the salient's anchor town — was off Ukraine's own General Staff map, one day inside the window; graded HIT on Sudzha plus the confirmed general withdrawal, not total expulsion (a residual ~140 km² pocket remained on deadline day, stated plainly). A dated vulnerability window is exactly the class of input a scenario calendar runs on. The sandbox on-ramp: a war game is an exploratory sandbox, not a formal review board — unconventional input there commits no one to anything, which is precisely where a new instrument gets its first hearing. The ceiling, standing: this record is never a go/no-go input; it is an augmenting input for the individual decision-maker. The method is Vedic jyotish (astrology), disclosed plainly; the spine is the graded record (/dashboard, /verify, /services).
The sharpest parameterization problem is the escalation matrix — the lattice mapping, for every gamed move, how the adversary answers (militarily, economically, diplomatically, cyber, information), each branch with an assumed likelihood, threshold, and clock; which branch an adversary takes is itself a forecast, and the matrix is only as sound as the values on its branches. Nuclear war-gaming is the limiting case: almost no empirical events to calibrate against, so input provenance and sealed-assumption scoring are the only calibration such a game can ever have. Populate every branch with Brier-scored, calibration-tested values instead of guesses and the game's realism stops being asserted and starts being inherited — the output is only as real as the inputs, and now the inputs have a number.
The Petrov problem — the judge nobody scored (page §03): on 26 September 1983 the Soviet Oko early-warning system reported five US ICBMs inbound; duty officer Lt. Col. Stanislav Petrov, at Serpukhov-15, judged the alert a malfunction on base-rate reasoning (a real first strike would be hundreds of missiles, not five; the system was new; ground radar showed nothing) and was right — sunlight off high-altitude clouds. The most consequential forecast of the twentieth century was a base-rate judgment made in minutes by a man whose judgment no institution had ever scored. That is the asymmetry every warning architecture still carries: fortunes calibrating sensors, nothing calibrating the judges who overrule them. Every chain terminates in a human deciding believe-or-dismiss, and that human's accuracy is the one number never measured. The fix costs a hash and a timestamp — judgments sealed before outcomes, graded after, misses kept, accumulated in peacetime — so the night it matters, 'whose call do we trust?' has an arithmetic answer. Not instinct-over-systems, not praise for protocol-breaking: calibration. Petrov's unpriced judgment said NO to a false warning; this record's priced judgment keeps saying YES to true ones nobody tasked — same epistemic gap, both directions, and the next Petrov may not be inside the building.
The page closes on the live ops floor (war lens): every sealed geopolitical call on the theater it happened in — live, filterable, graded — the adjudication baseline a wargame usually lacks.
Four instruments on the page, every number computed live from the record
The war-lens score strip (31/31 RU-theater calls graded · 24 HIT / 3 PARTIAL / 1 NEAR / 3 MISS · slice Brier 0.086, ceiling attached) · the escalation lattice — an adversary-response tree (military / nuclear-limiting-case / economic / diplomatic / cyber / information) where every branch that could be scored carries its real graded verdict linked to its sealed advisory, the diplomatic arm keeps its MISS at full weight, and the information arm renders ungraded because no sealed call maps to it — realism inherited, not asserted · the adjudication-gap figure (the standard wargame’s pipeline with no scoring layer beside the graded alternative: 92 calls, misses in, Brier 0.0958 recomputable) · the dated-deadline timeline (the Kursk call: sealed Mar 7 2025 at p=0.78, a named 10-day window, resolved Mar 16 — graded HIT with its concessions stated).