diff --git a/README.md b/README.md index 6be89fe..f8b5baa 100644 --- a/README.md +++ b/README.md @@ -48,7 +48,7 @@ packs or regenerate art on that machine: ```bash npm install # pulls classic-level, used to write the LevelDB packs npm run build # rebuilds packs/ from tools/content.mjs -npm run check # the fourteen guards; the build refuses to run if they fail +npm run check # the sixteen guards; the build refuses to run if they fail ``` To update a deployed server: `git pull` and restart Foundry. If the pull touches @@ -157,7 +157,7 @@ icons/ fonts/ art/ generated art packs/ built LevelDB compendia (committed — see Deploying) tools/ content.mjs the catalogue: skills, weapons, armour, gear, vehicles, NPCs - build-packs.mjs builds packs/ — runs the fourteen guards first and refuses on failure + build-packs.mjs builds packs/ — runs the sixteen guards first and refuses on failure make-icons.mjs draws all 231 icons rules-text.mjs generates the rules journal FROM rules.mjs mission.mjs the case generator @@ -178,11 +178,11 @@ stops being identity-equal to what it aliases, if a constant is re-declared as a literal, or if an exported rule has no spot-check. The packs used to be built under one set of numbers and played under another; this makes that impossible to ship. -The build runs all fourteen guards before it writes anything: +The build runs all sixteen guards before it writes anything: ``` -check-rules: OK — 110 rules, 43 files scanned, 3 aliases + 46 constant sets checked, 395 formulas verified +check-rules: OK — 110 rules, 46 files scanned, 3 aliases + 46 constant sets checked, 395 formulas verified check-kits: OK — 25 roles, 10 trades, 218 catalogue items, every kit key resolves, every posting can use what it carries, every loadout distinct check-lang: OK — en.json, 1018 keys, no leaf/branch collisions check-templates: OK — 19 templates compile @@ -195,6 +195,8 @@ check-lethality: OK — 47 creatures, every one fighting exactly as recorded aga check-focus: OK — focus fire measured against 31 packs, helps in all of them, 31 of them by more than their own noise check-firstblood: OK — the first disabling blow is worth 40 points of wipe rate in the four-player cut, 15.4 in the six-a-side line check-attackers: OK — three of the four-player cut disable on 11.1% of their attacks or better, and one of them on 0.1% +check-fight-tail: OK — a fight past 15 rounds wipes the party 1.4x as often in the four-player cut and 2.4x at six a side, and nothing was truncated +check-cited: OK — 23 cited figures across 13 scenario files resolve against 3 baselines check-bestiary: OK — 47 creatures, the document matches the game ``` diff --git a/docs/REVIEW_LOG.md b/docs/REVIEW_LOG.md index fb1b4ce..4ad67bf 100644 --- a/docs/REVIEW_LOG.md +++ b/docs/REVIEW_LOG.md @@ -6090,3 +6090,73 @@ cure, and that a later measurement which cares about the tail must pass its own no document cites is a maintenance obligation protecting no claim — and this suite has fourteen guards precisely because each one was built when something started being asserted. If the tail reaches a page, it gets a guard the same day. + +--- + +## R-272 — the tail reached a page, so it got a guard; and two guards to stop prose drifting + +R-271 ended "if the tail reaches a page, it gets a guard the same day." Tim asked for the +tail, the citation fix and the rules.mjs note together, so this is that day. Guards fifteen +and sixteen, and one comment. + +**check-fight-tail (15), on `tools/fight-tail.mjs`.** The suite could not see a long fight +because every guard in it averages. EXPOSURE printed medians for the same reason. + +The cap is the measurement here, so the tool passes its own — 400, not `runFight`'s default +40 — and `--update` refuses to record anything that reaches it. Proved by lowering it back +to 40: the cut loses 20 fights to the ceiling and the six-a-side 117, with 107 ending +neither won nor wiped, and the recorder stops rather than writing a truncated distribution. +That is R-271's defect turned into a refusal. + +The claim checks read a fresh measurement rather than the baseline, so `--update` cannot +silence them. Two claims, chosen because the obvious one does not discriminate — "the long +end is twice the median" is true of both encounters and so distinguishes nothing: + +- **Length predicts death.** A fight past fifteen rounds ends in a wipe materially more + often than a shorter one. A GM can act on that: a fight still running is not a stalemate, + it is one the party is losing slowly. +- **The six-a-side fight is the LONGER one**, median 14 against 10.7, which is R-270 showing + up as duration. A disabled fighter does not leave, they fight on 30 points down, so more + bodies means more people swinging badly for longer. The cut is shorter because it is + decisive, not because it is safer. + +**check-cited (16), on `tools/check-cited.mjs`.** check-firstblood and check-attackers catch +the game changing; neither reads the document. Re-record after a re-cast and the artifact +updates, the guard goes green, and the paragraph goes on printing the old number with a +citation under it saying where the new one lives — the citation making it worse, because it +tells a GM the figure is checked. + +So the citations are machine-readable: `**40**`, resolved +against the named JSON on every build. 25 of them now. It earned its place immediately by +failing three times on its first runs — a config keyed `column` that the prose called `line`, +two figures rounded from 32.3 to 32, and a vacuous pass when zero citations were found, which +is now fatal in its own right. + +**It also refuses to let a field be quoted at all.** `fight-tail.longest` is recorded for +drift and may not reach prose: same party, same seeds, same runs, and renaming a config from +`line` to `column` moved it from 71 to 90 rounds, because `seedFor` derives the stream from +the id. Measured across seven labels — median spread 0, p95 1, p99 3, **longest 21**. A +sample maximum reads exactly like a bound on the encounter and is a property of the label. +EXPOSURE states the long end as p99 instead, which is stable to about three rounds. + +The same discipline caught the deadlier ratio: 2.42 at six a side with a seed spread of 1.1, +forty-five per cent of its own value. The document states the direction of that effect and +explicitly declines to quote its size; `deadlierNoise` is recorded so the next reader can see +why. + +**rules.mjs.** A comment over `resolveLocationHit` recording that its two thresholds are +unrelated and usually agree. `disabled` is a fraction of the pool per location; `majorWound` +is ceil(hp/2) and feeds only `dyingLimitFor`; neither takes anybody out of a fight, which is +`conditionFor` at 2 hit points or a destroyed head. For a 10 hp human, leg, abdomen and chest +capacities are 5 and majorWoundFor is 5 — and those locations take 12 of 20 melee results and +15 of 20 ranged. Two readers independently reconstructed a rule that does not exist, checked +it against the log, and were confirmed by it, because the agreeing case is most of the hits. +The note says: if you are about to describe either threshold from watching a fight, test an +arm. It is the only place the difference is visible. + +**What this pass is really about.** Four of the defects found today were inside the guards +rather than the game. A guard is a claim we have agreed to stop checking, which is its whole +value and exactly why an assumption inside one is the least likely thing in the repo to be +questioned. Both new guards are therefore built to refuse rather than to record: one will not +write a censored measurement, the other will not let an unstable figure be cited. Neither can +be satisfied by re-recording. diff --git a/docs/scenarios/CLEAN_GROUND.md b/docs/scenarios/CLEAN_GROUND.md index efd0ae9..9bf6fe1 100644 --- a/docs/scenarios/CLEAN_GROUND.md +++ b/docs/scenarios/CLEAN_GROUND.md @@ -737,8 +737,8 @@ losing. Both of the first two are live in this scenario. | Measured | Party strikes first | Then wipes | If it does not | |---|---|---|---| - | **Three of the column vs the cut** | 50.3% (mean round 1.8) | **25.5%** | **65.5%** — a swing of **40** | - | Six of the column vs the full six | 54.1% (mean round 1.3) | 7.8% | 23.2% — a swing of 15.4 | + | **Three of the column vs the cut** | 50.3% (mean round 1.8) | **25.5%** | **65.5%** — a swing of **40** | + | Six of the column vs the full six | 54.1% (mean round 1.3) | 7.8% | 23.2% — a swing of 15.4 | First blood is very nearly a coin toss, and in the cut it is worth **forty points of wipe rate**. Against the full six it is worth a third of that, because one fighter is a sixth @@ -769,7 +769,8 @@ losing. Both of the first two are live in this scenario. second agent to connect in a round is worth more than the first — that is what `check-focus` measures across the bestiary, and it holds here. - **Recorded, and it agrees:** the three land **16 to 18%** of their attacks — - Bhattacharya 17.8, Renshaw 17.3, Braithwaite 16.2 — which is *above* the one-in-eight + Bhattacharya 17.8, Renshaw 17.3, + Braithwaite 16.2 — which is *above* the one-in-eight the arithmetic gives for a fresh dodge, precisely because most blows in a fight are not the first of their round. `tools/attackers-baseline.json`, held by `check-attackers`. - **This is the other half of why the cut is a different game.** Three attackers cannot @@ -780,6 +781,44 @@ losing. Both of the first two are live in this scenario. passed hand to hand with both parties off the ground rather than marksmanship. The column being uncatchable is the point; see their tactics line. +#### How long it runs, which is not the median + +**Every figure above is a middle, and a middle is not a plan.** The table prints median +rounds because that is what the suite could measure; it is the wrong number for a slot. + +| | Median | Past 15 rounds | Past 20 | Past 30 | 1 in 100 runs past | +|---|---|---|---|---|---| +| Three of the column vs the cut | 10.7 | 19.9% | 7.3% | 1.3% | 32.3 rounds | +| **Six of the column vs the full six** | **14** | **40.3%** | **21.6%** | 6.5% | **45.3** rounds | + +Three things a GM should take off that table. + +- **The six-a-side fight is the LONGER one**, which is the opposite of what more guns + suggests. It is R-270 showing up as duration: a disabled fighter does not leave, they + fight on at 30 points less, so more bodies means more people swinging badly for longer. + **The cut is shorter because it is decisive, not because it is safer.** +- **One six-a-side fight in five runs past twenty rounds, and one in a hundred past + forty-five.** Act Four is 45 minutes. That fight is the whole act and the Close is what + pays for it — the same warning point 2 makes about hollow men, now with the column's + numbers under it. *The last column is the honest way to state a worst case: the longest + fight in any one sample is not a bound and moves by a third of its value on a re-seed, so + `check-cited` refuses to let this document quote it.* +- **A fight still going at fifteen rounds is not a stalemate. It is one the party is losing + slowly.** Long fights end in a wipe materially more often than short ones, in both + encounters. *The size of that effect is not quoted here on purpose: at six a side it + varies by nearly half its own value across seeds, so only its direction is trustworthy.* + `check-fight-tail` holds the direction and fails the build if it reverses. + +**If the clock says the fight has gone long, it has already gone wrong.** Cut to the +consequence rather than rolling it out — the party is not going to turn it around, and +twenty more rounds of dice is how a convention slot dies. + +*Recorded in `tools/fight-tail-baseline.json`, re-checked on every build by +`check-fight-tail`; read it with `node tools/fight-tail.mjs`. Measured with the round +ceiling raised to 400 because the harness default of 40 truncates exactly this +distribution — at the default, one six-a-side fight in fifty is cut off rather than +finished, and the 99th percentile reads 40 because 40 is the wall.* + #### What to do instead of a fight - **Act Four's tension is a test, not a threat.** The question is which of two people is @@ -1064,7 +1103,8 @@ twice.* large part of why EXPOSURE says violence at four players is party-ending. Substituting Agyeman fixes the arithmetic as well as the lane. - **He is not sidelined in the fight — he is busy and ineffective.** He takes his - actions — **9.2 attacks across a fight, of which 1% land** (`check-attackers`) — and + actions — **9.2 attacks across a fight, of + which 1% land** (`check-attackers`) — and **one attack in a thousand disables anybody**. That plays worse at the table than being obviously out of it: a player who cannot act knows to do something else, and a player rolling once a round for an hour thinks he is fighting. **If a fight starts and Ashcroft is in it, give him something that is not diff --git a/package.json b/package.json index 965dae1..1f8c1f4 100644 --- a/package.json +++ b/package.json @@ -11,7 +11,7 @@ "play": "node tools/playthrough.mjs", "mj": "node tools/mj-queue.mjs", "simulate": "node tools/simulate.mjs", - "check": "bun tools/check-rules.mjs && bun tools/check-kits.mjs && bun tools/check-lang.mjs && bun tools/check-templates.mjs && bun tools/check-behaviour.mjs && bun tools/check-scenarios.mjs && bun tools/check-rollable.mjs && bun tools/check-creatures.mjs && bun tools/check-anatomy.mjs && bun tools/check-lethality.mjs && bun tools/check-focus.mjs && bun tools/check-firstblood.mjs && bun tools/check-attackers.mjs && bun tools/check-bestiary.mjs", + "check": "bun tools/check-rules.mjs && bun tools/check-kits.mjs && bun tools/check-lang.mjs && bun tools/check-templates.mjs && bun tools/check-behaviour.mjs && bun tools/check-scenarios.mjs && bun tools/check-rollable.mjs && bun tools/check-creatures.mjs && bun tools/check-anatomy.mjs && bun tools/check-lethality.mjs && bun tools/check-focus.mjs && bun tools/check-firstblood.mjs && bun tools/check-attackers.mjs && bun tools/check-fight-tail.mjs && bun tools/check-cited.mjs && bun tools/check-bestiary.mjs", "test": "bun run check", "readme": "bun tools/update-readme.mjs" }, diff --git a/rules.mjs b/rules.mjs index 800a9c9..e6fb555 100644 --- a/rules.mjs +++ b/rules.mjs @@ -552,6 +552,25 @@ export function locationFor(roll, speciesId = "baseline", mode = "ranged") { * Spec §3.6. Resolve a hit against a location. * A blow that meets the location's maximum disables it; twice that destroys it. * General hit points still take the damage, so the two systems agree. + * + * THE TWO THRESHOLDS HERE ARE UNRELATED, AND THEY USUALLY AGREE. `disabled` compares the + * running total against `locationMax` — locationMaxHp(totalHp, frac), a fraction of the + * pool, per location. `majorWound` compares THIS blow against `majorWoundThreshold` — + * majorWoundFor(hp), which is ceil(hp/2) and decides only how long a dying character + * lasts, via dyingLimitFor. Neither has anything to do with leaving the fight: that is + * conditionFor, at 2 hit points or a destroyed head. + * + * For a 10 hp human the leg, abdomen and chest capacities are 5 and majorWoundFor is also + * 5; at 12 hp both are 6. Those locations take 12 of 20 melee results and 15 of 20 ranged, + * so on most hits the two tests fire together and the narration prints them on one line. + * They diverge only on arms and head — 4 against 5, 5 against 6. + * + * Two people reading that output independently reconstructed the same rule that does not + * exist ("everyone goes down at half their maximum hit points"), checked it against the + * log, and were confirmed by it, because the agreeing case is 60-75% of hits by + * construction. It reached a published scenario and a review-log entry before anyone + * opened this file. R-270. If you are about to describe either threshold from watching a + * fight, test an ARM: it is the only place the difference is visible. */ export function resolveLocationHit({ damage, locationMax, locationTaken = 0, majorWoundThreshold }) { const d = Number(damage) || 0; diff --git a/tools/check-cited.mjs b/tools/check-cited.mjs new file mode 100644 index 0000000..099c053 --- /dev/null +++ b/tools/check-cited.mjs @@ -0,0 +1,137 @@ +/** + * A figure a scenario prints must still be the figure its artifact holds. + * + * check-firstblood and check-attackers both compare a baseline against a fresh run, so + * they catch the GAME changing. Neither of them reads the document. Re-record after a + * re-cast and the artifact updates, the guard goes green, and the paragraph goes on + * printing the old number with a citation under it saying where the new one lives. The + * citation makes it worse, not better: it tells a GM the figure is checked. + * + * So the citations are machine-readable. A figure copied out of an artifact is written + * + * a swing of **40** + * + * and this guard resolves every one of them against the named JSON. Copy a number wrong, + * or re-record and leave the prose alone, and the build stops with both values named. + * + * This is the check-bestiary arrangement one size down. That document is GENERATED from + * the game, so it cannot drift; a scenario is written by hand and cannot be, but the + * numbers inside it can be held to the same standard. + * + * node tools/check-cited.mjs check every citation in every scenario + * node tools/check-cited.mjs --list print what each one currently resolves to + */ +import { readFileSync, existsSync } from "node:fs"; +import path from "node:path"; +import { scenarioFiles } from "./check-scenarios.mjs"; + +const ROOT = path.resolve(path.dirname(new URL(import.meta.url).pathname.replace(/^\/([A-Za-z]:)/, "$1")), ".."); +const LIST = process.argv.slice(2).includes("--list"); + +/* Artifact name as written in a citation -> the file it means. Adding a baseline here is + what makes it citeable; nothing else in the document needs to know. */ +const ARTIFACTS = { + "first-blood": "tools/first-blood-baseline.json", + "attackers": "tools/attackers-baseline.json", + "fight-tail": "tools/fight-tail-baseline.json", + "lethality": "tools/lethality-baseline.json" +}; + +/* Fields that exist in an artifact but must never appear in prose. A sample maximum is + the clearest case: same party, same seeds, same runs, and renaming a config moved + fight-tail's `longest` from 71 to 90 rounds, because the stream is derived from the id. + It reads like a bound on the encounter and is a property of the label. Citing it would + be the n=1 mistake one level up, so the citation itself is refused. */ +const UNCITEABLE = { + longest: "a sample maximum, not a bound — it moves by a third of its value on a re-seed. " + + "Cite p99 instead: 'one fight in a hundred runs past N rounds' is stable to about three rounds." +}; + +const loaded = new Map(); +function artifact(name) { + if (loaded.has(name)) return loaded.get(name); + const rel = ARTIFACTS[name]; + if (!rel) return null; + const file = path.join(ROOT, rel); + const json = existsSync(file) ? JSON.parse(readFileSync(file, "utf8")) : null; + loaded.set(name, json); + return json; +} + +/** "configs.cut.swing" or "cut.swing" — the leading container is optional. */ +function resolve(json, dotted) { + const direct = dotted.split(".").reduce((o, k) => (o == null ? o : o[k]), json); + if (direct !== undefined) return direct; + for (const container of ["configs", "agents"]) { + const v = dotted.split(".").reduce((o, k) => (o == null ? o : o[k]), json?.[container]); + if (v !== undefined) return v; + } + return undefined; +} + +/* The number immediately before the marker, allowing for markup and a trailing unit: + "**40**", "25.5%", "1.4x". Anchored to the end so it is the nearest one. */ +const CITE = /([\d]+(?:\.[\d]+)?)\s*(?:%|x)?\**\s*(?:—|-|–)?\s*/g; + +let checked = 0; +const problems = []; +const rows = []; + +/* Deliberately NOT scenarioText(): that strips HTML comments, which is where the + citations live. The raw file is the thing being checked. */ +for (const [rel, abs] of scenarioFiles()) { + const text = readFileSync(abs, "utf8"); + for (const m of text.matchAll(CITE)) { + const [, printed, name, dotted] = m; + checked++; + const json = artifact(name); + if (!json) { + problems.push(` ${rel}: cites "${name}", which is not a known artifact ` + + `(${Object.keys(ARTIFACTS).join(", ")})`); + continue; + } + const held = resolve(json, dotted); + if (held === undefined) { + problems.push(` ${rel}: cites ${name} ${dotted}, which that artifact does not hold`); + continue; + } + const leaf = dotted.split(".").pop(); + if (UNCITEABLE[leaf]) { + problems.push(` ${rel}: cites ${name} ${dotted}, which must not be quoted — ${UNCITEABLE[leaf]}`); + continue; + } + rows.push([rel, `${name} ${dotted}`, printed, String(held)]); + if (Number(printed) !== Number(held)) { + problems.push(` ${rel}: prints ${printed} but ${ARTIFACTS[name]} holds ${held} ` + + `(${dotted}). Re-recording updated the artifact and left the prose behind — fix the sentence.`); + } + } +} + +if (LIST) { + for (const [f, src, printed, held] of rows) { + console.log(` ${f.padEnd(32)} ${src.padEnd(28)} prints ${printed.padStart(6)} holds ${held}`); + } + process.exit(0); +} + +/* A guard that passes because it found nothing to check is the defect this suite has now + been caught by four times, so it is not allowed here: the citations exist, and a run + that cannot see them means the markers were reformatted away, not that the prose is + clean. */ +if (!checked) { + console.error("check-cited: FAILED — no cited figures found at all. The scenarios carry " + + "`` markers next to every measured number; finding none " + + "means they have been stripped or reformatted, not that there is nothing to check."); + process.exit(1); +} + +if (problems.length) { + console.error("check-cited: FAILED — a scenario prints a figure its artifact no longer holds"); + problems.forEach(p => console.error(p)); + process.exit(1); +} + +const artifacts = new Set(rows.map(r => r[1].split(" ")[0])); +console.log(`check-cited: OK — ${checked} cited figures across ${scenarioFiles().length} scenario files ` + + `resolve against ${artifacts.size} baselines`); diff --git a/tools/check-fight-tail.mjs b/tools/check-fight-tail.mjs new file mode 100644 index 0000000..68b30b3 --- /dev/null +++ b/tools/check-fight-tail.mjs @@ -0,0 +1,27 @@ +/** + * CLEAN GROUND tells a GM how long a fight takes, and every figure it had was a median. + * + * A median is what the suite could see: fourteen guards averaged, so none of them could + * have said that one six-a-side fight in five runs past twenty rounds, or that a fight + * still going at fifteen is one the party is losing. Those are the sentences EXPOSURE now + * prints, and this is what keeps them true. + * + * Thin on purpose: the measurement lives in fight-tail.mjs, which is also the tool a + * reader runs, so the guard and the reading cannot drift apart. + */ +import { spawnSync } from "node:child_process"; + +const r = spawnSync(process.execPath, [new URL("fight-tail.mjs", import.meta.url).pathname, "--check"], + { encoding: "utf8" }); + +if (r.status !== 0) { + process.stdout.write(r.stdout ?? ""); + process.stderr.write(r.stderr ?? ""); + console.error("check-fight-tail: FAILED — these fights no longer run the length CLEAN GROUND " + + "says they do. Read `node tools/fight-tail.mjs` before re-recording."); + process.exit(1); +} +// update-readme reads the line beginning with this file's own name. +const m = r.stdout.match(/wipes the party ([\d.]+)x as often in the cut and ([\d.]+)x at six a side/) ?? []; +console.log(`check-fight-tail: OK — a fight past 15 rounds wipes the party ${m[1] ?? "?"}x as often in ` + + `the four-player cut and ${m[2] ?? "?"}x at six a side, and nothing was truncated`); diff --git a/tools/fight-tail-baseline.json b/tools/fight-tail-baseline.json new file mode 100644 index 0000000..13caac3 --- /dev/null +++ b/tools/fight-tail-baseline.json @@ -0,0 +1,68 @@ +{ + "note": "Generated by tools/fight-tail.mjs --update. Do not edit by hand.", + "runs": 2000, + "seeds": [ + 11, + 4242, + 90210 + ], + "cap": 400, + "longRound": 15, + "deadlierBar": 1.15, + "configs": { + "cut": { + "label": "four of the roster vs three of the column", + "party": [ + "ashcroft", + "bhattacharya", + "renshaw", + "braithwaite" + ], + "creature": "quiet_neighbours_npc", + "count": 3, + "median": 10.7, + "p75": 14.3, + "p90": 19, + "p95": 22.3, + "p99": 32.3, + "longest": 67, + "over15": 19.9, + "over20": 7.3, + "over30": 1.3, + "wipeIfLong": 58.1, + "wipeIfShort": 42.5, + "longShare": 19.9, + "deadlierNoise": 0.1, + "capped": 0, + "undecided": 0 + }, + "column": { + "label": "six of the roster vs six of the column", + "party": [ + "ashcroft", + "bhattacharya", + "renshaw", + "braithwaite", + "pollard", + "okonkwo" + ], + "creature": "quiet_neighbours_npc", + "count": 6, + "median": 14, + "p75": 19.7, + "p90": 27, + "p95": 32.7, + "p99": 45.3, + "longest": 90, + "over15": 40.3, + "over20": 21.6, + "over30": 6.5, + "wipeIfLong": 23.2, + "wipeIfShort": 9.6, + "longShare": 40.3, + "deadlierNoise": 1.1, + "capped": 0, + "undecided": 0 + } + } +} diff --git a/tools/fight-tail.mjs b/tools/fight-tail.mjs new file mode 100644 index 0000000..758cc32 --- /dev/null +++ b/tools/fight-tail.mjs @@ -0,0 +1,276 @@ +/** + * fight-tail — how long these fights run, and what a median hides. + * + * Fourteen guards measured this scenario's combat and not one could have told a GM that a + * twenty-round fight was ordinary, because every one of them averages. EXPOSURE prints + * median rounds, and a median is the least useful number for a convention slot: the GM is + * not planning for the typical fight, they are planning for the one that eats the act. + * + * node tools/fight-tail.mjs read the recorded distribution + * node tools/fight-tail.mjs --check compare against the baseline + * node tools/fight-tail.mjs --update re-record it + * + * THE CAP IS THE MEASUREMENT HERE, which is why this tool passes its own. + * + * runFight stops at maxRounds and scores the result by who is standing, so a fight that + * would have run longer is recorded as exactly maxRounds and counts as neither a win nor a + * wipe. At the 40-round default that censors the top of the distribution — the six-a-side + * 99th percentile reads 40 because 40 is the wall, not because the fights end there. For + * every other guard that is a rounding error on a mean. For this one it is the finding, so + * the cap is set far past any observed fight and the check refuses to record a baseline if + * anything reaches it. R-271 documented the censoring; this tool is the one that cannot + * tolerate it. + * + * TWO CLAIMS, because the obvious one does not discriminate. "The long end is roughly twice + * the median" is true of both encounters and so tells a GM nothing about which is which. + * What is worth knowing is that LENGTH PREDICTS DEATH — a fight still going at fifteen + * rounds is not a stalemate, it is one the party is losing slowly — and that the SIX-A-SIDE + * fight is the longer one, which is counterintuitive until R-270: more bodies means more of + * them fighting on at reduced skill instead of dropping. The cut is shorter because it is + * decisive, not because it is safer. + * + * "Past N rounds" means strictly more than N. Stated because a >= count reads five points + * higher at N=15 and the two are easy to confuse in prose. + */ +import { readFile, writeFile } from "node:fs/promises"; +import { existsSync } from "node:fs"; +import path from "node:path"; +import { NPCS } from "./content.mjs"; +import { ROSTER } from "./roster.mjs"; +import { runFight, seedFor, makeRng } from "./simulate.mjs"; +import { castAndCut } from "./declared-cast.mjs"; + +const ROOT = path.resolve(path.dirname(new URL(import.meta.url).pathname.replace(/^\/([A-Za-z]:)/, "$1")), ".."); +const BASELINE = path.join(ROOT, "tools", "fight-tail-baseline.json"); + +const argv = process.argv.slice(2); +const UPDATE = argv.includes("--update"); +const CHECK = argv.includes("--check"); + +const RUNS = 2000; +const SEEDS = [11, 4242, 90210]; +/* Far past the longest fight ever observed here (71 rounds). Not runFight's default of 40, + which truncates precisely what this tool measures. */ +const CAP = 400; +/* The round by which a convention act is gone whatever the party size, so the long/short + split is an absolute number rather than a per-config percentile. */ +const LONG = 15; + +const { line: LINE, cut: CUT } = castAndCut("fight-tail"); +const CONFIGS = [ + { id: "cut", label: "four of the roster vs three of the column", party: CUT, creature: "quiet_neighbours_npc", count: 3 }, + { id: "column", label: "six of the roster vs six of the column", party: LINE, creature: "quiet_neighbours_npc", count: 6 } +]; + +/* How much likelier a long fight must be to end in a wipe before the claim counts as true. + Measured at 1.37x in the cut (seed noise 0.1) and 2.42x at six a side (seed noise 1.1 — + forty-five per cent of the value, which is why no prose quotes that ratio). The bar sits + at roughly twice the cut's noise below the cut's figure, so it asserts the SIGN of the + relationship and nothing finer: the point is to catch it disappearing, not to pin a + number three seeds cannot support. */ +const DEADLIER = 1.15; + +const roster = keys => keys.map(k => { + const found = ROSTER.find(r => r.key === k || r.key === "pc_" + k); + if (!found) { console.error(`fight-tail: "${k}" is not on the duty roster`); process.exit(1); } + return found; +}); +const creature = key => { + const spec = NPCS.find(n => n.key === key); + if (!spec) { console.error(`fight-tail: no creature keyed "${key}"`); process.exit(1); } + return spec; +}; + +const mean = a => a.reduce((x, y) => x + y, 0) / a.length; +const r1 = n => Number(n.toFixed(1)); + +/** One sweep. Lengths are kept whole rather than averaged, which is the entire point. */ +function sweep(party, foes, seed, runs) { + const rounds = []; + let capped = 0, undecided = 0; + const longF = { n: 0, wiped: 0 }, shortF = { n: 0, wiped: 0 }; + for (let i = 0; i < runs; i++) { + const r = runFight(makeRng(seed + i * 2654435761), party, foes, { maxRounds: CAP }); + rounds.push(r.rounds); + if (r.rounds >= CAP) capped++; + if (!r.wiped && !r.won) undecided++; + const bucket = r.rounds > LONG ? longF : shortF; + bucket.n++; + if (r.wiped) bucket.wiped++; + } + rounds.sort((a, b) => a - b); + const q = p => rounds[Math.min(rounds.length - 1, Math.floor(p * rounds.length))]; + const over = n => (100 * rounds.filter(x => x > n).length) / rounds.length; + return { + median: q(0.5), p75: q(0.75), p90: q(0.9), p95: q(0.95), p99: q(0.99), + longest: rounds[rounds.length - 1], + over15: over(15), over20: over(20), over30: over(30), + wipeIfLong: (100 * longF.wiped) / Math.max(longF.n, 1), + wipeIfShort: (100 * shortF.wiped) / Math.max(shortF.n, 1), + longShare: (100 * longF.n) / Math.max(longF.n + shortF.n, 1), + capped, undecided + }; +} + +function measure(runs = RUNS) { + const out = {}; + for (const cfg of CONFIGS) { + const party = roster(cfg.party); + const foes = Array(cfg.count).fill(creature(cfg.creature)); + const each = SEEDS.map(s => sweep(party, foes, seedFor(s, cfg.id), runs)); + const avg = f => r1(mean(each.map(f))); + out[cfg.id] = { + label: cfg.label, party: cfg.party, creature: cfg.creature, count: cfg.count, + median: avg(e => e.median), p75: avg(e => e.p75), p90: avg(e => e.p90), + p95: avg(e => e.p95), p99: avg(e => e.p99), + /* A single order statistic from one sample and the least stable number here: same + party, same seeds, same runs, and renaming this config from "line" to "column" + moved it 71 -> 90, because seedFor derives the stream from the id. Measured range + across seven labels: median 0, p95 1, p99 3, longest 21. Recorded for exact drift + only — check-cited refuses to let it be quoted, because a sample maximum reads + exactly like a limit. */ + longest: Math.max(...each.map(e => e.longest)), + over15: avg(e => e.over15), over20: avg(e => e.over20), over30: avg(e => e.over30), + wipeIfLong: avg(e => e.wipeIfLong), wipeIfShort: avg(e => e.wipeIfShort), + longShare: avg(e => e.longShare), + /* Seed-to-seed spread of the deadlier ratio. Recorded because renaming this config + re-seeded it via seedFor and moved the six-a-side figure from 2.8 to 2.4 — three + seeds are not enough to quote this ratio to one decimal, only to assert its sign. */ + deadlierNoise: r1(Math.max(...each.map(e => e.wipeIfLong / Math.max(e.wipeIfShort, 1e-9))) + - Math.min(...each.map(e => e.wipeIfLong / Math.max(e.wipeIfShort, 1e-9)))), + capped: each.reduce((a, e) => a + e.capped, 0), + undecided: each.reduce((a, e) => a + e.undecided, 0) + }; + } + return out; +} + +/** Refuse to record or trust a censored measurement. */ +function assertUncensored(m, when) { + const bad = Object.entries(m).filter(([, c]) => c.capped > 0 || c.undecided > 0); + if (!bad.length) return; + console.error(`fight-tail: FAILED — ${when}: ${bad.map(([id, c]) => + `${id} had ${c.capped} fights reach the ${CAP}-round ceiling and ${c.undecided} that ended ` + + `neither won nor wiped`).join("; ")}. The top of the distribution is truncated, so every ` + + `percentile above it is a wall rather than a measurement. Raise CAP.`); + process.exit(1); +} + +if (UPDATE) { + const m = measure(); + assertUncensored(m, "refusing to record"); + await writeFile(BASELINE, JSON.stringify({ + note: "Generated by tools/fight-tail.mjs --update. Do not edit by hand.", + runs: RUNS, seeds: SEEDS, cap: CAP, longRound: LONG, deadlierBar: DEADLIER, + configs: m + }, null, 2) + "\n", "utf8"); + console.log(`fight-tail: baseline recorded — ${CONFIGS.length} encounters, ${RUNS} runs x ${SEEDS.length} seeds, nothing censored`); + process.exit(0); +} + +if (!existsSync(BASELINE)) { + console.error("fight-tail: no baseline. Run with --update to record one."); + process.exit(1); +} +const base = JSON.parse(await readFile(BASELINE, "utf8")); + +if (!CHECK) { + console.log(`\nhow long the fight runs — ${base.runs} runs x ${base.seeds.length} seeds, ` + + `nothing truncated (ceiling ${base.cap})\n`); + for (const c of Object.values(base.configs)) { + console.log(` ${c.label}`); + console.log(` median ${c.median} · 75th ${c.p75} · 90th ${c.p90} · 95th ${c.p95} · 99th ${c.p99}`); + console.log(` past 15 rounds ${c.over15}% · past 20 ${c.over20}% · past 30 ${c.over30}%`); + console.log(` the honest long-end figure is the 99th percentile, ${c.p99} — stable to about ` + + `three rounds across streams.`); + console.log(` (longest in this sample ${c.longest}, which is NOT a bound: the maximum of ` + + `6000 fights moves across a 21-round range on`); + console.log(` a re-seed and is recorded for drift only. Do not cite it.)`); + console.log(` a fight past 15 rounds wipes the party ${c.wipeIfLong}% of the time, ` + + `against ${c.wipeIfShort}% for shorter ones — ${r1(c.wipeIfLong / c.wipeIfShort)}x`); + console.log(""); + } + console.log(` The six-a-side fight is the LONGER one (median ${base.configs.column.median} against ` + + `${base.configs.cut.median}). More bodies means more of them fighting on at reduced skill`); + console.log(` instead of dropping. The cut is shorter because it is decisive, not because it is safer.\n`); + process.exit(0); +} + +/* ---- --check ------------------------------------------------------------------ */ + +if (base.runs !== RUNS || base.seeds?.join() !== SEEDS.join() + || base.cap !== CAP || base.longRound !== LONG) { + console.error(`fight-tail: FAILED — the baseline was recorded under different conditions ` + + `(${base.runs} runs, seeds ${base.seeds?.join(",")}, ceiling ${base.cap}, long at ${base.longRound}). ` + + `Re-record it with --update.`); + process.exit(1); +} +for (const cfg of CONFIGS) { + const was = base.configs[cfg.id]; + if (!was) { console.error(`fight-tail: FAILED — nothing recorded for "${cfg.id}". Re-record with --update.`); process.exit(1); } + if (was.party.join() !== cfg.party.join() || was.creature !== cfg.creature || was.count !== cfg.count) { + const gone = was.party.filter(k => !cfg.party.includes(k)); + const added = cfg.party.filter(k => !was.party.includes(k)); + console.error(`fight-tail: FAILED — ${cfg.id}: the scenario now casts different people` + + (gone.length ? ` — no longer ${gone.join(", ")}` : "") + + (added.length ? `, now ${added.join(", ")}` : "") + + `. The recorded distribution describes the old cast; re-record with --update, ` + + `do not revert the cast.`); + process.exit(1); + } +} + +/* An exact comparison is only honest if the measurement is exact. Prove it. */ +{ + const a = JSON.stringify(measure(150)); + const b = JSON.stringify(measure(150)); + if (a !== b) { + console.error("fight-tail: FAILED — the simulation is not deterministic, so an exact " + + "baseline cannot mean anything."); + process.exit(1); + } +} + +const now = measure(); +assertUncensored(now, "the measurement no longer fits under the ceiling"); + +const problems = []; +for (const [id, was] of Object.entries(base.configs)) { + const m = now[id]; + if (!m) { problems.push(` ${id}: no longer measured`); continue; } + for (const f of ["median", "p75", "p90", "p95", "p99", "longest", + "over15", "over20", "over30", "wipeIfLong", "wipeIfShort", "longShare", "deadlierNoise"]) { + if (m[f] !== was[f]) problems.push(` ${id}: ${f} ${was[f]} -> ${m[f]}`); + } +} +if (problems.length) { + console.error("fight-tail: FAILED — these fights no longer run the length they were recorded at"); + problems.forEach(p => console.error(p)); + console.error(" Read `node tools/fight-tail.mjs` before re-recording, and check whether " + + "EXPOSURE's prose about the long end still holds."); + process.exit(1); +} + +/* The two claims EXPOSURE is allowed to print, restated as tests. The drift comparison + above already catches every movement, so these assert only the sentences — and unlike + pinned figures they cannot be satisfied by re-recording, because --update rewrites the + numbers and leaves the shape. */ +const notDeadlier = Object.entries(now).filter(([, c]) => c.wipeIfLong / c.wipeIfShort < DEADLIER); +if (notDeadlier.length) { + console.error(`fight-tail: FAILED — ${notDeadlier.map(([id]) => id).join(", ")}: a long fight is no ` + + `longer at least ${DEADLIER}x likelier to end in a wipe than a short one, so "a fight still ` + + `going is one the party is losing slowly" is false. Rewrite the prose rather than re-recording.`); + process.exit(1); +} +if (now.column.median <= now.cut.median) { + console.error(`fight-tail: FAILED — the six-a-side fight is no longer the longer one ` + + `(median ${now.column.median} against the cut's ${now.cut.median}). EXPOSURE says more bodies ` + + `means a longer fight, not a safer one; that sentence is now wrong. Rewrite it rather than ` + + `re-recording.`); + process.exit(1); +} + +console.log(`fight-tail: OK — a fight past ${LONG} rounds wipes the party ` + + `${r1(now.cut.wipeIfLong / now.cut.wipeIfShort)}x as often in the cut and ` + + `${r1(now.column.wipeIfLong / now.column.wipeIfShort)}x at six a side, the six-a-side fight is the ` + + `longer one (median ${now.column.median} vs ${now.cut.median}), and nothing was truncated`); diff --git a/tools/update-readme.mjs b/tools/update-readme.mjs index 395cd9a..f9e95b3 100644 --- a/tools/update-readme.mjs +++ b/tools/update-readme.mjs @@ -37,7 +37,7 @@ const GUARDS = ["check-rules.mjs", "check-kits.mjs", "check-lang.mjs", "check-templates.mjs", "check-behaviour.mjs", "check-scenarios.mjs", "check-rollable.mjs", "check-creatures.mjs", "check-anatomy.mjs", "check-lethality.mjs", "check-focus.mjs", "check-firstblood.mjs", - "check-attackers.mjs", "check-bestiary.mjs"]; + "check-attackers.mjs", "check-fight-tail.mjs", "check-cited.mjs", "check-bestiary.mjs"]; /* The list above went stale the moment a guard was added without touching this file — which is what the comment above it already warned about, and which happened again