It imported NPCS directly, so once the bestiary moved into its own files it
measured 47 of 147 and printed OK. A guard blind to two thirds of what it
guards is not a guard, and this one failed silently — the worse kind.
Now reads collectSpecs(), the list check-creatures, make-portraits and
mj-queue already share. Scenario casts come in with the bestiary: they were
never measured, and they should have been, since the Act Three standoff
this repository publishes wipe rates for IS a scenario cast. 184 specs at
roughly 47ms each, about nine seconds.
Switched BEFORE re-recording, to prove the change moved nothing: the run
against the old baseline reported all 47 existing creatures fighting
exactly as recorded. The --update is then verifiably additive — 47 to 184,
137 added, 0 changed, 0 lost.
Still fail-open on arrival: an unbaselined creature is a note, not a
failure, so this printed "OK, 137 new" while 137 creatures were unguarded.
Left as the author wrote it rather than changed under a content drop, but
it is the same shape as the three classifier bugs in 00f6a1e and worth
closing.
First thing it found, invisible until now: march_stone is authored "it
kills perhaps one person a century" and measures 43.5% wipe, 2.72 of 4
down, over 22 rounds. naturalArmour 14 is the highest in the game — the
Apex tier is 7 — so the party cannot hurt it and it grinds them down. The
prose and the statblock describe different creatures.
All 23 guards pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
192 lines
9.7 KiB
JavaScript
192 lines
9.7 KiB
JavaScript
/**
|
|
* check-lethality — no creature may quietly become something else.
|
|
*
|
|
* The seven guards before this one check that content is WELL FORMED. None of them
|
|
* checks what it DOES. A creature can validate perfectly, pack perfectly, and have
|
|
* become twice as dangerous as it was last week because somebody adjusted a damage
|
|
* modifier, a hit-point formula or the armour on a service vest. That change is
|
|
* invisible in a diff of the creature, because the creature did not change.
|
|
*
|
|
* This is the guard the creature-forge plan called for, built as a regression test
|
|
* rather than as a set of hand-declared bands. Bands were the plan's design and they
|
|
* are the wrong one here: declaring "keepers: dangerous" on forty-six creatures means
|
|
* inventing forty-six judgements, and the judgement that matters is not "is this
|
|
* dangerous" but "is this the same as it was". A baseline answers that exactly, needs
|
|
* no authoring, and cannot be argued with.
|
|
*
|
|
* node tools/check-lethality.mjs compare against the committed baseline
|
|
* node tools/check-lethality.mjs --update re-record it (after a deliberate change)
|
|
* node tools/check-lethality.mjs --verbose print every creature, not just drift
|
|
*
|
|
* WHY A FIXED PARTY. --spread exists because one number cannot describe an encounter:
|
|
* the Act Three standoff runs 0% to 100% depending which four agents were picked. That
|
|
* is true and it is why this guard does NOT try to describe danger. It measures the
|
|
* same creature against the same four agents with the same seed, so the only thing that
|
|
* can move the number is a change to the rules or to the creature. The party below is
|
|
* deliberately mixed — two armed postings and two trades — so that a change affecting
|
|
* either kind of agent shows up, and it is frozen for reproducibility rather than
|
|
* chosen for realism.
|
|
*/
|
|
import { readFile, writeFile } from "node:fs/promises";
|
|
import { existsSync } from "node:fs";
|
|
import path from "node:path";
|
|
import { collectSpecs } from "./all-specs.mjs";
|
|
import { ROSTER } from "./roster.mjs";
|
|
import { measure, seedFor } from "./simulate.mjs";
|
|
|
|
const ROOT = path.resolve(path.dirname(new URL(import.meta.url).pathname.replace(/^\/([A-Za-z]:)/, "$1")), "..");
|
|
const BASELINE = path.join(ROOT, "tools", "lethality-baseline.json");
|
|
|
|
const argv = process.argv.slice(2);
|
|
const UPDATE = argv.includes("--update");
|
|
const VERBOSE = argv.includes("--verbose");
|
|
|
|
/* Frozen. Changing this invalidates every recorded number, so if it ever must change,
|
|
change it in the same commit as a --update and say why in the message. */
|
|
const PARTY_KEYS = ["pc_holloway", "pc_okonkwo", "pc_nkemdirim", "pc_ferriby"];
|
|
const RUNS = 2000;
|
|
const SEED = 11;
|
|
/* Recorded alongside the numbers, because changing HOW the seed is derived changes every
|
|
row without changing SEED itself — a baseline taken under the old shared-seed scheme
|
|
would otherwise be compared, silently and wrongly, against this one. */
|
|
const SEED_MODE = "per-creature";
|
|
|
|
/* NO TOLERANCE, deliberately. The reasoning it replaces sounded right — "200 runs of a
|
|
coin flip has a standard error around 3.5 points, so anything under 6 is inside the
|
|
noise" — and it is exactly backwards for this measurement. Sampling error would apply
|
|
if the runs were random; they are not. The guard used to allow six points of wipe
|
|
rate and a third of an agent before it complained, which sounds prudent and was why it
|
|
sat silent through the largest combat change the system has had: ddc4f99 routed every
|
|
ordinary blow through the hit location table, 26 of 46 creatures moved, and not one
|
|
of them moved far enough in a single number to trip the threshold. The baseline went
|
|
four commits describing a game nobody was playing.
|
|
|
|
A tolerance is for noise, and there is none here: same party, same seed, same counts,
|
|
and measure() builds its own generator per creature, so two recordings of unchanged
|
|
code are byte-identical. Anything that moves is a real change to the rules or to the
|
|
creature, which is exactly what this file exists to notice. `rounds` is compared too;
|
|
it was recorded and then ignored, so a creature could take a round longer to kill
|
|
forever without a word. */
|
|
|
|
const party = PARTY_KEYS.map(k => {
|
|
const found = ROSTER.find(r => r.key === k);
|
|
if (!found) {
|
|
console.error(`check-lethality: the frozen party names "${k}", which is not on the roster.`);
|
|
process.exit(1);
|
|
}
|
|
return found;
|
|
});
|
|
|
|
/* WHAT IS MEASURED.
|
|
*
|
|
* Every npc-kind spec in the repository, read through collectSpecs() — the one list that
|
|
* check-creatures, make-portraits and mj-queue already share. This file used to import
|
|
* NPCS from content.mjs directly, which was the same thing until the 100-creature
|
|
* bestiary landed in its own files: after that it measured 47 of 147 and reported OK.
|
|
* A guard that cannot see two thirds of the creatures it guards is not a guard, and it
|
|
* failed silently, which is worse than failing loudly.
|
|
*
|
|
* Scenario casts come in with them. They were never measured before and they should have
|
|
* been: the Act Three standoff this repository publishes wipe rates for IS a scenario
|
|
* cast, and a rules change reaches a ticket inspector exactly as it reaches a barghest.
|
|
* 184 specs at roughly 47ms each is about nine seconds. */
|
|
const SPECS = (await collectSpecs()).filter(r => r.kind === "npc").map(r => r.spec);
|
|
|
|
/** Solo, because a creature is the unit under test — counts are an encounter's business. */
|
|
const current = {};
|
|
for (const spec of SPECS) {
|
|
const r = measure(party, [spec], { runs: RUNS, seed: seedFor(SEED, spec.key) });
|
|
current[spec.key] = {
|
|
wipe: Number((r.wipeRate * 100).toFixed(1)),
|
|
down: Number(r.downMean.toFixed(2)),
|
|
rounds: r.roundsMedian
|
|
};
|
|
}
|
|
|
|
if (UPDATE) {
|
|
await writeFile(BASELINE, JSON.stringify({
|
|
note: "Generated by tools/check-lethality.mjs --update. Do not edit by hand.",
|
|
party: PARTY_KEYS, runs: RUNS, seed: SEED, seedMode: SEED_MODE,
|
|
creatures: current
|
|
}, null, 2) + "\n", "utf8");
|
|
console.log(`check-lethality: baseline recorded — ${Object.keys(current).length} creatures, `
|
|
+ `party ${PARTY_KEYS.map(k => k.replace(/^pc_/, "")).join(", ")}, ${RUNS} runs, `
|
|
+ `seed ${SEED} ${SEED_MODE}`);
|
|
process.exit(0);
|
|
}
|
|
|
|
if (!existsSync(BASELINE)) {
|
|
console.error("check-lethality: no baseline. Run with --update to record one.");
|
|
process.exit(1);
|
|
}
|
|
|
|
const base = JSON.parse(await readFile(BASELINE, "utf8"));
|
|
|
|
/* The recorded numbers mean nothing if they were taken under different conditions, and
|
|
a guard comparing incomparable numbers is worse than no guard. */
|
|
if (base.runs !== RUNS || base.seed !== SEED
|
|
|| base.seedMode !== SEED_MODE
|
|
|| base.party.join() !== PARTY_KEYS.join()) {
|
|
console.error("check-lethality: FAILED — the baseline was recorded under different conditions "
|
|
+ `(party ${base.party.join(",")}, ${base.runs} runs, seed ${base.seed} `
|
|
+ `${base.seedMode ?? "shared"}). Re-record it with --update.`);
|
|
process.exit(1);
|
|
}
|
|
|
|
/* Comparing exactly is only honest if the measurement IS exact. Prove it here rather
|
|
than trust it: one creature, measured twice, must come back identical. If a future
|
|
change reaches for Math.random or a Set iteration order, this says so instead of
|
|
letting the whole guard degrade into noise-chasing. */
|
|
{
|
|
const probe = SPECS[0];
|
|
const a = measure(party, [probe], { runs: RUNS, seed: seedFor(SEED, probe.key) });
|
|
const b = measure(party, [probe], { runs: RUNS, seed: seedFor(SEED, probe.key) });
|
|
const shape = r => JSON.stringify([r.wipeRate, r.downMean, r.roundsMedian, r.hurtMean]);
|
|
if (shape(a) !== shape(b)) {
|
|
console.error("check-lethality: FAILED — the simulation is not deterministic, so an exact "
|
|
+ `baseline cannot mean anything. ${probe.key} measured twice gave ${shape(a)} and ${shape(b)}.`);
|
|
process.exit(1);
|
|
}
|
|
}
|
|
|
|
const problems = [];
|
|
const added = [], removed = [];
|
|
|
|
for (const [key, now] of Object.entries(current)) {
|
|
const was = base.creatures[key];
|
|
if (!was) { added.push(key); continue; }
|
|
if (now.wipe === was.wipe && now.down === was.down && now.rounds === was.rounds) continue;
|
|
const dWipe = now.wipe - was.wipe, dDown = now.down - was.down;
|
|
const bits = [];
|
|
if (now.wipe !== was.wipe)
|
|
bits.push(`wiped ${was.wipe}% -> ${now.wipe}% (${dWipe >= 0 ? "+" : ""}${dWipe.toFixed(1)})`);
|
|
if (now.down !== was.down)
|
|
bits.push(`down ${was.down} -> ${now.down} (${dDown >= 0 ? "+" : ""}${dDown.toFixed(2)} of 4)`);
|
|
if (now.rounds !== was.rounds)
|
|
bits.push(`rounds ${was.rounds} -> ${now.rounds}`);
|
|
problems.push({ key, dWipe, text: `${key}: ${bits.join(", ")}` });
|
|
}
|
|
for (const key of Object.keys(base.creatures)) if (!(key in current)) removed.push(key);
|
|
|
|
if (VERBOSE) {
|
|
for (const [key, now] of Object.entries(current)) {
|
|
console.log(` ${key.padEnd(22)} ${String(now.wipe).padStart(5)}% wiped ${now.down.toFixed(2)} down ${now.rounds} rounds`);
|
|
}
|
|
}
|
|
|
|
if (added.length) console.log(` note: ${added.length} new creature(s) not in the baseline — ${added.join(", ")}`);
|
|
if (removed.length) console.log(` note: ${removed.length} creature(s) gone from content — ${removed.join(", ")}`);
|
|
|
|
if (problems.length) {
|
|
console.error(`check-lethality: FAILED — ${problems.length} of ${Object.keys(current).length} `
|
|
+ `creature(s) fight differently than recorded`);
|
|
problems.sort((a, b) => Math.abs(b.dWipe) - Math.abs(a.dWipe));
|
|
for (const p of problems) console.error(" " + p.text);
|
|
console.error("\n If this was deliberate, re-record with --update and say what changed in the commit.");
|
|
process.exit(1);
|
|
}
|
|
|
|
console.log(`check-lethality: OK — ${Object.keys(current).length} creatures, every one fighting `
|
|
+ `exactly as recorded against the frozen party`
|
|
+ `${added.length ? `, ${added.length} new` : ""}`);
|