Files
RingBRP/tools/check-packs.mjs
slaguru666andClaude Opus 5 a5c4378521 R-294: EXPOSURE's tables measured into an artifact, and 58.7% was 74.9%
Desk pass 9 found the scenario's most consequential table held by nothing:
~30 sampled figures, no baseline, catchable only by re-running the exact
command. lethality-baseline.json does not cover them — it measures creatures
SOLO against a frozen four-agent party that is not this cast.

Building the measurer found the figures were also wrong.

simulate.mjs's CLI calls runFight with no options, so every published figure
was measured at the default 40-round ceiling — and runFight does not report
truncation, it scores whoever is standing when the loop stops. Measured at
400:

  line.column6      14.8% -> 15.2%    34/2000 truncated
  line.hollow6col2   7.8% ->  8.3%   129
  line.hollow6col3  28.7% -> 30.1%   264
  cut.hollow6       58.7% -> 74.9%   578   <- 29% of runs never finished

Fourteen points on the single most alarming number in the case, and the one
v0.16 added specifically to warn four-player tables. fight-tail learned this
in R-270 and carries an assertUncensored; EXPOSURE's own tables never got
one. The longest fight at cap 400 is 120 rounds and 400 vs 2000 are
identical, so the cap is comfortable rather than merely sufficient.

tools/pack-tables.mjs measures all 17 configs against the DERIVED cast via
castAndCut, using measure()'s exact discipline — one rng threaded through
every run, not a reseed per run, because reseeding is a different stream and
would not reproduce the published table. It refuses to report or record a
truncated sweep.

tools/check-packs.mjs holds the baseline against the game, so the pair is not
a loop: check-cited holds the prose against the record, this holds the record
against the harness. Hard claims read `now` and never `base`: nothing
truncated, the party equals the declared cast, config floor, and more of the
same creature may not make the party safer. Figures compare exactly, since
the runs are deterministic.

Verified by breaking it: a drifted baseline, CAP lowered to 40, and a removed
config each turn it red, and --update refuses outright rather than recording
a truncated sweep. The removed-config test first passed for a bad reason —
MIN_CONFIGS is a floor and 16 clears it — so a dropped-config check was added
and re-tested on a row in no monotonic chain. All files restored
byte-identical after each probe.

EXPOSURE's three tables and the nine prose figures around them are rebuilt
from the artifact with citation markers; the multi-seed stability claims were
re-measured too (the cut is 74.9/72.7/75.0/73.5 across four seeds, not
58.7/60.5/60.3/59.8). CLEAN GROUND v0.20.

npm run check: 20 guards pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 16:32:35 +01:00

130 lines
6.7 KiB
JavaScript

/**
* EXPOSURE's tables, held against the harness that produced them.
*
* node tools/check-packs.mjs hold the claim
* node tools/check-packs.mjs --list print what it found
* node tools/pack-tables.mjs --update re-record after a deliberate change
*
* check-cited holds the PROSE against the baseline. This holds the BASELINE against the
* game, and without it the pair is a loop: re-record, and the prose can be corrected to
* match a figure nobody has checked came from anywhere.
*
* THE HARD CLAIMS READ `now` AND NEVER `base`, so --update cannot silence them — R-270's
* arrangement in check-fight-tail, which is the only part of that guard that cannot be
* argued with. Chief among them is the ceiling: these tables were published for versions at
* runFight's default 40 rounds, which truncated 578 of 2000 runs in the worst config and
* understated its wipe rate by fourteen points. A guard that could be talked out of that
* check would be worth nothing.
*/
import { readFileSync, existsSync } from "node:fs";
import path from "node:path";
import { measureAll, censoredIn, CONFIGS, BASELINE, CAP, RUNS } from "./pack-tables.mjs";
import { castAndCut } from "./declared-cast.mjs";
const ROOT = path.resolve(path.dirname(new URL(import.meta.url).pathname.replace(/^\/([A-Za-z]:)/, "$1")), "..");
const LIST = process.argv.slice(2).includes("--list");
const MIN_CONFIGS = 14;
const FIELDS = ["hurt", "down", "wiped", "deaths", "median"];
if (!existsSync(BASELINE)) {
console.error(`check-packs: FAILED — no baseline at ${path.relative(ROOT, BASELINE)}. `
+ `Record one with: node tools/pack-tables.mjs --update`);
process.exit(1);
}
const base = JSON.parse(readFileSync(BASELINE, "utf8"));
const now = measureAll();
const problems = [];
const notes = [];
/* ── HARD CLAIMS ───────────────────────────────────────────────────────────────────── */
const censored = censoredIn(now);
if (censored.length) {
problems.push(` ${censored.length} config(s) hit the ${CAP}-round ceiling: ${censored.join(", ")}. `
+ `A truncated fight is scored by whoever is standing when the loop stops, so the wipe rate `
+ `is understated and nothing says so. This is the defect that produced the 58.7% these `
+ `tables carried for three versions when the real figure was 74.9%. Raise CAP, re-record.`);
} else {
notes.push(`nothing truncated at ${CAP} rounds (longest ${Math.max(...CONFIGS.map(c => now[c.party][c.id].longest))})`);
}
if (CONFIGS.length < MIN_CONFIGS) {
problems.push(` only ${CONFIGS.length} configs, and EXPOSURE prints at least ${MIN_CONFIGS} rows. `
+ `A shrunken config list would let this guard report full agreement over a table it is no `
+ `longer measuring.`);
}
const { line, cut } = castAndCut("check-packs");
for (const [which, want] of [["line", line], ["cut", cut]]) {
const got = now.party[which];
if (got.join(",") !== want.join(",")) {
problems.push(` the ${which} party measured here is ${got.join(", ")}, and the scenario declares `
+ `${want.join(", ")}. These figures would be a real measurement of a fight this case does not cast.`);
}
}
if (!problems.length) notes.push(`measured against the declared cast: ${line.join(", ")}`);
/* More of the same creature must not make the party safer. Not a style rule — EXPOSURE's
"the curve between six and ten is almost vertical" is an argument about this ordering,
and a violation means the harness has broken rather than that the document is wrong. */
for (const [which, ids] of [["line", ["hollow1", "hollow3", "hollow6"]], ["line", ["column3", "column6", "column10", "column15"]],
["cut", ["hollow3", "hollow6"]], ["cut", ["column3", "column6"]]]) {
for (let i = 1; i < ids.length; i++) {
const a = now[which][ids[i - 1]], b = now[which][ids[i]];
if (b.wiped < a.wiped) {
problems.push(` ${which}.${ids[i]} (${b.n} of them) wipes ${b.wiped}% and ${which}.${ids[i - 1]} `
+ `(${a.n}) wipes ${a.wiped}%. More of the same creature made the party safer, which is the `
+ `harness misbehaving, not a finding.`);
}
}
}
/* ── AGREEMENT WITH THE RECORD — deterministic, so exact ───────────────────────────── */
let drift = 0;
for (const c of CONFIGS) {
const b = base[c.party]?.[c.id], n = now[c.party][c.id];
if (!b) { problems.push(` ${c.party}.${c.id} is measured here and absent from the baseline. Re-record.`); continue; }
for (const f of FIELDS) {
if (b[f] !== n[f]) {
drift++;
problems.push(` ${c.party}.${c.id}.${f}: the game now says ${n[f]}, the baseline holds ${b[f]}. `
+ `These runs are deterministic at seed ${now.seed}, so this is a real change in the harness `
+ `or the cast, not noise. Re-record with --update once you know which, and fix EXPOSURE's `
+ `prose in the same commit — check-cited will not let you forget.`);
}
}
}
/* The mirror of the check above, and the one a config-list edit needs. Dropping a config
leaves the loop above iterating over what remains and agreeing with itself; coverage
falls and every figure still matches. MIN_CONFIGS is a floor, not a ratchet, so it
cannot see a single row going missing. */
const dropped = [];
for (const which of ["line", "cut"])
for (const id of Object.keys(base[which] ?? {}))
if (!CONFIGS.some(c => c.party === which && c.id === id)) dropped.push(`${which}.${id}`);
if (dropped.length) {
problems.push(` ${dropped.length} config(s) in the record are no longer measured: ${dropped.join(", ")}. `
+ `Coverage fell and every remaining figure still agreed, which is how this guard would go on `
+ `reporting success over a smaller table than the one EXPOSURE prints. Re-record deliberately `
+ `if a row was really retired.`);
}
if (!drift && !dropped.length) notes.push(`${CONFIGS.length} configs x ${FIELDS.length} figures agree exactly with the record`);
if (LIST) {
for (const c of CONFIGS) {
const r = now[c.party][c.id];
console.log(` ${(c.party + "." + c.id).padEnd(20)} ${c.label.padEnd(38)} hurt ${r.hurt} down ${r.down} wiped ${r.wiped}% deaths ${r.deaths} median ${r.median}`);
}
notes.forEach(n => console.log(` ${n}`));
problems.forEach(p => console.log(p));
process.exit(0);
}
if (problems.length) {
console.error("check-packs: FAILED — EXPOSURE's tables no longer describe the fight the game runs");
problems.forEach(p => console.error(p));
process.exit(1);
}
console.log(`check-packs: OK — ${CONFIGS.length} pack fights, ${RUNS} runs each at seed ${now.seed}, `
+ `nothing truncated at ${CAP} rounds, every figure matching the record`);