tools/playthrough.mjs exists to show a reader why a measured number is what it is. I cited it to another session — read seed 2, see the wipe the 45% row is made of — and the reproduction failed in front of them. They reported different outcomes at both seeds and guessed the cause correctly from outside: the two tools were not drawing from the same stream. They were not. simulate.mjs used a private mulberry32 makeRng; playthrough.mjs had its own LCG written to look like it. Both deterministic, both reproducible alone, and "seed 2" named a different fight in each — which breaks the only thing the tool is for. Its own comment claimed a seed here names the same fight there. check-focus carried a third copy of that LCG, so the two guards described the same game with different dice. makeRng is exported and both files use it. A playthrough seed is now exactly the first fight of simulate.mjs --seed <n>: --runs 1 --seed 2 and the playthrough give 9 rounds, 4 of 4 down, 1 dead, both. The other half was my citation rather than the code: the command I sent omitted --mode, so it plays both targeting arms and prints two fights. They read the last line, I quoted the first. The summary line now names the arm. check-focus re-recorded under the shared stream; figures move a point or two. What it buys is that the bimodality analysis now reproduces the published means exactly — 2.31, 3.03, 0.40, 0.42 against the four rows CLEAN GROUND publishes. Under the old LCG it agreed to within a decimal, which looked like corroboration and was two experiments landing near each other. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
204 lines
9.4 KiB
JavaScript
204 lines
9.4 KiB
JavaScript
/**
|
|
* check-focus — the bestiary's advice must keep being true.
|
|
*
|
|
* R-258 put a tactic on a page a GM reads, measured in a harness nobody guards. R-259
|
|
* found it worth nothing when run through the simulator. R-260 found the replacement
|
|
* advice right for the wrong reason. Each time the correction was an assertion in a
|
|
* document, which is the same arrangement that let ddc4f99 describe a game nobody was
|
|
* playing: a number written down with nothing checking it.
|
|
*
|
|
* So the claim is recorded as an artifact and re-measured on every build, exactly the
|
|
* way check-lethality does it. The bestiary reads the artifact rather than restating it,
|
|
* so the page, the guard and the simulator cannot disagree.
|
|
*
|
|
* node tools/check-focus.mjs compare against the committed baseline
|
|
* node tools/check-focus.mjs --update re-record it (after a deliberate change)
|
|
* node tools/check-focus.mjs --verbose print every creature, not just drift
|
|
*
|
|
* WHAT IS MEASURED. For each creature, one pack size, and two arms: every agent picking
|
|
* a target at random, versus every agent putting their attacks into the same creature.
|
|
* The difference is what focus fire is worth against that creature.
|
|
*
|
|
* WHY A RECORDED PACK SIZE. A fight the party wins 100% of the time, or loses 100% of
|
|
* the time, cannot show an effect of any size — three of the first four packs measured
|
|
* in R-259 were pinned like that and said nothing. --update picks, for each creature,
|
|
* the size whose spread-fire rate sits nearest an even fight, and RECORDS it. The guard
|
|
* then re-measures at that size rather than choosing again, so a creature that changes
|
|
* shows up as drift instead of quietly moving to a different question.
|
|
*/
|
|
import { readFile, writeFile } from "node:fs/promises";
|
|
import { existsSync } from "node:fs";
|
|
import path from "node:path";
|
|
import { NPCS } from "./content.mjs";
|
|
import { ROSTER } from "./roster.mjs";
|
|
import { runFight, seedFor, makeRng } from "./simulate.mjs";
|
|
|
|
const ROOT = path.resolve(path.dirname(new URL(import.meta.url).pathname.replace(/^\/([A-Za-z]:)/, "$1")), "..");
|
|
const BASELINE = path.join(ROOT, "tools", "focus-baseline.json");
|
|
|
|
const argv = process.argv.slice(2);
|
|
const UPDATE = argv.includes("--update");
|
|
const VERBOSE = argv.includes("--verbose");
|
|
|
|
/* The same frozen party as check-lethality, for the same reason and so the two guards
|
|
describe the same game. */
|
|
const PARTY_KEYS = ["pc_holloway", "pc_okonkwo", "pc_nkemdirim", "pc_ferriby"];
|
|
/* 1000 rather than 500 on evidence, not preference: at 500 runs one row's gain no longer
|
|
clears its own seed-to-seed noise and the claim-check below fails, because the page's
|
|
"without exception" is not supported by that much sampling. The build pays 8s for a
|
|
sentence it is entitled to print. */
|
|
const RUNS = 1000;
|
|
const SEEDS = [11, 4242, 90210];
|
|
const SEED_MODE = "per-creature";
|
|
const SIZES = [2, 3, 4, 5, 6];
|
|
const SCAN_RUNS = 400;
|
|
/* A pack whose spread-fire rate falls outside this cannot move enough to measure, and a
|
|
row that cannot move is a row that would pass this guard whatever happened to it. */
|
|
const MEASURABLE = [15, 85];
|
|
|
|
const party = PARTY_KEYS.map(k => {
|
|
const found = ROSTER.find(r => r.key === k);
|
|
if (!found) {
|
|
console.error(`check-focus: the frozen party names "${k}", which is not on the roster.`);
|
|
process.exit(1);
|
|
}
|
|
return found;
|
|
});
|
|
|
|
/* One generator per arm, derived from the creature's own key, so inserting a creature
|
|
cannot perturb its neighbours — R-252, learned the hard way. */
|
|
/* R-265: the harness's own generator, not a second one of the same shape. This file used
|
|
to roll its own LCG, so the two guards described the same game with different dice and
|
|
no figure here could be reproduced from simulate.mjs. */
|
|
const rngFor = makeRng;
|
|
const winRate = (foes, targets, seed) => {
|
|
let won = 0;
|
|
for (let i = 0; i < RUNS; i++) {
|
|
if (runFight(rngFor(seed + i * 2654435761), party, foes, { partyTargets: targets }).won) won++;
|
|
}
|
|
return (100 * won) / RUNS;
|
|
};
|
|
const mean = a => a.reduce((x, y) => x + y, 0) / a.length;
|
|
|
|
const measureAt = (spec, n) => {
|
|
const foes = Array(n).fill(spec);
|
|
const seeds = SEEDS.map(s => seedFor(s, spec.key));
|
|
const spread = seeds.map(s => winRate(foes, "random", s));
|
|
const focus = seeds.map(s => winRate(foes, "focus", s));
|
|
return {
|
|
n,
|
|
spread: Number(mean(spread).toFixed(1)),
|
|
focus: Number(mean(focus).toFixed(1)),
|
|
gain: Number((mean(focus) - mean(spread)).toFixed(1)),
|
|
noise: Number((Math.max(...spread) - Math.min(...spread)).toFixed(1))
|
|
};
|
|
};
|
|
|
|
/* --update only: choose the pack size that makes the question answerable. */
|
|
const pickSize = spec => {
|
|
let best = null;
|
|
for (const n of SIZES) {
|
|
const w = winRate(Array(n).fill(spec), "random", seedFor(SEEDS[0], spec.key), SCAN_RUNS);
|
|
if (!best || Math.abs(w - 50) < Math.abs(best.w - 50)) best = { n, w };
|
|
if (w < 5) break; // already hopeless; more of them is worse
|
|
}
|
|
return best;
|
|
};
|
|
|
|
if (UPDATE) {
|
|
const creatures = {};
|
|
let measurable = 0;
|
|
for (const spec of NPCS) {
|
|
const pick = pickSize(spec);
|
|
if (pick.w < MEASURABLE[0] || pick.w > MEASURABLE[1]) {
|
|
creatures[spec.key] = { n: pick.n, pinned: Number(pick.w.toFixed(1)) };
|
|
continue;
|
|
}
|
|
creatures[spec.key] = measureAt(spec, pick.n);
|
|
measurable++;
|
|
}
|
|
await writeFile(BASELINE, JSON.stringify({
|
|
note: "Generated by tools/check-focus.mjs --update. Do not edit by hand.",
|
|
party: PARTY_KEYS, runs: RUNS, seeds: SEEDS, seedMode: SEED_MODE,
|
|
measurableBand: MEASURABLE, creatures
|
|
}, null, 2) + "\n", "utf8");
|
|
console.log(`check-focus: baseline recorded — ${measurable} of ${NPCS.length} creatures measurable, `
|
|
+ `${RUNS} runs x ${SEEDS.length} seeds, seed ${SEEDS[0]} ${SEED_MODE}`);
|
|
process.exit(0);
|
|
}
|
|
|
|
if (!existsSync(BASELINE)) {
|
|
console.error("check-focus: no baseline. Run with --update to record one.");
|
|
process.exit(1);
|
|
}
|
|
const base = JSON.parse(await readFile(BASELINE, "utf8"));
|
|
|
|
if (base.runs !== RUNS || base.seeds?.join() !== SEEDS.join()
|
|
|| base.seedMode !== SEED_MODE || base.party.join() !== PARTY_KEYS.join()) {
|
|
console.error("check-focus: FAILED — the baseline was recorded under different conditions "
|
|
+ `(party ${base.party.join(",")}, ${base.runs} runs, seeds ${base.seeds?.join(",")} `
|
|
+ `${base.seedMode ?? "shared"}). Re-record it with --update.`);
|
|
process.exit(1);
|
|
}
|
|
|
|
/* Exact comparison is only honest if the measurement is exact. Prove it, don't assume. */
|
|
{
|
|
const probe = NPCS.find(s => base.creatures[s.key]?.gain !== undefined);
|
|
if (probe) {
|
|
const a = JSON.stringify(measureAt(probe, base.creatures[probe.key].n));
|
|
const b = JSON.stringify(measureAt(probe, base.creatures[probe.key].n));
|
|
if (a !== b) {
|
|
console.error("check-focus: FAILED — the simulation is not deterministic, so an exact "
|
|
+ `baseline cannot mean anything. ${probe.key} measured twice gave ${a} and ${b}.`);
|
|
process.exit(1);
|
|
}
|
|
}
|
|
}
|
|
|
|
const problems = [], added = [], removed = [];
|
|
let checked = 0, helped = 0, aboveNoise = 0;
|
|
|
|
for (const spec of NPCS) {
|
|
const was = base.creatures[spec.key];
|
|
if (!was) { added.push(spec.key); continue; }
|
|
if (was.pinned !== undefined) continue; // recorded as unmeasurable; nothing to compare
|
|
const now = measureAt(spec, was.n);
|
|
checked++;
|
|
if (now.gain > 0) helped++;
|
|
if (now.gain > now.noise) aboveNoise++;
|
|
const same = now.spread === was.spread && now.focus === was.focus && now.gain === was.gain;
|
|
if (!same) {
|
|
problems.push(` ${spec.key}: ${was.n} of them — spread ${was.spread}% -> ${now.spread}%, `
|
|
+ `focus ${was.focus}% -> ${now.focus}%, focus fire worth ${was.gain} -> ${now.gain}`);
|
|
} else if (VERBOSE) {
|
|
console.log(` ${spec.key.padEnd(24)} ${was.n}x spread ${was.spread}% focus ${was.focus}% +${was.gain}`);
|
|
}
|
|
}
|
|
for (const key of Object.keys(base.creatures)) if (!NPCS.some(s => s.key === key)) removed.push(key);
|
|
|
|
if (problems.length || added.length || removed.length) {
|
|
console.error(`check-focus: FAILED — ${problems.length} creature(s) answer differently than recorded`);
|
|
problems.forEach(p => console.error(p));
|
|
if (added.length) console.error(` not in the baseline: ${added.join(", ")} — re-record with --update`);
|
|
if (removed.length) console.error(` gone from the bestiary: ${removed.join(", ")} — re-record with --update`);
|
|
process.exit(1);
|
|
}
|
|
|
|
/* The claim the bestiary makes, restated as a test rather than as prose: concentrating a
|
|
round's attacks is supposed to help against every pack whose fight is in doubt. If it
|
|
ever stops doing that, the page is wrong and should be rewritten, not re-recorded. */
|
|
/* Two different claims, and only one of them is the page's. The bestiary states the
|
|
above-noise COUNT from this artifact, so it cannot overstate that however the number
|
|
moves — R-264 dropped it from 31 of 31 to 28 of 30 and the page rewrote itself. What
|
|
the page does assert outright is that concentrating fire helps, and that is what fails
|
|
here. A failure means rewrite the advice; it is not a baseline to re-record. */
|
|
if (helped !== checked) {
|
|
console.error(`check-focus: FAILED — focus fire helped only ${helped} of ${checked} packs. `
|
|
+ `docs/BESTIARY.md advises concentrating fire against every pack whose fight is in `
|
|
+ `doubt, and that is no longer true. Rewrite the advice rather than re-recording.`);
|
|
process.exit(1);
|
|
}
|
|
|
|
console.log(`check-focus: OK — focus fire measured against ${checked} packs, helps in all of them`
|
|
+ `, ${aboveNoise} of them by more than their own noise`);
|