Files
RingBRP/tools/check-focus.mjs
T
slaguru666andClaude Opus 5 fa484071cd R-281: a ratchet, because "helps in all of them" survives the advice decaying
R-280 left the claim satisfiable by a gain of 0.2. The artifact now records how
many measurable packs clear their own noise -- reliable: {aboveNoise: 36, of: 38}
-- and the check refuses if that share falls. It may rise freely.

A ratchet rather than a threshold: any threshold here would be a number I chose,
and choosing one just under the current value is what produced MEASURABLE =
[15,85]. A share rather than a count, so widening admission cannot pay it off.

Proved three ways in worktrees: making focus fire actively bad fires the drift
check first, which is correct; making it unreliable and re-recording fires the
older claim at 37 of 38; and claiming a better past, 38 of 38, is refused by the
ratchet itself. The ratchet bites exactly where the old claim does not -- between
"still helps everywhere" and "helps as reliably as it did", which is where a slow
degradation lives.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 14:34:10 +01:00

259 lines
13 KiB
JavaScript

/**
* check-focus — the bestiary's advice must keep being true.
*
* R-258 put a tactic on a page a GM reads, measured in a harness nobody guards. R-259
* found it worth nothing when run through the simulator. R-260 found the replacement
* advice right for the wrong reason. Each time the correction was an assertion in a
* document, which is the same arrangement that let ddc4f99 describe a game nobody was
* playing: a number written down with nothing checking it.
*
* So the claim is recorded as an artifact and re-measured on every build, exactly the
* way check-lethality does it. The bestiary reads the artifact rather than restating it,
* so the page, the guard and the simulator cannot disagree.
*
* node tools/check-focus.mjs compare against the committed baseline
* node tools/check-focus.mjs --update re-record it (after a deliberate change)
* node tools/check-focus.mjs --verbose print every creature, not just drift
*
* WHAT IS MEASURED. For each creature, one pack size, and two arms: every agent picking
* a target at random, versus every agent putting their attacks into the same creature.
* The difference is what focus fire is worth against that creature.
*
* WHY A RECORDED PACK SIZE. A fight the party wins 100% of the time, or loses 100% of
* the time, cannot show an effect of any size — three of the first four packs measured
* in R-259 were pinned like that and said nothing. --update picks, for each creature,
* the size whose spread-fire rate sits nearest an even fight, and RECORDS it. The guard
* then re-measures at that size rather than choosing again, so a creature that changes
* shows up as drift instead of quietly moving to a different question.
*/
import { readFile, writeFile } from "node:fs/promises";
import { existsSync } from "node:fs";
import path from "node:path";
import { NPCS } from "./content.mjs";
import { ROSTER } from "./roster.mjs";
import { runFight, seedFor, makeRng } from "./simulate.mjs";
const ROOT = path.resolve(path.dirname(new URL(import.meta.url).pathname.replace(/^\/([A-Za-z]:)/, "$1")), "..");
const BASELINE = path.join(ROOT, "tools", "focus-baseline.json");
const argv = process.argv.slice(2);
const UPDATE = argv.includes("--update");
const VERBOSE = argv.includes("--verbose");
/* The same frozen party as check-lethality, for the same reason and so the two guards
describe the same game. */
const PARTY_KEYS = ["pc_holloway", "pc_okonkwo", "pc_nkemdirim", "pc_ferriby"];
/* 1000 rather than 500 on evidence, not preference: at 500 runs one row's gain no longer
clears its own seed-to-seed noise and the claim-check below fails, because the page's
"without exception" is not supported by that much sampling. The build pays 8s for a
sentence it is entitled to print. */
const RUNS = 1000;
const SEEDS = [11, 4242, 90210];
const SEED_MODE = "per-creature";
const SIZES = [2, 3, 4, 5, 6];
/* R-278. 400 was enough to find a size and not enough to find the SAME size twice. For
the redcap the criterion is near-tied — |win-50| is about 25 at two of them and 23 at
three — so a 400-run scan picked 2 instead of 3 roughly one seed in eight, and the
recorded gain moved 6.3 to 4.9 with it. Measured rather than guessed: the pick is still
unstable at 3000 and at 4000, and stable across twenty seeds at 6000. This is --update
only, so the build pays nothing; re-recording pays about half a minute. */
const SCAN_RUNS = 6000;
/* R-280. This was a band, [15, 85], and it was a PROXY for the question that matters:
can this fight move at all? A rate pinned at 0 or 100 cannot show an effect of any size,
so measuring one is pointless. The band answered that with two round numbers, and
the_arrears — 14.2% at the only size worth fighting — was excluded by three tenths of a
point while carrying a gain of 6.2 at 2.4x its own noise. A proxy that disagrees with
the thing it proxies for is worth replacing rather than tuning.
So the test is now direct and expressed in the same units as everything else here: a
fight is measurable if its spread-fire rate stands clear of BOTH ends by more than its
own seed-to-seed noise. That is the actual condition for an effect to be visible.
IT DELIBERATELY DOES NOT LOOK AT THE GAIN. Selecting the sizes where focus fire happens
to clear its noise would make this guard's own claim — that focus fire helps in every
measurable pack — true by construction, which is the worst thing a guard can be. Size is
still chosen on nearest-an-even-fight, and admission is still decided without reference
to the effect being measured. */
const measurable = row => row.spread - row.noise > 0 && row.spread + row.noise < 100;
const party = PARTY_KEYS.map(k => {
const found = ROSTER.find(r => r.key === k);
if (!found) {
console.error(`check-focus: the frozen party names "${k}", which is not on the roster.`);
process.exit(1);
}
return found;
});
/* One generator per arm, derived from the creature's own key, so inserting a creature
cannot perturb its neighbours — R-252, learned the hard way. */
/* R-265: the harness's own generator, not a second one of the same shape. This file used
to roll its own LCG, so the two guards described the same game with different dice and
no figure here could be reproduced from simulate.mjs. */
const rngFor = makeRng;
/* R-278: `runs` was not a parameter. pickSize has always passed SCAN_RUNS as a fourth
argument to a three-argument function, so JavaScript discarded it and every scan ran at
RUNS instead — the constant was declared, documented, passed and never read, which is
the defect this project keeps finding in itself, sitting inside a guard. */
const winRate = (foes, targets, seed, runs = RUNS) => {
let won = 0;
for (let i = 0; i < runs; i++) {
if (runFight(rngFor(seed + i * 2654435761), party, foes, { partyTargets: targets }).won) won++;
}
return (100 * won) / runs;
};
const mean = a => a.reduce((x, y) => x + y, 0) / a.length;
const measureAt = (spec, n) => {
const foes = Array(n).fill(spec);
const seeds = SEEDS.map(s => seedFor(s, spec.key));
const spread = seeds.map(s => winRate(foes, "random", s));
const focus = seeds.map(s => winRate(foes, "focus", s));
return {
n,
spread: Number(mean(spread).toFixed(1)),
focus: Number(mean(focus).toFixed(1)),
gain: Number((mean(focus) - mean(spread)).toFixed(1)),
noise: Number((Math.max(...spread) - Math.min(...spread)).toFixed(1))
};
};
/* --update only: choose the pack size that makes the question answerable. */
const pickSize = spec => {
let best = null;
for (const n of SIZES) {
const w = winRate(Array(n).fill(spec), "random", seedFor(SEEDS[0], spec.key), SCAN_RUNS);
if (!best || Math.abs(w - 50) < Math.abs(best.w - 50)) best = { n, w };
if (w < 5) break; // already hopeless; more of them is worse
}
return best;
};
if (UPDATE) {
const creatures = {};
let measurableCount = 0, aboveNoiseCount = 0;
for (const spec of NPCS) {
const pick = pickSize(spec);
/* Measured first, then admitted or not: the noise this test needs is a product of the
measurement, so unlike the band it cannot be decided from the scan alone. */
const row = measureAt(spec, pick.n);
if (!measurable(row)) {
creatures[spec.key] = { n: pick.n, pinned: row.spread };
continue;
}
creatures[spec.key] = row;
measurableCount++;
if (row.gain > row.noise) aboveNoiseCount++;
}
await writeFile(BASELINE, JSON.stringify({
note: "Generated by tools/check-focus.mjs --update. Do not edit by hand.",
party: PARTY_KEYS, runs: RUNS, seeds: SEEDS, seedMode: SEED_MODE,
admission: "spread clear of 0 and 100 by more than its own noise",
/* R-281. The ratchet the claim check holds to: how many of the measurable packs
showed a gain clearing their own noise when this was recorded. Focus fire may
get more reliable and may not quietly get less. */
reliable: { aboveNoise: aboveNoiseCount, of: measurableCount },
creatures
}, null, 2) + "\n", "utf8");
console.log(`check-focus: baseline recorded — ${measurableCount} of ${NPCS.length} creatures measurable, `
+ `${RUNS} runs x ${SEEDS.length} seeds, seed ${SEEDS[0]} ${SEED_MODE}`);
process.exit(0);
}
if (!existsSync(BASELINE)) {
console.error("check-focus: no baseline. Run with --update to record one.");
process.exit(1);
}
const base = JSON.parse(await readFile(BASELINE, "utf8"));
if (base.runs !== RUNS || base.seeds?.join() !== SEEDS.join()
|| base.seedMode !== SEED_MODE || base.party.join() !== PARTY_KEYS.join()) {
console.error("check-focus: FAILED — the baseline was recorded under different conditions "
+ `(party ${base.party.join(",")}, ${base.runs} runs, seeds ${base.seeds?.join(",")} `
+ `${base.seedMode ?? "shared"}). Re-record it with --update.`);
process.exit(1);
}
/* Exact comparison is only honest if the measurement is exact. Prove it, don't assume. */
{
const probe = NPCS.find(s => base.creatures[s.key]?.gain !== undefined);
if (probe) {
const a = JSON.stringify(measureAt(probe, base.creatures[probe.key].n));
const b = JSON.stringify(measureAt(probe, base.creatures[probe.key].n));
if (a !== b) {
console.error("check-focus: FAILED — the simulation is not deterministic, so an exact "
+ `baseline cannot mean anything. ${probe.key} measured twice gave ${a} and ${b}.`);
process.exit(1);
}
}
}
const problems = [], added = [], removed = [];
let checked = 0, helped = 0, aboveNoise = 0;
for (const spec of NPCS) {
const was = base.creatures[spec.key];
if (!was) { added.push(spec.key); continue; }
if (was.pinned !== undefined) continue; // recorded as unmeasurable; nothing to compare
const now = measureAt(spec, was.n);
checked++;
if (now.gain > 0) helped++;
if (now.gain > now.noise) aboveNoise++;
const same = now.spread === was.spread && now.focus === was.focus && now.gain === was.gain;
if (!same) {
problems.push(` ${spec.key}: ${was.n} of them — spread ${was.spread}% -> ${now.spread}%, `
+ `focus ${was.focus}% -> ${now.focus}%, focus fire worth ${was.gain} -> ${now.gain}`);
} else if (VERBOSE) {
console.log(` ${spec.key.padEnd(24)} ${was.n}x spread ${was.spread}% focus ${was.focus}% +${was.gain}`);
}
}
for (const key of Object.keys(base.creatures)) if (!NPCS.some(s => s.key === key)) removed.push(key);
if (problems.length || added.length || removed.length) {
console.error(`check-focus: FAILED — ${problems.length} creature(s) answer differently than recorded`);
problems.forEach(p => console.error(p));
if (added.length) console.error(` not in the baseline: ${added.join(", ")} — re-record with --update`);
if (removed.length) console.error(` gone from the bestiary: ${removed.join(", ")} — re-record with --update`);
process.exit(1);
}
/* The claim the bestiary makes, restated as a test rather than as prose: concentrating a
round's attacks is supposed to help against every pack whose fight is in doubt. If it
ever stops doing that, the page is wrong and should be rewritten, not re-recorded. */
/* Two different claims, and only one of them is the page's. The bestiary states the
above-noise COUNT from this artifact, so it cannot overstate that however the number
moves — R-264 dropped it from 31 of 31 to 28 of 30 and the page rewrote itself. What
the page does assert outright is that concentrating fire helps, and that is what fails
here. A failure means rewrite the advice; it is not a baseline to re-record. */
if (helped !== checked) {
console.error(`check-focus: FAILED — focus fire helped only ${helped} of ${checked} packs. `
+ `docs/BESTIARY.md advises concentrating fire against every pack whose fight is in `
+ `doubt, and that is no longer true. Rewrite the advice rather than re-recording.`);
process.exit(1);
}
/* R-281. "Helps in all of them" is satisfied by a gain of 0.2, which is true and thin: it
survives the advice becoming useless everywhere it is not already decisive. So the
RELIABILITY is recorded and ratcheted — the share of measurable packs where the gain
clears its own noise may rise and may not quietly fall.
A ratchet rather than a threshold, because any threshold here would be a number I chose.
This one is the measurement itself, and it fails only on a real change in the game. Note
the direction: this cannot be satisfied by admitting more packs, since it is a share. */
const wasReliable = base.reliable;
if (wasReliable && wasReliable.of > 0) {
const wasShare = wasReliable.aboveNoise / wasReliable.of;
const nowShare = aboveNoise / Math.max(checked, 1);
if (nowShare < wasShare) {
console.error(`check-focus: FAILED — focus fire clears its own noise in ${aboveNoise} of `
+ `${checked} packs, down from ${wasReliable.aboveNoise} of ${wasReliable.of}. The advice `
+ `still helps everywhere, and it has become less reliably worth taking than when it was `
+ `recorded. That is a change in the game, not a baseline to refresh: find what moved `
+ `before re-recording, and rewrite the page if the advice is now weaker than it reads.`);
process.exit(1);
}
}
console.log(`check-focus: OK — focus fire measured against ${checked} packs, helps in all of them`
+ `, ${aboveNoise} of them by more than their own noise`);