collectSpecs() in tools/all-specs.mjs is the single place that decides
which specs exist. Nine modules were importing NPCS and PREGENS straight
from content.mjs instead, so each carried its own idea of the population
and measured a different subset of the game. check-lethality was the
clearest case: it scored 47 creatures and reported OK for all 184.
check-seam (guard 24, first in the suite) scans every module and fails if
anything but all-specs.mjs names NPCS or PREGENS in a content.mjs import.
The nine violators are re-pointed at the seam.
Re-recording check-focus's baseline against the full 184 raised it from
47 packs to 138 and surfaced one creature focus fire does not help: the
dun cow. Rather than re-record that away, the guard now requires the
advice section of BESTIARY.md to name every such exception, and
bestiary.mjs generates the sentence.
The first version of that check asked whether the name appeared anywhere
in BESTIARY.md, which every creature's own heading satisfies — deleting
the exception sentence still passed. It reads only the
"Shooting at something that moves" section now, and the mutation test
fails as it should.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
R-280 widened admission from a [15,85] band to "the spread rate stands clear of
both ends by more than its own noise", which admits fights that are nearly
settled -- the_choir at 99.3%, the_stanchion at 6% -- as long as their noise is
smaller still. The page went on saying "whose outcome was ever in doubt", which
was a fair description of the band and is a loose one of the rule.
It now says "whose odds leave room for a difference to show", and the provenance
line prints the admission rule itself, read from the artifact rather than
paraphrased, so the page cannot drift from the guard again. The rule is phrased
as a clause in check-focus for that reason.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
R-280 left the claim satisfiable by a gain of 0.2. The artifact now records how
many measurable packs clear their own noise -- reliable: {aboveNoise: 36, of: 38}
-- and the check refuses if that share falls. It may rise freely.
A ratchet rather than a threshold: any threshold here would be a number I chose,
and choosing one just under the current value is what produced MEASURABLE =
[15,85]. A share rather than a count, so widening admission cannot pay it off.
Proved three ways in worktrees: making focus fire actively bad fires the drift
check first, which is correct; making it unreliable and re-recording fires the
older claim at 37 of 38; and claiming a better past, 38 of 38, is refused by the
ratchet itself. The ratchet bites exactly where the old claim does not -- between
"still helps everywhere" and "helps as reliably as it did", which is where a slow
degradation lives.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
MEASURABLE = [15,85] proxied for "can this fight move at all". The direct form,
in the units the guard already uses: the spread rate must stand clear of both
ends by more than its own noise. Deliberately blind to the gain -- admitting the
sizes where focus fire clears its noise would make the guard's claim true by
construction. Size is still picked on nearest-an-even-fight.
31 measurable became 38. the_arrears returns at 6.3, and switchboard arrives at
8.1 -- the second largest gain in the artifact, thrown away for being one point
past a round number. Four of the eight carry effects larger than most rows the
band already admitted.
Two of them do not clear their own noise: the_choir has 0.7 points of headroom
and used 0.2, the_stanchion has six and used 0.2. Above-noise falls 31/31 to
36/38 and the bestiary prints 36. That is two measurements reporting no
detectable effect, which the band suppressed by refusing to take them.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Asked to raise the scan until the redcap's pack size stopped flipping. Measured
the threshold -- unstable at 3000 and 4000, stable across twenty seeds at 6000 --
raised it, and the re-record took fifteen seconds, which was impossible.
winRate takes three parameters and pickSize passed SCAN_RUNS as a fourth.
JavaScript discards it, so every scan has always run at RUNS and SCAN_RUNS has
never been read by anything. The fix I was asked to make was inert in the same
way as the thing it was fixing.
winRate takes runs now. The redcap is still n=3 with gain 6.3, arrived at stably
rather than luckily; the_arrears drops out of the measurable band at an honest
scan, 31 packs to 30; the_committee moves 2 to 6 and stays pinned. Claim check
still passes, bestiary regenerated.
The only signal was a number being too small. A fifteen-second re-record is good
news, and good news is what nobody investigates.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
41 statblocks carry a POWER in their tactics and simulate.mjs read none of them,
so a redcap that ignores the cumulative defence penalty has been measured as a
creature that tires -- in check-lethality, in check-focus, and in every figure
published about it. The defect was not that the powers were unimplemented, it was
that nothing said they were not.
powers.mjs classifies all 41: 2 wired, 14 notSimulable with a stated reason, 25
not fight rules. check-powers refuses an unclassified POWER and refuses a
notSimulable without a reason -- and it does not test that the harness imports a
power, it fights the creature with and without and requires the two to disagree.
Moved: the courier 9.4% to 1.0% wiped (it attacks at half while carrying), the
redcap 0.7% to 0.9% (small, because these fights rarely spend a second defence).
ARGENT AND GULES was wired and then un-wired: it tripled the supporter's wipe rate
to 75.2% because the harness has no ground and applied the borough-ground condition
unconditionally. Same reason THE PULL is not wired. I had wired one and refused the
other on identical facts.
Lethality and focus re-recorded, bestiary regenerated.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
tools/playthrough.mjs exists to show a reader why a measured number is what it is. I cited
it to another session — read seed 2, see the wipe the 45% row is made of — and the
reproduction failed in front of them. They reported different outcomes at both seeds and
guessed the cause correctly from outside: the two tools were not drawing from the same
stream.
They were not. simulate.mjs used a private mulberry32 makeRng; playthrough.mjs had its own
LCG written to look like it. Both deterministic, both reproducible alone, and "seed 2"
named a different fight in each — which breaks the only thing the tool is for. Its own
comment claimed a seed here names the same fight there. check-focus carried a third copy of
that LCG, so the two guards described the same game with different dice.
makeRng is exported and both files use it. A playthrough seed is now exactly the first
fight of simulate.mjs --seed <n>: --runs 1 --seed 2 and the playthrough give 9 rounds, 4 of
4 down, 1 dead, both.
The other half was my citation rather than the code: the command I sent omitted --mode, so
it plays both targeting arms and prints two fights. They read the last line, I quoted the
first. The summary line now names the arm.
check-focus re-recorded under the shared stream; figures move a point or two. What it buys
is that the bimodality analysis now reproduces the published means exactly — 2.31, 3.03,
0.40, 0.42 against the four rows CLEAN GROUND publishes. Under the old LCG it agreed to
within a decimal, which looked like corroboration and was two experiments landing near each
other.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Reading a narrated supporter fight showed it being hit on armR and head. It has wings; it
has never had arms. Three defects, all in the pipeline every published number comes from.
1. It located hits by species, not body plan — simulate.mjs read defender.species where
the game writes spec.bodyPlan ?? spec.species ?? "baseline" into speciesProfile
(build-packs.mjs:294). The barghest, kelpie and church grim were fought on two legs
with arms, and the supporter with no wings, so the fight its tactics call the one the
agents can win could not occur in a measured fight.
2. Armour was one scalar for the whole creature, and the harness's locations had no
armour field. Bare wings, the grounded halving and the vital exemption — R-254 and
R-255 — were invisible to every number. Worn armour was summed the same way, which put
a stab vest on a cleaner's head.
3. Nothing was ever grounded: a wing could be ruined and the creature kept flying.
Fixed in the game's order — locate, then apply what that location carries, every term
imported from rules.mjs. A ruined wing calls groundedPlanFor and the wounds carry across
by severity through remapLocationDamage, the function _preUpdate uses.
A bug of mine no guard would have caught: remapLocationDamage returns { damage, moved,
rescaled } and my first draft passed the whole object where a damage map was expected, so
every wound a creature carried was forgiven the moment it came down. check-lethality would
have passed it — fewer wounds means a longer fight, which reads as a number moving, and
this commit moves numbers. Found by probing a landing by hand.
20 of 47 creatures moved, 18 deadlier and 2 less. The supporter goes 69.3% -> 25.1% wiped,
3.38 -> 2.16 down, second deadliest to fourth: it was being measured as a 30-hit-point
creature in uniform armour 9 that could not be grounded. The small rises elsewhere are the
party's armour no longer covering locations it never protected.
check-focus then caught the page overclaiming, which is what it is for: focus fire still
helps in all 30 but only 28 clear their own noise where 31 of 31 did. The guard was
asserting more than the page needs — the bestiary prints that count from the artifact and
cannot overstate it — so it now checks only what the page asserts outright, and the page
rewrote itself to "in 28 of them".
Both baselines re-recorded. Minor version, not a patch: the published numbers changed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
R-258 put a tactic on a GM-facing page from an unguarded harness, R-259 found it worth
nothing, R-260 found the replacement right for the wrong reason. Three corrections, each
landing as prose with nothing checking it — the arrangement that let ddc4f99 describe a
game nobody was playing.
tools/focus-baseline.json records, per creature, the pack size and the spread-fire and
focus-fire win rates. check-focus re-measures all of it on every build and compares
exactly, with the party, seeds, run count and seeding scheme recorded alongside so numbers
taken under different conditions are refused rather than compared. The bestiary READS the
artifact instead of restating it, and check-bestiary refuses a page that has fallen behind
it. Page, guard and simulator cannot disagree.
The pack size is recorded rather than re-chosen: a fight at 0% or 100% cannot show an
effect, and a guard that picked again each run would let a changed creature move quietly
to a different question and pass. 31 of 47 creatures land in the measurable band; the
other 16 are recorded as pinned, with the rate that pinned them.
It guards the claim as well as the numbers. The page says focus fire helps in every fight
in doubt; check-focus fails if any row's gain reaches zero or stops clearing its own
noise. That failure means rewrite the page, not re-record the baseline.
The run count was chosen by evidence. 1000 x 3 seeds costs 8.5s and takes the suite from
2.5s to 13.7s. I tried 500 to halve it and the claim-check failed — at 500 runs one row
no longer clears its noise, so "without exception" is not supported by that much
sampling. Recording twice at 1000 gives byte-identical files.
Negative-tested four ways, all firing: a creature quietly made nimbler (redcap dodge
75 -> 85, spread 40.6% -> 27.6%), a baseline under different seeds, a hand-edited page,
and the claim failing at 500 runs.
Guard eleven (check-rollable) arrived from another session mid-build; this is twelve.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>