Commit Graph
4 Commits
Author SHA1 Message Date
slaguru666andClaude Opus 5 97fe2704e8 Measure creatures over 2000 runs, not 200 (R-253)
R-252 established that a recorded figure carried several points of slack
at 200 runs and that only more runs narrows it. This is that.

Measured by taking every creature under four different base seeds and
looking at how far the four answers sat apart, which is the noise a GM
reading the bestiary is unknowingly trusting:

                   mean spread    worst creature
  200 runs          1.28 points     11.5 points
  2000 runs         0.36 points      2.7 points

The worst case is what mattered. One creature's published wipe rate
could be eleven points from the same creature rolled with different
dice, on a figure printed in docs/BESTIARY.md for somebody to plan an
evening around. Under three now.

It corrected 19 of 47 published figures, mean 0.35 points: the arrears
6% -> 8.9%, the supporter 44% -> 46.8%, the nuckelavee 13.5% -> 16%.
Each was a sampling artefact printed as a property.

The cost is the part worth recording, because it is why this was never
done. I guessed "maybe a minute of build time" when I suggested it,
which was wrong by a factor of thirty and would have been a fair reason
to say no:

  check-lethality alone   0.29s -> 1.88s
  whole nine-guard suite           2.55s

BESTIARY.md regenerates from the baseline, so its run count follows
automatically; its header now also says the seeds are derived per
creature, since "at seed 11" stopped being the whole truth in R-252.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 22:33:40 +01:00
slaguru666andClaude Opus 5 2dbea88d4b Seed the lethality sim per creature (R-252)
Each creature now derives its own seed from its key, FNV-1a mixed into
the base seed, instead of every one of them being measured against seed
11.

This is the second reading of the request and the one I had not tested.
What I checked before was position independence — whether a creature's
numbers depend on its neighbours — and that was already true and still
is: inserting a creature ahead of the barghest changes 0 of 47 existing
rows under the new scheme, same as the old. What was NOT true is that
each creature had its own dice. All forty-seven faced the same two
hundred sequences.

Measured, because the argument for doing it is better than the result:

                                  shared 11   per-creature
  mean wipe rate across bestiary     5.67%        5.67%
  mean agents down                   0.538        0.522
  rows changed                         --        34 of 47

Seed 11 was not biasing the book. There was no systematic luck to
remove, and the aggregate is unmoved to two decimal places. What the
change buys is decorrelation: the error in each row no longer comes from
the same draw as every other row.

The useful number fell out of the comparison rather than the change.
Individual creatures moved up to four points of wipe rate purely from
being handed different dice — the courier 3.5% -> 7.5%, the long walker
66.5% -> 62.5%, quarantine unit 84.5% -> 88.5%. That is the sampling
noise inside any single recorded figure at 200 runs, and it means these
numbers are an exact regression fingerprint and a loose description of a
creature at the same time. Only more runs narrows the second; more seeds
does not. R-252 says so on the page.

seedMode is recorded alongside the numbers and checked, because changing
how a seed is derived moves every row without changing SEED itself. A
baseline from the old scheme is now refused rather than compared against
this one silently and wrongly — verified by running the new code against
the old file before re-recording.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 22:28:45 +01:00
slaguru666andClaude Opus 5 4cfacd47fe check-lethality compares exactly, because it already seeded per creature (R-97)
I asked for this on a false premise of my own: I reported that adding a
creature shifted a shared random stream and perturbed every other
creature's recorded number. That is not true and the code never did it.
measure() builds its own mulberry32 from the seed on every call, and
check-lethality calls it once per creature, so a creature's numbers do
not depend on its neighbours or its position. Measured rather than
argued: inserting a creature ahead of the barghest changes 0 of 47
existing entries. The file's own claim — "the only thing that can move
the number is a change to the rules or to the creature" — was accurate
all along, and my last commit message says otherwise. It is wrong.

The real cause, found by replaying each commit against the baseline as
committed at 4b71859:

  4b71859  baseline recorded            0 of 46 differ
  322389b  bestiary                     0 of 46 differ
  ddc4f99  hit locations reach combat  26 of 46 differ   <-- here
  a90c4f3 .. e5dc9b5                   26 of 46 differ

ddc4f99 routed every ordinary blow through the hit location table. That
is the largest change the combat system has had and it moved 26 of 46
creatures, which is correct and expected. What is not correct is that
nobody noticed for four commits: each creature moved by one or two
points, the guard allowed six, and it reported OK while describing a
game nobody was playing.

So the tolerance goes. It exists for sampling noise and there is no
sampling noise here — same party, same seed, same counts, and two
recordings of unchanged code are byte-identical. Anything that moves is
a real change, which is the entire point of the file. `rounds` is now
compared too; it was recorded and then never read, so a creature could
take a round longer to kill forever without a word.

Because exactness only means something if the measurement is exact, the
guard now proves it instead of assuming it: one creature measured twice
must come back identical, and it says so plainly if a future change
reaches for Math.random.

Negative-tested. A 2% change to locationMaxHp now trips 8 creatures at
+1.5 and +0.5 points of wipe rate — every one of which the old tolerance
would have passed. Breaking determinism is caught and named.

Also fixes update-readme, which advertised 8 guards while the build ran
9: check-anatomy was added without touching the list, which is precisely
what the comment above that list already warned had happened once. The
list is no longer trusted — it is checked against the `check` script in
package.json, and refuses to write a README advertising a different set
than the build runs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 22:22:19 +01:00
slaguru666andClaude Opus 5 4b718595fe check-lethality: guard eight, and the README stops advertising a subset
Seven guards checked that content is WELL FORMED. None checked what it DOES, so a
change to a damage modifier, a hit-point formula or the armour value on a service vest
could double a creature's lethality without touching one line of that creature - and
nothing in the build would notice, because the creature did not change.

Built as a regression test rather than the hand-declared bands the creature-forge plan
described. Bands are the wrong shape here: declaring "keepers: dangerous" across
forty-six creatures means inventing forty-six judgements, and after --spread it is
clear the interesting question is not "is this dangerous" - that has no single answer -
but "is this the same as it was". A recorded baseline answers exactly that, needs no
authoring, and cannot be argued with.

Every creature is measured solo against a FROZEN party of four, two armed postings and
two trades, at a fixed seed, so the only thing that can move a number is a change to
the rules or to the creature. Tolerances are 6 points of wipe rate and 0.35 agents,
which is outside the noise floor of 200 runs - a guard that cries wolf gets deleted.

Verified by tampering: told the baseline the nuckelavee was harmless and the guard
caught it at +16.5 points and +1.67 agents down, exit 1. Runs in half a second, so it
joins the pre-build checks rather than being something to remember to run.

Also: the README's guard block was a hardcoded list of four while the suite was seven.
It had silently stopped mentioning every guard added after it was written, including
check-creatures and check-scenarios. Now all eight, and the prose says eight.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 21:20:09 +01:00