32eaee1eb047c7b77eef654c19d6a234c1f5a2c5
4
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
97fe2704e8 |
Measure creatures over 2000 runs, not 200 (R-253)
R-252 established that a recorded figure carried several points of slack
at 200 runs and that only more runs narrows it. This is that.
Measured by taking every creature under four different base seeds and
looking at how far the four answers sat apart, which is the noise a GM
reading the bestiary is unknowingly trusting:
mean spread worst creature
200 runs 1.28 points 11.5 points
2000 runs 0.36 points 2.7 points
The worst case is what mattered. One creature's published wipe rate
could be eleven points from the same creature rolled with different
dice, on a figure printed in docs/BESTIARY.md for somebody to plan an
evening around. Under three now.
It corrected 19 of 47 published figures, mean 0.35 points: the arrears
6% -> 8.9%, the supporter 44% -> 46.8%, the nuckelavee 13.5% -> 16%.
Each was a sampling artefact printed as a property.
The cost is the part worth recording, because it is why this was never
done. I guessed "maybe a minute of build time" when I suggested it,
which was wrong by a factor of thirty and would have been a fair reason
to say no:
check-lethality alone 0.29s -> 1.88s
whole nine-guard suite 2.55s
BESTIARY.md regenerates from the baseline, so its run count follows
automatically; its header now also says the seeds are derived per
creature, since "at seed 11" stopped being the whole truth in R-252.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
2dbea88d4b |
Seed the lethality sim per creature (R-252)
Each creature now derives its own seed from its key, FNV-1a mixed into
the base seed, instead of every one of them being measured against seed
11.
This is the second reading of the request and the one I had not tested.
What I checked before was position independence — whether a creature's
numbers depend on its neighbours — and that was already true and still
is: inserting a creature ahead of the barghest changes 0 of 47 existing
rows under the new scheme, same as the old. What was NOT true is that
each creature had its own dice. All forty-seven faced the same two
hundred sequences.
Measured, because the argument for doing it is better than the result:
shared 11 per-creature
mean wipe rate across bestiary 5.67% 5.67%
mean agents down 0.538 0.522
rows changed -- 34 of 47
Seed 11 was not biasing the book. There was no systematic luck to
remove, and the aggregate is unmoved to two decimal places. What the
change buys is decorrelation: the error in each row no longer comes from
the same draw as every other row.
The useful number fell out of the comparison rather than the change.
Individual creatures moved up to four points of wipe rate purely from
being handed different dice — the courier 3.5% -> 7.5%, the long walker
66.5% -> 62.5%, quarantine unit 84.5% -> 88.5%. That is the sampling
noise inside any single recorded figure at 200 runs, and it means these
numbers are an exact regression fingerprint and a loose description of a
creature at the same time. Only more runs narrows the second; more seeds
does not. R-252 says so on the page.
seedMode is recorded alongside the numbers and checked, because changing
how a seed is derived moves every row without changing SEED itself. A
baseline from the old scheme is now refused rather than compared against
this one silently and wrongly — verified by running the new code against
the old file before re-recording.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
4cfacd47fe |
check-lethality compares exactly, because it already seeded per creature (R-97)
I asked for this on a false premise of my own: I reported that adding a creature shifted a shared random stream and perturbed every other creature's recorded number. That is not true and the code never did it. measure() builds its own mulberry32 from the seed on every call, and check-lethality calls it once per creature, so a creature's numbers do not depend on its neighbours or its position. Measured rather than argued: inserting a creature ahead of the barghest changes 0 of 47 existing entries. The file's own claim — "the only thing that can move the number is a change to the rules or to the creature" — was accurate all along, and my last commit message says otherwise. It is wrong. The real cause, found by replaying each commit against the baseline as committed at |
||
|
|
4b718595fe |
check-lethality: guard eight, and the README stops advertising a subset
Seven guards checked that content is WELL FORMED. None checked what it DOES, so a change to a damage modifier, a hit-point formula or the armour value on a service vest could double a creature's lethality without touching one line of that creature - and nothing in the build would notice, because the creature did not change. Built as a regression test rather than the hand-declared bands the creature-forge plan described. Bands are the wrong shape here: declaring "keepers: dangerous" across forty-six creatures means inventing forty-six judgements, and after --spread it is clear the interesting question is not "is this dangerous" - that has no single answer - but "is this the same as it was". A recorded baseline answers exactly that, needs no authoring, and cannot be argued with. Every creature is measured solo against a FROZEN party of four, two armed postings and two trades, at a fixed seed, so the only thing that can move a number is a change to the rules or to the creature. Tolerances are 6 points of wipe rate and 0.35 agents, which is outside the noise floor of 200 runs - a guard that cries wolf gets deleted. Verified by tampering: told the baseline the nuckelavee was harmless and the guard caught it at +16.5 points and +1.67 agents down, exit 1. Runs in half a second, so it joins the pre-build checks rather than being something to remember to run. Also: the README's guard block was a hardcoded list of four while the suite was seven. It had silently stopped mentioning every guard added after it was written, including check-creatures and check-scenarios. Now all eight, and the prose says eight. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |