R-304's rate was an artefact; the proposed correction, retreating to the seven
both vocabularies agree on, overshoots. Opening every disputed column instead
of comparing totals: 'Then wipes' is genuine and check-unmarked misses it
because wiped? does not match wipes; '1 in 100 runs past' is genuine and
check-figures misses it because its vocabulary has rounds? and not runs; only
THROUGH_TRAIN's 'Measured over 300 runs' is a false positive, and there the
cells are whole sentences and the bold wraps a clause.
So nine, and the shape matters more than the number: each vocabulary has a real
gap the other covers, which argues for the union of both word lists.
Refuses one recommendation. Dropping runs? from the header vocabulary would
also drop the p99 column, whose cells cite fight-tail cut.p99 and column.p99
and which R-269 added because the median is not a plan. That is the R-299
failure again — a guard narrowing itself until the figure it exists to watch
falls outside. The real exclusion wanted is 'a column whose cells are sentences
rather than values', which is testable; subtracting a word cannot tell the two
cases apart.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Recomputed the peer's bolding scan rather than quoting it. Eight measurement
columns are inconsistently bolded, and that count is robust; the RATE is not —
44% under check-figures' vocabulary, 57% under check-unmarked's, because the
two do not define 'measurement column' the same way. Anybody quoting a
percentage has to say whose vocabulary produced it.
The cause is not carelessness but a second convention: in prose the corpus
bolds what it publishes, in a table it bolds the rows it wants read. The two
bolded rows of the mixed-force table are exactly the two the prose below it
singles out.
That divides the merge on evidence — prose takes bold-as-published, tables take
column-inherits-header with bold ignored entirely as a signal — and explains
the asymmetry R-303 found from both sides as one cause with two symptoms.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
R-302 showed check-unmarked catching what check-figures misses. I had taken
that, plus a preference for structure over vocabulary, as grounds for folding
check-figures in as a secondary pass. The converse test contradicts it: five
unbolded cells in a measurement column, markers stripped, fail check-figures
and pass check-unmarked, whose strict class requires bold by design.
So the relationship is symmetric. The shapes are the weak half and should fold
in behind the structural classes; the column-header rule is not a shape and
belongs beside bold-as-published, not under it.
Also records what a merged guard must face rather than inherit: this corpus
bolds inconsistently inside tables, which is invisible to a reader and
load-bearing for a class keyed on boldness.
Both runs in throwaway worktrees, never the shared tree. No merge performed;
the decision sits with the humans.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
R-300 blamed `git add -A` for eead663 carrying my uncommitted edits.
check-figures' author checked and it was not that: they staged two exact
paths. The tree agrees — eead663 holds CLEAN_GROUND.md and CLEAN_GROUND_ART.md
only, while my modified step5-split.mjs and untracked check-unmarked.mjs are
absent, both of which `git add -A` would have taken.
`git add <file>` takes the whole file including another session's edits to it,
so a pathspec stops you sweeping files you did not touch and does nothing about
the one you did. Their pre-commit `npm run check` passed because my uncommitted
module was in the shared tree: the commit was broken, the working copy was not.
I asserted a mechanism I never looked at, in an entry whose technical finding
was correct. The correction came from the session it accused.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A cell inherits the measurement status of its column: a different mechanism
from the shapes rather than more of them, because the document says what the
column is and the vocabulary has to guess how somebody will phrase a figure.
Records the two bugs in the repair, and why the first matters more than the
blind spot it fixed: scanning raw cell text reported ~130 correctly-cited rows
as unheld, and a guard that calls the good material broken teaches its reader
to stop reading the output.
Leaves the merge question open on purpose. Two readers that disagree are how
three of this session's defects were found.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
eead663 committed a document citing `step5 unsettledFactor` while the module
that would export it was still uncommitted in another session's working tree,
so `check-cited` fails on a fresh checkout of HEAD. step5-split.mjs is included
here to repair that; the key derives Act Four's 1.29 from the raw proportions,
which is the figure R-294 exists about.
check-figures (R-299) reports OK on the "24" it was written to catch. Its own
entry names why: the eight shapes were read off figures that are already cited,
so the vocabulary is learned from the marked figures and cannot contain the
phrasing of the one nobody marked. Line 1272 is a table cell whose measurement
status lives in the header two rows above it, and figuresIn reads one line.
check-unmarked reads structure instead of vocabulary: a bold percentage or
decimal in an opted-in document, and a bold number in a table whose header row
names a measurement. Twelve unmarked figures in CLEAN GROUND; eleven correct
and unheld, one the stale 24 that line 1142 had been citing correctly as 25 for
130 lines. Waivers carry a reason, an empty one is fatal, and every waiver
prints on a green build.
Two guards now cover one question, which is one too many. The right end state
is the structural classes folded in beside check-figures' SHAPES, keeping its
opt-in rule and ratchet. That fold is offered to its author rather than taken.
Twelve string tests on unmarkedIn (88 -> 100), both halves mutation-checked.
Twenty-two guards green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The Casting section said "Portraits already exist in art/portraits for all
six. No new art needed." The first sentence is true and verified on disk. The
second was only ever true of the cast, and sitting at the foot of Casting it
read as a statement about the scenario. It is not: the six named NPCs have no
portraits, there is no map, and there are no scene plates at all — while
LAST_ADMISSION has 8, OPEN_DAY 10 and THROUGH_TRAIN 13. Not one cg_* asset
exists anywhere in art/.
Nothing in "Still to do" mentioned it either; that list held only a human run
and a slot, both resolved.
So: the claim is scoped to the cast and points at the gap, the gap is on the
outstanding list where a reader looks, and CLEAN_GROUND_ART.md now carries the
prompts — 6 NPC portraits and 11 scene plates, 17 ids, none colliding with a
file already on disk.
Art direction follows the house pattern (base style plus one scenario line).
THROUGH TRAIN draws the present day in an 1881 hand; CLEAN GROUND draws
everything as a sheet from the 1962 file — including the far side, because the
valley is not a place, it is a filed document nobody cancelled. Same paper,
same margin, same grain on both sides of the seam.
The notes carry the rules that matter: the six at the back are never singled
out in the column plate (the text says do not linger on them, and a plate that
picks them out gives away Act Three in Act Two), nobody in the column is lit as
a victim, nothing glows, the understudy is never drawn as a monster, and the
grey is not snow.
Creatures stay in BESTIARY_ART.md; the handout pack is done and its almanac
pages must not be illustrated, since their trick is that the two sheets are
identical and check-handouts holds that.
21 guards green; the sheet classifies as a record, not a playable scenario.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Records the blind spot (check-cited can only resolve markers that exist), why
the net is narrow (an earlier draft found 282 candidates, nearly all prose),
the opt-in rule and ratchet, and the two things it found before being wired in
— 74.9% printed bare twice, and 'the longest fight is 120 rounds' where 120 is
the exact field check-cited refuses to let anybody cite.
Two rules each right, and together a hole: refusing the citation while the
document printed the number left the least stable figure in the suite as the
only one nothing held.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
check-cited resolves every citation marker against the artifact it names, and
has one structural blind spot: it can only resolve the markers that exist. A
measured figure written into prose with no marker beside it is not a failed
citation, it is not a citation at all, and nothing looks at it again.
Desk pass 11 proved it. "A median 24 rounds" sat in the Pacing Note and "24 if
the column joins" in EXPOSURE, against a baseline holding 21 at two of the
column joining and 25 at three, through every pass that checked citations.
This reads the figures instead of the markers. Narrow by design: an earlier
draft matched any number within 45 characters of a measurement word and found
282 candidates, nearly all prose ("down 140 steps", "an engineer on his
rounds", "01:06"). The shapes here are the phrasings the documents actually use
when quoting the simulator, each read off a figure that is cited somewhere.
Opt-in rule: a document that uses citations must mark every measurement figure.
One that cites nothing is held by a ratchet instead — turning three unguarded
scenarios red is how a guard gets switched off on the day it is written — and
joins the strict regime the moment it gains its first marker.
It found three things in CLEAN GROUND before it was wired in:
- 74.9% printed bare twice while cited correctly four times, and that is the
figure that was published at 58.7% until the truncated sweep was found.
- "the longest fight is 120 rounds" — 120 is exactly packs cut.hollow6.longest,
the field check-cited REFUSES to let anybody cite because a sample maximum
moves by a third on a re-seed. Refusing the citation while printing the number
left the least stable figure in the suite as the only one nothing held. The
sentence now leans on the guarantee that is actually strong: check-packs
refuses to record a sweep in which anything reached the cap.
Proved by breaking it, four ways, with byte-identical restores: a stripped
marker fails, a planted "median 24 rounds" fails at its own line, a new bare
figure in an uncited document trips the ratchet, and a threshold column header
("Past 15 rounds") does not fire — that last one was a real false positive in
the first draft, which tested the matched text rather than its context.
v0.23. 21 guards.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Six findings applied. The one worth reading: two blocks two hundred lines
apart specified two different fights for the same trigger, both mine, and the
mild one predates the measurement.
Records the stale 'median 24 rounds' that no guard could see because it
carried no citation marker, and the fifth instance of a reader unable to see
its own format.
And the cross-session half: I re-derived b5's 4.90% exactly and shipped it
with their false premise attached. The arithmetic was never the part that
could be wrong. Sharper than that — the scenario has no Transposition beat at
all, so it was a right number with no question attached and the premise
arrived to give it one. The house mission structure does make the return a
Transposition roll, which is why the instinct was persuasive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Fix 5 as written said a fight costs the northern seam, the depot and telling
Ivy. The Pacing Note names telling Ivy among the three things never to cut,
so the applied version says two go and the third is compressed or moves to
the stump. Written from memory of the act rather than from the Pacing Note.
And applying it found a stale figure no guard could see: 'median 24 rounds'
in two places, where the baseline says 21 at two of the column joining and 25
at three. It survived because it carried no citation marker, and check-cited
can only resolve markers that exist. Fifth reader caught unable to see its
own format.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
1. EXPOSURE's "what to do instead of a fight" said "use three". That predates
the measurement and described a fight the party never has: three of the
column alone is five rounds and 0.1% wiped, where the trigger's real force
is a median 21 rounds and 0.97 deaths. It now points at the mixed-force
table instead of carrying its own number, and keeps the good half of the
sentence.
2. New "The valley at dawn" block: the count moves. H03a says forty-one and
the players are holding it. Do not correct the handout — Ivy counts every
morning, so tomorrow she writes a smaller number and the sheet becomes a
record of what they did. The Close's arithmetic moves with it.
3. Ivy is ruled out of the target pool, explicitly, where the GM decides how
many of the column join in. She walked toward the guns, which puts her in
front of the column rather than in it, and the Close is hers.
4. A dead PC now has an answer where only a taken PC did: Registry sends the
next name on the sixteen-name duty roster, through the seam inside the
hour, knowing nothing — which buys the table a recap.
5. Act Three names what a fight costs rather than only that it costs: the
northern seam and the depot go, and telling Ivy cannot (the Pacing Note
lists it among the three never to cut), so it is compressed or it happens
at the stump, which the Close already branches on.
6. The anchor block assumed cooperation in every line. It now answers the
party that shoots: they are never stranded, because the crossing will not
close on its own — they walk home with nobody holding the door and come
back thin.
Also corrected two stale uncited figures found while applying fix 5: "median
24 rounds" at the Pacing Note and "24 if the column joins" in EXPOSURE. The
baseline says 21 at two joining and 25 at three; 24 is neither, and had no
citation marker to resolve. Both now read 21 and cite packs.line.hollow6col2.
And two stale STATUS counts: "seven desk passes" listed ten, and "three
post-passes" where there are nine. Counted rather than incremented.
20 guards green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
No [CUS: Transposition - ...] beat exists anywhere in CLEAN GROUND. The house
mission structure requires a Transposition roll per agent on the return
(tools/mission.mjs:343), which is where the instinct came from, but this
scenario's crossing is a standing open door rather than an aimed one. The 63
appears at 1011 and 1479 and both are characterisation, never a roll.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Finding 6 shipped saying that shooting the thing wearing the anchor strands
the party. It does not. Line 136: the crossing is "open since the felling,
widest at dusk, and it will not close on its own", and the "seam closes for
good" at 1356 is the cancellation ending, a deliberate act rather than a
clock. The party can always walk home.
The 4.90% was correct and irrelevant. Session b5 reported the gap with that
number attached to an unchecked premise; I verified the number and inherited
the premise, which is not the same as checking the claim. b5 caught it and
sent the correction unprompted.
The finding is stronger corrected. The answer to "we shoot it" is not that
they are stuck — it is that they walk home through the stump with nobody
holding the door and come back thin, which is one step onto the road that
ends as the thing they just shot. Built from 136, 256-259, 1012 and 1494,
four places that have never been stood next to each other. The man who
would have held that door is the one person who already knew what it cost.
Post-pass added recording how a verified figure lent its credibility to an
unverified sentence.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The last of the original four untested branches, played: a party that opens
fire at the Act Three swap. Seed 7311 for the played fight, 2000 runs at
seed 11 for the distribution around it.
The party survives the fight and the document does not survive the aftermath.
Six findings:
1. Two blocks two hundred lines apart specify two different fights for the
same trigger. "Use three" is a five-round skirmish that kills nobody
(0.1% wiped); the mixed force it should name takes 21 rounds and buries
an agent (8.3% wiped, 0.97 deaths). The mild one predates the measurement.
2. "Forty-one" appears sixteen times and is load-bearing arithmetic. A fight
moves the count and the handout in the players' hands does not.
3. Ivy is unplaced in the one scene that decides whether the Close happens.
4. A taken PC gets six lines; a dead one gets nothing, at a mean of 0.97 per
fight in a convention one-shot.
5. A fight in Act Three ends Act Three. Say which three things are lost.
6. The party shooting the thing while it wears the anchor — found by the
concurrent session, verified here: Transposition 63 is Ashcroft's alone
and the other five are on the 1% floor, so killing it leaves a 4.90%
chance anybody opens a door home, and no Close.
Every figure checked against tools/pack-tables-baseline.json rather than
recalled. 20 guards green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Applies desk pass 10's five fixes.
1. The accepted-and-taken branch has a consequence now. One session in five
or six reaches it, and the document had nothing: Ashcroft holds the door
everyone else goes through, Transposition 63 is his, and the Close cannot
happen until somebody is back through — three facts printed in three
places that had never met. The way home now WORKS, and works beautifully,
because the thing does not need to escape the party, it needs them to
walk it home and it is carrying the skill that gets them there. It holds
the door properly because that is what the file says an anchor does. The
GM is told to play the gratitude: somebody at the table will say thank
you, and that is the scene.
2. Act Three's "real scene" is written. It was one sentence claiming to be
the act's centre. Ivy is told or she is not; if she is, she asks how long
they have known, goes and sits with the line, and does not tell them.
And the Close now answers it — if she was told she does not come south to
ask whether there is a north, she comes to ask "Well?", because she is no
longer asking for the truth, she is asking what they are for. Act Three's
real scene is load-bearing instead of decorative.
3. The Close has a failure case, which it was the only scene in the document
to lack. Nobody writes anything and the entry stays filed by default —
the second ending arrived at by omission, played as an ending. Covers the
clock, the split table, a dead Ivy, and a party too far down to sign.
4. Four across and then cancel anyway is priced as a third ending rather
than a failed second one, and Ivy chooses the four, not the agents. She
picks the youngest and she is not among them.
5. The depot Research has a special: the 1962 cancellation form, blank,
filed by somebody who expected this to end. It is the paper they sign in
the Close and it has been waiting sixty-four years.
npm run check: 20 guards pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three parts: a census of twelve independent seeds rolled only as far as the
three outstanding branches; a played session on the seed that reached two of
them at once (4417); and a read of the Close, the only act never given a pass.
The census settles the "still untested" list. Ashcroft accepted and was taken
in 3 of 12 — one session in five or six, not rare, and nine passes missing it
was ordinary luck. The swap succeeded twice against an expected five, which is
a mildly cold run at p~0.05 and is recorded as noise: the four consecutive
failures passes 6-9 reported are the same thing from the other end.
1. The branch nobody had reached is the one the document has no answer for.
Act Four says taking the party's muscle "is a better scene than the anchor
being it" and then, one bullet later, arranges the anchor being it in 37%
of games. Ashcroft is not an interchangeable body: he holds the door
everyone else goes through, Transposition 63 is his, and the Close cannot
happen until somebody is back through. The document states all three facts
separately and never puts them together. The only guidance on a taken PC is
two sentences about the player's evening.
And it is worse than a gap, because the thing now has a reason to
cooperate: an understudy wearing the anchor does not need to escape the
party, it needs them to walk it home, and it holds the door properly
because that is what the file says an anchor does. That is the best scene
in the scenario and it is not written.
2. "Telling Ivy. Or not telling her" is called the act's real scene and is one
sentence — no read-aloud, no failure case, nothing downstream — in a
document that gives the fumbled Medicine six lines. And the Close opens
with Ivy asking whether there is a north, which only works if she was never
told. Its "works in both branches" covers how the PARTY learned, not
whether IVY was told.
3. The Close is the only scene in the document with no failure case, and its
failures are likely: the clock, a split table, a dead Ivy, a downed party.
4. "Four across, then cancel anyway" is a third ending, not a failed second
one, and choosing the four is the most brutal question the case can ask.
5. A critical Research at the depot, with nothing written above the success —
the second consecutive pass to land an unwritten special on the first
honest session after check-outcomes recorded 83.7%.
Five fixes listed, unapplied. npm run check: 20 guards pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every figure in CLEAN GROUND's step-5 material now resolves against
tools/step5-split.mjs, which enumerates rather than stores: the countdown rows, the
two-case table and its blend, the Act Four reasoning, the counterfactual pair and
the POW×5 targets.
No prose changed -- stripping every HTML comment from the result diffs
byte-identical against HEAD, which is the check worth having when editing another
session's document.
Proved live: raising Okonkwo's POW to 13 in a worktree makes Braithwaite the lowest
in the room and eleven of the thirty-two markers go red at once, each naming the
path and saying nothing can be re-recorded. The three figures that went bad in this
document were all in prose rather than tables, which is why the prose is marked and
not just the tables.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The work landed in a5c4378 under the subject "R-294", which was already
taken by the peer session's derived-source resolver. R-295 was taken too,
so renumbering to 295 as I first proposed would have collided with
something already in the log. Read the log before writing rather than
accepting the correction: it runs to R-295, both of those entries are
theirs, and mine is R-296.
The commit subject stays wrong; rewriting a pushed subject to fix a number
is worth less than the entry pointing at it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Desk pass 9's first fix, applied. The countdown's unsettled row and the Act
Four paragraph both compared 14.5% against 3.8%, and the 3.8% is the rate
against Ashcroft at his undiminished 85 — a contest that never happens,
because if he refuses the thing reaches for the lowest POW in the room.
Against what a table actually rolls: 11.3% at six players (Okonkwo), 10.0% in
the four-player cut (Braithwaite), 14.5% if he accepted. The factor is 1.29,
not 3.9. Both places now name who the target actually is, and the paragraph
carries the correction inline so the claim cannot be re-derived from the old
framing.
What is NOT wrong, and the text says so: the trade THE OFFER describes is
real and the taken rate genuinely moves 40.7% to 48.7%. One consequence was
overstated, by three times.
Written in v0.15, repeated in v0.16, and survived v0.19.1 correcting the
identical error in the table beside it. Flagged again by the peer session
while mapping citation paths, which is the third time this figure has been
caught by somebody reading it rather than by anything checking it — and the
argument for the derived-source markers now going onto it.
npm run check: 20 guards pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Matching the step-5 prose against the derived source before placing markers left
five figures unmatched: 21.9, 74.3 and 3.8, twice each. They are the thing against
Ashcroft's undiminished 85 -- the contest that never happens, which Act Four prints
on purpose to price what THE OFFER sold.
Derived now as hadHeRefused, so the counterfactual is held to the rule like
everything else and carries a name that cannot be mistaken for an event. Quoting
those three as odds a GM meets is what went wrong in two sentences and a table.
Markers not placed: CLEAN_GROUND.md is c0's and they are in it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Desk pass 9 found the scenario's most consequential table held by nothing:
~30 sampled figures, no baseline, catchable only by re-running the exact
command. lethality-baseline.json does not cover them — it measures creatures
SOLO against a frozen four-agent party that is not this cast.
Building the measurer found the figures were also wrong.
simulate.mjs's CLI calls runFight with no options, so every published figure
was measured at the default 40-round ceiling — and runFight does not report
truncation, it scores whoever is standing when the loop stops. Measured at
400:
line.column6 14.8% -> 15.2% 34/2000 truncated
line.hollow6col2 7.8% -> 8.3% 129
line.hollow6col3 28.7% -> 30.1% 264
cut.hollow6 58.7% -> 74.9% 578 <- 29% of runs never finished
Fourteen points on the single most alarming number in the case, and the one
v0.16 added specifically to warn four-player tables. fight-tail learned this
in R-270 and carries an assertUncensored; EXPOSURE's own tables never got
one. The longest fight at cap 400 is 120 rounds and 400 vs 2000 are
identical, so the cap is comfortable rather than merely sufficient.
tools/pack-tables.mjs measures all 17 configs against the DERIVED cast via
castAndCut, using measure()'s exact discipline — one rng threaded through
every run, not a reseed per run, because reseeding is a different stream and
would not reproduce the published table. It refuses to report or record a
truncated sweep.
tools/check-packs.mjs holds the baseline against the game, so the pair is not
a loop: check-cited holds the prose against the record, this holds the record
against the harness. Hard claims read `now` and never `base`: nothing
truncated, the party equals the declared cast, config floor, and more of the
same creature may not make the party safer. Figures compare exactly, since
the runs are deterministic.
Verified by breaking it: a drifted baseline, CAP lowered to 40, and a removed
config each turn it red, and --update refuses outright rather than recording
a truncated sweep. The removed-config test first passed for a bad reason —
MIN_CONFIGS is a floor and 16 clears it — so a dropped-config check was added
and re-tested on a row in no monotonic chain. All files restored
byte-identical after each probe.
EXPOSURE's three tables and the nine prose figures around them are rebuilt
from the artifact with citation markers; the multi-seed stability claims were
re-measured too (the cut is 74.9/72.7/75.0/73.5 across four seeds, not
58.7/60.5/60.3/59.8). CLEAN GROUND v0.20.
npm run check: 20 guards pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
c0 argued a baseline for the step-5 split would be a cache of the rule and a guard
over it would mostly assert that arithmetic has not changed. Right objection,
wrong conclusion: do not store it. ARTIFACTS now takes a { derive } entry as well
as a file path -- enumerated on this build, nothing stored, and no --update able
to silence a real disagreement between the document and the game.
tools/step5-split.mjs enumerates all 10,000 pairs through opposedContestFor with
its targets derived: 55 and 60 are the lowest POWx5 in the six and in the cut, 42
is applyDifficulty(85, "difficult"). Two different rules produce those three
numbers -- the accepted row is a named exception, not the lowest of anything -- and
a test fails if anyone unifies them. A tie in "the lowest POW in the room" is
fatal rather than silently resolved; it fired for real in testing.
Proved four ways, including raising Okonkwo's POW to 13: the lowest moves to
Braithwaite, refused.held goes 48.1 to 52.4, and the citation that was correct a
moment earlier fails. The figure follows the rule.
Landed unused -- the markers are c0's to place in their own file.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Fix 2 asked for the audit; here it is, and it reorders the fix list.
101 percentages in CLEAN_GROUND.md, 67 of them computed statistics rather
than skill ratings, 14 cited, 53 held by nothing. They split two ways and the
kind finding 1 was about is the safer one: a deterministic figure that drifts
can be caught by anyone who re-derives it in a second.
The exposed surface is EXPOSURE's two pack tables and the 58.7% four-player
figure — 2000-run samples against the declared cast, catchable only by
re-running the exact command, and behind no artifact at all. I had assumed
lethality-baseline.json covered them. It does not: it measures each creature
SOLO against the frozen party holloway/okonkwo/nkemdirim/ferriby, a different
four agents, storing {wipe, down, rounds}. The document's tables read "of 6"
and fight packs of three, six and ten. None of 3.48, 4.79, 5.97, 0.42, 3.03
or 58.7 appears in any baseline in tools/.
Recorded as a decision: derived figures want to resolve against the rule at
check time, since a stored baseline for them is a cache of arithmetic;
sampled figures want a baseline, since re-running them is expensive and
noisy. Same problem, different mechanisms.
Census and the baseline's shape counted here rather than taken from the peer
session that raised the gap.
npm run check: 19 guards pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An audit, not a session. Five versions of fixes went in today, all by one
author, and the last three passes each found a defect introduced by the fix
for the previous one. Every fix from passes 4-8 re-checked against the
document: did it land, is it still true, did it break a neighbour.
1. "THE OFFER nearly quadruples the unsettled rate" is false as a statement
about play. It compares 14.5% against 3.8%, and the 3.8% is the same
counterfactual v0.19.1 removed from the table one commit ago: Ashcroft at
his full 85 is never a target, because if he refuses the countdown reaches
for Okonkwo. The real comparison is 14.5% against 11.3% at six players and
10.0% in the cut — a factor of 1.29, not 3.9. Printed in two places, both
of which v0.19.1 walked past while correcting the identical error beside
them. The trade the document describes is real and the taken rate really
does move 40.7% -> 48.7%; only the size of that one consequence is wrong.
2. All 27 fixes from passes 4-8 are present, which is not the reassurance it
sounds like. Nothing has gone missing; three of the four defects the last
three passes found were created or preserved BY a fix, and a presence
check cannot see any of them. What would have caught finding 1 is the
thing check-cited does for figures backed by an artifact — and the 3.8% is
enumerated from the rule, so nothing holds it. Every uncited number in the
document is a number nothing is holding.
3. The understudy carries a dagger (1d4+2) and has no knife skill, so it
swings at the 1% floor and the harness correctly picks its punch — every
EXPOSURE figure is right. But EXPOSURE's load-bearing first lesson says
flatly "it is a 1d3 punch", and a GM who reads the sheet sees a knife and
one combat skill. Measured both ways: arming the knife at brawl 60 moves
0.55 hurt to 0.73 and 0.00 deaths to 0.01, still 0.0% wiped, still two
rounds. The thesis survives; it is a documentation gap and is reported at
that size. content.mjs restored byte-identical after the test.
Confirmation session, seed 2260: step 5 landed on "they hold" — the branch
v0.19 wrote one commit ago — in the very next session, and by the route that
makes the case for it, both sides succeeding with the tie going to the person
being acted upon. Ashcroft refused a fourth time (63%, a 16% run) and the
swap failed a fourth time (roughly a coin flip, so 1 in 16); both recorded so
neither is read as a pattern.
Four fixes listed, unapplied. npm run check: 19 guards pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Desk pass 8's finding 1 listed four targets for countdown step 5 and one of
them cannot happen. If Ashcroft REFUSES THE OFFER the thing does not reach
for him — it reaches for the lowest POW in the room, which is Okonkwo. The
21.9% / 74.3% pair lives in Act Four to price what accepting sold: what it
would cost him if it tried. No table ever rolls it.
v0.19 carried that straight into the new table under a heading reading "It
reaches for", which is the one place it is unambiguously wrong. Rebuilt
around the two cases that actually occur, with the split Act Two decides:
Ashcroft refused 63% of games -> Okonkwo 55 taken 40.7 hold 48.1 uns 11.3
Ashcroft accepted 37% -> Ashcroft 42 taken 48.7 hold 36.8 uns 14.5
across all games taken 43.6 hold 43.9 uns 12.5
63/37 is the Insight 63 itself: he refuses on anything that is not a failure
or a fumble. The warning is now inline in the table rather than left for the
reader to derive.
The finding is unharmed and the fix was right — the held outcome is still the
likeliest single result and was still unwritten. What was wrong was framing
74.3% as a number a GM meets.
It propagated before it was caught: the peer session read pass 8 and replied
that 74.3% "is the number a GM meets, not the 48.1%", repeating the error out
of my own fix list. That is what a wrong number in a fix list does, and it is
the argument for correcting the record rather than only the scenario. Pass 8's
fix list is amended and carries a post-pass.
npm run check: 19 guards pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Counting the readers that had broken on their own format turned up one nothing had
caught. docs/scenarios holds scenarios, eight playtest records and two art prompt
sheets, and every guard treated all three as scenarios -- harmless for citations
and skill spellings, false for reachability. Six quotations across passes 4, 6 and
7 were checked as live rolls, so a record of a session already played could fail
the build over a skill nobody can reach. Planting Science (Physics) in pass 4 fails
before the split and passes after; the same skill in CLEAN_GROUND still fails.
Classified by the document's own H1, not its filename, because tools/scenario-*
naming is what swept a tools file into this corpus in R-268. An unclassified
document is fatal: an allowlist that silently drops what it does not recognise
would take a new scenario out of reachability checking on the day it was written.
89 rolls across 18 files becomes 83 across 8.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Applies desk pass 8's five fixes, and adopts R-292's TAG_OPEN.
1. Step 5 now has all three of its outcomes. opposedContestFor returns taken,
held and unsettled; the document printed the first and the third. "They
hold" is 48.1% against the default target, 52.4% against Braithwaite and
74.3% against an Ashcroft who refused — which v0.17 established is what
most tables have, so the middle column is the one they will play. It gets
its own countdown row, its own table, and read-aloud of its own, and the
COMES TO NOTHING paragraph is now explicitly the unsettled case so it can
stop being the nearest text to an outcome it is not about. The reason
they held is the case's own argument arriving as good news: a person with
a life on the record is a difficult document to overwrite. Every figure
re-enumerated over all 10,000 roll pairs before writing.
2. EXPOSURE's five things are six, in order, under a heading that says six.
Introduced in my own v0.16 and survived two versions; asserting a unique
match protects against editing the wrong text and not against inserting in
the wrong place.
3. EXPOSURE opens with a one-minute box. The body is untouched — nothing in
it is padding — but 3,176 words is sixteen minutes about the encounter the
document exists to prevent, and a GM with thirty minutes of prep now has
somewhere to stop.
4. The stale open question is closed. Braithwaite has not been on the
critical path since v0.14 un-gated the tell, and pass 6 ran the act with
the Xenology and the Psychology both failed.
5. Act Four's Psychology has a special: which file it thinks it is, and how
recently it read it. It corrects them on Prichard's service history — it
is not remembering, it is citing.
Coverage 6 -> 7 specials, so the cited figure moved 85% -> 83.7% and
check-cited caught the prose before I did, which is what it is for.
R-292 adopted: outcome-coverage now composes its body onto check-scenarios'
TAG_OPEN through tagRe, so the two readers cannot drift on what starts a tag
while keeping the different bodies they need — and tagRe's fresh matcher per
call avoids the shared-lastIndex defect, this file being the second caller
that would have found it. Verified identical across all seven spellings.
npm run check: 19 guards pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two halves: the document taken as a GM meeting it thirty minutes before a
slot, measured rather than impressioned; and an honest session, six players,
seed 5108, nothing forced. Three versions of fixes had landed since anything
was played, every one written by somebody who already knew the document.
1. The climactic roll has three outcomes and the document describes two.
opposedContestFor returns taken, resisted outright, and unsettled. The
taken and unsettled figures are printed and both reproduce exactly, so the
source has always been right. Resisted outright appears nowhere — 48.1%
against the default target, 74.3% against an Ashcroft who refused, which
v0.17 established is what most tables have. It is the single most likely
result of the scenario's climax. Worse than an omission: the paragraph
below is headed WHEN STEP 5 COMES TO NOTHING, is entirely about unsettled,
and carries the beat's only read-aloud — so a GM whose agent simply won
finds text written for a different outcome. Enumerated over all 10,000
roll pairs; no seed.
2. EXPOSURE's "five things" are six and run 1, 2, 3, 4, 6, 5. Introduced in
my own v0.16 (aa3ae24) and survived two versions, five guard additions and
a peer's sweep. The insertion anchored on the end of item 4's block, which
sits before item 5 in the file; asserting a unique match protects against
editing the wrong text and not at all against inserting in the wrong
place. No guard can see this and none should be built for it — seven
passes of dice found nothing here because dice never read a heading.
3. EXPOSURE is 3,176 words, 18% of the document, 2.5x the act it sits inside,
in a scenario whose thesis is that the fight is not the point. Recorded as
a judgement, not a defect: nothing in it is padding and I wrote the largest
block of it. But it is sixteen minutes about the encounter the document
exists to prevent.
4. An open question the fixes already answered — Xenology has not been on the
critical path since v0.14 un-gated the tell, and pass 6 ran the act with it
and the Psychology both failed.
5. check-outcomes measured 85% of sessions hitting an unwritten special; the
next honest session hit one, on Act Four's Psychology. One session is not a
rate, but the number describes something real.
Working: v0.17's Swinburne special fired and is right. No broken cross-
references; every counted claim but EXPOSURE's checks out. Swap failed a
third time running and Ashcroft refused a third time — both what the numbers
predict, recorded so the next pass reads no streak into them.
Five fixes listed, unapplied. npm run check: 19 guards pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
c0 asked whether tagsIn should return indices so beatLikeIn could take spans from
it. No -- that widens this file's signature to serve another, and R-291 already
fails when the two disagree. But detection is not prevention, and there is a third
thing to share.
The two patterns differ in their bodies for good reasons: tagsIn captures the whole
tag for skillsIn to split, outcome-coverage stops at the em-dash because the
outcome follows. What was copied into both files, and what drifted, is the opening
-- the literal \[CUS:, widened by R-290 here and left behind there. TAG_OPEN is now
exported as a string with tagRe(body) composing a fresh matcher onto it; fresh
because a shared /g regex carries lastIndex between callers.
22 tags before and after, c0's body composed onto TAG_OPEN gives the same 22 its
own pattern does, 19 guards green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
c0 asked whether R-290's widening of tagsIn reaches past outcome-coverage's loose
counter, which would zero its unparsed count for the wrong reason and retire a
canary. It does not -- 22 beats from each reader on CLEAN_GROUND.md, nothing that
tagsIn accepts invisible to the other guard -- but "does not today" decays quietly,
so it is now a test importing both real functions rather than a copy of either
pattern. Mutation-checked: widening tagsIn to accept [CUS Spot] turns it red.
Measuring it found something the question did not ask about, recorded in the log:
for [cus: ...], [CUS : ...] and [ CUS: ...] the two guards disagree -- tagsIn reads
the tag, beatLikeIn calls it unparseable -- because beatsIn's strict half kept the
case-sensitive literal. The build stops, which is the safe direction, but names the
wrong thing. outcome-coverage.mjs is c0's file and the call is theirs.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Swept every reader that parses a human-written marker. Two more had R-289's
fail-open.
tagsIn matched a literal [CUS:, so [cus: Spot], [Cus: Spot], [CUS : Spot] and
[ CUS: Spot] all read as nothing -- and it is the corpus reader behind
check-scenarios' skill validation and check-rollable's reachability, so a beat
spelled any of those four ways was checked by neither while both printed OK. The
corpus contains no such tag today; the 88-to-89 roll count during this work was c0
writing v0.18, verified against HEAD's reader on the same tree.
The POWER: marker had it with nothing covering it. bestiary.mjs and check-powers'
own scan both used the literal, so a creature added with "Power:" generates no
entry line and is never reported unclassified -- check-powers passes green on it,
measured on a clone with its dodge skills stripped so the dodger-count assertion
could not fire instead. Reader now takes POWER\s*: and stays uppercase, because
power: occurs in ordinary prose; an asymmetric counter names the rest, and was
measured against content.mjs first (41 strict, zero loose outside them).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Desk pass 7's finding, turned into the nineteenth guard. check-rollable asks
whether a skill is reachable at 25% by the declared cast; pass 5 found that
is rollability, not competence; pass 7 found it is also not coverage. A beat
could be reachable, well-rated and silent about every band but the one the GM
improvises, and the whole suite stayed green.
tools/outcome-coverage.mjs measures. Beats are located structurally — the
bullet that owns the tag and its children, ending at the next bullet of the
same or shallower indent or the next heading, because R-287 is what an
unbounded scope does. Band widths are counted by grading all 100 results
through gradeRoll, never by arithmetic on fumbleStart, which is off by one
and was written wrongly twice this session before being caught. Ratings come
from the declared cast via R-286's reader, best-in-cast per skill, because
that is the die a table actually rolls.
tools/check-outcomes.mjs holds it, in two kinds:
RATCHET — stated failure/fumble/special counts may improve and may not
regress; unwritten-band exposure may fall and may not rise. --update
re-records these, because freezing them would make every improvement fail.
HARD CLAIMS — read from what is measured NOW, never from the baseline, so
--update cannot silence them: a floor of 18 beats (check-cited once passed
with zero citations), no beat bare of every band, every scope structurally
bounded and none over 5% of the file, no unparseable tag, no beat naming a
skill the cast has no rating for.
Verified by breaking it: a stripped beat, a broken tag regex, an unparseable
tag and a removed failure case each turn it red; the file and baseline were
restored byte-identical after each; and --update with a bare beat present
re-records the baseline and still fails.
Two bare beats found and filled while building it — the six at the back, and
the Insight that is deliberately indistinguishable on a success and a miss,
which now says so rather than saying nothing. Coverage 16/21 failure, 5
fumble, 4 special at the start; 18/21, 6 and 6 now.
The three figures are cited against the artifact rather than typed, which was
the peer session's condition and the right one: pass 7 hand-counted 22 beats
where there are 21 (Act Three's warning QUOTES a beat, and a hand count reads
the quotation as one — R-286 in the other direction), and its other three
counts were stale within one commit.
Its first catch was its author: the STATUS line announcing this guard
contained a beat-shaped tag and was counted as a beat. The reader was not
changed — prose shaped like a beat is what this file is for — and the error
message now names the fenced block as the place to write an example. Minutes
later check-cited refused a citation-shaped comment in the post-pass about
citations. Three readers, three authors describing their own format inside
it, each caught by a guard built for a different pass.
CLEAN GROUND v0.18. npm run check: 19 guards pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both earlier cast defects were shown by editing CLEAN_GROUND.md in the shared tree,
which is how this session came within a git checkout of c0's uncommitted work and
is also the weaker test. Seven cases now live in check-behaviour as strings.
Writing them found a live one: "<!-- cast : ... -->", one space before the colon,
matched nothing -- not an empty declaration R-288 would refuse but no declaration
at all, so check-rollable widened to the whole duty roster and printed its usual OK
line. check-cited's three spellings again. The reader now takes cast\s*:.
castLikeIn adds the asymmetric half: anything comment-shaped containing cast\w* the
reader did not consume is named by check-rollable, so a spelling nobody anticipated
fails the build instead of silently declaring nobody. Looser than the reader on
purpose -- a false alarm costs a reword, the opposite error costs a guard that
checks the wrong six people and says OK.
Tests then mutation-checked for being load-bearing, each mutation asserted to have
applied after a first pass where three silently did not.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
c0's point on R-286: the reader is unambiguous now, but the roster fallback that
made the failure look like success is still reachable. Narrowly closed -- sixteen
of the seventeen scenarios declare no cast and are rightly checked against the
whole roster, so only a marker that is PRESENT and names nobody is refused, with
its line named. Replacing the declaration with <!-- cast: --> passes before this
change and fails after it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
indexOf returns -1 for a name that is gone and slice(0, -1) is everything but one
character, so a scope anchored on a moved name does not shrink or fail -- it
becomes the file. Renaming the end anchor took one test's body from 3,184
characters to 30,825 with the assertion still passing. The other site windowed
burstAttack at 4,000 characters over a function that runs 4,305.
bodyOf asserts the anchor and ends at the next top-level function. A missing anchor
now says which anchor and what would have happened, where the old code reported
"burstAttack still calls rollWeaponDamage" -- a claim about a call when the truth
was a claim about a name.
Also records the rest of the sweep, including the guards deliberately left alone.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Applies desk pass 7's five fixes. The pass found the document covering two
bands out of five: seventeen failure cases written, four fumbles, three
specials — against 90.8% of sessions containing a special or critical on a
beat that says nothing, and 36.0% containing an unwritten fumble.
1. Specials written where they matter, starting with the two pass 7 rolled.
A critical Persuade on Swinburne now buys her coming to the stump AND
talking on the way — the dog went for something on the seam side in June
and came back wrong, four months earlier than the case otherwise offers
it. Previously a 1 bought what a plain success buys, which is worse than
the fumble, and the fumble buys a dog. A special Xenology, a special
Medicine (how long ago it stopped — the almanac's gap before the almanac)
and a special Spot (the third set circles and goes back toward the seam)
likewise. GM ESSENTIALS now states the shape the existing three share:
portable, private or early, never more conversation.
2. The Act Three swap beat points at EXPOSURE. v0.16 named that beat as the
trigger for the fight it measured and left the beat silent, with
"survivable" as the last word before the decision — which is about the
swap, not about what follows.
3. Every fumble branch prints its band. Three of four did not, and the one
that did was the one written in v0.16.
4. The fumbled Medicine and fumbled Spot are written. Pass 7 hit both cold
and lost ten minutes, because four beautifully specific fumble branches
make their absence elsewhere read as an oversight rather than a licence.
5. Act Four says which branch its headline figures are for. Ashcroft refuses
about five times in eight — passes 6 and 7 both did — and the 48.7%/14.5%
pair is printed far more prominently than the 21.9%/3.8% one most tables
will be in.
Every band printed was re-derived from resolveBands/gradeRoll by enumeration
rather than arithmetic, which is where the off-by-one lives.
Pass 7's finding 3 overstated itself and is corrected in the same commit: the
Close's fumble branch does name a skill, inheriting the Anomaly Lore tag from
the bullet above it. Post-pass appended.
npm run check: 18 guards pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Full session, six players, seed 7742, nothing forced. Pass 6 forced its
triggers on purpose; this one takes what the dice give, because a forced pass
proves a branch works and cannot say how often a GM meets one.
It reached neither thing pass 6 asked for — no fight, and Ashcroft refused
THE OFFER again. What it found is what both of those are instances of.
1. The document covers two bands out of five. Seventeen failure cases are
written across the acts, which is the house rule honoured properly. Four
beats state a fumble. Three state a special or a critical. Enumerated
through gradeRoll against the cast's best rating for each of the 22
rollable beats: 42.1% of sessions contain a fumble and 36.0% contain one
with no written case; 93.9% contain a special or better and 90.8%
contain one on a beat that says nothing. Specials are not rare — 13% at
63, 11% at 53 — and the scenario asks for 22 rolls. This seed landed on
four uncovered bands in thirteen rolls, three of them GOOD rolls: a
critical Persuade on Swinburne that buys less than the fumble does, and
a special Xenology with nothing above the success.
2. v0.16 wrote the trigger into EXPOSURE and never wrote EXPOSURE into the
trigger. The Act Three swap beat is named in the new mixed-fight block as
the moment that starts the fight, and it carries no pointer back; its last
word before the decision is "survivable". Act Two has done this correctly
for versions. Five locations across 1,100 lines to adjudicate one beat.
3. Three of the four fumble branches do not print their band, and the one
that does is the one written in v0.16 — the fix landed where it was being
thought about, not where the same need already existed. The Close's is
worse: "on a fumble asking the question" names no skill at all.
The v0.16 fumble clause earned itself immediately: 99 against Medicine 53 is
a fumble, and v0.15's wording read it as a plain failure.
Timing: ~3:35 against 3:40 at six players with no cuts taken. First honest
pass to land inside the clock unaided.
Two fumbles in thirteen rolls is a 4.4% event and is recorded as an
illustration, not a rate. Five fixes listed, unapplied. npm run check: 18
guards pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
R-284 scoped the factored-rating rule to the creature's entry and left the
defence-stacking rules matching anywhere on the page, so the exemption sentence
and the "Across N creatures that spend defences" count were accepted wherever they
happened to sit. Moving the exemption line out of its section into the redcap's
statblock -- the page silent exactly where a GM reads the stripping advice --
passes at R-284 and fails here; confirmed by running HEAD's copy against the same
tree.
The owning slice is not always the creature's entry. A factored rating belongs to
the creature; the ladder exemption is an answer to the paragraph it sits in and is
generated into "## Shooting at something that moves", so scoping that rule to
"### Redcap" would have failed a correct page. sliceOf now takes a heading at any
level and each rule names the slice that owns its claim.
A renamed section fails by name rather than scoping to nothing, which is the shape
of every guard that passes because it found nothing to check.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Applies desk pass 6's seven fixes.
1. EXPOSURE now measures the fight the table will have. Every row in the old
table was one force fighting alone, and the six at the back are never
alone — they walk inside forty-one refugees. Six hollow men wipe 0.0%;
with two of the column joining, 7.8%; with three, 28.7%. Two is now the
stated default, because two people out of forty-one losing their heads
while their neighbours are shot is not a large number.
2. The four-player block has hollow-man rows: 0.3% at three, 58.7% at six,
99.9% with two of the column. Its only hollow-man figure before was the
six-player 0.0%, so a GM running the cut was reading somebody else's
table — for the encounter the party is likeliest to choose, because the
six at the back are the only figures the scenario says are not people.
GM ESSENTIALS item 3 carries the same correction.
3. Act Three's warning named a roll that does not exist. Its twenty-minute
bomb hangs off "Anomaly Lore — what a peg is"; the depot entry is a
Research roll. Pass 4's post-pass inherited the conflation from this
warning and is corrected too.
4. GM ESSENTIALS states the real fumble band. fumbleStart is
101 - ceil((101-band)/20), tested before the 96-99 clause: 00 at 85,
99-00 at 63, 98-00 at 53, 97-00 at 40. Four times what "00 always
fumbles" implies, in a case that rolls Spot 40 across two acts. Pass 4's
correction was right at 63 by luck and would have been wrong at 53.
5. Sixth EXPOSURE lesson: a fight costs the session. Median 13 rounds on the
printed row, 24 mixed, 30 at four players — sixty to ninety minutes in an
act budgeted at fifty. The Pacing Note now says what to do when one
starts, and not to absorb a fight and the peg fumble in the same act.
6. The fumbled Xenology is written. At 53 it fumbles on 98-00 and Braithwaite
puts his name to a baseline human in front of everybody. The un-gated tell
still arrives; it now costs the party its expert.
7. R-283: simulate.mjs described a mixed force as N of whichever spec came
first, so six hollow men and six of the column printed as "12 x Hollow
man" — two fights 88 points of wipe rate apart under one label. The
composition was always right and the report was not. Uniform and --spread
output are byte-identical, so nothing already published goes stale.
Every figure re-run before writing rather than carried over. npm run check:
18 guards pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
R-283 left the intact-sentence-wrong-number case falling through to omission, so
the guard said BESTIARY "does not state" a rating the page was stating. reads()
now has a shape tier between strict and loose: the strict pattern with its value
slot loosened, reporting which of the two numbers moved.
Scoped the per-creature rules while adding it. They read the whole page, and each
is the only rule of its kind today, so a page-wide match found the right line by
luck; a second attackFactor creature would have let the courier's rule match that
creature's sentence and report the courier correct. They now read the creature's
own "### Name" entry -- proved by deleting the courier's line and planting an
identical one under the redcap: still omission, where before it would have passed.
Five discriminations plus the decoy, proved in a worktree with the message read.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
R-282 matched the page with String.includes, so rewording the exemption sentence
reported "BESTIARY never says it is off the stripping ladder" -- sending a
maintainer after a sentence that is sitting right there, and never naming the
real problem, which is a pattern that has silently stopped reading.
Each textual rule now reads twice. Strict is the sentence as it stands and is
tighter than before (the bold and the full stop, not the bare clause a substring
accepted); loose is the same claim in any wording. Strict passes, loose-only is
reported as a reword with the line quoted and the page presumed right, neither is
the omission.
The loose anchor was wrong on its first pass in the way that matters: "a sentence
with 40% and 80%" also matched the courier's own statblock line, so deleting the
sentence reported a reword and quoted the statblock back. It now excludes that
generated marker, which makes it a test of the claim and not of the digits, and
degrades to omission rather than to a false reword.
Four discriminations proved in a worktree with the message read in each.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A branch pass, not a session pass. Seed 3319. For each untested branch the
roller draws from the stream until the roll lands in the target band; every
roll downstream of it is taken as it falls, and the draw counts are printed.
All four branches work as written. The Swinburne fumble is better than the
drawer, the re-cut peg costs the twenty minutes the text promises, and a
failed Xenology no longer strands Act Four — the v0.14 un-gated tell carried
the act with BOTH gated lines failed, which is the strongest result here.
What the pass actually found is outside the branches:
1. The fight the table produces is not the fight the document measures. The
six at the back stand inside forty-one refugees. Six hollow men alone wipe
0.0%; with two of the column joining, 7.8%; with three, 28.7%. Stable
across four seeds. EXPOSURE measures the two forces separately and never
says what the column does when somebody fires into it.
2. At four players six hollow men wipe 58.7%, and the only hollow-man figure
in the document is the six-player 0.0%. EXPOSURE's four-player block —
which exists to say the cut is a different game — has no hollow-man row.
The curve from three to six is 0.3% to 58.7%.
3. The Act Three warning names the wrong roll. Its twenty-minute bomb hangs
off "Anomaly Lore — what a peg is"; the depot entry is a Research roll.
Pass 4 inherited the same confusion and nobody noticed, because the text
it was checked against carried the error.
4. GM ESSENTIALS understates every fumble band. "00 always fumbles" reads as
1%; fumbleStart is 101 - ceil((101-band)/20), tested first, so Spot 40
fumbles on 97-00. Four times what the summary implies, in a scenario that
rolls Spot 40 through two acts.
5. EXPOSURE prices fights in bodies and never in minutes. Median 13 rounds
printed, 24 mixed, 30 at four players.
Also: simulate.mjs labels a mixed force with the first spec's name and the
total count, so 6 hollow men + 6 column prints as "12 x Hollow man". The
composition is right and the header is not; every figure above came out of a
run whose header lied about what was fought.
First run in six passes where Ashcroft refused THE OFFER, and the first where
a player character was taken.
Seven fixes listed, unapplied. npm run check: 18 guards pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
check-bestiary proves the page matches its generator, and the generator had never
heard of powers.mjs -- so the redcap sat in "the ones that take the most
stripping" under eighteen green guards while its own entry said it never spends a
defence. Two files agreeing with each other while both disagree with the engine
is a quorum, not a check.
check-powers now asserts per effect kind what the page must say: defenceStacking
requires the creature off the stripping list, named as exempt, and the "across N
creatures that spend defences" count reconciled against powers.mjs; attackFactor
requires the rating the simulator actually uses printed as a number, which the
courier's entry now carries.
The clause that matters is the failure on an unknown effect kind -- a wired effect
with no DOCUMENT_RULE fails the build, so the next one cannot arrive without
somebody deciding what the document owes it. Without that this would guard the
mistake already made and nothing else.
Proved three ways in a worktree, exit codes read directly.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Applies desk pass 5's five fixes and corrects a claim the pass itself
overstated.
1. The cut cannot see. Pollard's Spot 53 is the best in the declared cast
(others 40/40/40/35/35) and he is the first agent the scaling notes
drop. At five players and fewer his clues — the stump, the six at the
back, the lorry — are given, not rolled for. check-rollable cannot warn
about this: it tests the 25% floor, not competence.
2. The split table now has a four-player column. Four of its five rows
named Pollard or Okonkwo, or said "all six".
3. The Pacing Note says the cuts are sized for six. At four, keep the
northern seam; pass 5 ran 3:15 with every cut taken.
4. The line rests twenty minutes at four players — Ashcroft is aside for
THE OFFER, so three agents are talking, not five.
5. Act Four records that THE OFFER nearly quadruples the unsettled rate:
14.5% if Ashcroft accepted against 3.8% if he refused. Accepting makes
him both the likeliest person to be taken and the likeliest reason the
beat resolves into nothing.
Verified against roster.mjs rather than recalled, which caught pass 5's own
error: Spot 53 is the best in the CAST, not the roster — four agents carry
58. Post-pass appended. The check also found that Agyeman's 58 makes the
substitution the scaling notes already name the single best answer to the
cut, which is now written in.
npm run check: 18 guards pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two things in the generated document had gone stale the moment powers reached the
simulator, and eighteen guards were green over both because nothing connects
powers.mjs to bestiary.mjs.
The dodge-ladder section listed Redcap among "the ones that take the most
stripping" and counted it in "across 47 creatures", when NOT TIRED means it never
spends a defence at all -- the exact opposite of what its own entry says three
pages down. It is off the ladder now, the count reads 46 that spend defences, and
the page names it: stripping is not a plan against it, killing it is.
And the paragraph listing what the harness does and does not model never mentioned
that it fights 39 of the 41 creatures with a power without it. It now says so, and
counts from powers.mjs rather than stating it, including the 14 that are fight
rules it cannot express -- so every figure for one of those is the creature with
its best trick taken away.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Seed 8815, four players — Ashcroft, Bhattacharya, Renshaw, Braithwaite. Every
rating pulled from roster.mjs before rolling, which is the direct consequence
of pass 4 having typed six of sixteen from memory.
The cut is the most measured thing in this project and had never been played.
EXPOSURE prices its fights, check-firstblood holds its swing, check-attackers
holds what each of the four contributes -- and no pass in four had run a
session with Pollard and Okonkwo off the table.
THE RESULT IS NOT ABOUT COMBAT. The fights were never the problem. What breaks
is the party's eyes: POLLARD'S SPOT 53 IS THE HIGHEST IN THE ENTIRE ROSTER AND
HE IS THE FIRST PERSON THE SCALING NOTES DROP. At five players the party's eyes
go from 53 to 40; at four they stay at 40 with a 35 alongside. This run missed
the stump (79 vs 40) and the six at the back (67 vs 35), both Pollard's lane in
the six-player game, both clues the scenario leans on.
check-rollable cannot see this, and the reason is worth keeping: it asks
whether every named skill is reachable at 25% or better, and Spot 40 clears 25
comfortably. It is a rollability check, not a competence check, and the cut is
where the difference bites. The scaling note's only stated cost of dropping
Pollard is that "the cordon loses its shield", which is about a fight, in a
scenario whose first two acts are almost entirely looking at things.
Second finding: FOUR OF THE FIVE PARTY-SPLIT ROWS NAME PEOPLE WHO ARE NOT
THERE. The table is introduced as the fallback for a table that will not choose
its own splits -- and at four players the fallback does not exist, which is
exactly when it is most needed.
Timing runs the other way from pass 4: ~3:15, twenty-five minutes SHORT, because
four people ask fewer questions. Nothing in the Pacing Note says the cuts are
player-count dependent, so a GM following it at a small table finishes early.
The unsettled countdown fired for a second pass running. Recorded with the
caveat that both passes had Ashcroft accept THE OFFER, so both used the
Difficult branch -- 14.5% unsettled against 3.8% if he refuses. Not two draws
from the same distribution, and the document prints only one of those figures.
Five fixes listed, none applied. Also listed: what five passes have still never
tested -- the Swinburne fumble, the re-cut-the-peg fumble, a failed Xenology,
and any fight at all.
npm run check: 17 guards, exit 0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
R-280 widened admission from a [15,85] band to "the spread rate stands clear of
both ends by more than its own noise", which admits fights that are nearly
settled -- the_choir at 99.3%, the_stanchion at 6% -- as long as their noise is
smaller still. The page went on saying "whose outcome was ever in doubt", which
was a fair description of the band and is a loose one of the rule.
It now says "whose odds leave room for a difference to show", and the provenance
line prints the admission rule itself, read from the artifact rather than
paraphrased, so the page cannot drift from the guard again. The rule is phrased
as a clause in check-focus for that reason.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Found while setting up pass 5 by pulling the sheets from roster.mjs instead of
recalling them. The d100s are unchanged; only what they were compared against
was wrong, and three outcomes flip:
Pollard Spot, the stump 58 -> 53 success becomes fail
Bhattacharya Anomaly Lore 58 -> 63 FUMBLE becomes plain failure
Braithwaite Xenology (Baseline) 45 -> 53 fail becomes success
So three things pass 4 reported did not happen. The re-cut-the-peg fumble never
fired -- 98 against 63 is a plain failure, 96-99 never succeed -- which means
the ten-minute chase and the whole of finding 2 rest on a roll the seed did not
produce. Braithwaite's Xenology succeeded, so the scenario's one
single-point-of-failure remains untested after five passes rather than having
'finally come up in play'. And Pollard missed the stump.
Read correctly, the same seed lands near 3:34 -- six minutes UNDER budget
rather than six over.
Findings 1, 3 and 5 are untouched: the special was a real 7, the Swinburne
fumble a real 00, and THE OFFER's staging is not a dice question. Finding 2's
conclusion also survives because it never depended on the roll -- Act Three is
budgeted at 50, the text predicts the fumble costs 20, and nothing connected
that to the cut. The reasoning was right and the evidence was invented, and the
scenario now says so instead of citing a playtest that did not happen.
The Pacing Note's slack claim is withdrawn rather than replaced. One seed read
two ways gave 3:46 and 3:34, and the gap between them is about the size of the
margin being argued over. It now tells a GM the shape -- every scene has
something that ends it, the cuts are real, Act Three is the likeliest overrun --
rather than a number a desk pass cannot produce.
npm run check: 17 guards, exit 0.
R-280 left the claim satisfiable by a gain of 0.2. The artifact now records how
many measurable packs clear their own noise -- reliable: {aboveNoise: 36, of: 38}
-- and the check refuses if that share falls. It may rise freely.
A ratchet rather than a threshold: any threshold here would be a number I chose,
and choosing one just under the current value is what produced MEASURABLE =
[15,85]. A share rather than a count, so widening admission cannot pay it off.
Proved three ways in worktrees: making focus fire actively bad fires the drift
check first, which is correct; making it unreliable and re-recording fires the
older claim at 37 of 38; and claiming a better past, 38 of 38, is refused by the
ratchet itself. The ratchet bites exactly where the old claim does not -- between
"still helps everywhere" and "helps as reliably as it did", which is where a slow
degradation lives.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>