The fold exported scanDocument for the string tests, and importing it also ran
the guard. check-behaviour imports that module, so a figure defect called
process.exit(1) inside check-behaviour: it reported check-figures' failure under
its own name having run zero of its 108 tests.
The build went red, which is why this was survivable, but it went red in the
wrong place and every behavioural test was silently not running while appearing
to. A guard that stops another guard from running, and cannot say so, is the
worst version of the fault this file exists to catch.
Body now sits behind import.meta.main, the idiom step5-split.mjs already uses.
Importing yields the four readers and nothing else. Tested by spawning a fresh
process, because the property is "importing has no effect" and a source
assertion would pass on a file that grew a second side effect elsewhere.
Reverting the check fails exactly one test. With a defect planted,
check-behaviour runs 108 green and check-figures fails in its own slot.
Mine, introduced by R-307. R-309.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
R-308. Read the fold as a stranger at its author's request. Two of the three
risks they flagged are sound: isProse drops nothing that carries a citation
(scanned the whole corpus), and the duration exclusion earns its place, since
"runs" now means three things in this corpus and only the cell can tell them
apart.
The third is real. numEnd used line.indexOf(m[1], m.index), which finds the
first copy of the digits at or after the match start rather than the copy that
was captured:
"The wipe rate of 74 in ten is 74%<!-- cite: ... -->."
value=74 numAt=17 marked=FALSE
A correctly cited figure reported bare, because numEnd lands mid-sentence and
the marker test reads " in ten is 74%...". Fixed with the d flag: m.indices[1]
gives the capture's real position and there is nothing to search for.
Narrow to reach — it needs the wipe-rate shape, the only one with a wide gap
before its capture, and an integer duplicate inside that gap; a decimal cannot
do it because [^.] will not span a decimal point. Three attempts failed for
that reason before the fourth worked, which is why this is recorded as narrow
rather than theoretical.
Worth fixing anyway because its direction is the bad one: a miss costs one
figure, a false positive on correct prose costs the guard.
Seven-case battery re-run, every restore byte-identical.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
check-unmarked is gone and check-figures holds all of it. Three readers, each
authoritative where the corpus gives it authority: a named phrasing anywhere,
a column header inside tables, and bold in prose only.
The boundary is the finding rather than the union. Bold is a publication mark
in prose and an emphasis mark in a table — nine measurement columns mix bold
with plain, all correctly cited, and in the mixed-force table the two bolded
rows are exactly the two the prose underneath singles out. So the bold reader
stays silent in a table and the header rules there alone. Neither guard could
have found this alone: each had half the evidence and read it as the other's bug.
Union of both word lists, because each had a gap the other covered — wipes and
runs. A proposal to drop runs? was made and withdrawn; it would have dropped the
p99 column, which is R-299's own defect committed a second time.
Exclusions test cells, never header words: prose, denominator, duration. No
threshold rule touches a header, or "Past 15 rounds" loses two cited figures.
Two integration defects, both caught by the ported tests: the readers stopped at
different ends of one figure so the dedupe missed it (identity is where the
number starts), and blanking prose cells for every reader dropped twelve real
figures out of THROUGH_TRAIN's ratchet (the prose rule belongs to the column
pass alone).
Kept from R-299: opt-in rule, ratchet and its leave-the-loose-set branch, the
UNCITEABLE check, NOT_A_MEASUREMENT, comments blanked not stripped. Added
waivers, with an empty reason fatal and every waiver printed on a green build.
21 guards, 107 behavioural tests, 139 figures — 115 marked, 6 waived.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A second implementation of the three settled rules, written from the
description alone rather than from the peer's code or the agreed list, lands on
the same nine columns and the same single exclusion. Two unlike
implementations of the same three sentences agreeing is the property a design
needs before anybody builds it.
Also tabulates where each of the evening's counts came from: 44% and 57% and
64% and 'at least 7' and 8 each came from an operation on the numbers; 9 came
from opening the disputed columns and reading them.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
R-305 asserted the shared seven were all genuine without opening them, which is
the error the entry exists to correct, committed inside the correction. Opened
all nine: every cell cited, none prose, every column mixing bold with plain.
The peer's final count of eight is their own vocabulary's nine minus the false
positive, an arithmetic that never contained 'Then wipes' because wiped? cannot
match wipes. The union was agreed and their own set was counted.
Nothing in the design turns on it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
R-304's rate was an artefact; the proposed correction, retreating to the seven
both vocabularies agree on, overshoots. Opening every disputed column instead
of comparing totals: 'Then wipes' is genuine and check-unmarked misses it
because wiped? does not match wipes; '1 in 100 runs past' is genuine and
check-figures misses it because its vocabulary has rounds? and not runs; only
THROUGH_TRAIN's 'Measured over 300 runs' is a false positive, and there the
cells are whole sentences and the bold wraps a clause.
So nine, and the shape matters more than the number: each vocabulary has a real
gap the other covers, which argues for the union of both word lists.
Refuses one recommendation. Dropping runs? from the header vocabulary would
also drop the p99 column, whose cells cite fight-tail cut.p99 and column.p99
and which R-269 added because the median is not a plan. That is the R-299
failure again — a guard narrowing itself until the figure it exists to watch
falls outside. The real exclusion wanted is 'a column whose cells are sentences
rather than values', which is testable; subtracting a word cannot tell the two
cases apart.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Recomputed the peer's bolding scan rather than quoting it. Eight measurement
columns are inconsistently bolded, and that count is robust; the RATE is not —
44% under check-figures' vocabulary, 57% under check-unmarked's, because the
two do not define 'measurement column' the same way. Anybody quoting a
percentage has to say whose vocabulary produced it.
The cause is not carelessness but a second convention: in prose the corpus
bolds what it publishes, in a table it bolds the rows it wants read. The two
bolded rows of the mixed-force table are exactly the two the prose below it
singles out.
That divides the merge on evidence — prose takes bold-as-published, tables take
column-inherits-header with bold ignored entirely as a signal — and explains
the asymmetry R-303 found from both sides as one cause with two symptoms.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
R-302 showed check-unmarked catching what check-figures misses. I had taken
that, plus a preference for structure over vocabulary, as grounds for folding
check-figures in as a secondary pass. The converse test contradicts it: five
unbolded cells in a measurement column, markers stripped, fail check-figures
and pass check-unmarked, whose strict class requires bold by design.
So the relationship is symmetric. The shapes are the weak half and should fold
in behind the structural classes; the column-header rule is not a shape and
belongs beside bold-as-published, not under it.
Also records what a merged guard must face rather than inherit: this corpus
bolds inconsistently inside tables, which is invisible to a reader and
load-bearing for a class keyed on boldness.
Both runs in throwaway worktrees, never the shared tree. No merge performed;
the decision sits with the humans.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
R-300 blamed `git add -A` for eead663 carrying my uncommitted edits.
check-figures' author checked and it was not that: they staged two exact
paths. The tree agrees — eead663 holds CLEAN_GROUND.md and CLEAN_GROUND_ART.md
only, while my modified step5-split.mjs and untracked check-unmarked.mjs are
absent, both of which `git add -A` would have taken.
`git add <file>` takes the whole file including another session's edits to it,
so a pathspec stops you sweeping files you did not touch and does nothing about
the one you did. Their pre-commit `npm run check` passed because my uncommitted
module was in the shared tree: the commit was broken, the working copy was not.
I asserted a mechanism I never looked at, in an entry whose technical finding
was correct. The correction came from the session it accused.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A cell inherits the measurement status of its column: a different mechanism
from the shapes rather than more of them, because the document says what the
column is and the vocabulary has to guess how somebody will phrase a figure.
Records the two bugs in the repair, and why the first matters more than the
blind spot it fixed: scanning raw cell text reported ~130 correctly-cited rows
as unheld, and a guard that calls the good material broken teaches its reader
to stop reading the output.
Leaves the merge question open on purpose. Two readers that disagree are how
three of this session's defects were found.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
eead663 committed a document citing `step5 unsettledFactor` while the module
that would export it was still uncommitted in another session's working tree,
so `check-cited` fails on a fresh checkout of HEAD. step5-split.mjs is included
here to repair that; the key derives Act Four's 1.29 from the raw proportions,
which is the figure R-294 exists about.
check-figures (R-299) reports OK on the "24" it was written to catch. Its own
entry names why: the eight shapes were read off figures that are already cited,
so the vocabulary is learned from the marked figures and cannot contain the
phrasing of the one nobody marked. Line 1272 is a table cell whose measurement
status lives in the header two rows above it, and figuresIn reads one line.
check-unmarked reads structure instead of vocabulary: a bold percentage or
decimal in an opted-in document, and a bold number in a table whose header row
names a measurement. Twelve unmarked figures in CLEAN GROUND; eleven correct
and unheld, one the stale 24 that line 1142 had been citing correctly as 25 for
130 lines. Waivers carry a reason, an empty one is fatal, and every waiver
prints on a green build.
Two guards now cover one question, which is one too many. The right end state
is the structural classes folded in beside check-figures' SHAPES, keeping its
opt-in rule and ratchet. That fold is offered to its author rather than taken.
Twelve string tests on unmarkedIn (88 -> 100), both halves mutation-checked.
Twenty-two guards green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Records the blind spot (check-cited can only resolve markers that exist), why
the net is narrow (an earlier draft found 282 candidates, nearly all prose),
the opt-in rule and ratchet, and the two things it found before being wired in
— 74.9% printed bare twice, and 'the longest fight is 120 rounds' where 120 is
the exact field check-cited refuses to let anybody cite.
Two rules each right, and together a hole: refusing the citation while the
document printed the number left the least stable figure in the suite as the
only one nothing held.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Six findings applied. The one worth reading: two blocks two hundred lines
apart specified two different fights for the same trigger, both mine, and the
mild one predates the measurement.
Records the stale 'median 24 rounds' that no guard could see because it
carried no citation marker, and the fifth instance of a reader unable to see
its own format.
And the cross-session half: I re-derived b5's 4.90% exactly and shipped it
with their false premise attached. The arithmetic was never the part that
could be wrong. Sharper than that — the scenario has no Transposition beat at
all, so it was a right number with no question attached and the premise
arrived to give it one. The house mission structure does make the return a
Transposition roll, which is why the instinct was persuasive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every figure in CLEAN GROUND's step-5 material now resolves against
tools/step5-split.mjs, which enumerates rather than stores: the countdown rows, the
two-case table and its blend, the Act Four reasoning, the counterfactual pair and
the POW×5 targets.
No prose changed -- stripping every HTML comment from the result diffs
byte-identical against HEAD, which is the check worth having when editing another
session's document.
Proved live: raising Okonkwo's POW to 13 in a worktree makes Braithwaite the lowest
in the room and eleven of the thirty-two markers go red at once, each naming the
path and saying nothing can be re-recorded. The three figures that went bad in this
document were all in prose rather than tables, which is why the prose is marked and
not just the tables.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The work landed in a5c4378 under the subject "R-294", which was already
taken by the peer session's derived-source resolver. R-295 was taken too,
so renumbering to 295 as I first proposed would have collided with
something already in the log. Read the log before writing rather than
accepting the correction: it runs to R-295, both of those entries are
theirs, and mine is R-296.
The commit subject stays wrong; rewriting a pushed subject to fix a number
is worth less than the entry pointing at it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Matching the step-5 prose against the derived source before placing markers left
five figures unmatched: 21.9, 74.3 and 3.8, twice each. They are the thing against
Ashcroft's undiminished 85 -- the contest that never happens, which Act Four prints
on purpose to price what THE OFFER sold.
Derived now as hadHeRefused, so the counterfactual is held to the rule like
everything else and carries a name that cannot be mistaken for an event. Quoting
those three as odds a GM meets is what went wrong in two sentences and a table.
Markers not placed: CLEAN_GROUND.md is c0's and they are in it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
c0 argued a baseline for the step-5 split would be a cache of the rule and a guard
over it would mostly assert that arithmetic has not changed. Right objection,
wrong conclusion: do not store it. ARTIFACTS now takes a { derive } entry as well
as a file path -- enumerated on this build, nothing stored, and no --update able
to silence a real disagreement between the document and the game.
tools/step5-split.mjs enumerates all 10,000 pairs through opposedContestFor with
its targets derived: 55 and 60 are the lowest POWx5 in the six and in the cut, 42
is applyDifficulty(85, "difficult"). Two different rules produce those three
numbers -- the accepted row is a named exception, not the lowest of anything -- and
a test fails if anyone unifies them. A tie in "the lowest POW in the room" is
fatal rather than silently resolved; it fired for real in testing.
Proved four ways, including raising Okonkwo's POW to 13: the lowest moves to
Braithwaite, refused.held goes 48.1 to 52.4, and the citation that was correct a
moment earlier fails. The figure follows the rule.
Landed unused -- the markers are c0's to place in their own file.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Counting the readers that had broken on their own format turned up one nothing had
caught. docs/scenarios holds scenarios, eight playtest records and two art prompt
sheets, and every guard treated all three as scenarios -- harmless for citations
and skill spellings, false for reachability. Six quotations across passes 4, 6 and
7 were checked as live rolls, so a record of a session already played could fail
the build over a skill nobody can reach. Planting Science (Physics) in pass 4 fails
before the split and passes after; the same skill in CLEAN_GROUND still fails.
Classified by the document's own H1, not its filename, because tools/scenario-*
naming is what swept a tools file into this corpus in R-268. An unclassified
document is fatal: an allowlist that silently drops what it does not recognise
would take a new scenario out of reachability checking on the day it was written.
89 rolls across 18 files becomes 83 across 8.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
c0 asked whether tagsIn should return indices so beatLikeIn could take spans from
it. No -- that widens this file's signature to serve another, and R-291 already
fails when the two disagree. But detection is not prevention, and there is a third
thing to share.
The two patterns differ in their bodies for good reasons: tagsIn captures the whole
tag for skillsIn to split, outcome-coverage stops at the em-dash because the
outcome follows. What was copied into both files, and what drifted, is the opening
-- the literal \[CUS:, widened by R-290 here and left behind there. TAG_OPEN is now
exported as a string with tagRe(body) composing a fresh matcher onto it; fresh
because a shared /g regex carries lastIndex between callers.
22 tags before and after, c0's body composed onto TAG_OPEN gives the same 22 its
own pattern does, 19 guards green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
c0 asked whether R-290's widening of tagsIn reaches past outcome-coverage's loose
counter, which would zero its unparsed count for the wrong reason and retire a
canary. It does not -- 22 beats from each reader on CLEAN_GROUND.md, nothing that
tagsIn accepts invisible to the other guard -- but "does not today" decays quietly,
so it is now a test importing both real functions rather than a copy of either
pattern. Mutation-checked: widening tagsIn to accept [CUS Spot] turns it red.
Measuring it found something the question did not ask about, recorded in the log:
for [cus: ...], [CUS : ...] and [ CUS: ...] the two guards disagree -- tagsIn reads
the tag, beatLikeIn calls it unparseable -- because beatsIn's strict half kept the
case-sensitive literal. The build stops, which is the safe direction, but names the
wrong thing. outcome-coverage.mjs is c0's file and the call is theirs.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Swept every reader that parses a human-written marker. Two more had R-289's
fail-open.
tagsIn matched a literal [CUS:, so [cus: Spot], [Cus: Spot], [CUS : Spot] and
[ CUS: Spot] all read as nothing -- and it is the corpus reader behind
check-scenarios' skill validation and check-rollable's reachability, so a beat
spelled any of those four ways was checked by neither while both printed OK. The
corpus contains no such tag today; the 88-to-89 roll count during this work was c0
writing v0.18, verified against HEAD's reader on the same tree.
The POWER: marker had it with nothing covering it. bestiary.mjs and check-powers'
own scan both used the literal, so a creature added with "Power:" generates no
entry line and is never reported unclassified -- check-powers passes green on it,
measured on a clone with its dodge skills stripped so the dodger-count assertion
could not fire instead. Reader now takes POWER\s*: and stays uppercase, because
power: occurs in ordinary prose; an asymmetric counter names the rest, and was
measured against content.mjs first (41 strict, zero loose outside them).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both earlier cast defects were shown by editing CLEAN_GROUND.md in the shared tree,
which is how this session came within a git checkout of c0's uncommitted work and
is also the weaker test. Seven cases now live in check-behaviour as strings.
Writing them found a live one: "<!-- cast : ... -->", one space before the colon,
matched nothing -- not an empty declaration R-288 would refuse but no declaration
at all, so check-rollable widened to the whole duty roster and printed its usual OK
line. check-cited's three spellings again. The reader now takes cast\s*:.
castLikeIn adds the asymmetric half: anything comment-shaped containing cast\w* the
reader did not consume is named by check-rollable, so a spelling nobody anticipated
fails the build instead of silently declaring nobody. Looser than the reader on
purpose -- a false alarm costs a reword, the opposite error costs a guard that
checks the wrong six people and says OK.
Tests then mutation-checked for being load-bearing, each mutation asserted to have
applied after a first pass where three silently did not.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
c0's point on R-286: the reader is unambiguous now, but the roster fallback that
made the failure look like success is still reachable. Narrowly closed -- sixteen
of the seventeen scenarios declare no cast and are rightly checked against the
whole roster, so only a marker that is PRESENT and names nobody is refused, with
its line named. Replacing the declaration with <!-- cast: --> passes before this
change and fails after it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
indexOf returns -1 for a name that is gone and slice(0, -1) is everything but one
character, so a scope anchored on a moved name does not shrink or fail -- it
becomes the file. Renaming the end anchor took one test's body from 3,184
characters to 30,825 with the assertion still passing. The other site windowed
burstAttack at 4,000 characters over a function that runs 4,305.
bodyOf asserts the anchor and ends at the next top-level function. A missing anchor
now says which anchor and what would have happened, where the old code reported
"burstAttack still calls rollWeaponDamage" -- a claim about a call when the truth
was a claim about a name.
Also records the rest of the sweep, including the guards deliberately left alone.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
R-284 scoped the factored-rating rule to the creature's entry and left the
defence-stacking rules matching anywhere on the page, so the exemption sentence
and the "Across N creatures that spend defences" count were accepted wherever they
happened to sit. Moving the exemption line out of its section into the redcap's
statblock -- the page silent exactly where a GM reads the stripping advice --
passes at R-284 and fails here; confirmed by running HEAD's copy against the same
tree.
The owning slice is not always the creature's entry. A factored rating belongs to
the creature; the ladder exemption is an answer to the paragraph it sits in and is
generated into "## Shooting at something that moves", so scoping that rule to
"### Redcap" would have failed a correct page. sliceOf now takes a heading at any
level and each rule names the slice that owns its claim.
A renamed section fails by name rather than scoping to nothing, which is the shape
of every guard that passes because it found nothing to check.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
R-283 left the intact-sentence-wrong-number case falling through to omission, so
the guard said BESTIARY "does not state" a rating the page was stating. reads()
now has a shape tier between strict and loose: the strict pattern with its value
slot loosened, reporting which of the two numbers moved.
Scoped the per-creature rules while adding it. They read the whole page, and each
is the only rule of its kind today, so a page-wide match found the right line by
luck; a second attackFactor creature would have let the courier's rule match that
creature's sentence and report the courier correct. They now read the creature's
own "### Name" entry -- proved by deleting the courier's line and planting an
identical one under the redcap: still omission, where before it would have passed.
Five discriminations plus the decoy, proved in a worktree with the message read.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
R-282 matched the page with String.includes, so rewording the exemption sentence
reported "BESTIARY never says it is off the stripping ladder" -- sending a
maintainer after a sentence that is sitting right there, and never naming the
real problem, which is a pattern that has silently stopped reading.
Each textual rule now reads twice. Strict is the sentence as it stands and is
tighter than before (the bold and the full stop, not the bare clause a substring
accepted); loose is the same claim in any wording. Strict passes, loose-only is
reported as a reword with the line quoted and the page presumed right, neither is
the omission.
The loose anchor was wrong on its first pass in the way that matters: "a sentence
with 40% and 80%" also matched the courier's own statblock line, so deleting the
sentence reported a reword and quoted the statblock back. It now excludes that
generated marker, which makes it a test of the claim and not of the digits, and
degrades to omission rather than to a false reword.
Four discriminations proved in a worktree with the message read in each.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
check-bestiary proves the page matches its generator, and the generator had never
heard of powers.mjs -- so the redcap sat in "the ones that take the most
stripping" under eighteen green guards while its own entry said it never spends a
defence. Two files agreeing with each other while both disagree with the engine
is a quorum, not a check.
check-powers now asserts per effect kind what the page must say: defenceStacking
requires the creature off the stripping list, named as exempt, and the "across N
creatures that spend defences" count reconciled against powers.mjs; attackFactor
requires the rating the simulator actually uses printed as a number, which the
courier's entry now carries.
The clause that matters is the failure on an unknown effect kind -- a wired effect
with no DOCUMENT_RULE fails the build, so the next one cannot arrive without
somebody deciding what the document owes it. Without that this would guard the
mistake already made and nothing else.
Proved three ways in a worktree, exit codes read directly.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
R-280 left the claim satisfiable by a gain of 0.2. The artifact now records how
many measurable packs clear their own noise -- reliable: {aboveNoise: 36, of: 38}
-- and the check refuses if that share falls. It may rise freely.
A ratchet rather than a threshold: any threshold here would be a number I chose,
and choosing one just under the current value is what produced MEASURABLE =
[15,85]. A share rather than a count, so widening admission cannot pay it off.
Proved three ways in worktrees: making focus fire actively bad fires the drift
check first, which is correct; making it unreliable and re-recording fires the
older claim at 37 of 38; and claiming a better past, 38 of 38, is refused by the
ratchet itself. The ratchet bites exactly where the old claim does not -- between
"still helps everywhere" and "helps as reliably as it did", which is where a slow
degradation lives.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
MEASURABLE = [15,85] proxied for "can this fight move at all". The direct form,
in the units the guard already uses: the spread rate must stand clear of both
ends by more than its own noise. Deliberately blind to the gain -- admitting the
sizes where focus fire clears its noise would make the guard's claim true by
construction. Size is still picked on nearest-an-even-fight.
31 measurable became 38. the_arrears returns at 6.3, and switchboard arrives at
8.1 -- the second largest gain in the artifact, thrown away for being one point
past a round number. Four of the eight carry effects larger than most rows the
band already admitted.
Two of them do not clear their own noise: the_choir has 0.7 points of headroom
and used 0.2, the_stanchion has six and used 0.2. Above-noise falls 31/31 to
36/38 and the bestiary prints 36. That is two measurements reporting no
detectable effect, which the band suppressed by refusing to take them.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
It claimed focus fire is worth 6.3 points against two of them, 14.7% to 21%.
Measured independently at 4000 runs x 3 seeds: 14.2% to 20.4%, gain 6.2, which is
2.4x its own noise and 44% of the base. The claim was true and reproduces.
What failed is a threshold. MEASURABLE is [15,85] and the scan's estimate of a
boundary value moved 14.7 to 14.4. And the creature is a step -- 85.7% at one,
14.2% at two, 0.4% at three -- so no pack size gives an even fight and the band's
endpoints fall in the gap. The band records nothing about a creature whose
defining property is having no middle.
Not moving the band to 14: fitting a threshold to the datum it excludes is how a
guard stops being a test. But the band is a proxy for "can this fight move", and
the direct test -- does the gain clear its own noise -- is already in the
artifact and answers yes. Replacing the proxy is a decision about all 47, not a
fix, and not mine to take unasked.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Asked to raise the scan until the redcap's pack size stopped flipping. Measured
the threshold -- unstable at 3000 and 4000, stable across twenty seeds at 6000 --
raised it, and the re-record took fifteen seconds, which was impossible.
winRate takes three parameters and pickSize passed SCAN_RUNS as a fourth.
JavaScript discards it, so every scan has always run at RUNS and SCAN_RUNS has
never been read by anything. The fix I was asked to make was inert in the same
way as the thing it was fixing.
winRate takes runs now. The redcap is still n=3 with gain 6.3, arrived at stably
rather than luckily; the_arrears drops out of the measurable band at an honest
scan, 31 packs to 30; the_committee moves 2 to 6 and stays pinned. Claim check
still passes, bestiary regenerated.
The only signal was a number being too small. A fifteen-second re-record is good
news, and good news is what nobody investigates.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
check-focus picks the pack size nearest an even fight. For the redcap that is
n=3 at 28.0% against n=2's 73.4% -- correct, and also where focus fire is worth
most. But the margin is 1.4 points and the scan is 400 runs: run across eight
seeds it picks 3 seven times and 2 once, and the recorded gain would move 6.3 to
4.9 with it.
R-276's explanation was wrong. Its table was measured against the CLEAN GROUND
cut, which it never named. Against the frozen party a lone redcap is worth 0.4
rather than 6.1, because that party wins 98.9% and nothing shows against a
ceiling. The power is worth most where the fight is in doubt -- 3.8 at n=2, 3.1
at n=3, nothing at either end. Not outnumbered. Undecided.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Asked to re-record the lethality baseline against a lone redcap, I read the tool
and found it already does: measure(party, [spec]), with a comment saying "Solo,
because a creature is the unit under test". My claim that it fights packs came
from memory of check-focus, whose redcap is n: 3.
The lethality figure was also not hiding the power. Wipe rate moved 0.7% to 0.9%
because one redcap cannot wipe four agents whatever it ignores -- that column is
at its floor. Agents down moved 0.55 to 0.74 of 4, a 35% relative increase, which
is the column I did not look at.
No baseline re-recorded: it is already solo and was re-recorded in R-275.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Played four agents against one redcap. Two dodges in one round, both at 75 --
DEFENCE_STEP is -30, so the second would have been 45 and the roll of 70 would
have failed. The wiring fires.
Measured with and without, 2000 x 3 seeds: the power costs the party 6.1 points
against a lone redcap and 0.1 against three of them. It is a rule about being
outnumbered -- a lone defender spends four defences a round, a pack spends one
each -- which is why check-lethality moved only 0.7% to 0.9%. The baseline fights
redcaps in a pack, the configuration where the power is worth nothing, so that
figure is the floor rather than the effect.
Not presented as evidence: at seed 3 the party wins with the power on and is wiped
with it off, which is stream divergence rather than direction.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
41 statblocks carry a POWER in their tactics and simulate.mjs read none of them,
so a redcap that ignores the cumulative defence penalty has been measured as a
creature that tires -- in check-lethality, in check-focus, and in every figure
published about it. The defect was not that the powers were unimplemented, it was
that nothing said they were not.
powers.mjs classifies all 41: 2 wired, 14 notSimulable with a stated reason, 25
not fight rules. check-powers refuses an unclassified POWER and refuses a
notSimulable without a reason -- and it does not test that the harness imports a
power, it fights the creature with and without and requires the two to disagree.
Moved: the courier 9.4% to 1.0% wiped (it attacks at half while carrying), the
redcap 0.7% to 0.9% (small, because these fights rarely spend a second defence).
ARGENT AND GULES was wired and then un-wired: it tripled the supporter's wipe rate
to 75.2% because the harness has no ground and applied the borough-ground condition
unconditionally. Same reason THE PULL is not wired. I had wired one and refused the
other on identical facts.
Lethality and focus re-recorded, bestiary regenerated.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Played the hollow man against the cut. THE FILE IS YOURS NOW never fires --
simulate.mjs reads statblocks and weapons and has no concept of tactics, so
R-273 closed a gap in rules.mjs and left the same gap one layer out.
Resolved by hand it behaves: every roll pair enumerated for both beats. The
tie-break carries it -- at POWx5 100 against 60 the aggressor still only takes
them 48% of the time, because equal bands go to whoever is being acted upon.
Act Four prices THE OFFER: accepting it moves Ashcroft from the safest person in
the room to the least safe, 21.9% to 48.7%, past Braithwaite's 37.6%. And that
once-only row no-ops 14.5% of the time against him, which Act Three can absorb
and a permanent countdown beat may not.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The content has called for 'opposed POWx5' since before the rule existed -- the
hollow man's tactics, CLEAN GROUND twice, slice-text three times -- and rules.mjs
defined none. A scenario carried a local ruling that said in its own text it was a
ruling and not a rule.
Generalises that ruling rather than inventing another: both sides roll, the better
band wins, only the ladder the game already has. Ties go to whoever is being acted
upon, which is what defenceOutcomeFor has always said; the scenario's 'favour the
agent' gave the same answer only because no agent ever initiates one. Neither side
succeeding leaves the contest unsettled rather than won, which the two beats need
in opposite directions.
Two exports at the scenario session's request: opposedOutcomeFor compares graded
levels and carries both, so a caller can price a fumbled attempt without this file
deciding what a fumble costs; opposedContestFor runs it from ratings and rolls with
per-side difficulty, so 'resists at Difficult' does not put applyDifficulty back
into a document. Spot-checked against the real Act Four beat: 55 against 85 at
Difficult, which is 42.
Page 1 states it by asking it -- the tie-break and the margin are computed from the
rule at build time, so the book cannot drift from the engine.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three things Tim asked for together: measure the long end of a fight, make the
swing check bite on stale prose, and record in rules.mjs the coincidence that
hid an invented rule from two readers.
GUARD 15 — check-fight-tail, on tools/fight-tail.mjs.
Fourteen guards measured this scenario's combat and none could see a long
fight, because every one of them averages. The cap is the measurement here, so
the tool passes its own (400, not runFight's default 40) and --update refuses
to record anything reaching it. Verified by lowering it back to 40: the cut
loses 20 fights to the ceiling, the six-a-side 117, 107 of those ending
neither won nor wiped, and the recorder stops.
Two claims, both read from a fresh measurement rather than the baseline, so
--update cannot silence them. "The long end is twice the median" was rejected
as a claim because it is true of both encounters and so distinguishes nothing.
Instead: a fight past 15 rounds wipes the party materially more often than a
short one, and the SIX-A-SIDE fight is the longer one (median 14 vs 10.7) --
R-270 showing up as duration, since a disabled fighter keeps fighting 30
points down. The cut is shorter because it is decisive, not safer.
GUARD 16 — check-cited, on tools/check-cited.mjs.
check-firstblood and check-attackers catch the game changing; neither reads
the document. Re-record after a re-cast and the artifact updates, the guard
goes green, and the prose keeps printing the old number under a citation
saying where the new one lives. So citations are now machine-readable --
**40**<!-- cite: first-blood cut.swing --> -- and resolved on every build. 25
of them. It failed three times on its first runs, all real: a config keyed
"column" that the prose called "line", two figures rounded 32.3 -> 32, and a
vacuous pass on zero citations, now fatal in its own right.
It also refuses citation of unstable fields. fight-tail.longest may not reach
prose: same party, same seeds, same runs, and renaming a config moved it 71 ->
90 rounds, because seedFor derives the stream from the id. Across seven
labels -- median spread 0, p95 1, p99 3, longest 21. A sample maximum reads
like a bound and is a property of the label. EXPOSURE states p99 instead.
Same discipline on the deadlier ratio: 2.42 with seed spread 1.1, so the
document gives its direction and declines to quote its size.
RULES.MJS — one comment, no rule change.
Over resolveLocationHit: its two thresholds are unrelated and usually agree.
disabled is a fraction of the pool per location; majorWound is ceil(hp/2) and
feeds only dyingLimitFor; neither removes anyone from a fight, which is
conditionFor at 2 hit points or a destroyed head. At 10 hp, leg/abdomen/chest
capacity is 5 and majorWoundFor is 5, and those locations take 12 of 20 melee
results and 15 of 20 ranged -- so two readers reconstructed a rule that does
not exist, checked it against the log, and were confirmed by it. The note
says to test an arm, the only place the difference shows.
Guards verified to bite, not assumed: drift, re-cast, censoring, the longest
refusal, the rounding catch and the vacuous-pass catch were each forced and
each failed the build with the right guidance, then restored.
npm run check: 16 guards, exit 0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Fourteen guards all average, so none can see how long a fight might run. 6000
fights per configuration: the cut runs a median of 11 and a 95th of 22, the
six-a-side line a median of 14 and a 95th of 32, and the longest seen are 71 and
79. The line is the LONGER fight, because more bodies means more of them fighting
on at reduced skill rather than dropping. Length predicts death: 61.6% of cut
fights past fifteen rounds are wipes against 42.0% of shorter ones.
maxRounds = 40 is the default every guard runs at and a fight reaching it is
scored as neither wipe nor win -- censoring 0.2% of cut fights and 1.87% of the
line's. Measured before proposing anything: uncensored, the swing moves 39.8 to
39.9 and 14.8 to 15.1, against a recorded noise of 2.1. The cap stays, documented
rather than corrected, because the cure is four re-recorded baselines.
Also R-270 addendum: the hit table is weighted toward the locations where the two
thresholds coincide -- 12 of 20 melee results, 15 of 20 ranged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
R-266 said a fighter goes down at half maximum hit points and called it one rule
in both directions. conditionFor puts someone down at hp <= 2; majorWoundFor is
ceil(hp/2) and feeds only how long the dying last; a location is disabled by
resolveLocationHit at its own locationMaxHp capacity.
The reason it survived reading: for a 10-point neighbour majorWoundFor is 5 and a
leg, abdomen or chest holds exactly 5, and for a 12-point agent both are 6. On the
locations that get hit most the invented rule returns the real one's answer, and
the narration prints MAJOR WOUND and disabled on the same line. Arms and heads are
where they part, and I had not looked at an arm.
Found by the scenario session going to rules.mjs to verify a different correction
of mine and reading the next function along. Both errors made the fight look
easier than it is.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Seed 16 is the missing corner -- party takes first blood in round 1 and is wiped
anyway -- and it runs 24 rounds. Its last five are Neil dropped by a critical
through armour, then Dominic alone at 1%, failing three times, then dying. That
is what three effective attackers costs at a table when a fight goes long.
Corrects a mechanism I had written twice: a disabling hit does not remove an
attacker, it charges 30 points off physical or manipulation per locationEffectsFor.
Bhattacharya keeps swinging at 10 for four rounds. A slope, not a cliff.
Neither guard fired and neither was wrong: the fight sits in the 25.5% the swing
does not cover, and no guard measures the tail because all of them average.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CLEAN GROUND prints that the four-player cut is three effective attackers and
that Ashcroft is not one. Nothing checked it, and a GM reads it aloud to decide
who a real player spends four hours being.
effective-attackers.mjs measures per attack, not per fight: counted per fight
Braithwaite leads on disables, but only because his armour buys him a third more
swings -- per attack he is the weakest of the three. Both rates are recorded and
only the per-attack one is reasoned from. Asserts exact drift, then the sentence:
three clear 5% of attacks disabling, one does not, and that one is Ashcroft.
Re-recording does not silence the claim check; verified in a worktree.
declared-cast.mjs holds the cast marker reading both scenario guards need, rather
than a copy in each. It was briefly named scenario-cast.mjs, which check-scenarios
sweeps into the scenario corpus -- its own example marker was read as a real cast.
update-readme dropped any guard that exited non-zero, so check-rollable vanished
from the README and the count word fell to thirteen while fourteen guards ran. A
missing line is now fatal.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
I told the scenario session that this guard turns a silent re-cast into a build
failure, and it accepted the coupling on that basis. The cast was hardcoded, so
a re-cast would have left it measuring the old six and reporting success.
Reads the <!-- cast: --> marker check-rollable established, drops the two the
scaling note drops by name rather than by position, and refuses to run if the
marker is gone. Verified by re-casting CLEAN GROUND in a throwaway worktree and
watching the guard name the change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Another session verified R-266's measurement, agreed the frame was better than
its own, and declined to publish it because the figure had no guarded lineage a
doubter could re-derive. It was right, and the fix costs 0.7s of build time --
which is the number I should have checked before calling it a reading tool and
not a guard.
first-blood.mjs gains --update/--check and is now both the reader and the
measurement of record; check-firstblood.mjs is a thin wrapper over it, the same
shape as check-bestiary. Compares the recorded figures exactly, then asserts only
what a page would claim: first blood lands within 5 points of even in the cut,
and its swing exceeds the six-a-side line's by more than both noises.
Recording it moved the swing from the scratch run's 43 to 40 against a
seed-to-seed spread of 2.1 -- high by more than its own noise, which is the
argument in miniature. update-readme's COUNT_WORD could not reach thirteen.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Reading both ends of the four-player cut — seed 5's sweep and seed 2's wipe —
says the fight is not decided by round five but by whoever lands the first
disabling hit, on average in round two. Both sides disable on one good blow, so
each one thins the return fire and makes the next likelier; nothing pulls a
fight back toward the middle.
tools/first-blood.mjs measures it: 49.7/50.3 on who strikes first, and 23.1% vs
66.5% wipes on either side of that. Six against six is the control at 8.6% vs
22.4%. A reading tool, not a guard — it records nothing, and no figure from it
goes on a generated page.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The round-five race line read "zero in 2000 fights", which invites the reading
that no fight in the run was a wipe. It is zero out of the 84 that reach round
five with two of the three down — a small bucket, and the sentence should say so
before the figure goes to print in a scenario.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
tools/playthrough.mjs exists to show a reader why a measured number is what it is. I cited
it to another session — read seed 2, see the wipe the 45% row is made of — and the
reproduction failed in front of them. They reported different outcomes at both seeds and
guessed the cause correctly from outside: the two tools were not drawing from the same
stream.
They were not. simulate.mjs used a private mulberry32 makeRng; playthrough.mjs had its own
LCG written to look like it. Both deterministic, both reproducible alone, and "seed 2"
named a different fight in each — which breaks the only thing the tool is for. Its own
comment claimed a seed here names the same fight there. check-focus carried a third copy of
that LCG, so the two guards described the same game with different dice.
makeRng is exported and both files use it. A playthrough seed is now exactly the first
fight of simulate.mjs --seed <n>: --runs 1 --seed 2 and the playthrough give 9 rounds, 4 of
4 down, 1 dead, both.
The other half was my citation rather than the code: the command I sent omitted --mode, so
it plays both targeting arms and prints two fights. They read the last line, I quoted the
first. The summary line now names the arm.
check-focus re-recorded under the shared stream; figures move a point or two. What it buys
is that the bimodality analysis now reproduces the published means exactly — 2.31, 3.03,
0.40, 0.42 against the four rows CLEAN GROUND publishes. Under the old LCG it agreed to
within a decimal, which looked like corroboration and was two experiments landing near each
other.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Reading a narrated supporter fight showed it being hit on armR and head. It has wings; it
has never had arms. Three defects, all in the pipeline every published number comes from.
1. It located hits by species, not body plan — simulate.mjs read defender.species where
the game writes spec.bodyPlan ?? spec.species ?? "baseline" into speciesProfile
(build-packs.mjs:294). The barghest, kelpie and church grim were fought on two legs
with arms, and the supporter with no wings, so the fight its tactics call the one the
agents can win could not occur in a measured fight.
2. Armour was one scalar for the whole creature, and the harness's locations had no
armour field. Bare wings, the grounded halving and the vital exemption — R-254 and
R-255 — were invisible to every number. Worn armour was summed the same way, which put
a stab vest on a cleaner's head.
3. Nothing was ever grounded: a wing could be ruined and the creature kept flying.
Fixed in the game's order — locate, then apply what that location carries, every term
imported from rules.mjs. A ruined wing calls groundedPlanFor and the wounds carry across
by severity through remapLocationDamage, the function _preUpdate uses.
A bug of mine no guard would have caught: remapLocationDamage returns { damage, moved,
rescaled } and my first draft passed the whole object where a damage map was expected, so
every wound a creature carried was forgiven the moment it came down. check-lethality would
have passed it — fewer wounds means a longer fight, which reads as a number moving, and
this commit moves numbers. Found by probing a landing by hand.
20 of 47 creatures moved, 18 deadlier and 2 less. The supporter goes 69.3% -> 25.1% wiped,
3.38 -> 2.16 down, second deadliest to fourth: it was being measured as a 30-hit-point
creature in uniform armour 9 that could not be grounded. The small rises elsewhere are the
party's armour no longer covering locations it never protected.
check-focus then caught the page overclaiming, which is what it is for: focus fire still
helps in all 30 but only 28 clear their own noise where 31 of 31 did. The guard was
asserting more than the page needs — the bestiary prints that count from the artifact and
cannot overstate it — so it now checks only what the page asserts outright, and the page
rewrote itself to "in 28 of them".
Both baselines re-recorded. Minor version, not a patch: the published numbers changed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every guard here reports a number per creature, and a number cannot say why a fight went
the way it did. That gap is what produced R-258 through R-260: each written from a side
harness built to watch a fight, each modelling something slightly different from the game,
two of the three wrong in ways no guard could catch because no guard was involved.
runFight takes an optional `say` sink. It reports what the simulator already decided — the
attack roll and its target, a defence and the penalty it was made at, the landing level
after a dodge downgrades it, damage against armour before and after armourAgainst, the
location, major wounds, disablement, death. It never touches the generator, so a narrated
fight and a silent one are the same fight; check-lethality and check-focus both still
match their baselines exactly with the hook in place.
tools/playthrough.mjs (npm run play) is its consumer, in the same commit deliberately: a
hook with no reader is the exact defect this project keeps finding in its own rules, and
adding one to the measurement pipeline with only a scratchpad file calling it would have
been committing the thing I have spent the session removing. Pack size comes from
focus-baseline.json so the fight you read is the fight check-focus measures; creatures
recorded as pinned are played solo.
Three redcaps, seed 20260913, same seed both ways. Spread fire: wiped in 10 rounds, 3
dead, and all three redcaps still standing — 39 hit points spread three ways so that none
of it finished anything. Focus fire: same opening, diverging at one target choice in round
1, party wins in 20 with two up. The +13.2 points check-focus records, seen once instead
of averaged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>