The fold exported scanDocument for the string tests, and importing it also ran
the guard. check-behaviour imports that module, so a figure defect called
process.exit(1) inside check-behaviour: it reported check-figures' failure under
its own name having run zero of its 108 tests.
The build went red, which is why this was survivable, but it went red in the
wrong place and every behavioural test was silently not running while appearing
to. A guard that stops another guard from running, and cannot say so, is the
worst version of the fault this file exists to catch.
Body now sits behind import.meta.main, the idiom step5-split.mjs already uses.
Importing yields the four readers and nothing else. Tested by spawning a fresh
process, because the property is "importing has no effect" and a source
assertion would pass on a file that grew a second side effect elsewhere.
Reverting the check fails exactly one test. With a defect planted,
check-behaviour runs 108 green and check-figures fails in its own slot.
Mine, introduced by R-307. R-309.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
R-308. Read the fold as a stranger at its author's request. Two of the three
risks they flagged are sound: isProse drops nothing that carries a citation
(scanned the whole corpus), and the duration exclusion earns its place, since
"runs" now means three things in this corpus and only the cell can tell them
apart.
The third is real. numEnd used line.indexOf(m[1], m.index), which finds the
first copy of the digits at or after the match start rather than the copy that
was captured:
"The wipe rate of 74 in ten is 74%<!-- cite: ... -->."
value=74 numAt=17 marked=FALSE
A correctly cited figure reported bare, because numEnd lands mid-sentence and
the marker test reads " in ten is 74%...". Fixed with the d flag: m.indices[1]
gives the capture's real position and there is nothing to search for.
Narrow to reach — it needs the wipe-rate shape, the only one with a wide gap
before its capture, and an integer duplicate inside that gap; a decimal cannot
do it because [^.] will not span a decimal point. Three attempts failed for
that reason before the fourth worked, which is why this is recorded as narrow
rather than theoretical.
Worth fixing anyway because its direction is the bad one: a miss costs one
figure, a false positive on correct prose costs the guard.
Seven-case battery re-run, every restore byte-identical.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
check-unmarked is gone and check-figures holds all of it. Three readers, each
authoritative where the corpus gives it authority: a named phrasing anywhere,
a column header inside tables, and bold in prose only.
The boundary is the finding rather than the union. Bold is a publication mark
in prose and an emphasis mark in a table — nine measurement columns mix bold
with plain, all correctly cited, and in the mixed-force table the two bolded
rows are exactly the two the prose underneath singles out. So the bold reader
stays silent in a table and the header rules there alone. Neither guard could
have found this alone: each had half the evidence and read it as the other's bug.
Union of both word lists, because each had a gap the other covered — wipes and
runs. A proposal to drop runs? was made and withdrawn; it would have dropped the
p99 column, which is R-299's own defect committed a second time.
Exclusions test cells, never header words: prose, denominator, duration. No
threshold rule touches a header, or "Past 15 rounds" loses two cited figures.
Two integration defects, both caught by the ported tests: the readers stopped at
different ends of one figure so the dedupe missed it (identity is where the
number starts), and blanking prose cells for every reader dropped twelve real
figures out of THROUGH_TRAIN's ratchet (the prose rule belongs to the column
pass alone).
Kept from R-299: opt-in rule, ratchet and its leave-the-loose-set branch, the
UNCITEABLE check, NOT_A_MEASUREMENT, comments blanked not stripped. Added
waivers, with an empty reason fatal and every waiver printed on a green build.
21 guards, 107 behavioural tests, 139 figures — 115 marked, 6 waived.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
R-300 was right, and it was right with a live defect rather than an argument:
a bare **24** in a table whose own header row says "Median rounds", reported
OK by this guard, against a baseline holding 25. The scan read one line at a
time, so measurement status living two lines above was invisible.
Fixed by reading the header. A cell now inherits the measurement status of its
column — which is a claim the DOCUMENT makes, rather than one this file's
vocabulary has to anticipate. That is the narrow repair, and it is deliberately
a different mechanism from the shapes: the shapes guess at phrasing, the header
does not have to.
The general point in R-300 still stands and the file now says so where the
shapes are defined: a vocabulary learned from the marked figures cannot contain
the phrasing of the figure nobody marked. check-unmarked attacks that from the
other end, treating bold as the corpus's own mark of a published figure.
Two bugs found while testing, both mine, both caught before commit:
- A citation marker is full of digits and none of them are figures. "packs
line.hollow6.hurt" holds a 6; "fight-tail cut.over15" holds a 15. Scanning
raw cell text reported roughly 130 correctly-cited rows as unheld — the
best-marked tables in the corpus. Comments are now blanked rather than
removed, so every offset still points at the right character.
- "3.48 of 6" — the 6 is the party size the mean is out of, not a measurement.
Coverage goes from 25 figures, 7 marked, to 120 figures, 102 marked.
Proved again by breaking it, five ways, every restore byte-identical: the
historical bare **24** fails at its own line, a stripped prose marker fails, a
planted "median 24 rounds" fails, a new bare figure in an uncited document
trips the ratchet, and neither a threshold column header nor any of the 102
cited cells fires.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
check-cited resolves every citation marker against the artifact it names, and
has one structural blind spot: it can only resolve the markers that exist. A
measured figure written into prose with no marker beside it is not a failed
citation, it is not a citation at all, and nothing looks at it again.
Desk pass 11 proved it. "A median 24 rounds" sat in the Pacing Note and "24 if
the column joins" in EXPOSURE, against a baseline holding 21 at two of the
column joining and 25 at three, through every pass that checked citations.
This reads the figures instead of the markers. Narrow by design: an earlier
draft matched any number within 45 characters of a measurement word and found
282 candidates, nearly all prose ("down 140 steps", "an engineer on his
rounds", "01:06"). The shapes here are the phrasings the documents actually use
when quoting the simulator, each read off a figure that is cited somewhere.
Opt-in rule: a document that uses citations must mark every measurement figure.
One that cites nothing is held by a ratchet instead — turning three unguarded
scenarios red is how a guard gets switched off on the day it is written — and
joins the strict regime the moment it gains its first marker.
It found three things in CLEAN GROUND before it was wired in:
- 74.9% printed bare twice while cited correctly four times, and that is the
figure that was published at 58.7% until the truncated sweep was found.
- "the longest fight is 120 rounds" — 120 is exactly packs cut.hollow6.longest,
the field check-cited REFUSES to let anybody cite because a sample maximum
moves by a third on a re-seed. Refusing the citation while printing the number
left the least stable figure in the suite as the only one nothing held. The
sentence now leans on the guarantee that is actually strong: check-packs
refuses to record a sweep in which anything reached the cap.
Proved by breaking it, four ways, with byte-identical restores: a stripped
marker fails, a planted "median 24 rounds" fails at its own line, a new bare
figure in an uncited document trips the ratchet, and a threshold column header
("Past 15 rounds") does not fire — that last one was a real false positive in
the first draft, which tested the matched text rather than its context.
v0.23. 21 guards.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>