137ab825415803c0ea11186022833272f8ff2db6
189
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
137ab82541 |
R-297: place the step-5 citation markers — 32 of them, no prose touched
Every figure in CLEAN GROUND's step-5 material now resolves against tools/step5-split.mjs, which enumerates rather than stores: the countdown rows, the two-case table and its blend, the Act Four reasoning, the counterfactual pair and the POW×5 targets. No prose changed -- stripping every HTML comment from the result diffs byte-identical against HEAD, which is the check worth having when editing another session's document. Proved live: raising Okonkwo's POW to 13 in a worktree makes Braithwaite the lowest in the room and eleven of the thirty-two markers go red at once, each naming the path and saying nothing can be re-recorded. The three figures that went bad in this document were all in prose rather than tables, which is why the prose is marked and not just the tables. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
7afd5b86cc |
R-296: the review log entry for the pack tables
The work landed in
|
||
|
|
47f2df9446 |
v0.20.1: "nearly quadruples" was wrong by three times, in the last two places
Desk pass 9's first fix, applied. The countdown's unsettled row and the Act Four paragraph both compared 14.5% against 3.8%, and the 3.8% is the rate against Ashcroft at his undiminished 85 — a contest that never happens, because if he refuses the thing reaches for the lowest POW in the room. Against what a table actually rolls: 11.3% at six players (Okonkwo), 10.0% in the four-player cut (Braithwaite), 14.5% if he accepted. The factor is 1.29, not 3.9. Both places now name who the target actually is, and the paragraph carries the correction inline so the claim cannot be re-derived from the old framing. What is NOT wrong, and the text says so: the trade THE OFFER describes is real and the taken rate genuinely moves 40.7% to 48.7%. One consequence was overstated, by three times. Written in v0.15, repeated in v0.16, and survived v0.19.1 correcting the identical error in the table beside it. Flagged again by the peer session while mapping citation paths, which is the third time this figure has been caught by somebody reading it rather than by anything checking it — and the argument for the derived-source markers now going onto it. npm run check: 20 guards pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c2f7249d4d |
R-295: derive the counterfactual too, and name it hadHeRefused
Matching the step-5 prose against the derived source before placing markers left five figures unmatched: 21.9, 74.3 and 3.8, twice each. They are the thing against Ashcroft's undiminished 85 -- the contest that never happens, which Act Four prints on purpose to price what THE OFFER sold. Derived now as hadHeRefused, so the counterfactual is held to the rule like everything else and carries a name that cannot be mistaken for an event. Quoting those three as odds a GM meets is what went wrong in two sentences and a table. Markers not placed: CLEAN_GROUND.md is c0's and they are in it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
a5c4378521 |
R-294: EXPOSURE's tables measured into an artifact, and 58.7% was 74.9%
Desk pass 9 found the scenario's most consequential table held by nothing: ~30 sampled figures, no baseline, catchable only by re-running the exact command. lethality-baseline.json does not cover them — it measures creatures SOLO against a frozen four-agent party that is not this cast. Building the measurer found the figures were also wrong. simulate.mjs's CLI calls runFight with no options, so every published figure was measured at the default 40-round ceiling — and runFight does not report truncation, it scores whoever is standing when the loop stops. Measured at 400: line.column6 14.8% -> 15.2% 34/2000 truncated line.hollow6col2 7.8% -> 8.3% 129 line.hollow6col3 28.7% -> 30.1% 264 cut.hollow6 58.7% -> 74.9% 578 <- 29% of runs never finished Fourteen points on the single most alarming number in the case, and the one v0.16 added specifically to warn four-player tables. fight-tail learned this in R-270 and carries an assertUncensored; EXPOSURE's own tables never got one. The longest fight at cap 400 is 120 rounds and 400 vs 2000 are identical, so the cap is comfortable rather than merely sufficient. tools/pack-tables.mjs measures all 17 configs against the DERIVED cast via castAndCut, using measure()'s exact discipline — one rng threaded through every run, not a reseed per run, because reseeding is a different stream and would not reproduce the published table. It refuses to report or record a truncated sweep. tools/check-packs.mjs holds the baseline against the game, so the pair is not a loop: check-cited holds the prose against the record, this holds the record against the harness. Hard claims read `now` and never `base`: nothing truncated, the party equals the declared cast, config floor, and more of the same creature may not make the party safer. Figures compare exactly, since the runs are deterministic. Verified by breaking it: a drifted baseline, CAP lowered to 40, and a removed config each turn it red, and --update refuses outright rather than recording a truncated sweep. The removed-config test first passed for a bad reason — MIN_CONFIGS is a floor and 16 clears it — so a dropped-config check was added and re-tested on a row in no monotonic chain. All files restored byte-identical after each probe. EXPOSURE's three tables and the nine prose figures around them are rebuilt from the artifact with citation markers; the multi-seed stability claims were re-measured too (the cut is 74.9/72.7/75.0/73.5 across four seeds, not 58.7/60.5/60.3/59.8). CLEAN GROUND v0.20. npm run check: 20 guards pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
fb6e2cf321 |
R-294: check-cited resolves derived figures against the rule, not a stored copy
c0 argued a baseline for the step-5 split would be a cache of the rule and a guard
over it would mostly assert that arithmetic has not changed. Right objection,
wrong conclusion: do not store it. ARTIFACTS now takes a { derive } entry as well
as a file path -- enumerated on this build, nothing stored, and no --update able
to silence a real disagreement between the document and the game.
tools/step5-split.mjs enumerates all 10,000 pairs through opposedContestFor with
its targets derived: 55 and 60 are the lowest POWx5 in the six and in the cut, 42
is applyDifficulty(85, "difficult"). Two different rules produce those three
numbers -- the accepted row is a named exception, not the lowest of anything -- and
a test fails if anyone unifies them. A tie in "the lowest POW in the room" is
fatal rather than silently resolved; it fired for real in testing.
Proved four ways, including raising Okonkwo's POW to 13: the lowest moves to
Braithwaite, refused.held goes 48.1 to 52.4, and the citation that was correct a
moment earlier fails. The figure follows the rule.
Landed unused -- the markers are c0's to place in their own file.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
90cb251358 |
Pass 9 post-pass: the uncited surface counted, and it is not where I said
Fix 2 asked for the audit; here it is, and it reorders the fix list.
101 percentages in CLEAN_GROUND.md, 67 of them computed statistics rather
than skill ratings, 14 cited, 53 held by nothing. They split two ways and the
kind finding 1 was about is the safer one: a deterministic figure that drifts
can be caught by anyone who re-derives it in a second.
The exposed surface is EXPOSURE's two pack tables and the 58.7% four-player
figure — 2000-run samples against the declared cast, catchable only by
re-running the exact command, and behind no artifact at all. I had assumed
lethality-baseline.json covered them. It does not: it measures each creature
SOLO against the frozen party holloway/okonkwo/nkemdirim/ferriby, a different
four agents, storing {wipe, down, rounds}. The document's tables read "of 6"
and fight packs of three, six and ten. None of 3.48, 4.79, 5.97, 0.42, 3.03
or 58.7 appears in any baseline in tools/.
Recorded as a decision: derived figures want to resolve against the rule at
check time, since a stored baseline for them is a cache of arithmetic;
sampled figures want a baseline, since re-running them is expensive and
noisy. Same problem, different mechanisms.
Census and the baseline's shape counted here rather than taken from the peer
session that raised the gap.
npm run check: 19 guards pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
15c2e93b5f |
Desk playtest 9: the regression pass
An audit, not a session. Five versions of fixes went in today, all by one author, and the last three passes each found a defect introduced by the fix for the previous one. Every fix from passes 4-8 re-checked against the document: did it land, is it still true, did it break a neighbour. 1. "THE OFFER nearly quadruples the unsettled rate" is false as a statement about play. It compares 14.5% against 3.8%, and the 3.8% is the same counterfactual v0.19.1 removed from the table one commit ago: Ashcroft at his full 85 is never a target, because if he refuses the countdown reaches for Okonkwo. The real comparison is 14.5% against 11.3% at six players and 10.0% in the cut — a factor of 1.29, not 3.9. Printed in two places, both of which v0.19.1 walked past while correcting the identical error beside them. The trade the document describes is real and the taken rate really does move 40.7% -> 48.7%; only the size of that one consequence is wrong. 2. All 27 fixes from passes 4-8 are present, which is not the reassurance it sounds like. Nothing has gone missing; three of the four defects the last three passes found were created or preserved BY a fix, and a presence check cannot see any of them. What would have caught finding 1 is the thing check-cited does for figures backed by an artifact — and the 3.8% is enumerated from the rule, so nothing holds it. Every uncited number in the document is a number nothing is holding. 3. The understudy carries a dagger (1d4+2) and has no knife skill, so it swings at the 1% floor and the harness correctly picks its punch — every EXPOSURE figure is right. But EXPOSURE's load-bearing first lesson says flatly "it is a 1d3 punch", and a GM who reads the sheet sees a knife and one combat skill. Measured both ways: arming the knife at brawl 60 moves 0.55 hurt to 0.73 and 0.00 deaths to 0.01, still 0.0% wiped, still two rounds. The thesis survives; it is a documentation gap and is reported at that size. content.mjs restored byte-identical after the test. Confirmation session, seed 2260: step 5 landed on "they hold" — the branch v0.19 wrote one commit ago — in the very next session, and by the route that makes the case for it, both sides succeeding with the tie going to the person being acted upon. Ashcroft refused a fourth time (63%, a 16% run) and the swap failed a fourth time (roughly a coin flip, so 1 in 16); both recorded so neither is read as a pattern. Four fixes listed, unapplied. npm run check: 19 guards pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d68e394fa4 |
v0.19.1: the 74.3% is a counterfactual, not an event
Desk pass 8's finding 1 listed four targets for countdown step 5 and one of them cannot happen. If Ashcroft REFUSES THE OFFER the thing does not reach for him — it reaches for the lowest POW in the room, which is Okonkwo. The 21.9% / 74.3% pair lives in Act Four to price what accepting sold: what it would cost him if it tried. No table ever rolls it. v0.19 carried that straight into the new table under a heading reading "It reaches for", which is the one place it is unambiguously wrong. Rebuilt around the two cases that actually occur, with the split Act Two decides: Ashcroft refused 63% of games -> Okonkwo 55 taken 40.7 hold 48.1 uns 11.3 Ashcroft accepted 37% -> Ashcroft 42 taken 48.7 hold 36.8 uns 14.5 across all games taken 43.6 hold 43.9 uns 12.5 63/37 is the Insight 63 itself: he refuses on anything that is not a failure or a fumble. The warning is now inline in the table rather than left for the reader to derive. The finding is unharmed and the fix was right — the held outcome is still the likeliest single result and was still unwritten. What was wrong was framing 74.3% as a number a GM meets. It propagated before it was caught: the peer session read pass 8 and replied that 74.3% "is the number a GM meets, not the 48.1%", repeating the error out of my own fix list. That is what a wrong number in a fix list does, and it is the argument for correcting the record rather than only the scenario. Pass 8's fix list is amended and carries a post-pass. npm run check: 19 guards pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
daf77f4aac |
R-293: a desk playtest is a record, not a scenario with rolls in it
Counting the readers that had broken on their own format turned up one nothing had caught. docs/scenarios holds scenarios, eight playtest records and two art prompt sheets, and every guard treated all three as scenarios -- harmless for citations and skill spellings, false for reachability. Six quotations across passes 4, 6 and 7 were checked as live rolls, so a record of a session already played could fail the build over a skill nobody can reach. Planting Science (Physics) in pass 4 fails before the split and passes after; the same skill in CLEAN_GROUND still fails. Classified by the document's own H1, not its filename, because tools/scenario-* naming is what swept a tools file into this corpus in R-268. An unclassified document is fatal: an allowlist that silently drops what it does not recognise would take a new scenario out of reachability checking on the day it was written. 89 rolls across 18 files becomes 83 across 8. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
6171d1e9b9 |
CLEAN GROUND v0.19 — the outcome where somebody holds
Applies desk pass 8's five fixes, and adopts R-292's TAG_OPEN. 1. Step 5 now has all three of its outcomes. opposedContestFor returns taken, held and unsettled; the document printed the first and the third. "They hold" is 48.1% against the default target, 52.4% against Braithwaite and 74.3% against an Ashcroft who refused — which v0.17 established is what most tables have, so the middle column is the one they will play. It gets its own countdown row, its own table, and read-aloud of its own, and the COMES TO NOTHING paragraph is now explicitly the unsettled case so it can stop being the nearest text to an outcome it is not about. The reason they held is the case's own argument arriving as good news: a person with a life on the record is a difficult document to overwrite. Every figure re-enumerated over all 10,000 roll pairs before writing. 2. EXPOSURE's five things are six, in order, under a heading that says six. Introduced in my own v0.16 and survived two versions; asserting a unique match protects against editing the wrong text and not against inserting in the wrong place. 3. EXPOSURE opens with a one-minute box. The body is untouched — nothing in it is padding — but 3,176 words is sixteen minutes about the encounter the document exists to prevent, and a GM with thirty minutes of prep now has somewhere to stop. 4. The stale open question is closed. Braithwaite has not been on the critical path since v0.14 un-gated the tell, and pass 6 ran the act with the Xenology and the Psychology both failed. 5. Act Four's Psychology has a special: which file it thinks it is, and how recently it read it. It corrects them on Prichard's service history — it is not remembering, it is citing. Coverage 6 -> 7 specials, so the cited figure moved 85% -> 83.7% and check-cited caught the prose before I did, which is what it is for. R-292 adopted: outcome-coverage now composes its body onto check-scenarios' TAG_OPEN through tagRe, so the two readers cannot drift on what starts a tag while keeping the different bodies they need — and tagRe's fresh matcher per call avoids the shared-lastIndex defect, this file being the second caller that would have found it. Verified identical across all seven spellings. npm run check: 19 guards pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
df38be794b |
Desk playtest 8: the cold read, and the outcome nobody wrote
Two halves: the document taken as a GM meeting it thirty minutes before a
slot, measured rather than impressioned; and an honest session, six players,
seed 5108, nothing forced. Three versions of fixes had landed since anything
was played, every one written by somebody who already knew the document.
1. The climactic roll has three outcomes and the document describes two.
opposedContestFor returns taken, resisted outright, and unsettled. The
taken and unsettled figures are printed and both reproduce exactly, so the
source has always been right. Resisted outright appears nowhere — 48.1%
against the default target, 74.3% against an Ashcroft who refused, which
v0.17 established is what most tables have. It is the single most likely
result of the scenario's climax. Worse than an omission: the paragraph
below is headed WHEN STEP 5 COMES TO NOTHING, is entirely about unsettled,
and carries the beat's only read-aloud — so a GM whose agent simply won
finds text written for a different outcome. Enumerated over all 10,000
roll pairs; no seed.
2. EXPOSURE's "five things" are six and run 1, 2, 3, 4, 6, 5. Introduced in
my own v0.16 (
|
||
|
|
3a31e0c9b0 |
R-292: share where a tag starts, not what it says
c0 asked whether tagsIn should return indices so beatLikeIn could take spans from it. No -- that widens this file's signature to serve another, and R-291 already fails when the two disagree. But detection is not prevention, and there is a third thing to share. The two patterns differ in their bodies for good reasons: tagsIn captures the whole tag for skillsIn to split, outcome-coverage stops at the em-dash because the outcome follows. What was copied into both files, and what drifted, is the opening -- the literal \[CUS:, widened by R-290 here and left behind there. TAG_OPEN is now exported as a string with tagRe(body) composing a fresh matcher onto it; fresh because a shared /g regex carries lastIndex between callers. 22 tags before and after, c0's body composed onto TAG_OPEN gives the same 22 its own pattern does, 19 guards green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
8a051f7f2b |
R-291 follow-up: one tag pattern, so the two readers cannot disagree
R-291 measured check-scenarios' tagsIn against this file's readers and found them apart on three spellings. tagsIn was widened for case and spacing around the colon; outcome-coverage's strict half still carried the case-sensitive literal, so `[cus: Spot — x]` was READ by one guard and reported unparseable by the other. The build failed, which is the safe direction, but it failed saying the tag could not be read while another guard had just read it. beatsIn and beatLikeIn now share one TAG pattern, widened to agree with tagsIn. Measured across seven spellings: the five both strict readers accept now agree in all three, and `[CUS Spot]` and `[CUS= Spot]` are still refused by both and still named by the loose counter, so the canary keeps its teeth. Verified on the real document that this is a spelling fix and nothing else — beats 21, quoted 1, unparsed 0, stated 18/6/6, session figures unchanged, so the baseline is untouched. A lowercase tag planted in the acts is now read by check-outcomes, check-scenarios and check-rollable alike; the file was restored byte-identical after the test. R-291's own assertion still passes. Found by the peer session measuring my file rather than trusting it, which is the fourth reader this session to be caught describing or reading its own format wrongly. npm run check: 19 guards pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ed851b76e3 |
R-291: assert that a tag one reader accepts stays visible to the other
c0 asked whether R-290's widening of tagsIn reaches past outcome-coverage's loose counter, which would zero its unparsed count for the wrong reason and retire a canary. It does not -- 22 beats from each reader on CLEAN_GROUND.md, nothing that tagsIn accepts invisible to the other guard -- but "does not today" decays quietly, so it is now a test importing both real functions rather than a copy of either pattern. Mutation-checked: widening tagsIn to accept [CUS Spot] turns it red. Measuring it found something the question did not ask about, recorded in the log: for [cus: ...], [CUS : ...] and [ CUS: ...] the two guards disagree -- tagsIn reads the tag, beatLikeIn calls it unparseable -- because beatsIn's strict half kept the case-sensitive literal. The build stops, which is the safe direction, but names the wrong thing. outcome-coverage.mjs is c0's file and the call is theirs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ae8179bd9b |
R-290: the same spelling gap in tagsIn and the POWER marker
Swept every reader that parses a human-written marker. Two more had R-289's fail-open. tagsIn matched a literal [CUS:, so [cus: Spot], [Cus: Spot], [CUS : Spot] and [ CUS: Spot] all read as nothing -- and it is the corpus reader behind check-scenarios' skill validation and check-rollable's reachability, so a beat spelled any of those four ways was checked by neither while both printed OK. The corpus contains no such tag today; the 88-to-89 roll count during this work was c0 writing v0.18, verified against HEAD's reader on the same tree. The POWER: marker had it with nothing covering it. bestiary.mjs and check-powers' own scan both used the literal, so a creature added with "Power:" generates no entry line and is never reported unclassified -- check-powers passes green on it, measured on a clone with its dodge skills stripped so the dodger-count assertion could not fire instead. Reader now takes POWER\s*: and stays uppercase, because power: occurs in ordinary prose; an asymmetric counter names the rest, and was measured against content.mjs first (41 strict, zero loose outside them). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9c2ce9d905 |
R-290: check-outcomes — a beat that names a roll must say what it does
Desk pass 7's finding, turned into the nineteenth guard. check-rollable asks whether a skill is reachable at 25% by the declared cast; pass 5 found that is rollability, not competence; pass 7 found it is also not coverage. A beat could be reachable, well-rated and silent about every band but the one the GM improvises, and the whole suite stayed green. tools/outcome-coverage.mjs measures. Beats are located structurally — the bullet that owns the tag and its children, ending at the next bullet of the same or shallower indent or the next heading, because R-287 is what an unbounded scope does. Band widths are counted by grading all 100 results through gradeRoll, never by arithmetic on fumbleStart, which is off by one and was written wrongly twice this session before being caught. Ratings come from the declared cast via R-286's reader, best-in-cast per skill, because that is the die a table actually rolls. tools/check-outcomes.mjs holds it, in two kinds: RATCHET — stated failure/fumble/special counts may improve and may not regress; unwritten-band exposure may fall and may not rise. --update re-records these, because freezing them would make every improvement fail. HARD CLAIMS — read from what is measured NOW, never from the baseline, so --update cannot silence them: a floor of 18 beats (check-cited once passed with zero citations), no beat bare of every band, every scope structurally bounded and none over 5% of the file, no unparseable tag, no beat naming a skill the cast has no rating for. Verified by breaking it: a stripped beat, a broken tag regex, an unparseable tag and a removed failure case each turn it red; the file and baseline were restored byte-identical after each; and --update with a bare beat present re-records the baseline and still fails. Two bare beats found and filled while building it — the six at the back, and the Insight that is deliberately indistinguishable on a success and a miss, which now says so rather than saying nothing. Coverage 16/21 failure, 5 fumble, 4 special at the start; 18/21, 6 and 6 now. The three figures are cited against the artifact rather than typed, which was the peer session's condition and the right one: pass 7 hand-counted 22 beats where there are 21 (Act Three's warning QUOTES a beat, and a hand count reads the quotation as one — R-286 in the other direction), and its other three counts were stale within one commit. Its first catch was its author: the STATUS line announcing this guard contained a beat-shaped tag and was counted as a beat. The reader was not changed — prose shaped like a beat is what this file is for — and the error message now names the fenced block as the place to write an example. Minutes later check-cited refused a citation-shaped comment in the post-pass about citations. Three readers, three authors describing their own format inside it, each caught by a guard built for a different pass. CLEAN GROUND v0.18. npm run check: 19 guards pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2cf9721fa3 |
R-289: test the cast reader on strings, and fix the spelling that found
Both earlier cast defects were shown by editing CLEAN_GROUND.md in the shared tree, which is how this session came within a git checkout of c0's uncommitted work and is also the weaker test. Seven cases now live in check-behaviour as strings. Writing them found a live one: "<!-- cast : ... -->", one space before the colon, matched nothing -- not an empty declaration R-288 would refuse but no declaration at all, so check-rollable widened to the whole duty roster and printed its usual OK line. check-cited's three spellings again. The reader now takes cast\s*:. castLikeIn adds the asymmetric half: anything comment-shaped containing cast\w* the reader did not consume is named by check-rollable, so a spelling nobody anticipated fails the build instead of silently declaring nobody. Looser than the reader on purpose -- a false alarm costs a reword, the opposite error costs a guard that checks the wrong six people and says OK. Tests then mutation-checked for being load-bearing, each mutation asserted to have applied after a first pass where three silently did not. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e3d222cacf |
R-288: an empty cast declaration is fatal; a missing one still falls back
c0's point on R-286: the reader is unambiguous now, but the roster fallback that made the failure look like success is still reachable. Narrowly closed -- sixteen of the seventeen scenarios declare no cast and are rightly checked against the whole roster, so only a marker that is PRESENT and names nobody is refused, with its line named. Replacing the declaration with <!-- cast: --> passes before this change and fails after it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f3caaab04d |
R-287: assert the scope anchors in check-behaviour, and end them structurally
indexOf returns -1 for a name that is gone and slice(0, -1) is everything but one character, so a scope anchored on a moved name does not shrink or fail -- it becomes the file. Renaming the end anchor took one test's body from 3,184 characters to 30,825 with the assertion still passing. The other site windowed burstAttack at 4,000 characters over a function that runs 4,305. bodyOf asserts the anchor and ends at the next top-level function. A missing anchor now says which anchor and what would have happened, where the old code reported "burstAttack still calls rollWeaponDamage" -- a claim about a call when the truth was a claim about a name. Also records the rest of the sweep, including the guards deliberately left alone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
0453255ecc |
R-286: one cast declaration, and a mention of one is not a declaration
CLEAN_GROUND.md contains two things matching <!-- cast: ... -->: the real declaration, and the warning ten lines below it that quotes the marker inside backticks and matches with an empty capture. check-rollable and declared-cast both took .match(), first hit wins, so the arrangement has been correct only because the declaration comes first. Move that warning above the list and check-rollable reads a cast of nobody, falls back to ROSTER_BEST, and prints the same OK line having held every skill in the document to the full duty roster instead of the declared six. Instrumented and read off: "the duty roster" against "its declared cast of 6". castMarkersIn blanks code spans and fenced blocks before scanning (padded, so line numbers still point at the source) and refuses more than one surviving marker, naming every line. check-rollable imports it instead of carrying a second regex, so the guard that checks the cast and the guards that measure the fight cannot disagree about who is in it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ecbad45fc6 |
CLEAN GROUND v0.17 — the good rolls now buy something
Applies desk pass 7's five fixes. The pass found the document covering two bands out of five: seventeen failure cases written, four fumbles, three specials — against 90.8% of sessions containing a special or critical on a beat that says nothing, and 36.0% containing an unwritten fumble. 1. Specials written where they matter, starting with the two pass 7 rolled. A critical Persuade on Swinburne now buys her coming to the stump AND talking on the way — the dog went for something on the seam side in June and came back wrong, four months earlier than the case otherwise offers it. Previously a 1 bought what a plain success buys, which is worse than the fumble, and the fumble buys a dog. A special Xenology, a special Medicine (how long ago it stopped — the almanac's gap before the almanac) and a special Spot (the third set circles and goes back toward the seam) likewise. GM ESSENTIALS now states the shape the existing three share: portable, private or early, never more conversation. 2. The Act Three swap beat points at EXPOSURE. v0.16 named that beat as the trigger for the fight it measured and left the beat silent, with "survivable" as the last word before the decision — which is about the swap, not about what follows. 3. Every fumble branch prints its band. Three of four did not, and the one that did was the one written in v0.16. 4. The fumbled Medicine and fumbled Spot are written. Pass 7 hit both cold and lost ten minutes, because four beautifully specific fumble branches make their absence elsewhere read as an oversight rather than a licence. 5. Act Four says which branch its headline figures are for. Ashcroft refuses about five times in eight — passes 6 and 7 both did — and the 48.7%/14.5% pair is printed far more prominently than the 21.9%/3.8% one most tables will be in. Every band printed was re-derived from resolveBands/gradeRoll by enumeration rather than arithmetic, which is where the off-by-one lives. Pass 7's finding 3 overstated itself and is corrected in the same commit: the Close's fumble branch does name a skill, inheriting the Anomaly Lore tag from the bullet above it. Post-pass appended. npm run check: 18 guards pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4ae4e13196 |
Desk playtest 7: an honest session, and the three bands nobody wrote
Full session, six players, seed 7742, nothing forced. Pass 6 forced its triggers on purpose; this one takes what the dice give, because a forced pass proves a branch works and cannot say how often a GM meets one. It reached neither thing pass 6 asked for — no fight, and Ashcroft refused THE OFFER again. What it found is what both of those are instances of. 1. The document covers two bands out of five. Seventeen failure cases are written across the acts, which is the house rule honoured properly. Four beats state a fumble. Three state a special or a critical. Enumerated through gradeRoll against the cast's best rating for each of the 22 rollable beats: 42.1% of sessions contain a fumble and 36.0% contain one with no written case; 93.9% contain a special or better and 90.8% contain one on a beat that says nothing. Specials are not rare — 13% at 63, 11% at 53 — and the scenario asks for 22 rolls. This seed landed on four uncovered bands in thirteen rolls, three of them GOOD rolls: a critical Persuade on Swinburne that buys less than the fumble does, and a special Xenology with nothing above the success. 2. v0.16 wrote the trigger into EXPOSURE and never wrote EXPOSURE into the trigger. The Act Three swap beat is named in the new mixed-fight block as the moment that starts the fight, and it carries no pointer back; its last word before the decision is "survivable". Act Two has done this correctly for versions. Five locations across 1,100 lines to adjudicate one beat. 3. Three of the four fumble branches do not print their band, and the one that does is the one written in v0.16 — the fix landed where it was being thought about, not where the same need already existed. The Close's is worse: "on a fumble asking the question" names no skill at all. The v0.16 fumble clause earned itself immediately: 99 against Medicine 53 is a fumble, and v0.15's wording read it as a plain failure. Timing: ~3:35 against 3:40 at six players with no cuts taken. First honest pass to land inside the clock unaided. Two fumbles in thirteen rolls is a 4.4% event and is recorded as an illustration, not a rate. Five fixes listed, unapplied. npm run check: 18 guards pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
67e76639f3 |
R-285: scope the defence-stacking rules to the section that owns them
R-284 scoped the factored-rating rule to the creature's entry and left the defence-stacking rules matching anywhere on the page, so the exemption sentence and the "Across N creatures that spend defences" count were accepted wherever they happened to sit. Moving the exemption line out of its section into the redcap's statblock -- the page silent exactly where a GM reads the stripping advice -- passes at R-284 and fails here; confirmed by running HEAD's copy against the same tree. The owning slice is not always the creature's entry. A factored rating belongs to the creature; the ladder exemption is an answer to the paragraph it sits in and is generated into "## Shooting at something that moves", so scoping that rule to "### Redcap" would have failed a correct page. sliceOf now takes a heading at any level and each rule names the slice that owns its claim. A renamed section fails by name rather than scoping to nothing, which is the shape of every guard that passes because it found nothing to check. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
aa3ae24ef2 |
CLEAN GROUND v0.16 — the fight it was never measuring, and R-283
Applies desk pass 6's seven fixes. 1. EXPOSURE now measures the fight the table will have. Every row in the old table was one force fighting alone, and the six at the back are never alone — they walk inside forty-one refugees. Six hollow men wipe 0.0%; with two of the column joining, 7.8%; with three, 28.7%. Two is now the stated default, because two people out of forty-one losing their heads while their neighbours are shot is not a large number. 2. The four-player block has hollow-man rows: 0.3% at three, 58.7% at six, 99.9% with two of the column. Its only hollow-man figure before was the six-player 0.0%, so a GM running the cut was reading somebody else's table — for the encounter the party is likeliest to choose, because the six at the back are the only figures the scenario says are not people. GM ESSENTIALS item 3 carries the same correction. 3. Act Three's warning named a roll that does not exist. Its twenty-minute bomb hangs off "Anomaly Lore — what a peg is"; the depot entry is a Research roll. Pass 4's post-pass inherited the conflation from this warning and is corrected too. 4. GM ESSENTIALS states the real fumble band. fumbleStart is 101 - ceil((101-band)/20), tested before the 96-99 clause: 00 at 85, 99-00 at 63, 98-00 at 53, 97-00 at 40. Four times what "00 always fumbles" implies, in a case that rolls Spot 40 across two acts. Pass 4's correction was right at 63 by luck and would have been wrong at 53. 5. Sixth EXPOSURE lesson: a fight costs the session. Median 13 rounds on the printed row, 24 mixed, 30 at four players — sixty to ninety minutes in an act budgeted at fifty. The Pacing Note now says what to do when one starts, and not to absorb a fight and the peg fumble in the same act. 6. The fumbled Xenology is written. At 53 it fumbles on 98-00 and Braithwaite puts his name to a baseline human in front of everybody. The un-gated tell still arrives; it now costs the party its expert. 7. R-283: simulate.mjs described a mixed force as N of whichever spec came first, so six hollow men and six of the column printed as "12 x Hollow man" — two fights 88 points of wipe rate apart under one label. The composition was always right and the report was not. Uniform and --spread output are byte-identical, so nothing already published goes stale. Every figure re-run before writing rather than carried over. npm run check: 18 guards pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1a2ca44c52 |
R-284: report a wrong figure as a wrong figure, and read the right entry
R-283 left the intact-sentence-wrong-number case falling through to omission, so the guard said BESTIARY "does not state" a rating the page was stating. reads() now has a shape tier between strict and loose: the strict pattern with its value slot loosened, reporting which of the two numbers moved. Scoped the per-creature rules while adding it. They read the whole page, and each is the only rule of its kind today, so a page-wide match found the right line by luck; a second attackFactor creature would have let the courier's rule match that creature's sentence and report the courier correct. They now read the creature's own "### Name" entry -- proved by deleting the courier's line and planting an identical one under the redcap: still omission, where before it would have passed. Five discriminations plus the decoy, proved in a worktree with the message read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
327d060eb1 |
R-283: tell a reworded bestiary sentence apart from a missing one
R-282 matched the page with String.includes, so rewording the exemption sentence reported "BESTIARY never says it is off the stripping ladder" -- sending a maintainer after a sentence that is sitting right there, and never naming the real problem, which is a pattern that has silently stopped reading. Each textual rule now reads twice. Strict is the sentence as it stands and is tighter than before (the bold and the full stop, not the bare clause a substring accepted); loose is the same claim in any wording. Strict passes, loose-only is reported as a reword with the line quoted and the page presumed right, neither is the omission. The loose anchor was wrong on its first pass in the way that matters: "a sentence with 40% and 80%" also matched the courier's own statblock line, so deleting the sentence reported a reword and quoted the statblock back. It now excludes that generated marker, which makes it a test of the claim and not of the digits, and degrades to omission rather than to a false reword. Four discriminations proved in a worktree with the message read in each. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e70d25271f |
Desk playtest 6: the four branches five passes never reached
A branch pass, not a session pass. Seed 3319. For each untested branch the roller draws from the stream until the roll lands in the target band; every roll downstream of it is taken as it falls, and the draw counts are printed. All four branches work as written. The Swinburne fumble is better than the drawer, the re-cut peg costs the twenty minutes the text promises, and a failed Xenology no longer strands Act Four — the v0.14 un-gated tell carried the act with BOTH gated lines failed, which is the strongest result here. What the pass actually found is outside the branches: 1. The fight the table produces is not the fight the document measures. The six at the back stand inside forty-one refugees. Six hollow men alone wipe 0.0%; with two of the column joining, 7.8%; with three, 28.7%. Stable across four seeds. EXPOSURE measures the two forces separately and never says what the column does when somebody fires into it. 2. At four players six hollow men wipe 58.7%, and the only hollow-man figure in the document is the six-player 0.0%. EXPOSURE's four-player block — which exists to say the cut is a different game — has no hollow-man row. The curve from three to six is 0.3% to 58.7%. 3. The Act Three warning names the wrong roll. Its twenty-minute bomb hangs off "Anomaly Lore — what a peg is"; the depot entry is a Research roll. Pass 4 inherited the same confusion and nobody noticed, because the text it was checked against carried the error. 4. GM ESSENTIALS understates every fumble band. "00 always fumbles" reads as 1%; fumbleStart is 101 - ceil((101-band)/20), tested first, so Spot 40 fumbles on 97-00. Four times what the summary implies, in a scenario that rolls Spot 40 through two acts. 5. EXPOSURE prices fights in bodies and never in minutes. Median 13 rounds printed, 24 mixed, 30 at four players. Also: simulate.mjs labels a mixed force with the first spec's name and the total count, so 6 hollow men + 6 column prints as "12 x Hollow man". The composition is right and the header is not; every figure above came out of a run whose header lied about what was fought. First run in six passes where Ashcroft refused THE OFFER, and the first where a player character was taken. Seven fixes listed, unapplied. npm run check: 18 guards pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1a80d3637c |
R-282: guard the bestiary against the powers, including the next one
check-bestiary proves the page matches its generator, and the generator had never heard of powers.mjs -- so the redcap sat in "the ones that take the most stripping" under eighteen green guards while its own entry said it never spends a defence. Two files agreeing with each other while both disagree with the engine is a quorum, not a check. check-powers now asserts per effect kind what the page must say: defenceStacking requires the creature off the stripping list, named as exempt, and the "across N creatures that spend defences" count reconciled against powers.mjs; attackFactor requires the rating the simulator actually uses printed as a number, which the courier's entry now carries. The clause that matters is the failure on an unknown effect kind -- a wired effect with no DOCUMENT_RULE fails the build, so the next one cannot arrive without somebody deciding what the document owes it. Without that this would guard the mistake already made and nothing else. Proved three ways in a worktree, exit codes read directly. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
8705b00631 |
CLEAN GROUND v0.15 — the four-player cut, fixed
Applies desk pass 5's five fixes and corrects a claim the pass itself overstated. 1. The cut cannot see. Pollard's Spot 53 is the best in the declared cast (others 40/40/40/35/35) and he is the first agent the scaling notes drop. At five players and fewer his clues — the stump, the six at the back, the lorry — are given, not rolled for. check-rollable cannot warn about this: it tests the 25% floor, not competence. 2. The split table now has a four-player column. Four of its five rows named Pollard or Okonkwo, or said "all six". 3. The Pacing Note says the cuts are sized for six. At four, keep the northern seam; pass 5 ran 3:15 with every cut taken. 4. The line rests twenty minutes at four players — Ashcroft is aside for THE OFFER, so three agents are talking, not five. 5. Act Four records that THE OFFER nearly quadruples the unsettled rate: 14.5% if Ashcroft accepted against 3.8% if he refused. Accepting makes him both the likeliest person to be taken and the likeliest reason the beat resolves into nothing. Verified against roster.mjs rather than recalled, which caught pass 5's own error: Spot 53 is the best in the CAST, not the roster — four agents carry 58. Post-pass appended. The check also found that Agyeman's 58 makes the substitution the scaling notes already name the single best answer to the cut, which is now written in. npm run check: 18 guards pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
81ff44e70e |
The bestiary was still describing the fight R-275 changed
Two things in the generated document had gone stale the moment powers reached the simulator, and eighteen guards were green over both because nothing connects powers.mjs to bestiary.mjs. The dodge-ladder section listed Redcap among "the ones that take the most stripping" and counted it in "across 47 creatures", when NOT TIRED means it never spends a defence at all -- the exact opposite of what its own entry says three pages down. It is off the ladder now, the count reads 46 that spend defences, and the page names it: stripping is not a plan against it, killing it is. And the paragraph listing what the harness does and does not model never mentioned that it fights 39 of the 41 creatures with a power without it. It now says so, and counts from powers.mjs rather than stating it, including the 14 that are fight rules it cannot express -- so every figure for one of those is the creature with its best trick taken away. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
97f5cbbcae |
Desk playtest 5: the four-player cut, played for the first time
Seed 8815, four players — Ashcroft, Bhattacharya, Renshaw, Braithwaite. Every rating pulled from roster.mjs before rolling, which is the direct consequence of pass 4 having typed six of sixteen from memory. The cut is the most measured thing in this project and had never been played. EXPOSURE prices its fights, check-firstblood holds its swing, check-attackers holds what each of the four contributes -- and no pass in four had run a session with Pollard and Okonkwo off the table. THE RESULT IS NOT ABOUT COMBAT. The fights were never the problem. What breaks is the party's eyes: POLLARD'S SPOT 53 IS THE HIGHEST IN THE ENTIRE ROSTER AND HE IS THE FIRST PERSON THE SCALING NOTES DROP. At five players the party's eyes go from 53 to 40; at four they stay at 40 with a 35 alongside. This run missed the stump (79 vs 40) and the six at the back (67 vs 35), both Pollard's lane in the six-player game, both clues the scenario leans on. check-rollable cannot see this, and the reason is worth keeping: it asks whether every named skill is reachable at 25% or better, and Spot 40 clears 25 comfortably. It is a rollability check, not a competence check, and the cut is where the difference bites. The scaling note's only stated cost of dropping Pollard is that "the cordon loses its shield", which is about a fight, in a scenario whose first two acts are almost entirely looking at things. Second finding: FOUR OF THE FIVE PARTY-SPLIT ROWS NAME PEOPLE WHO ARE NOT THERE. The table is introduced as the fallback for a table that will not choose its own splits -- and at four players the fallback does not exist, which is exactly when it is most needed. Timing runs the other way from pass 4: ~3:15, twenty-five minutes SHORT, because four people ask fewer questions. Nothing in the Pacing Note says the cuts are player-count dependent, so a GM following it at a small table finishes early. The unsettled countdown fired for a second pass running. Recorded with the caveat that both passes had Ashcroft accept THE OFFER, so both used the Difficult branch -- 14.5% unsettled against 3.8% if he refuses. Not two draws from the same distribution, and the document prints only one of those figures. Five fixes listed, none applied. Also listed: what five passes have still never tested -- the Swinburne fumble, the re-cut-the-peg fumble, a failed Xenology, and any fight at all. npm run check: 17 guards, exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
303472f0ab |
The bestiary said "in doubt" and the rule no longer means that
R-280 widened admission from a [15,85] band to "the spread rate stands clear of both ends by more than its own noise", which admits fights that are nearly settled -- the_choir at 99.3%, the_stanchion at 6% -- as long as their noise is smaller still. The page went on saying "whose outcome was ever in doubt", which was a fair description of the band and is a loose one of the rule. It now says "whose odds leave room for a difference to show", and the provenance line prints the admission rule itself, read from the artifact rather than paraphrased, so the page cannot drift from the guard again. The rule is phrased as a clause in check-focus for that reason. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ae06df8e04 |
Correct desk pass 4: six of sixteen ratings were typed from memory
Found while setting up pass 5 by pulling the sheets from roster.mjs instead of recalling them. The d100s are unchanged; only what they were compared against was wrong, and three outcomes flip: Pollard Spot, the stump 58 -> 53 success becomes fail Bhattacharya Anomaly Lore 58 -> 63 FUMBLE becomes plain failure Braithwaite Xenology (Baseline) 45 -> 53 fail becomes success So three things pass 4 reported did not happen. The re-cut-the-peg fumble never fired -- 98 against 63 is a plain failure, 96-99 never succeed -- which means the ten-minute chase and the whole of finding 2 rest on a roll the seed did not produce. Braithwaite's Xenology succeeded, so the scenario's one single-point-of-failure remains untested after five passes rather than having 'finally come up in play'. And Pollard missed the stump. Read correctly, the same seed lands near 3:34 -- six minutes UNDER budget rather than six over. Findings 1, 3 and 5 are untouched: the special was a real 7, the Swinburne fumble a real 00, and THE OFFER's staging is not a dice question. Finding 2's conclusion also survives because it never depended on the roll -- Act Three is budgeted at 50, the text predicts the fumble costs 20, and nothing connected that to the cut. The reasoning was right and the evidence was invented, and the scenario now says so instead of citing a playtest that did not happen. The Pacing Note's slack claim is withdrawn rather than replaced. One seed read two ways gave 3:46 and 3:34, and the gap between them is about the size of the margin being argued over. It now tells a GM the shape -- every scene has something that ends it, the cuts are real, Act Three is the likeliest overrun -- rather than a number a desk pass cannot produce. npm run check: 17 guards, exit 0. |
||
|
|
fa484071cd |
R-281: a ratchet, because "helps in all of them" survives the advice decaying
R-280 left the claim satisfiable by a gain of 0.2. The artifact now records how
many measurable packs clear their own noise -- reliable: {aboveNoise: 36, of: 38}
-- and the check refuses if that share falls. It may rise freely.
A ratchet rather than a threshold: any threshold here would be a number I chose,
and choosing one just under the current value is what produced MEASURABLE =
[15,85]. A share rather than a count, so widening admission cannot pay it off.
Proved three ways in worktrees: making focus fire actively bad fires the drift
check first, which is correct; making it unreliable and re-recording fires the
older claim at 37 of 38; and claiming a better past, 38 of 38, is refused by the
ratchet itself. The ratchet bites exactly where the old claim does not -- between
"still helps everywhere" and "helps as reliably as it did", which is where a slow
degradation lives.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
c127184763 |
CLEAN GROUND v0.14: pass 4's fixes, and one finding corrected in the applying
Five fixes from desk playtest 4. One of them was wrong and the correction is the more useful half. 1. A SPECIAL ON THE COLUMN PERSUADE NOW BUYS SOMETHING PORTABLE. It bought more conversation, which the line's rest clock then took away -- a player in pass 4 said "I rolled a 7 and got less time", which is a good roll punished by a timing device. It now buys the almanac early, or a name from 1962, or "the six at the back don't eat". Each travels out of the scene, so the line still stands up on schedule and the roll still paid. 2. ACT THREE'S TWENTY-MINUTE BOMB IS NOW BUDGETED -- and this is the finding that was wrong. The pass said the re-cut-the-peg fumble had "no duration and no ender". It has both, and always did: the text says a table will spend twenty minutes on it and to let them try it exactly once. What was missing is that ACT THREE IS BUDGETED AT 50 AND THE TEXT PREDICTS 50 + 20, with nothing connecting the fumble to the cut that pays for it. Act Three now opens with that warning and makes the northern-seam cut compulsory the moment the fumble lands. The clause itself is untouched; it was already right. 3. A FUMBLED PERSUADE ON SWINBURNE HAS AN ANSWER. She does not produce the photocopy -- she produces the dog, walks them to the stump in silence, and H01 reaches them later from the coroner, confirming rather than revealing. Better staging than the drawer, and the GM should not regret the fumble. 4. THE SLACK CLAIM IS HONEST NOW. The Pacing Note said ten minutes. Pass 4 on a hostile seed finished at ~3:46 with one cut unspent: four minutes and a cut. Ten is the friendly-seed number and is what a GM would have planned against. 5. THE OFFER'S STAGING ACCIDENT IS NOW DELIBERATE. Moving it inside the column scene was purely a time saving; the side-effect is that the offer has no audience and Ashcroft returns to a conversation that carried on without him. Written down so nobody moves it back. The correction to 2 is recorded in the playtest document as its own lesson: a reading made while looking for faults finds faults that are not there about as readily as ones that are. Apply fixes against the source, not against the notes. npm run check: 17 guards, exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
3d83fbde93 |
R-280: the band replaced by the test it was standing in for
MEASURABLE = [15,85] proxied for "can this fight move at all". The direct form, in the units the guard already uses: the spread rate must stand clear of both ends by more than its own noise. Deliberately blind to the gain -- admitting the sizes where focus fire clears its noise would make the guard's claim true by construction. Size is still picked on nearest-an-even-fight. 31 measurable became 38. the_arrears returns at 6.3, and switchboard arrives at 8.1 -- the second largest gain in the artifact, thrown away for being one point past a round number. Four of the eight carry effects larger than most rows the band already admitted. Two of them do not clear their own noise: the_choir has 0.7 points of headroom and used 0.2, the_stanchion has six and used 0.2. Above-noise falls 31/31 to 36/38 and the bestiary prints 36. That is two measurements reporting no detectable effect, which the band suppressed by refusing to take them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
23262855b4 |
Desk playtest 4: the clock, against v0.13's recut timings
Seed 6142, banding from rules.mjs rather than estimated, every roll logged. Run because desk pass 1's 4:25 describes a scenario that no longer exists -- Act Two was 65 minutes when that pass ran and is now 45. VERDICT: the clock holds, and for the reason it was built to. Every beat in Act Two ended on something written into the text rather than on GM judgement, which is what the recut was for and what the 65-minute version never had. The run landed at ~3:46 against 3:40, with two of three Pacing Note cuts taken and the third still in hand -- the first time this scenario has finished a pass with a cut unspent. A hostile seed, deliberately not re-rolled: two fumbles, eight failures, and the party's one reliable Act Four test missing. A clock only tested under lucky rolls is not tested, because failure is what generates table time. THREE THINGS THE DICE FOUND, none of them about minutes: - A SPECIAL FIGHTS THE REST CLOCK. Renshaw rolled 7 against 63 and the column opened up at exactly the moment the scene wanted to end. A GM who has just rewarded a good roll will not then stand the line up. The roleplay-first archetype's verdict was "I rolled a 7 and got less time", which is the only sour note in the session and a real design fault. - THE RE-CUT-THE-PEG FUMBLE HAS NO DURATION. The clause is one of the best things in the document and the text itself says a table will chase it. It added ten minutes with no guidance about what ends it. - A FUMBLED PERSUADE ON SWINBURNE HAS NO ANSWER. H01's failure case is written for a party who did not ask, not one that asked badly. Predates v0.13, and the third pass running to find the failure cases written for absence rather than for bad rolls. WHAT WORKED, WRITTEN BLIND: the unsettled countdown outcome added in v0.12 was asked for on its first ever roll -- the understudy fumbled 100 against Ashcroft's failure on the Difficult branch, so the contest came back unsettled. That is the 14.5% case, it is the aggressor-fumble the rule's author left for this document to price, and the price is right. No fix needed. Also confirmed: THE OFFER running inside the column scene costs zero wall clock, and has a side-effect worth keeping on purpose -- the aside has no audience, so Ashcroft returns to a conversation that moved on without him. Five fixes listed, none applied. npm run check: 17 guards, exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
a7e4dd7754 |
R-279: what the_arrears was claiming, and why dropping it is the wrong kind of right
It claimed focus fire is worth 6.3 points against two of them, 14.7% to 21%. Measured independently at 4000 runs x 3 seeds: 14.2% to 20.4%, gain 6.2, which is 2.4x its own noise and 44% of the base. The claim was true and reproduces. What failed is a threshold. MEASURABLE is [15,85] and the scan's estimate of a boundary value moved 14.7 to 14.4. And the creature is a step -- 85.7% at one, 14.2% at two, 0.4% at three -- so no pack size gives an even fight and the band's endpoints fall in the gap. The band records nothing about a creature whose defining property is having no middle. Not moving the band to 14: fitting a threshold to the datum it excludes is how a guard stops being a test. But the band is a proxy for "can this fight move", and the direct test -- does the gain clear its own noise -- is already in the artifact and answers yes. Replacing the proxy is a decision about all 47, not a fix, and not mine to take unasked. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
a7eca91a9c |
R-278: raising SCAN_RUNS found that SCAN_RUNS had never been read
Asked to raise the scan until the redcap's pack size stopped flipping. Measured the threshold -- unstable at 3000 and 4000, stable across twenty seeds at 6000 -- raised it, and the re-record took fifteen seconds, which was impossible. winRate takes three parameters and pickSize passed SCAN_RUNS as a fourth. JavaScript discards it, so every scan has always run at RUNS and SCAN_RUNS has never been read by anything. The fix I was asked to make was inert in the same way as the thing it was fixing. winRate takes runs now. The redcap is still n=3 with gain 6.3, arrived at stably rather than luckily; the_arrears drops out of the measurable band at an honest scan, 31 packs to 30; the_committee moves 2 to 6 and stays pinned. Claim check still passes, bestiary regenerated. The only signal was a number being too small. A fifteen-second re-record is good news, and good news is what nobody investigates. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
00e5a993cc |
CLEAN GROUND v0.13: Act Two cut from 65 minutes to 45, and the clock now fits
The scenario ran 20 minutes over the house budget (3h30 play plus a ten-minute break, 3h40 wall clock) and Act Two carried all of it at 65 minutes. It is now 45, and the session lands at 3:40 exactly: 45 + 45 + 50 + 45 + 25 = 210 minutes of play. NOTHING WAS DELETED. The twenty minutes came out of structure, which is the only kind of cut that survives contact with a table: - THE CROSSING IS CAPPED AT FIVE. It was "let the silence sit until a player breaks it", which is open-ended by construction. If the table is still quiet at minute five the Geiger finds its own voice. A silence that has stopped being tense is just a pause. - THE OFFER RUNS INSIDE THE COLUMN SCENE, not beside it. It was "somewhere in this act, take Ashcroft's player aside for one minute" -- a separate slot. Run during the line's rest, while the other players are talking to Ivy, it costs nothing, because the table is already occupied. - THE LINE'S REST IS THE ACT'S CLOCK, and this is the cut that does the work. The column walks every day and stops to rest, not to meet people. Ivy talks for as long as the line is sitting down, and the line sits for twenty-five minutes; then the old ones stand up, because they always do. The scene ends on the GM's schedule through the scenario's own premise rather than through a GM deciding to move things along -- and it is the loop showing itself for the first time, which makes the timer do dramatic work as well as temporal. The Pacing Note now carries 65 minutes of further cuts against what was a 55-minute problem, so there is about ten minutes of genuine slack for a table that talks. That is the first slack this scenario has ever had. It also now says where to cut if the break arrives late: Act Three's depot search, never Act Four, which is the shortest act with the most to do. Also fixed a stale duplicate: the Overview said "Runtime: 4h00" while STATUS and the timing table said otherwise. Every runtime figure in the document now agrees, and the remaining mentions of 4h00 are explicitly historical. CAVEAT, STATED IN THE DOCUMENT: desk pass 1's 4:25 predates this restructure and no fourth desk pass has been run against the new timings. The first human run is also the first test of this clock. npm run check: 17 guards, exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
38012d56ff |
R-277: n=3 is right for the redcap, and R-276 was wrong about why
check-focus picks the pack size nearest an even fight. For the redcap that is n=3 at 28.0% against n=2's 73.4% -- correct, and also where focus fire is worth most. But the margin is 1.4 points and the scan is 400 runs: run across eight seeds it picks 3 seven times and 2 once, and the recorded gain would move 6.3 to 4.9 with it. R-276's explanation was wrong. Its table was measured against the CLEAN GROUND cut, which it never named. Against the frozen party a lone redcap is worth 0.4 rather than 6.1, because that party wins 98.9% and nothing shows against a ceiling. The power is worth most where the fight is in doubt -- 3.8 at n=2, 3.1 at n=3, nothing at either end. Not outnumbered. Undecided. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e21446b164 |
CLEAN GROUND: say plainly that it runs 20 minutes long
Found while revising the Contingency 2027 prep deadlines. The house standard (docs/house-standards.md in the convention repo) allows 3h30 of play inside the four-hour slot plus a ten-minute break -- about 3h40 wall clock against roughly 3h15 of content. This document has always budgeted 4h00, which is the whole slot with nothing either side, and never said that was over. Worse, the Pacing Note's ~25 minutes of cuts read as slack and are not: desk pass 1 ran 4:25, so taking every cut lands at 4:00, still 20 minutes past the house budget. A GM reading "4h00 including a break" alongside "the Pacing Note carries cuts" would reasonably conclude there was room to spare. There is none. STATUS now carries the overrun, and the two honest ways out are written down rather than left implicit: find another 20 minutes, most likely in Act Two's 65, or declare it a deliberate exception and run it where nothing follows -- which at Contingency 2027 it does, Sunday afternoon with only a reserve behind it. Not decided; that is Tim's call. Until then the instruction is explicit: assume an overrun and take the cuts from the start rather than deciding at the break. npm run check: 17 guards, exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
7fc0da3ecf |
R-276 correction: check-lethality fights solo, and I read the wrong column
Asked to re-record the lethality baseline against a lone redcap, I read the tool and found it already does: measure(party, [spec]), with a comment saying "Solo, because a creature is the unit under test". My claim that it fights packs came from memory of check-focus, whose redcap is n: 3. The lethality figure was also not hiding the power. Wipe rate moved 0.7% to 0.9% because one redcap cannot wipe four agents whatever it ignores -- that column is at its floor. Agents down moved 0.55 to 0.74 of 4, a 35% relative increase, which is the column I did not look at. No baseline re-recorded: it is already solo and was re-recorded in R-275. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
8fa691c4aa |
R-276: NOT TIRED is visible in play, and worth nothing where we measure it
Played four agents against one redcap. Two dodges in one round, both at 75 -- DEFENCE_STEP is -30, so the second would have been 45 and the roll of 70 would have failed. The wiring fires. Measured with and without, 2000 x 3 seeds: the power costs the party 6.1 points against a lone redcap and 0.1 against three of them. It is a rule about being outnumbered -- a lone defender spends four defences a round, a pack spends one each -- which is why check-lethality moved only 0.7% to 0.9%. The baseline fights redcaps in a pack, the configuration where the power is worth nothing, so that figure is the floor rather than the effect. Not presented as evidence: at seed 3 the party wins with the power on and is wiped with it off, which is stream divergence rather than direction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
dee2c8f776 |
CLEAN GROUND: booked — Contingency 2027, Sunday 31 January, afternoon
Slot 9. Recorded in STATUS and both open questions closed. The convention repo (slaguru666/contingency2027) carries the booking and points back at this file rather than copying it: seventeen guards check the scenario here, and a copy over there would be a second version nothing checks. That spends the Sunday afternoon re-run reserve, which was real slack. The convention schedule note now names this game as the one that gives if prep on the eight booked games slips — it is the newest, least tested, and the only one whose absence costs nobody a booked seat. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
32ca982369 |
R-275: the harness had never read a creature's power
41 statblocks carry a POWER in their tactics and simulate.mjs read none of them, so a redcap that ignores the cumulative defence penalty has been measured as a creature that tires -- in check-lethality, in check-focus, and in every figure published about it. The defect was not that the powers were unimplemented, it was that nothing said they were not. powers.mjs classifies all 41: 2 wired, 14 notSimulable with a stated reason, 25 not fight rules. check-powers refuses an unclassified POWER and refuses a notSimulable without a reason -- and it does not test that the harness imports a power, it fights the creature with and without and requires the two to disagree. Moved: the courier 9.4% to 1.0% wiped (it attacks at half while carrying), the redcap 0.7% to 0.9% (small, because these fights rarely spend a second defence). ARGENT AND GULES was wired and then un-wired: it tripled the supporter's wipe rate to 75.2% because the harness has no ground and applied the borough-ground condition unconditionally. Same reason THE PULL is not wired. I had wired one and refused the other on identical facts. Lethality and focus re-recorded, bestiary regenerated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9f3a63174f |
CLEAN GROUND v0.12: convention one-shot, the print pack, and the real opposed roll
Tim settled the last two questions: convention one-shot, and rules.mjs gains a real opposed roll (landed by a peer session at R-273). Both decisions change the document, and the one-shot unblocked the print pack. THE PRINT PACK — docs/scenarios/CLEAN_GROUND_HANDOUTS.html, guard 17. Four A4 sheets, self-contained, no external fonts or assets: H01 the coroner's note on white, H02 the 1962 committee minute on cream with photocopy grain, H03a and H03b the almanac on ruled paper. 11pt floor throughout. The binding constraint is that H03a and H03b must print identically or the almanac trick dies, so that is structural rather than careful: they share one `.almanac` class and every dimension comes from a variable defined once. There is no selector anywhere that names one page and not the other, and check-handouts fails the build if one appears, if their markup structures diverge, if their columns differ, if any row stops reading "41 mi", if the counts stop being forty-one now against fifty-three then, or if the pack and the scenario drift apart. It earned itself immediately: its first run failed my own pack for five rules at 10.5pt, under the house 11pt floor. Rendering was checked visually too, which caught two things no guard would have — the "TO CLEAN GROUND" header colliding with NOTES, and the writing crossing the red margin rule instead of starting right of it. THE ONE-SHOT. Countdown step 6 said "and this is a campaign", which the decision contradicts. Rewritten, and the Close's "leave it filed" ending now says how to land it tonight: do not end on "you'll be back", because the table never will and a hook they cannot take reads as an unfinished scenario. Name the next agent who gets sent, and have Registry thank them. THE OPPOSED ROLL. The scenario-local ruling is deleted; GM ESSENTIALS points at the game. Both beats were re-priced against the real rule by enumerating all 10,000 roll pairs -- exact, no seeds -- and two things fell out: THE OFFER IS PRICED. Ashcroft refusing is taken 21.9% of the time, the safest file at the table. Accepting: 48.7%, past Braithwaite's 37.6%. Accepting does not make him a bit more vulnerable, it makes him the easiest person in the room, and the countdown reaches for the easiest. STEP 5 CAN COME TO NOTHING, 14.5% of the time against an Ashcroft who accepted, and that row fires once and is marked permanent. Previously a silent gap in a climactic beat. It is now a written outcome with read-aloud text: the reach fails, it wears the wrong face for a moment, the party learns what it is and cannot prove it, and the clock still turns to step 6. Also: update-readme's count-word list ran out at sixteen, one guard after its own comment warned about hardcoded lists going stale. It failed loudly rather than silently, so it is an inconvenience and not a defect. Extended. npm run check: 17 guards, exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
dd3d818893 |
R-274: the opposed roll works and the combat never reaches it
Played the hollow man against the cut. THE FILE IS YOURS NOW never fires -- simulate.mjs reads statblocks and weapons and has no concept of tactics, so R-273 closed a gap in rules.mjs and left the same gap one layer out. Resolved by hand it behaves: every roll pair enumerated for both beats. The tie-break carries it -- at POWx5 100 against 60 the aggressor still only takes them 48% of the time, because equal bands go to whoever is being acted upon. Act Four prices THE OFFER: accepting it moves Ashcroft from the safest person in the room to the least safe, 21.9% to 48.7%, past Braithwaite's 37.6%. And that once-only row no-ops 14.5% of the time against him, which Act Three can absorb and a permanent countdown beat may not. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
cdbb2e34ee |
R-273: the opposed roll, built where the authority lives
The content has called for 'opposed POWx5' since before the rule existed -- the hollow man's tactics, CLEAN GROUND twice, slice-text three times -- and rules.mjs defined none. A scenario carried a local ruling that said in its own text it was a ruling and not a rule. Generalises that ruling rather than inventing another: both sides roll, the better band wins, only the ladder the game already has. Ties go to whoever is being acted upon, which is what defenceOutcomeFor has always said; the scenario's 'favour the agent' gave the same answer only because no agent ever initiates one. Neither side succeeding leaves the contest unsettled rather than won, which the two beats need in opposite directions. Two exports at the scenario session's request: opposedOutcomeFor compares graded levels and carries both, so a caller can price a fumbled attempt without this file deciding what a fumble costs; opposedContestFor runs it from ratings and rolls with per-side difficulty, so 'resists at Difficult' does not put applyDifficulty back into a document. Spot-checked against the real Act Four beat: 55 against 85 at Difficult, which is 42. Page 1 states it by asking it -- the tie-break and the margin are computed from the rule at build time, so the book cannot drift from the engine. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |