e80f7bb564efd60264dcca0aaf9bc832a7519c05
227
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e80f7bb564 |
detention-tables: measure STANDING ORDER's Act Two fight, and hold it
A second scenario gets pack-tables' arrangement: four configs (Teague plus one, three or five wardens) against the declared six and the four-player cut, 2000 runs at seed 11, cap 400, recorded to detention-baseline.json, citeable as "detention", and --check added to npm run check. standingorder joins all-specs' SCENARIOS, and simulate.mjs now reads that list instead of keeping its own copy, so a new cast is fightable the day it is registered. The lethality, focus and pack baselines are unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
0fa21263b4 |
STANDING ORDER: an Adventure compendium, built from the document
New pack ringbrp.standingorder, in the Scenarios folder. It holds the GM journal (13 pages, one per section), the three player handouts, nine plates to show, six faces and the Bellhouse section map as a scene. No cast actors, because nobody in the case has a stat block. The GM pages are read from docs/scenarios/STANDING_ORDER.md at build time rather than transcribed, so the compendium cannot drift from the document the desk passes ran against. STATUS and Open questions are left out as authoring notes; STATUS's "not obvious" list leads page 1. Doc fixes found on the way: the Act Two heading still said ~55 min against a 75-minute budget; STATUS said four passes and "no desk playtest run" after five; the runtime open question still called 45/55/50 unbacked. figures-baseline: generator enrolled at 0; STANDING_ORDER.md leaves the uncited list because it cites. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e780181e07 |
Rebuild the rules pack: R-273's opposed-roll section never reached it
|
||
|
|
1d42fb2444 |
STANDING ORDER v0.7: pass 5 verifies the retry fix, and catches my own error
A narrow verification pass rather than a session: does Act Two land where v0.6 budgets it now that a failed filed-clue roll costs a scene instead of ten minutes? Seed 4471. THE RETRY FIX WORKS AND DOES NOT DO WHAT v0.6 ASSUMED. Forced both filed clues to fail — about one table in seven, since each is a 37% miss at these ratings. Under the old rule that was twenty minutes and a party standing in Registry waiting; under the new one Pennyfeather fetches each while they do something else and it costs nothing, while the roll still means something. So the fix removes up to twenty minutes of VARIANCE. It does not shorten the typical case, and Act Two uncut is 75 minutes whether or not anybody fails a filed roll. AND v0.6 CONTAINED AN ARITHMETIC CONTRADICTION I PUT THERE. The Runtime row budgeted Act Two at 65 — pass 4's figure, measured with nineteen minutes of cuts taken — while the Pacing Note two screens later said the correct number of cuts to plan for is zero. Both cannot be true; it was a ten-minute overrun written in on purpose. The three prior passes reconcile exactly (75 uncut, and pass 4's 66 is 75 − 19 + 10), so the number was never in doubt, only which configuration it belonged to. I measured under one configuration, changed the configuration, and kept the number. Act Two is now budgeted at its uncut base of 75. Acts of 45 / 75 / 50 is 2h50 of play, 3h00 wall with the break, in a 3h15 slot — fifteen minutes of real slack rather than twenty-five claimed ones. The cuts stay optional, which keeps Kit Marlow's scene, which pass 4 showed is the price of the likeliest ending. The timing section now records that 75 is the uncut base three passes agree on, and that the retry fix moved the worst case and not the typical one, so nobody re-derives the same error the next time the slot moves. All twenty-one guards pass. Five passes; what is left is human beings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1a6c1474da |
STANDING ORDER v0.6: the slot is 3h15
Decided rather than worked around. Four passes measured Act Two's floor at about 66 minutes against a budget that said 55; the alternative was twelve minutes out of the sick bay, which is where Revelation 3's best route lives. The slot moved, because 3h00 was an assumption and the sick bay is not. 2h50 of play in three acts of 45 / 65 / 50 plus the break, in a 3h15 slot. Act Two is budgeted at the 66 pass 4 actually measured with every cut taken, not at the 56 the retry fix predicts, because nothing has measured the fix. The twenty-five minutes of slack are deliberate: the house standard says a scenario landing to the minute with zero slack is a fail for a table of strangers, and every figure came off a desk pass run by a GM who already knew the document. THE CUTS ARE NO LONGER DEFAULTS, which is the real answer to pass 4's finding rather than a workaround for it. They existed because Act Two had to lose twenty minutes it did not have. At 3h15 the correct number to plan for is zero — and that retires the trap where the cheapest default cut hollowed out the price of the likeliest ending. AND THE REBUDGET EXPOSED AN OLDER DEFECT. The Countdown put step 3 at +1h20 and step 4 at +2h00, measured from the blast door — so under every budget this scenario has ever had, both fired AFTER Act Two ended. Both are backstops for Essential revelations: step 3 is the second road to the sum and step 4 is Revelation 3's last resort. They were arriving too late to back anything up. The whole table is re-timed inside the act at +15 / +35 / +45 / +60, the reference point is stated, and Act One's Stealth roll moves step 2 within that window instead of pushing Teague past the end of the act. The act checkpoints were trailing their own acts too — the signing one left five minutes for the decision. A checkpoint is a point you can still correct from, so the sum is wanted at 1:40 with twenty minutes of Act Two left, and Act Three's is ten minutes in. Running times are printed so they mean something. All twenty-one guards pass. What is left is human beings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4148d13768 |
STANDING ORDER v0.5: desk playtest 4, the clock with the cuts taken
The pass three passes had not run: an honest session at six players, no contrarian choices and no forced branches, with the Pacing Note's cuts taken exactly as written. Seed 6620. Nineteen minutes cut — pipe room folded, Kit Marlow to the corridor, census dropped. THE CUTS DID NOT CLOSE THE GAP, AND ONE RETRY UNDID MOST OF THEM. Act Two asks for six rolls whose failure case was "ten minutes and a second attempt", roughly two fail in an average session, and a single failed Research handed back ten of the nineteen minutes that had just been cut. The Pacing Note's arithmetic was sized against the act's length and never against its variance, so it could not hold. A failed roll on a filed clue now costs a SCENE instead — Pennyfeather fetches it while the party do something else — which keeps the roll meaningful and costs nothing on the clock. THE CHEAPEST CUT GUTS THE LIKELIEST ENDING. The Pacing Note priced moving Kit Marlow to a corridor at "the drawings". This pass took that cut and then reached the narrow row of ending 3 — the likeliest one, since it needs a single dossier item — whose entire price is that the sixty-one stay unrecorded. Kit is the sixty-one made into a person and her own entry says she must be a person before she is a price. The cut is re-costed honestly, and the corridor version now has to do her one job: she shows them a drawing and asks whether she got the blue right. The blanket Act One failure case named Joan for all seven clues, and Joan does not go down the adit. Clues 6 and 7 now have their own: time, never access, because the door has never been locked. Joan herself is at the farm AND walks up with them, so the moor walk has somebody in it. And 13b came off the dice after a 98 lost the reason the proper channel failed — the thematic centre of the case — which is now written at the bottom of Pennyfeather's own sheet. Act Two's floor is about 66 minutes against a budget of 55, measured across three passes at 75 uncut, 70 uncut and 66 fully cut. The retry fix should recover most of that and nothing has measured it. THE REMAINING QUESTION IS A BUDGET DECISION AND IT IS TIM'S: move the slot to 3h15, or take twelve minutes out of the sick bay and lose Revelation 3's best route. Written up in Open questions and deliberately not decided here. All twenty-one guards pass. Four passes, every content defect they found fixed. What is left is human beings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
3db28885f0 |
STANDING ORDER v0.4: desk playtest 3, player chair, contrarian
The cadence's third pass: a full session choosing contrarily at every fork, forced to the one ending nothing had reached. Seed 3307. The party skipped the scenario's best NPC, told the antagonist the truth before anyone else, refused the hospitality, split so the wrong specialist was in every room, took a resident to the surface, and cancelled. BOTH OF PASS 1'S HEADLINE FIXES WORK, AND ONE WORKS BETTER THAN WRITTEN. The pen held above the paper converted the ending from pass 2's ninety-second reflex into a real argument about whether a narrow amendment they could no longer reach beat a cancellation they could. And Pennyfeather's spoken cost landed on a party who had failed the birth register, so the sixty-one arrived as new information at the moment of decision — which is better than knowing in advance, and clue 14 is demoted to Supporting on the strength of it. THE SUM HAD ONE ROUTE AND NO FALLBACK. Pass 1 found the Revelation 3 fallback circular and it was repaired with three roads; nobody then asked whether the other essential revelations had the same shape. The contrarian split put Insight 40 in Registry instead of 63, one roll failed, and the party cancelled the Standing Order having never learned why it was urgent. Worse, the clue asked for a roll that Pennyfeather's own roster line contradicted — she hands the sheet to anybody who asks her a straight question. The roster line wins: the sum is no longer behind a roll at all, the Insight now buys what she did about it, and the establishment return she posts at Countdown step 3 is a second road for a table that never thinks to ask her anything. Fixing an instance is not fixing a class. TWO OF THREE HOOKS ROUTED THE PARTY AROUND THE ACT'S ENGINE. Only one hook involves the farm, and a professional team with a grid reference drives past it — so Act One ran fifteen minutes short with nobody in it, and both its Essential clues lost their stated failure case, which is Joan pointing at them. She is now at the vent head, where the woman in her own entry would be anyway, and no route into the act can miss her. Also: the pen pause was staged only for a Teague who stood down and now prints the hostile version; a photograph brought back down is answered; refusing the tea is answered; and an uninformed cancellation is written as the ending it deserves to be, since a party who do the right thing by accident is the best possible close for a case about people doing the wrong thing correctly. All twenty-one guards pass. Three passes and none of them a normal table: GM chair with bad dice, a branch stress, and a deliberately contrarian run. Pass 3 landed on the 150-minute budget only because one act collapsed and the other spent the difference. No pass has run Act Two with its cuts taken. That, and human beings, is what is left. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
bd1b6e71f2 |
STANDING ORDER v0.3: desk playtest 2, Act Three stress
The house cadence's second pass is a stress re-run of the weakest act. Pass 1 resolved Act Three in eight minutes of fifty and left three of its four branches untested, so this forces each in turn on seed 8143 — and runs the three creature encounters through tools/playthrough.mjs, the same runFight the lethality and focus baselines come from, against this scenario's own declared six rather than the frozen four. THE THREE DOSSIER SCENES CHANGED NOTHING. The party assembled one item of three and rolled 62 against 63 — exactly the roll they would have made with all three, because Borrowed Authority says items improve the credential and never the roll. Correct as a rule, catastrophic as staging: the act's best three scenes had no effect any player could see, and pass 1 could not find it because there the amendment failed and the question never arose. The items now buy scope instead. Same roll, three different worlds: a narrow amendment that stops the ash and leaves the site sealed, a lifted seal, or records above ground for the sixty-one so Kit Marlow can leave. ALL THREE CREATURES ARE INVISIBLE TO THE DICE. STERILISATION ORDER, THE CORRECTION and ENUMERATED are all classified outOfCombat in powers.mjs, so every measured number is a measurement of the creature's arms and legs and of nothing that makes it frightening. That is correct bookkeeping, not a defect in the guard, and it means the GM's text carries the whole threat at the climax. It did not. The overwriter now comes with four printed corrections specific to this case, and with the fact that force does not work: measured, one of them cannot meaningfully hurt anybody and five grind for forty rounds finishing nothing, so the end of the world was landing as an inconclusive scuffle. EXPOSURE wiped all six in fifteen rounds with five dead, without the sterilisation blast firing at all. The warning was right and understated. It now prints the six-round countdown that lived only in the bestiary entry, the argue-down as a named Borrowed Authority stretch, and what a failure costs — five rounds reaches the adit and does not reach the crèche. The best ending was three sentences, less than the failure branch beneath it, and now plays out properly: the tannoy, nobody cheering, Pennyfeather filing it, and Joan sweeping a yard that stays swept. The four-player cut was one sentence from disaster. Science (Biology) and Medicine are the only skills that vanish with Okonkwo and Nkemdirim, the cut note happened to reassign exactly those two clues, and check-rollable never saw any of it because the cast declaration names six and the cut is prose. Both clues are now written on Knowledge and Insight for every table size, so nothing load-bearing depends on an unguarded sentence. All twenty-one guards pass, and the document has entered check-figures' strict regime now that it carries citations. Still untested: ending 1, and with it Pennyfeather speaking the cost and the pen held above the paper — pass 1's two headline fixes. Pass 3 is the player chair, contrarian, and should cancel. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
07bc09129e |
art sheet: record what actually rendered, and why Joan did not
Eight plates and six of seven portraits are on disk. Joan Wetherlaw's portrait stalled three times — twice on the bridge's full 1800s window and once on deliberately reworded prompt text, which rules out both the prompt and deduplication. The thirteen before it returned in about forty-five seconds and everything after ~21:20 hung, which looks like fast GPU hours running out mid-batch. The sheet listed every stem as though it existed and named the plates by their pre-conversion stems rather than the webp files actually committed. Both fixed, with the retry instructions and the note that mj-gen exits 0 on a timeout, so the file is the thing to check and not the exit code. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
123e08612e |
STANDING ORDER v0.2: desk playtest 1, and its fix list applied
GM chair, all six cast pregens played as archetypes, every beat rolled on a seeded roller (mulberry32, seed 5291, rolls consumed in printed order) so the run is re-checkable rather than remembered. Fourteen plates and portraits too. The pass found three things worth the afternoon. REVELATION 3 WAS NEVER DELIVERED, AND ITS FALLBACK WAS CIRCULAR. The Psychology roll with Marlow failed, the Larkhall fallback failed, and the only route left was a file in Marlow's private quarters that nothing in the scenario told the players existed — findable only by knowing the thing it reveals. A fallback that requires the revelation it is a fallback for is not a fallback. The file now lives in the sick-bay day room where they already have reason to be, there is a second independent route through Pennyfeather and the 1974 return's CASUALTIES NIL, and the Vigil is an automatic backstop under all three. Psychology was the wrong skill anyway: the audit put the best in the cast at 40, a coin flip on the moral centre of the case, while the Casting table credited Renshaw with it as a strength. It is Insight now — 63, genuinely hers, and a better fit for a woman who is ashamed rather than confused. THE MOST OBVIOUS PLAYER MOVE HAD NO PRINTED ANSWER. The talker told the Controller there had been no war, four minutes in, and the document said nothing. Now printed in three mouths, plus what happens when they take a resident up to see the sky — which works, harms nobody, and solves nothing. ACT THREE RESOLVED IN EIGHT MINUTES OF A FIFTY-MINUTE ACT. The amendment failed and the party cancelled inside ninety seconds, because falling through to the end of the world is so much worse than the alternative that nobody deliberated — and they cancelled without knowing it kills the sixty-one children, because Pennyfeather was two corridors away. She now stands at the Controller's elbow and says the number out loud before the pen moves, and a failed amendment holds Sowerby's pen above the paper for one full round. Also: Act Two ran 75 against 55 and now has a real Pacing Note; the Act One Stealth roll bought nothing and now buys when Teague finds them; time inside versus outside is ruled in the text; the census is marked optional because missing it cost nothing; and Containment gets the one beat the heavy's 3/10 verdict asked for. All twenty-one guards still pass, and check-figures still records this document at zero bare figures. Untested: ending 2, a successful amendment, the quarantine unit, four players. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
70acc1e60b |
STANDING ORDER: a three-act post-apocalyptic crossing case
Two hundred and six people went underground in 1974 to govern the country after the war. There was no war. The Deputy Controller filed the drill as live, on purpose, to save the site a fortnight of arguing about whether it was allowed to start — and reality took the filing. The apocalypse in this one is clerical. At renewal 211 the site's population record exceeds the surviving population of the region it governs, and a record that cannot hold both resolves by correcting the region. Act Two shows the players the sum; Act Three is the decision, and all three endings cost something. Cancelling the order kills the sixty-one children born below, who have no record above ground. The text says so and does not offer the GM a way to make it painless. Deliberately not CLEAN GROUND. That case is a filed valley found by accident and resolved by cancellation; this is a filed institution the department built on purpose, and its best ending amends the order's scope rather than cancelling it — a Borrowed Authority problem, which is the setting's own thesis about arrangements beating fights. Three existing bestiary creatures, nothing new to build: the census counts in the Registry, the overwriter performs the correction if the renewal goes through, and the quarantine unit is the failure state you argue down. Cast declared as a different six from CLEAN GROUND's so the two cases exercise different sheets. All twenty-one guards pass: check-scenarios resolves 115 tags, check-rollable holds every route to the declared cast, and check-figures records the document at zero bare figures — it cites no measured number because it prints none, the quarantine unit's lethality included. The section drawing is hand-built SVG and must stay that way: every label on it is load-bearing and no generator holds legible text. Numbered callouts with a key beneath, per AFTERIMAGE v3.1. NOT playtested. Not once, desk or human. Every timing in it is a guess. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
fa30909dc9 |
handouts: the print rule zeroed the padding the slug depended on
Rendering the pack to PDF to print it showed the GM filing slug printed on top of the first line of three of the four sheets — both almanac tables and the minute's "Copy 3 of 4" line. The cause is one declaration. .slug is position:absolute at top:6mm, and what kept sheet content clear of it was .sheet's 18mm screen padding. The print block then said padding: 0, so in print — and only in print — content started at the very top and ran under the slug. On screen the pack looked perfect, which is why it survived being built, checked and shipped. Fixed by keeping a top padding in print: padding: 11mm 0 0. Side and bottom margins still come from @page, so nothing else moves. check-handouts could not see this and is not at fault for it: it holds the pack's STRUCTURE — the two almanac sheets sharing one class and one set of columns, every line reading 41 mi, nothing under 11pt — and a collision between an absolutely positioned element and the flow is a property of the rendering, not the markup. It took rendering the file and looking at the pages. The almanac trick itself is unaffected and was verified in the PDF: H03a and H03b print identically but for their dates and their counts, forty-one against fifty-three. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
fbe6288c94 |
CLEAN GROUND: re-roll the column plate, and record what the crossing costs
cg_05_column re-rolled and replaced. It now has Ivy a half-step ahead of the line with both hands held open and visible, which is the beat the read-aloud actually describes — "one woman walks forward with her hands held where you can see them". The line behind her is still drawn uniformly, so the six at the back remain indistinguishable, which is the one hard rule in the notes. cg_04_crossing is unchanged. Seven attempts did not beat the original: the plate has to be grey and dead and empty, and the generator supplies any two. The original's livestock are confirmed real rather than rocks — a crop of the hillside shows grazing animals and a barn — so the defect stands, recorded in STATUS rather than quietly kept. What the failures taught is now in CLEAN_GROUND_ART.md, because it is worth more than the plate: - Naming a thing to exclude it puts it in. "No birds, no sheep, no movement" and "the grey is dust and not snow" produced, respectively, livestock and snow. Both negations summoned what they forbade. - Steering off snow steers into summer: the moor comes back green and alive, which is worse than either. - A long --no list stalls the job outright. Four consecutive prompts with ten or more exclusions never produced a grid; the same prompt with four returned in two minutes. - Filter words costing a 30-minute timeout each: wound, dead ground, lifeless, smothered, and shot — including in "nobody in shot". Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
310d342767 |
CLEAN GROUND v0.25 — the artwork, generated and placed
Seventeen assets through the Midjourney bridge: six NPC portraits at 512x512 in art/portraits/cg_*.webp and eleven scene plates at 1024x682 in art/scenes/cg_*.webp, matching the dimensions every other scenario in the system already uses. CLEAN GROUND had none; LAST_ADMISSION has 8 plates, OPEN_DAY 10, THROUGH_TRAIN 13. The art direction holds: 1962 Ministry drawing-office hand, pen and ink with pencil shading, dyeline blue-grey wash, buff card grain and a ruled margin, so the far side is drawn on the same paper as the Registry corridor. Ivy reads as competent rather than frail, and cg_05_column obeys the one hard rule in the notes — the six at the back are drawn exactly like the other thirty-five. Three prompts had to be reworded around Midjourney's filter, which declines ephemerally and leaves mj-gen waiting out its full 30-minute timeout on silence. "An empty dog lead WOUND twice round one fist" cost 23 minutes before the pattern was recognised; "nobody in SHOT" and "nothing behind its face" were found by scanning the remaining prompts rather than by hitting them. Two misses recorded in STATUS rather than quietly kept: - cg_04_crossing appears to have livestock on the hillside, and the point of that scene is that every sheep is gone. - cg_05_column has no Ivy; the line is uniform where the text has one woman a half-step forward with her hands open. Neither stops the scenario running. Eight of the seventeen were inspected against their briefs in detail, including every plate carrying an explicit rule. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
fe92f6ceb2 |
CLEAN GROUND art sheet: square portraits, and the word photograph out of the tokens
Two corrections found by checking the repo before generating rather than after. The sheet said --ar 2:3 for portraits, copied from THROUGH TRAIN's tokens. Every raster portrait in art/portraits is 512x512, all sixteen, and every scene plate in art/scenes is 1024x682 across all three scenarios that have them. Foundry wants a square portrait. Corrected to --ar 1:1. And the token said 'composed as a 1962 personnel-file photograph'. The conceit is that the framing is a file record, but leaving the word photograph in the prompt fights --no photography and pulls the generation to photorealism, which is the house gotcha already written down from the Day One prompts. 'Composed square-on like a 1962 personnel record card' gets the framing without the word. --no photography, photorealism added to all three tokens. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d0db7954bc |
check-figures: the guard's body belongs behind the entry-point check
The fold exported scanDocument for the string tests, and importing it also ran the guard. check-behaviour imports that module, so a figure defect called process.exit(1) inside check-behaviour: it reported check-figures' failure under its own name having run zero of its 108 tests. The build went red, which is why this was survivable, but it went red in the wrong place and every behavioural test was silently not running while appearing to. A guard that stops another guard from running, and cannot say so, is the worst version of the fault this file exists to catch. Body now sits behind import.meta.main, the idiom step5-split.mjs already uses. Importing yields the four readers and nothing else. Tested by spawning a fresh process, because the property is "importing has no effect" and a source assertion would pass on a file that grew a second side effect elsewhere. Reverting the check fails exactly one test. With a defect planted, check-behaviour runs 108 green and check-figures fails in its own slot. Mine, introduced by R-307. R-309. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
bd9b3bdd27 |
check-figures: use the capture's position, not a search for its text
R-308. Read the fold as a stranger at its author's request. Two of the three
risks they flagged are sound: isProse drops nothing that carries a citation
(scanned the whole corpus), and the duration exclusion earns its place, since
"runs" now means three things in this corpus and only the cell can tell them
apart.
The third is real. numEnd used line.indexOf(m[1], m.index), which finds the
first copy of the digits at or after the match start rather than the copy that
was captured:
"The wipe rate of 74 in ten is 74%<!-- cite: ... -->."
value=74 numAt=17 marked=FALSE
A correctly cited figure reported bare, because numEnd lands mid-sentence and
the marker test reads " in ten is 74%...". Fixed with the d flag: m.indices[1]
gives the capture's real position and there is nothing to search for.
Narrow to reach — it needs the wipe-rate shape, the only one with a wide gap
before its capture, and an integer duplicate inside that gap; a decimal cannot
do it because [^.] will not span a decimal point. Three attempts failed for
that reason before the fourth worked, which is why this is recorded as narrow
rather than theoretical.
Worth fixing anyway because its direction is the bad one: a miss costs one
figure, a false positive on correct prose costs the guard.
Seven-case battery re-run, every restore byte-identical.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
b946a28312 |
check-figures: the fold — one guard, three readers, and the boundary between them
check-unmarked is gone and check-figures holds all of it. Three readers, each authoritative where the corpus gives it authority: a named phrasing anywhere, a column header inside tables, and bold in prose only. The boundary is the finding rather than the union. Bold is a publication mark in prose and an emphasis mark in a table — nine measurement columns mix bold with plain, all correctly cited, and in the mixed-force table the two bolded rows are exactly the two the prose underneath singles out. So the bold reader stays silent in a table and the header rules there alone. Neither guard could have found this alone: each had half the evidence and read it as the other's bug. Union of both word lists, because each had a gap the other covered — wipes and runs. A proposal to drop runs? was made and withdrawn; it would have dropped the p99 column, which is R-299's own defect committed a second time. Exclusions test cells, never header words: prose, denominator, duration. No threshold rule touches a header, or "Past 15 rounds" loses two cited figures. Two integration defects, both caught by the ported tests: the readers stopped at different ends of one figure so the dedupe missed it (identity is where the number starts), and blanking prose cells for every reader dropped twelve real figures out of THROUGH_TRAIN's ratchet (the prose rule belongs to the column pass alone). Kept from R-299: opt-in rule, ratchet and its leave-the-loose-set branch, the UNCITEABLE check, NOT_A_MEASUREMENT, comments blanked not stripped. Added waivers, with an empty reason fatal and every waiver printed on a green build. 21 guards, 107 behavioural tests, 139 figures — 115 marked, 6 waived. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4363b545a0 |
docs: R-306 — the merged design reproduces the answer from two implementations
A second implementation of the three settled rules, written from the description alone rather than from the peer's code or the agreed list, lands on the same nine columns and the same single exclusion. Two unlike implementations of the same three sentences agreeing is the property a design needs before anybody builds it. Also tabulates where each of the evening's counts came from: 44% and 57% and 64% and 'at least 7' and 8 each came from an operation on the numbers; 9 came from opening the disputed columns and reading them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4df8064559 |
docs: R-305 addendum — nine, now by inspection rather than assertion
R-305 asserted the shared seven were all genuine without opening them, which is the error the entry exists to correct, committed inside the correction. Opened all nine: every cell cited, none prose, every column mixing bold with plain. The peer's final count of eight is their own vocabulary's nine minus the false positive, an arithmetic that never contained 'Then wipes' because wiped? cannot match wipes. The union was agreed and their own set was counted. Nothing in the design turns on it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c59dc45117 |
docs: R-305 — the union, not the intersection: nine columns
R-304's rate was an artefact; the proposed correction, retreating to the seven both vocabularies agree on, overshoots. Opening every disputed column instead of comparing totals: 'Then wipes' is genuine and check-unmarked misses it because wiped? does not match wipes; '1 in 100 runs past' is genuine and check-figures misses it because its vocabulary has rounds? and not runs; only THROUGH_TRAIN's 'Measured over 300 runs' is a false positive, and there the cells are whole sentences and the bold wraps a clause. So nine, and the shape matters more than the number: each vocabulary has a real gap the other covers, which argues for the union of both word lists. Refuses one recommendation. Dropping runs? from the header vocabulary would also drop the p99 column, whose cells cite fight-tail cut.p99 and column.p99 and which R-269 added because the median is not a plan. That is the R-299 failure again — a guard narrowing itself until the figure it exists to watch falls outside. The real exclusion wanted is 'a column whose cells are sentences rather than values', which is testable; subtracting a word cannot tell the two cases apart. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ae72953a9d |
docs: R-304 — bold means two different things, and that decides the merge
Recomputed the peer's bolding scan rather than quoting it. Eight measurement columns are inconsistently bolded, and that count is robust; the RATE is not — 44% under check-figures' vocabulary, 57% under check-unmarked's, because the two do not define 'measurement column' the same way. Anybody quoting a percentage has to say whose vocabulary produced it. The cause is not carelessness but a second convention: in prose the corpus bolds what it publishes, in a table it bolds the rows it wants read. The two bolded rows of the mixed-force table are exactly the two the prose below it singles out. That divides the merge on evidence — prose takes bold-as-published, tables take column-inherits-header with bold ignored entirely as a signal — and explains the asymmetry R-303 found from both sides as one cause with two symptoms. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
5e3635ffed |
docs: R-303 — the two guards are disjoint in both directions
R-302 showed check-unmarked catching what check-figures misses. I had taken that, plus a preference for structure over vocabulary, as grounds for folding check-figures in as a secondary pass. The converse test contradicts it: five unbolded cells in a measurement column, markers stripped, fail check-figures and pass check-unmarked, whose strict class requires bold by design. So the relationship is symmetric. The shapes are the weak half and should fold in behind the structural classes; the column-header rule is not a shape and belongs beside bold-as-published, not under it. Also records what a merged guard must face rather than inherit: this corpus bolds inconsistently inside tables, which is invisible to a reader and load-bearing for a class keyed on boldness. Both runs in throwaway worktrees, never the shared tree. No merge performed; the decision sits with the humans. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e2a474b246 |
docs: R-302 — correcting R-300, the pathspec was not the fault
R-300 blamed `git add -A` for |
||
|
|
e8faac73db |
docs: R-301 — the narrow repair to R-299, and what R-300 got right
A cell inherits the measurement status of its column: a different mechanism from the shapes rather than more of them, because the document says what the column is and the vocabulary has to guess how somebody will phrase a figure. Records the two bugs in the repair, and why the first matters more than the blind spot it fixed: scanning raw cell text reported ~130 correctly-cited rows as unheld, and a guard that calls the good material broken teaches its reader to stop reading the output. Leaves the merge question open on purpose. Two readers that disagree are how three of this session's defects were found. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e9f7d458f0 |
check-figures: a cell inherits its column's header
R-300 was right, and it was right with a live defect rather than an argument: a bare **24** in a table whose own header row says "Median rounds", reported OK by this guard, against a baseline holding 25. The scan read one line at a time, so measurement status living two lines above was invisible. Fixed by reading the header. A cell now inherits the measurement status of its column — which is a claim the DOCUMENT makes, rather than one this file's vocabulary has to anticipate. That is the narrow repair, and it is deliberately a different mechanism from the shapes: the shapes guess at phrasing, the header does not have to. The general point in R-300 still stands and the file now says so where the shapes are defined: a vocabulary learned from the marked figures cannot contain the phrasing of the figure nobody marked. check-unmarked attacks that from the other end, treating bold as the corpus's own mark of a published figure. Two bugs found while testing, both mine, both caught before commit: - A citation marker is full of digits and none of them are figures. "packs line.hollow6.hurt" holds a 6; "fight-tail cut.over15" holds a 15. Scanning raw cell text reported roughly 130 correctly-cited rows as unheld — the best-marked tables in the corpus. Comments are now blanked rather than removed, so every offset still points at the right character. - "3.48 of 6" — the 6 is the party size the mean is out of, not a measurement. Coverage goes from 25 figures, 7 marked, to 120 figures, 102 marked. Proved again by breaking it, five ways, every restore byte-identical: the historical bare **24** fails at its own line, a stripped prose marker fails, a planted "median 24 rounds" fails, a new bare figure in an uncited document trips the ratchet, and neither a threshold column header nor any of the 102 cited cells fires. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
960fb6be27 |
check-unmarked: the figures a vocabulary cannot name, and a repair to HEAD
|
||
|
|
eead6633c8 |
CLEAN GROUND v0.24 — artwork prompt sheet, and an art claim that was too wide
The Casting section said "Portraits already exist in art/portraits for all six. No new art needed." The first sentence is true and verified on disk. The second was only ever true of the cast, and sitting at the foot of Casting it read as a statement about the scenario. It is not: the six named NPCs have no portraits, there is no map, and there are no scene plates at all — while LAST_ADMISSION has 8, OPEN_DAY 10 and THROUGH_TRAIN 13. Not one cg_* asset exists anywhere in art/. Nothing in "Still to do" mentioned it either; that list held only a human run and a slot, both resolved. So: the claim is scoped to the cast and points at the gap, the gap is on the outstanding list where a reader looks, and CLEAN_GROUND_ART.md now carries the prompts — 6 NPC portraits and 11 scene plates, 17 ids, none colliding with a file already on disk. Art direction follows the house pattern (base style plus one scenario line). THROUGH TRAIN draws the present day in an 1881 hand; CLEAN GROUND draws everything as a sheet from the 1962 file — including the far side, because the valley is not a place, it is a filed document nobody cancelled. Same paper, same margin, same grain on both sides of the seam. The notes carry the rules that matter: the six at the back are never singled out in the column plate (the text says do not linger on them, and a plate that picks them out gives away Act Three in Act Two), nobody in the column is lit as a victim, nothing glows, the understudy is never drawn as a monster, and the grey is not snow. Creatures stay in BESTIARY_ART.md; the handout pack is done and its almanac pages must not be illustrated, since their trick is that the two sheets are identical and check-handouts holds that. 21 guards green; the sheet classifies as a record, not a playable scenario. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9897d55b67 |
docs: R-299 — check-figures, the twenty-first guard
Records the blind spot (check-cited can only resolve markers that exist), why the net is narrow (an earlier draft found 282 candidates, nearly all prose), the opt-in rule and ratchet, and the two things it found before being wired in — 74.9% printed bare twice, and 'the longest fight is 120 rounds' where 120 is the exact field check-cited refuses to let anybody cite. Two rules each right, and together a hole: refusing the citation while the document printed the number left the least stable figure in the suite as the only one nothing held. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
91a104e682 |
check-figures: the twenty-first guard — figures with nothing holding them
check-cited resolves every citation marker against the artifact it names, and
has one structural blind spot: it can only resolve the markers that exist. A
measured figure written into prose with no marker beside it is not a failed
citation, it is not a citation at all, and nothing looks at it again.
Desk pass 11 proved it. "A median 24 rounds" sat in the Pacing Note and "24 if
the column joins" in EXPOSURE, against a baseline holding 21 at two of the
column joining and 25 at three, through every pass that checked citations.
This reads the figures instead of the markers. Narrow by design: an earlier
draft matched any number within 45 characters of a measurement word and found
282 candidates, nearly all prose ("down 140 steps", "an engineer on his
rounds", "01:06"). The shapes here are the phrasings the documents actually use
when quoting the simulator, each read off a figure that is cited somewhere.
Opt-in rule: a document that uses citations must mark every measurement figure.
One that cites nothing is held by a ratchet instead — turning three unguarded
scenarios red is how a guard gets switched off on the day it is written — and
joins the strict regime the moment it gains its first marker.
It found three things in CLEAN GROUND before it was wired in:
- 74.9% printed bare twice while cited correctly four times, and that is the
figure that was published at 58.7% until the truncated sweep was found.
- "the longest fight is 120 rounds" — 120 is exactly packs cut.hollow6.longest,
the field check-cited REFUSES to let anybody cite because a sample maximum
moves by a third on a re-seed. Refusing the citation while printing the number
left the least stable figure in the suite as the only one nothing held. The
sentence now leans on the guarantee that is actually strong: check-packs
refuses to record a sweep in which anything reached the cap.
Proved by breaking it, four ways, with byte-identical restores: a stripped
marker fails, a planted "median 24 rounds" fails at its own line, a new bare
figure in an uncited document trips the ratchet, and a threshold column header
("Past 15 rounds") does not fire — that last one was a real false positive in
the first draft, which tested the matched text rather than its context.
v0.23. 21 guards.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
a1f9e0dca2 |
docs: R-298 — desk pass 11, v0.22, and the cross-session error
Six findings applied. The one worth reading: two blocks two hundred lines apart specified two different fights for the same trigger, both mine, and the mild one predates the measurement. Records the stale 'median 24 rounds' that no guard could see because it carried no citation marker, and the fifth instance of a reader unable to see its own format. And the cross-session half: I re-derived b5's 4.90% exactly and shipped it with their false premise attached. The arithmetic was never the part that could be wrong. Sharper than that — the scenario has no Transposition beat at all, so it was a right number with no question attached and the premise arrived to give it one. The house mission structure does make the return a Transposition roll, which is why the instinct was persuasive. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
3af4c55f12 |
docs: playtest 11 — record what changed on the way into v0.22
Fix 5 as written said a fight costs the northern seam, the depot and telling Ivy. The Pacing Note names telling Ivy among the three things never to cut, so the applied version says two go and the third is compressed or moves to the stump. Written from memory of the act rather than from the Pacing Note. And applying it found a stale figure no guard could see: 'median 24 rounds' in two places, where the baseline says 21 at two of the column joining and 25 at three. It survived because it carried no citation marker, and check-cited can only resolve markers that exist. Fifth reader caught unable to see its own format. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
65ae156477 |
CLEAN GROUND v0.22 — pass 11's six fixes
1. EXPOSURE's "what to do instead of a fight" said "use three". That predates the measurement and described a fight the party never has: three of the column alone is five rounds and 0.1% wiped, where the trigger's real force is a median 21 rounds and 0.97 deaths. It now points at the mixed-force table instead of carrying its own number, and keeps the good half of the sentence. 2. New "The valley at dawn" block: the count moves. H03a says forty-one and the players are holding it. Do not correct the handout — Ivy counts every morning, so tomorrow she writes a smaller number and the sheet becomes a record of what they did. The Close's arithmetic moves with it. 3. Ivy is ruled out of the target pool, explicitly, where the GM decides how many of the column join in. She walked toward the guns, which puts her in front of the column rather than in it, and the Close is hers. 4. A dead PC now has an answer where only a taken PC did: Registry sends the next name on the sixteen-name duty roster, through the seam inside the hour, knowing nothing — which buys the table a recap. 5. Act Three names what a fight costs rather than only that it costs: the northern seam and the depot go, and telling Ivy cannot (the Pacing Note lists it among the three never to cut), so it is compressed or it happens at the stump, which the Close already branches on. 6. The anchor block assumed cooperation in every line. It now answers the party that shoots: they are never stranded, because the crossing will not close on its own — they walk home with nobody holding the door and come back thin. Also corrected two stale uncited figures found while applying fix 5: "median 24 rounds" at the Pacing Note and "24 if the column joins" in EXPOSURE. The baseline says 21 at two joining and 25 at three; 24 is neither, and had no citation marker to resolve. Both now read 21 and cite packs.line.hollow6col2. And two stale STATUS counts: "seven desk passes" listed ten, and "three post-passes" where there are nine. Counted rather than incremented. 20 guards green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
78e92a2dc8 |
docs: playtest 11 post-pass — the 4.90% answers a question the scenario never asks
No [CUS: Transposition - ...] beat exists anywhere in CLEAN GROUND. The house mission structure requires a Transposition roll per agent on the return (tools/mission.mjs:343), which is where the instinct came from, but this scenario's crossing is a standing open door rather than an aimed one. The 63 appears at 1011 and 1479 and both are characterisation, never a roll. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
a8f48df50e |
docs: playtest 11 — correct finding 6's premise
Finding 6 shipped saying that shooting the thing wearing the anchor strands the party. It does not. Line 136: the crossing is "open since the felling, widest at dusk, and it will not close on its own", and the "seam closes for good" at 1356 is the cancellation ending, a deliberate act rather than a clock. The party can always walk home. The 4.90% was correct and irrelevant. Session b5 reported the gap with that number attached to an unchecked premise; I verified the number and inherited the premise, which is not the same as checking the claim. b5 caught it and sent the correction unprompted. The finding is stronger corrected. The answer to "we shoot it" is not that they are stuck — it is that they walk home through the stump with nobody holding the door and come back thin, which is one step onto the road that ends as the thing they just shot. Built from 136, 256-259, 1012 and 1494, four places that have never been stood next to each other. The man who would have held that door is the one person who already knew what it cost. Post-pass added recording how a verified figure lent its credibility to an unverified sentence. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
cd3be7128f |
docs: desk playtest 11 — the fight the players choose
The last of the original four untested branches, played: a party that opens fire at the Act Three swap. Seed 7311 for the played fight, 2000 runs at seed 11 for the distribution around it. The party survives the fight and the document does not survive the aftermath. Six findings: 1. Two blocks two hundred lines apart specify two different fights for the same trigger. "Use three" is a five-round skirmish that kills nobody (0.1% wiped); the mixed force it should name takes 21 rounds and buries an agent (8.3% wiped, 0.97 deaths). The mild one predates the measurement. 2. "Forty-one" appears sixteen times and is load-bearing arithmetic. A fight moves the count and the handout in the players' hands does not. 3. Ivy is unplaced in the one scene that decides whether the Close happens. 4. A taken PC gets six lines; a dead one gets nothing, at a mean of 0.97 per fight in a convention one-shot. 5. A fight in Act Three ends Act Three. Say which three things are lost. 6. The party shooting the thing while it wears the anchor — found by the concurrent session, verified here: Transposition 63 is Ashcroft's alone and the other five are on the 1% floor, so killing it leaves a 4.90% chance anybody opens a door home, and no Close. Every figure checked against tools/pack-tables-baseline.json rather than recalled. 20 guards green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4b491b18cf |
CLEAN GROUND v0.21 — the anchor is the way home, and the Close has an ending
Applies desk pass 10's five fixes. 1. The accepted-and-taken branch has a consequence now. One session in five or six reaches it, and the document had nothing: Ashcroft holds the door everyone else goes through, Transposition 63 is his, and the Close cannot happen until somebody is back through — three facts printed in three places that had never met. The way home now WORKS, and works beautifully, because the thing does not need to escape the party, it needs them to walk it home and it is carrying the skill that gets them there. It holds the door properly because that is what the file says an anchor does. The GM is told to play the gratitude: somebody at the table will say thank you, and that is the scene. 2. Act Three's "real scene" is written. It was one sentence claiming to be the act's centre. Ivy is told or she is not; if she is, she asks how long they have known, goes and sits with the line, and does not tell them. And the Close now answers it — if she was told she does not come south to ask whether there is a north, she comes to ask "Well?", because she is no longer asking for the truth, she is asking what they are for. Act Three's real scene is load-bearing instead of decorative. 3. The Close has a failure case, which it was the only scene in the document to lack. Nobody writes anything and the entry stays filed by default — the second ending arrived at by omission, played as an ending. Covers the clock, the split table, a dead Ivy, and a party too far down to sign. 4. Four across and then cancel anyway is priced as a third ending rather than a failed second one, and Ivy chooses the four, not the agents. She picks the youngest and she is not among them. 5. The depot Research has a special: the 1962 cancellation form, blank, filed by somebody who expected this to end. It is the paper they sign in the Close and it has been waiting sixty-four years. npm run check: 20 guards pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2ee44cef79 |
Desk playtest 10: the Close, and the branch nine passes never reached
Three parts: a census of twelve independent seeds rolled only as far as the three outstanding branches; a played session on the seed that reached two of them at once (4417); and a read of the Close, the only act never given a pass. The census settles the "still untested" list. Ashcroft accepted and was taken in 3 of 12 — one session in five or six, not rare, and nine passes missing it was ordinary luck. The swap succeeded twice against an expected five, which is a mildly cold run at p~0.05 and is recorded as noise: the four consecutive failures passes 6-9 reported are the same thing from the other end. 1. The branch nobody had reached is the one the document has no answer for. Act Four says taking the party's muscle "is a better scene than the anchor being it" and then, one bullet later, arranges the anchor being it in 37% of games. Ashcroft is not an interchangeable body: he holds the door everyone else goes through, Transposition 63 is his, and the Close cannot happen until somebody is back through. The document states all three facts separately and never puts them together. The only guidance on a taken PC is two sentences about the player's evening. And it is worse than a gap, because the thing now has a reason to cooperate: an understudy wearing the anchor does not need to escape the party, it needs them to walk it home, and it holds the door properly because that is what the file says an anchor does. That is the best scene in the scenario and it is not written. 2. "Telling Ivy. Or not telling her" is called the act's real scene and is one sentence — no read-aloud, no failure case, nothing downstream — in a document that gives the fumbled Medicine six lines. And the Close opens with Ivy asking whether there is a north, which only works if she was never told. Its "works in both branches" covers how the PARTY learned, not whether IVY was told. 3. The Close is the only scene in the document with no failure case, and its failures are likely: the clock, a split table, a dead Ivy, a downed party. 4. "Four across, then cancel anyway" is a third ending, not a failed second one, and choosing the four is the most brutal question the case can ask. 5. A critical Research at the depot, with nothing written above the success — the second consecutive pass to land an unwritten special on the first honest session after check-outcomes recorded 83.7%. Five fixes listed, unapplied. npm run check: 20 guards pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
137ab82541 |
R-297: place the step-5 citation markers — 32 of them, no prose touched
Every figure in CLEAN GROUND's step-5 material now resolves against tools/step5-split.mjs, which enumerates rather than stores: the countdown rows, the two-case table and its blend, the Act Four reasoning, the counterfactual pair and the POW×5 targets. No prose changed -- stripping every HTML comment from the result diffs byte-identical against HEAD, which is the check worth having when editing another session's document. Proved live: raising Okonkwo's POW to 13 in a worktree makes Braithwaite the lowest in the room and eleven of the thirty-two markers go red at once, each naming the path and saying nothing can be re-recorded. The three figures that went bad in this document were all in prose rather than tables, which is why the prose is marked and not just the tables. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
7afd5b86cc |
R-296: the review log entry for the pack tables
The work landed in
|
||
|
|
47f2df9446 |
v0.20.1: "nearly quadruples" was wrong by three times, in the last two places
Desk pass 9's first fix, applied. The countdown's unsettled row and the Act Four paragraph both compared 14.5% against 3.8%, and the 3.8% is the rate against Ashcroft at his undiminished 85 — a contest that never happens, because if he refuses the thing reaches for the lowest POW in the room. Against what a table actually rolls: 11.3% at six players (Okonkwo), 10.0% in the four-player cut (Braithwaite), 14.5% if he accepted. The factor is 1.29, not 3.9. Both places now name who the target actually is, and the paragraph carries the correction inline so the claim cannot be re-derived from the old framing. What is NOT wrong, and the text says so: the trade THE OFFER describes is real and the taken rate genuinely moves 40.7% to 48.7%. One consequence was overstated, by three times. Written in v0.15, repeated in v0.16, and survived v0.19.1 correcting the identical error in the table beside it. Flagged again by the peer session while mapping citation paths, which is the third time this figure has been caught by somebody reading it rather than by anything checking it — and the argument for the derived-source markers now going onto it. npm run check: 20 guards pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c2f7249d4d |
R-295: derive the counterfactual too, and name it hadHeRefused
Matching the step-5 prose against the derived source before placing markers left five figures unmatched: 21.9, 74.3 and 3.8, twice each. They are the thing against Ashcroft's undiminished 85 -- the contest that never happens, which Act Four prints on purpose to price what THE OFFER sold. Derived now as hadHeRefused, so the counterfactual is held to the rule like everything else and carries a name that cannot be mistaken for an event. Quoting those three as odds a GM meets is what went wrong in two sentences and a table. Markers not placed: CLEAN_GROUND.md is c0's and they are in it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
a5c4378521 |
R-294: EXPOSURE's tables measured into an artifact, and 58.7% was 74.9%
Desk pass 9 found the scenario's most consequential table held by nothing: ~30 sampled figures, no baseline, catchable only by re-running the exact command. lethality-baseline.json does not cover them — it measures creatures SOLO against a frozen four-agent party that is not this cast. Building the measurer found the figures were also wrong. simulate.mjs's CLI calls runFight with no options, so every published figure was measured at the default 40-round ceiling — and runFight does not report truncation, it scores whoever is standing when the loop stops. Measured at 400: line.column6 14.8% -> 15.2% 34/2000 truncated line.hollow6col2 7.8% -> 8.3% 129 line.hollow6col3 28.7% -> 30.1% 264 cut.hollow6 58.7% -> 74.9% 578 <- 29% of runs never finished Fourteen points on the single most alarming number in the case, and the one v0.16 added specifically to warn four-player tables. fight-tail learned this in R-270 and carries an assertUncensored; EXPOSURE's own tables never got one. The longest fight at cap 400 is 120 rounds and 400 vs 2000 are identical, so the cap is comfortable rather than merely sufficient. tools/pack-tables.mjs measures all 17 configs against the DERIVED cast via castAndCut, using measure()'s exact discipline — one rng threaded through every run, not a reseed per run, because reseeding is a different stream and would not reproduce the published table. It refuses to report or record a truncated sweep. tools/check-packs.mjs holds the baseline against the game, so the pair is not a loop: check-cited holds the prose against the record, this holds the record against the harness. Hard claims read `now` and never `base`: nothing truncated, the party equals the declared cast, config floor, and more of the same creature may not make the party safer. Figures compare exactly, since the runs are deterministic. Verified by breaking it: a drifted baseline, CAP lowered to 40, and a removed config each turn it red, and --update refuses outright rather than recording a truncated sweep. The removed-config test first passed for a bad reason — MIN_CONFIGS is a floor and 16 clears it — so a dropped-config check was added and re-tested on a row in no monotonic chain. All files restored byte-identical after each probe. EXPOSURE's three tables and the nine prose figures around them are rebuilt from the artifact with citation markers; the multi-seed stability claims were re-measured too (the cut is 74.9/72.7/75.0/73.5 across four seeds, not 58.7/60.5/60.3/59.8). CLEAN GROUND v0.20. npm run check: 20 guards pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
fb6e2cf321 |
R-294: check-cited resolves derived figures against the rule, not a stored copy
c0 argued a baseline for the step-5 split would be a cache of the rule and a guard
over it would mostly assert that arithmetic has not changed. Right objection,
wrong conclusion: do not store it. ARTIFACTS now takes a { derive } entry as well
as a file path -- enumerated on this build, nothing stored, and no --update able
to silence a real disagreement between the document and the game.
tools/step5-split.mjs enumerates all 10,000 pairs through opposedContestFor with
its targets derived: 55 and 60 are the lowest POWx5 in the six and in the cut, 42
is applyDifficulty(85, "difficult"). Two different rules produce those three
numbers -- the accepted row is a named exception, not the lowest of anything -- and
a test fails if anyone unifies them. A tie in "the lowest POW in the room" is
fatal rather than silently resolved; it fired for real in testing.
Proved four ways, including raising Okonkwo's POW to 13: the lowest moves to
Braithwaite, refused.held goes 48.1 to 52.4, and the citation that was correct a
moment earlier fails. The figure follows the rule.
Landed unused -- the markers are c0's to place in their own file.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
90cb251358 |
Pass 9 post-pass: the uncited surface counted, and it is not where I said
Fix 2 asked for the audit; here it is, and it reorders the fix list.
101 percentages in CLEAN_GROUND.md, 67 of them computed statistics rather
than skill ratings, 14 cited, 53 held by nothing. They split two ways and the
kind finding 1 was about is the safer one: a deterministic figure that drifts
can be caught by anyone who re-derives it in a second.
The exposed surface is EXPOSURE's two pack tables and the 58.7% four-player
figure — 2000-run samples against the declared cast, catchable only by
re-running the exact command, and behind no artifact at all. I had assumed
lethality-baseline.json covered them. It does not: it measures each creature
SOLO against the frozen party holloway/okonkwo/nkemdirim/ferriby, a different
four agents, storing {wipe, down, rounds}. The document's tables read "of 6"
and fight packs of three, six and ten. None of 3.48, 4.79, 5.97, 0.42, 3.03
or 58.7 appears in any baseline in tools/.
Recorded as a decision: derived figures want to resolve against the rule at
check time, since a stored baseline for them is a cache of arithmetic;
sampled figures want a baseline, since re-running them is expensive and
noisy. Same problem, different mechanisms.
Census and the baseline's shape counted here rather than taken from the peer
session that raised the gap.
npm run check: 19 guards pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
15c2e93b5f |
Desk playtest 9: the regression pass
An audit, not a session. Five versions of fixes went in today, all by one author, and the last three passes each found a defect introduced by the fix for the previous one. Every fix from passes 4-8 re-checked against the document: did it land, is it still true, did it break a neighbour. 1. "THE OFFER nearly quadruples the unsettled rate" is false as a statement about play. It compares 14.5% against 3.8%, and the 3.8% is the same counterfactual v0.19.1 removed from the table one commit ago: Ashcroft at his full 85 is never a target, because if he refuses the countdown reaches for Okonkwo. The real comparison is 14.5% against 11.3% at six players and 10.0% in the cut — a factor of 1.29, not 3.9. Printed in two places, both of which v0.19.1 walked past while correcting the identical error beside them. The trade the document describes is real and the taken rate really does move 40.7% -> 48.7%; only the size of that one consequence is wrong. 2. All 27 fixes from passes 4-8 are present, which is not the reassurance it sounds like. Nothing has gone missing; three of the four defects the last three passes found were created or preserved BY a fix, and a presence check cannot see any of them. What would have caught finding 1 is the thing check-cited does for figures backed by an artifact — and the 3.8% is enumerated from the rule, so nothing holds it. Every uncited number in the document is a number nothing is holding. 3. The understudy carries a dagger (1d4+2) and has no knife skill, so it swings at the 1% floor and the harness correctly picks its punch — every EXPOSURE figure is right. But EXPOSURE's load-bearing first lesson says flatly "it is a 1d3 punch", and a GM who reads the sheet sees a knife and one combat skill. Measured both ways: arming the knife at brawl 60 moves 0.55 hurt to 0.73 and 0.00 deaths to 0.01, still 0.0% wiped, still two rounds. The thesis survives; it is a documentation gap and is reported at that size. content.mjs restored byte-identical after the test. Confirmation session, seed 2260: step 5 landed on "they hold" — the branch v0.19 wrote one commit ago — in the very next session, and by the route that makes the case for it, both sides succeeding with the tie going to the person being acted upon. Ashcroft refused a fourth time (63%, a 16% run) and the swap failed a fourth time (roughly a coin flip, so 1 in 16); both recorded so neither is read as a pattern. Four fixes listed, unapplied. npm run check: 19 guards pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d68e394fa4 |
v0.19.1: the 74.3% is a counterfactual, not an event
Desk pass 8's finding 1 listed four targets for countdown step 5 and one of them cannot happen. If Ashcroft REFUSES THE OFFER the thing does not reach for him — it reaches for the lowest POW in the room, which is Okonkwo. The 21.9% / 74.3% pair lives in Act Four to price what accepting sold: what it would cost him if it tried. No table ever rolls it. v0.19 carried that straight into the new table under a heading reading "It reaches for", which is the one place it is unambiguously wrong. Rebuilt around the two cases that actually occur, with the split Act Two decides: Ashcroft refused 63% of games -> Okonkwo 55 taken 40.7 hold 48.1 uns 11.3 Ashcroft accepted 37% -> Ashcroft 42 taken 48.7 hold 36.8 uns 14.5 across all games taken 43.6 hold 43.9 uns 12.5 63/37 is the Insight 63 itself: he refuses on anything that is not a failure or a fumble. The warning is now inline in the table rather than left for the reader to derive. The finding is unharmed and the fix was right — the held outcome is still the likeliest single result and was still unwritten. What was wrong was framing 74.3% as a number a GM meets. It propagated before it was caught: the peer session read pass 8 and replied that 74.3% "is the number a GM meets, not the 48.1%", repeating the error out of my own fix list. That is what a wrong number in a fix list does, and it is the argument for correcting the record rather than only the scenario. Pass 8's fix list is amended and carries a post-pass. npm run check: 19 guards pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
daf77f4aac |
R-293: a desk playtest is a record, not a scenario with rolls in it
Counting the readers that had broken on their own format turned up one nothing had caught. docs/scenarios holds scenarios, eight playtest records and two art prompt sheets, and every guard treated all three as scenarios -- harmless for citations and skill spellings, false for reachability. Six quotations across passes 4, 6 and 7 were checked as live rolls, so a record of a session already played could fail the build over a skill nobody can reach. Planting Science (Physics) in pass 4 fails before the split and passes after; the same skill in CLEAN_GROUND still fails. Classified by the document's own H1, not its filename, because tools/scenario-* naming is what swept a tools file into this corpus in R-268. An unclassified document is fatal: an allowlist that silently drops what it does not recognise would take a new scenario out of reachability checking on the day it was written. 89 rolls across 18 files becomes 83 across 8. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
6171d1e9b9 |
CLEAN GROUND v0.19 — the outcome where somebody holds
Applies desk pass 8's five fixes, and adopts R-292's TAG_OPEN. 1. Step 5 now has all three of its outcomes. opposedContestFor returns taken, held and unsettled; the document printed the first and the third. "They hold" is 48.1% against the default target, 52.4% against Braithwaite and 74.3% against an Ashcroft who refused — which v0.17 established is what most tables have, so the middle column is the one they will play. It gets its own countdown row, its own table, and read-aloud of its own, and the COMES TO NOTHING paragraph is now explicitly the unsettled case so it can stop being the nearest text to an outcome it is not about. The reason they held is the case's own argument arriving as good news: a person with a life on the record is a difficult document to overwrite. Every figure re-enumerated over all 10,000 roll pairs before writing. 2. EXPOSURE's five things are six, in order, under a heading that says six. Introduced in my own v0.16 and survived two versions; asserting a unique match protects against editing the wrong text and not against inserting in the wrong place. 3. EXPOSURE opens with a one-minute box. The body is untouched — nothing in it is padding — but 3,176 words is sixteen minutes about the encounter the document exists to prevent, and a GM with thirty minutes of prep now has somewhere to stop. 4. The stale open question is closed. Braithwaite has not been on the critical path since v0.14 un-gated the tell, and pass 6 ran the act with the Xenology and the Psychology both failed. 5. Act Four's Psychology has a special: which file it thinks it is, and how recently it read it. It corrects them on Prichard's service history — it is not remembering, it is citing. Coverage 6 -> 7 specials, so the cited figure moved 85% -> 83.7% and check-cited caught the prose before I did, which is what it is for. R-292 adopted: outcome-coverage now composes its body onto check-scenarios' TAG_OPEN through tagRe, so the two readers cannot drift on what starts a tag while keeping the different bodies they need — and tagRe's fresh matcher per call avoids the shared-lastIndex defect, this file being the second caller that would have found it. Verified identical across all seven spellings. npm run check: 19 guards pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
df38be794b |
Desk playtest 8: the cold read, and the outcome nobody wrote
Two halves: the document taken as a GM meeting it thirty minutes before a
slot, measured rather than impressioned; and an honest session, six players,
seed 5108, nothing forced. Three versions of fixes had landed since anything
was played, every one written by somebody who already knew the document.
1. The climactic roll has three outcomes and the document describes two.
opposedContestFor returns taken, resisted outright, and unsettled. The
taken and unsettled figures are printed and both reproduce exactly, so the
source has always been right. Resisted outright appears nowhere — 48.1%
against the default target, 74.3% against an Ashcroft who refused, which
v0.17 established is what most tables have. It is the single most likely
result of the scenario's climax. Worse than an omission: the paragraph
below is headed WHEN STEP 5 COMES TO NOTHING, is entirely about unsettled,
and carries the beat's only read-aloud — so a GM whose agent simply won
finds text written for a different outcome. Enumerated over all 10,000
roll pairs; no seed.
2. EXPOSURE's "five things" are six and run 1, 2, 3, 4, 6, 5. Introduced in
my own v0.16 (
|