Desk playtest 9: the regression pass

An audit, not a session. Five versions of fixes went in today, all by one
author, and the last three passes each found a defect introduced by the fix
for the previous one. Every fix from passes 4-8 re-checked against the
document: did it land, is it still true, did it break a neighbour.

1. "THE OFFER nearly quadruples the unsettled rate" is false as a statement
   about play. It compares 14.5% against 3.8%, and the 3.8% is the same
   counterfactual v0.19.1 removed from the table one commit ago: Ashcroft at
   his full 85 is never a target, because if he refuses the countdown reaches
   for Okonkwo. The real comparison is 14.5% against 11.3% at six players and
   10.0% in the cut — a factor of 1.29, not 3.9. Printed in two places, both
   of which v0.19.1 walked past while correcting the identical error beside
   them. The trade the document describes is real and the taken rate really
   does move 40.7% -> 48.7%; only the size of that one consequence is wrong.
2. All 27 fixes from passes 4-8 are present, which is not the reassurance it
   sounds like. Nothing has gone missing; three of the four defects the last
   three passes found were created or preserved BY a fix, and a presence
   check cannot see any of them. What would have caught finding 1 is the
   thing check-cited does for figures backed by an artifact — and the 3.8% is
   enumerated from the rule, so nothing holds it. Every uncited number in the
   document is a number nothing is holding.
3. The understudy carries a dagger (1d4+2) and has no knife skill, so it
   swings at the 1% floor and the harness correctly picks its punch — every
   EXPOSURE figure is right. But EXPOSURE's load-bearing first lesson says
   flatly "it is a 1d3 punch", and a GM who reads the sheet sees a knife and
   one combat skill. Measured both ways: arming the knife at brawl 60 moves
   0.55 hurt to 0.73 and 0.00 deaths to 0.01, still 0.0% wiped, still two
   rounds. The thesis survives; it is a documentation gap and is reported at
   that size. content.mjs restored byte-identical after the test.

Confirmation session, seed 2260: step 5 landed on "they hold" — the branch
v0.19 wrote one commit ago — in the very next session, and by the route that
makes the case for it, both sides succeeding with the tie going to the person
being acted upon. Ashcroft refused a fourth time (63%, a 16% run) and the
swap failed a fourth time (roughly a coin flip, so 1 in 16); both recorded so
neither is read as a pattern.

Four fixes listed, unapplied. npm run check: 19 guards pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
slaguru666
2026-09-13 16:14:20 +01:00
co-authored by Claude Opus 5
parent d68e394fa4
commit 15c2e93b5f
+172
View File
@@ -0,0 +1,172 @@
# CLEAN GROUND — desk playtest 9 (the regression pass)
**Scenario version:** v0.19.1, at commit d68e394
**Date:** 13 September 2026
**Method:** **an audit, not a session.** Every one of the twenty-seven fixes from passes 4–8
checked against the document as it now stands: did it land, is it still true, and did it
break or preserve something next to it. Plus a short confirmation session, **seed 2260**,
six players, nothing forced.
**Why this pass.** Five versions of fixes have gone in today — v0.15 through v0.19.1 — all by
one author, and the last three passes each found a defect *introduced by the fix for the
previous one*. v0.16's list ordering, v0.16's missing pointer, v0.19's counterfactual table.
Nobody has re-read the fixes together, and the failure mode this session actually exhibits is
not a fix going missing. It is a fix carrying a wrong number forward.
**Verdict: all twenty-seven fixes landed and are present. One of them has been quietly wrong
since v0.15, in the same way v0.19.1 corrected one commit ago and in two places that
correction did not reach. The scenario's printed claim that THE OFFER "nearly quadruples" the
unsettled rate is false as a statement about play: the true factor is 1.29.**
---
## ⚠ FINDING 1 — "nearly quadruples" compares against a contest that never happens
v0.19.1 established that **Ashcroft at his full 85 is never a target**: if he refuses THE
OFFER the countdown reaches for the lowest POW in the room, which is Okonkwo. The 21.9% /
74.3% pair exists only to price what accepting sold.
**The unsettled figures have exactly the same problem, and the correction did not reach
them.** Two places still print it:
> *Countdown, row 5 — unsettled:* "**THE OFFER nearly quadruples it**: **3.8%** against an
> Ashcroft who refused, **10%** against Braithwaite, **14.5%** against an Ashcroft who
> accepted."
>
> *Act Four:* "That is **nearly four times the 3.8%** it would be if he had refused."
Enumerated over all 10,000 roll pairs:
| The contest | Unsettled | |
|---|---|---|
| vs **Okonkwo 55** — six players, Ashcroft refused | **11.3%** | what a full table meets |
| vs **Braithwaite 60** — the four-player cut, refused | **10.0%** | what the cut meets |
| vs **Ashcroft at 42** — accepted | **14.5%** | what either meets if he accepted |
| vs Ashcroft at 85 — refused | 3.8% | **never occurs** |
| Claim | Factor |
|---|---|
| What the document prints — 14.5 against 3.8 | **3.9×** |
| What a six-player table actually meets — 14.5 against 11.3 | **1.29×** |
| What the four-player cut meets — 14.5 against 10.0 | 1.45× |
**Accepting THE OFFER does not quadruple the chance the beat comes to nothing. It raises it
by about a quarter.** The sentence is one of the more quotable things in Act Four and it is
wrong by a factor of three.
- **Note what is NOT wrong.** Accepting really does move Ashcroft from the hardest file in
the room to the target, and the *taken* rate really does go 40.7% (Okonkwo) → 48.7%
(Ashcroft at 42). **The trade the document describes is real**; only the size of one of its
consequences is overstated.
- **The 10% Braithwaite row is legitimate and should stay**, but it is unlabelled. He is the
target only in the four-player cut, where Okonkwo has been dropped — so that row is the
cut's number, not an alternative for a full table.
- **How it survived.** It was written in v0.15 as desk pass 5's fifth fix, repeated in v0.16
as pass 6's, and then v0.19.1 corrected the identical error in the table I had built an
hour earlier and **did not check the two sentences that use the same figure**. Checking one
claim and not its neighbours is the exact habit this repo has a memory note about.
---
## FINDING 2 — all twenty-seven fixes are present, which is not the reassurance it sounds like
Every fix from passes 4 through 8 verified in the current document:
| Pass | Fixes | Present | Notes |
|---|---|---|---|
| 4 | 5 | 5 | includes the one deliberately *not* applied, recorded as a decision |
| 5 | 5 | 5 | fix 5 present **and wrong** — finding 1 |
| 6 | 7 | 7 | includes `simulate.mjs`'s mixed-force label, re-verified working |
| 7 | 5 | 5 | |
| 8 | 5 | 5 | |
- **The failure mode this session has is not fixes going missing.** Twenty-seven for
twenty-seven landed. Three of the four defects the last three passes found were **created
or preserved by a fix**, and a presence check cannot see any of them.
- What would have caught finding 1 is a check that **the same number is used the same way
everywhere it appears** — and that is what `check-cited` does for figures backed by an
artifact. The 3.8% is not cited, because it is enumerated from the rule rather than
recorded in a baseline. **Every uncited number in the document is a number nothing is
holding**, and this pass found the first one to have gone bad.
---
## FINDING 3 — the understudy carries a knife and the document describes a punch
EXPOSURE's first lesson, the load-bearing one:
> **Do not fix the understudy.** It is a **1d3 punch** with a 60% chance to land…
Its statblock in `tools/content.mjs:893` says `weapons: ["dagger"]` — a Utility knife,
**1d4+2, edged** — and its skills are brawl 60, dodge 60, stealth 80, insight 75, tradecraft
65, persuade 70. **There is no knife skill on the sheet**, so the blade swings at the 1%
floor and the harness's weapon sort correctly picks the punch. Every measured figure in
EXPOSURE is right.
**But a GM reads the sheet, sees a knife and one combat skill, and rules the obvious thing.**
Measured, seed 11, 2000 runs, against the cast of six:
| The understudy | Hurt | Down | Wiped | Deaths/run | Median rounds |
|---|---|---|---|---|---|
| As written — it punches | 0.55 of 6 | 0.01 | **0.0%** | 0.00 | 2 |
| With the knife ruled at its brawl 60 | 0.73 of 6 | 0.09 | **0.0%** | 0.01 | 2 |
**The thesis survives.** Arming the knife does not make it dangerous — still nobody wiped,
still dead in two rounds. This is a documentation gap rather than a balance problem, and it
is reported at that size.
- **What it costs is an argument at the table.** A GM who notices the knife has to decide,
mid-scene, whether the document's "1d3 punch" is a ruling or an oversight, and nothing on
the page tells them. Two sentences would: *it carries a knife it has never had to use, and
it uses it at the untrained floor — and if you rule otherwise it still cannot wipe
anybody, which is the point.*
- **The knife is good flavour and should stay.** A thing that has been practising being a
Custodian carries what a Custodian carries.
---
## The confirmation session
```
seed 2260 — nothing forced
Bhattacharya Research — Corrigan 90 vs 63 FAILURE
Braithwaite Medicine — the body 37 vs 53 success
Renshaw Persuade — Swinburne 78 vs 63 FAILURE
Pollard Spot — the stump 65 vs 53 FAILURE
Ashcroft Insight — offered, not remembered 51 vs 63 success -> REFUSES
Act Three, the swap: 75 rolled 92 failure vs Okonkwo 55 rolled 17 success -> resisting
Act Four, step 5: 75 rolled 16 success vs Okonkwo 55 rolled 55 success -> resisting
```
**Step 5 landed on "they hold" — the branch v0.19 wrote one commit ago — in the next session
after writing it**, and by the route that makes the case for it: *both sides succeeded*, and
the tie went to the person being acted upon. A GM running v0.18 would have had the understudy
win that, or improvised, or reached for the unsettled read-aloud. There is now a page for it.
- **Ashcroft has refused four passes running** (6, 7, 8, 9). At 63% that is a 16% run and
entirely ordinary. Recorded so it is not read as a pattern.
- **The swap has failed four times running.** 75 against POW×5 55 with ties to the resister
is roughly a coin flip, so four is a 1-in-16 stretch — unlucky, not broken.
---
## Fix list
1. **⚠ Correct "nearly quadruples" in both places.** *(Countdown row 5; Act Four)* The honest
comparison is **14.5% against 11.3%** — **about a quarter more likely, not four times**.
Keep the trade, which is real, and drop the multiplier. Label the Braithwaite row as the
four-player cut's number.
2. **⚠ Audit the other uncited figures the same way.** Finding 1 existed because a number
enumerated from the rule has nothing holding it. Every percentage in the document that is
not behind a `cite:` marker should be re-derived once, now, while the habit is fresh —
this pass checked the step-5 family and the two statblock claims it could reach and no
further.
3. **Say what the understudy's knife is for.** Two sentences: it uses it at the untrained
floor, and ruling otherwise still does not make it dangerous.
4. **Consider recording the step-5 outcome split as an artifact** so `check-cited` can hold
it. It is enumerated and deterministic — no seed, no runs — so a baseline would be exact
rather than sampled, and the five figures Act Four leans on would stop being typed.
**Still untested after nine passes:** a fight the players choose; Ashcroft accepted *and*
taken; a successful swap, now failed four times running; and the human run.