Desk playtest 9: the regression pass
An audit, not a session. Five versions of fixes went in today, all by one author, and the last three passes each found a defect introduced by the fix for the previous one. Every fix from passes 4-8 re-checked against the document: did it land, is it still true, did it break a neighbour. 1. "THE OFFER nearly quadruples the unsettled rate" is false as a statement about play. It compares 14.5% against 3.8%, and the 3.8% is the same counterfactual v0.19.1 removed from the table one commit ago: Ashcroft at his full 85 is never a target, because if he refuses the countdown reaches for Okonkwo. The real comparison is 14.5% against 11.3% at six players and 10.0% in the cut — a factor of 1.29, not 3.9. Printed in two places, both of which v0.19.1 walked past while correcting the identical error beside them. The trade the document describes is real and the taken rate really does move 40.7% -> 48.7%; only the size of that one consequence is wrong. 2. All 27 fixes from passes 4-8 are present, which is not the reassurance it sounds like. Nothing has gone missing; three of the four defects the last three passes found were created or preserved BY a fix, and a presence check cannot see any of them. What would have caught finding 1 is the thing check-cited does for figures backed by an artifact — and the 3.8% is enumerated from the rule, so nothing holds it. Every uncited number in the document is a number nothing is holding. 3. The understudy carries a dagger (1d4+2) and has no knife skill, so it swings at the 1% floor and the harness correctly picks its punch — every EXPOSURE figure is right. But EXPOSURE's load-bearing first lesson says flatly "it is a 1d3 punch", and a GM who reads the sheet sees a knife and one combat skill. Measured both ways: arming the knife at brawl 60 moves 0.55 hurt to 0.73 and 0.00 deaths to 0.01, still 0.0% wiped, still two rounds. The thesis survives; it is a documentation gap and is reported at that size. content.mjs restored byte-identical after the test. Confirmation session, seed 2260: step 5 landed on "they hold" — the branch v0.19 wrote one commit ago — in the very next session, and by the route that makes the case for it, both sides succeeding with the tie going to the person being acted upon. Ashcroft refused a fourth time (63%, a 16% run) and the swap failed a fourth time (roughly a coin flip, so 1 in 16); both recorded so neither is read as a pattern. Four fixes listed, unapplied. npm run check: 19 guards pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
d68e394fa4
commit
15c2e93b5f
@@ -0,0 +1,172 @@
|
||||
# CLEAN GROUND — desk playtest 9 (the regression pass)
|
||||
|
||||
**Scenario version:** v0.19.1, at commit d68e394
|
||||
**Date:** 13 September 2026
|
||||
**Method:** **an audit, not a session.** Every one of the twenty-seven fixes from passes 4–8
|
||||
checked against the document as it now stands: did it land, is it still true, and did it
|
||||
break or preserve something next to it. Plus a short confirmation session, **seed 2260**,
|
||||
six players, nothing forced.
|
||||
|
||||
**Why this pass.** Five versions of fixes have gone in today — v0.15 through v0.19.1 — all by
|
||||
one author, and the last three passes each found a defect *introduced by the fix for the
|
||||
previous one*. v0.16's list ordering, v0.16's missing pointer, v0.19's counterfactual table.
|
||||
Nobody has re-read the fixes together, and the failure mode this session actually exhibits is
|
||||
not a fix going missing. It is a fix carrying a wrong number forward.
|
||||
|
||||
**Verdict: all twenty-seven fixes landed and are present. One of them has been quietly wrong
|
||||
since v0.15, in the same way v0.19.1 corrected one commit ago and in two places that
|
||||
correction did not reach. The scenario's printed claim that THE OFFER "nearly quadruples" the
|
||||
unsettled rate is false as a statement about play: the true factor is 1.29.**
|
||||
|
||||
---
|
||||
|
||||
## ⚠ FINDING 1 — "nearly quadruples" compares against a contest that never happens
|
||||
|
||||
v0.19.1 established that **Ashcroft at his full 85 is never a target**: if he refuses THE
|
||||
OFFER the countdown reaches for the lowest POW in the room, which is Okonkwo. The 21.9% /
|
||||
74.3% pair exists only to price what accepting sold.
|
||||
|
||||
**The unsettled figures have exactly the same problem, and the correction did not reach
|
||||
them.** Two places still print it:
|
||||
|
||||
> *Countdown, row 5 — unsettled:* "**THE OFFER nearly quadruples it**: **3.8%** against an
|
||||
> Ashcroft who refused, **10%** against Braithwaite, **14.5%** against an Ashcroft who
|
||||
> accepted."
|
||||
>
|
||||
> *Act Four:* "That is **nearly four times the 3.8%** it would be if he had refused."
|
||||
|
||||
Enumerated over all 10,000 roll pairs:
|
||||
|
||||
| The contest | Unsettled | |
|
||||
|---|---|---|
|
||||
| vs **Okonkwo 55** — six players, Ashcroft refused | **11.3%** | what a full table meets |
|
||||
| vs **Braithwaite 60** — the four-player cut, refused | **10.0%** | what the cut meets |
|
||||
| vs **Ashcroft at 42** — accepted | **14.5%** | what either meets if he accepted |
|
||||
| vs Ashcroft at 85 — refused | 3.8% | **never occurs** |
|
||||
|
||||
| Claim | Factor |
|
||||
|---|---|
|
||||
| What the document prints — 14.5 against 3.8 | **3.9×** |
|
||||
| What a six-player table actually meets — 14.5 against 11.3 | **1.29×** |
|
||||
| What the four-player cut meets — 14.5 against 10.0 | 1.45× |
|
||||
|
||||
**Accepting THE OFFER does not quadruple the chance the beat comes to nothing. It raises it
|
||||
by about a quarter.** The sentence is one of the more quotable things in Act Four and it is
|
||||
wrong by a factor of three.
|
||||
|
||||
- **Note what is NOT wrong.** Accepting really does move Ashcroft from the hardest file in
|
||||
the room to the target, and the *taken* rate really does go 40.7% (Okonkwo) → 48.7%
|
||||
(Ashcroft at 42). **The trade the document describes is real**; only the size of one of its
|
||||
consequences is overstated.
|
||||
- **The 10% Braithwaite row is legitimate and should stay**, but it is unlabelled. He is the
|
||||
target only in the four-player cut, where Okonkwo has been dropped — so that row is the
|
||||
cut's number, not an alternative for a full table.
|
||||
- **How it survived.** It was written in v0.15 as desk pass 5's fifth fix, repeated in v0.16
|
||||
as pass 6's, and then v0.19.1 corrected the identical error in the table I had built an
|
||||
hour earlier and **did not check the two sentences that use the same figure**. Checking one
|
||||
claim and not its neighbours is the exact habit this repo has a memory note about.
|
||||
|
||||
---
|
||||
|
||||
## FINDING 2 — all twenty-seven fixes are present, which is not the reassurance it sounds like
|
||||
|
||||
Every fix from passes 4 through 8 verified in the current document:
|
||||
|
||||
| Pass | Fixes | Present | Notes |
|
||||
|---|---|---|---|
|
||||
| 4 | 5 | 5 | includes the one deliberately *not* applied, recorded as a decision |
|
||||
| 5 | 5 | 5 | fix 5 present **and wrong** — finding 1 |
|
||||
| 6 | 7 | 7 | includes `simulate.mjs`'s mixed-force label, re-verified working |
|
||||
| 7 | 5 | 5 | |
|
||||
| 8 | 5 | 5 | |
|
||||
|
||||
- **The failure mode this session has is not fixes going missing.** Twenty-seven for
|
||||
twenty-seven landed. Three of the four defects the last three passes found were **created
|
||||
or preserved by a fix**, and a presence check cannot see any of them.
|
||||
- What would have caught finding 1 is a check that **the same number is used the same way
|
||||
everywhere it appears** — and that is what `check-cited` does for figures backed by an
|
||||
artifact. The 3.8% is not cited, because it is enumerated from the rule rather than
|
||||
recorded in a baseline. **Every uncited number in the document is a number nothing is
|
||||
holding**, and this pass found the first one to have gone bad.
|
||||
|
||||
---
|
||||
|
||||
## FINDING 3 — the understudy carries a knife and the document describes a punch
|
||||
|
||||
EXPOSURE's first lesson, the load-bearing one:
|
||||
|
||||
> **Do not fix the understudy.** It is a **1d3 punch** with a 60% chance to land…
|
||||
|
||||
Its statblock in `tools/content.mjs:893` says `weapons: ["dagger"]` — a Utility knife,
|
||||
**1d4+2, edged** — and its skills are brawl 60, dodge 60, stealth 80, insight 75, tradecraft
|
||||
65, persuade 70. **There is no knife skill on the sheet**, so the blade swings at the 1%
|
||||
floor and the harness's weapon sort correctly picks the punch. Every measured figure in
|
||||
EXPOSURE is right.
|
||||
|
||||
**But a GM reads the sheet, sees a knife and one combat skill, and rules the obvious thing.**
|
||||
Measured, seed 11, 2000 runs, against the cast of six:
|
||||
|
||||
| The understudy | Hurt | Down | Wiped | Deaths/run | Median rounds |
|
||||
|---|---|---|---|---|---|
|
||||
| As written — it punches | 0.55 of 6 | 0.01 | **0.0%** | 0.00 | 2 |
|
||||
| With the knife ruled at its brawl 60 | 0.73 of 6 | 0.09 | **0.0%** | 0.01 | 2 |
|
||||
|
||||
**The thesis survives.** Arming the knife does not make it dangerous — still nobody wiped,
|
||||
still dead in two rounds. This is a documentation gap rather than a balance problem, and it
|
||||
is reported at that size.
|
||||
|
||||
- **What it costs is an argument at the table.** A GM who notices the knife has to decide,
|
||||
mid-scene, whether the document's "1d3 punch" is a ruling or an oversight, and nothing on
|
||||
the page tells them. Two sentences would: *it carries a knife it has never had to use, and
|
||||
it uses it at the untrained floor — and if you rule otherwise it still cannot wipe
|
||||
anybody, which is the point.*
|
||||
- **The knife is good flavour and should stay.** A thing that has been practising being a
|
||||
Custodian carries what a Custodian carries.
|
||||
|
||||
---
|
||||
|
||||
## The confirmation session
|
||||
|
||||
```
|
||||
seed 2260 — nothing forced
|
||||
Bhattacharya Research — Corrigan 90 vs 63 FAILURE
|
||||
Braithwaite Medicine — the body 37 vs 53 success
|
||||
Renshaw Persuade — Swinburne 78 vs 63 FAILURE
|
||||
Pollard Spot — the stump 65 vs 53 FAILURE
|
||||
Ashcroft Insight — offered, not remembered 51 vs 63 success -> REFUSES
|
||||
|
||||
Act Three, the swap: 75 rolled 92 failure vs Okonkwo 55 rolled 17 success -> resisting
|
||||
Act Four, step 5: 75 rolled 16 success vs Okonkwo 55 rolled 55 success -> resisting
|
||||
```
|
||||
|
||||
**Step 5 landed on "they hold" — the branch v0.19 wrote one commit ago — in the next session
|
||||
after writing it**, and by the route that makes the case for it: *both sides succeeded*, and
|
||||
the tie went to the person being acted upon. A GM running v0.18 would have had the understudy
|
||||
win that, or improvised, or reached for the unsettled read-aloud. There is now a page for it.
|
||||
|
||||
- **Ashcroft has refused four passes running** (6, 7, 8, 9). At 63% that is a 16% run and
|
||||
entirely ordinary. Recorded so it is not read as a pattern.
|
||||
- **The swap has failed four times running.** 75 against POW×5 55 with ties to the resister
|
||||
is roughly a coin flip, so four is a 1-in-16 stretch — unlucky, not broken.
|
||||
|
||||
---
|
||||
|
||||
## Fix list
|
||||
|
||||
1. **⚠ Correct "nearly quadruples" in both places.** *(Countdown row 5; Act Four)* The honest
|
||||
comparison is **14.5% against 11.3%** — **about a quarter more likely, not four times**.
|
||||
Keep the trade, which is real, and drop the multiplier. Label the Braithwaite row as the
|
||||
four-player cut's number.
|
||||
2. **⚠ Audit the other uncited figures the same way.** Finding 1 existed because a number
|
||||
enumerated from the rule has nothing holding it. Every percentage in the document that is
|
||||
not behind a `cite:` marker should be re-derived once, now, while the habit is fresh —
|
||||
this pass checked the step-5 family and the two statblock claims it could reach and no
|
||||
further.
|
||||
3. **Say what the understudy's knife is for.** Two sentences: it uses it at the untrained
|
||||
floor, and ruling otherwise still does not make it dangerous.
|
||||
4. **Consider recording the step-5 outcome split as an artifact** so `check-cited` can hold
|
||||
it. It is enumerated and deterministic — no seed, no runs — so a baseline would be exact
|
||||
rather than sampled, and the five figures Act Four leans on would stop being typed.
|
||||
|
||||
**Still untested after nine passes:** a fight the players choose; Ashcroft accepted *and*
|
||||
taken; a successful swap, now failed four times running; and the human run.
|
||||
Reference in New Issue
Block a user