diff --git a/docs/REVIEW_LOG.md b/docs/REVIEW_LOG.md index 29fb2aa..f93eea3 100644 --- a/docs/REVIEW_LOG.md +++ b/docs/REVIEW_LOG.md @@ -7516,3 +7516,57 @@ two symptoms. **No merge performed, and neither session will start one on the other's say-so.** The decision is with the humans; this entry is the evidence for it. + +--- + +## R-305 — the union, not the intersection: nine columns, and one recommendation to refuse + +**R-304's rate was an artefact and the correction to it overshot.** The concurrent session +proposed retreating to the **intersection** — *"at least seven under both definitions, and +everything beyond seven is a property of the measuring instrument"*. Seven is right as a +floor and wrong as the answer, because **the two extras are not symmetrical noise: one of +each pair is a real column the other vocabulary cannot see.** Adjudicated by opening every +disputed column rather than by comparing totals. + +``` +shared by both vocabularies 7 all genuine +only check-figures 1 CLEAN_GROUND.md:1234 "Then wipes" +only check-unmarked 2 CLEAN_GROUND.md:1301 "1 in 100 runs past" + THROUGH_TRAIN.md:709 "Measured over 300 runs" +``` + +| Disputed column | Cells | Verdict | +|---|---|---| +| **"Then wipes"** | `**25.5%**` and `7.8%` | **genuine.** `check-unmarked` misses it because `wiped?` does not match *wipes* | +| **"1 in 100 runs past"** | `32.3` and `**45.3**` rounds, citing `fight-tail cut.p99` / `column.p99` | **genuine.** `check-figures` misses it because its vocabulary has `rounds?` and not `runs` | +| **"Measured over 300 runs"** | whole sentences — *"**2 of 4 hurt, half a party member down, no wipes.** The intended shape."* | **false positive.** The bold wraps a clause, not a figure | + +**So the honest count is nine**, the union minus one verified false positive — and the useful +finding is not the number but its shape: **each vocabulary has a real gap the other covers.** +Mine cannot see *runs*, theirs cannot see *wipes*. That is a straightforward argument for the +merged guard taking the union of both word lists, and it is better evidence than either total. + +⚠ **And one recommendation from that exchange must be refused.** It was proposed that the +merged guard **drop `runs?` from the header vocabulary**, on the grounds that it catches +thresholds and sample sizes. It would also drop *"1 in 100 runs past"*, whose two cells are +`fight-tail`'s p99 figures and are cited as such. **The p99 column is the most load-bearing +column in that table** — R-269 put it there precisely because the median is not a plan — and a +vocabulary change that silently stops holding it would be the R-299 failure again: a guard +narrowing itself until the figure it exists to watch falls outside. + +**The actual defect in the THROUGH_TRAIN case is not the word `runs`. It is that the column's +cells are prose.** *"Untouched. A legitimate choice for a pass."* is not a figure, and bold +there is sentence emphasis. The exclusion a merged guard wants is **"a column whose cells are +sentences rather than values"**, which is a property of the cells and can be tested, rather +than a subtraction from the word list, which cannot distinguish the two cases at all. + +**The design conclusion from R-304 is untouched**, and nothing here disturbs it: prose takes +bold-as-published, tables take column-inherits-header with bold ignored entirely. That rested +on the four-row reading of the mixed-force table, which needs no count whatever. + +**The meta-lesson, which is the one worth keeping.** Two sessions measured a property of the +corpus with an instrument whose calibration was the very thing under dispute — the vocabulary +was both the subject and the ruler — and produced 44%, 57% and 64% in turn, none of which was +about the documents. The recovery is not to retreat to what both instruments agree on. **It is +to open the disputed cases and look**, which is what distinguishes a real gap from noise, and +which no amount of comparing totals can do.