# Self-audit, v0.1, 28 September 2026

An adversarial pass over the draft asking what a hostile reviewer would say. Ten findings.
Six were defects and are fixed. Four are structural and are now disclosed rather than fixed,
because fixing them needs work that has not been done.

---

## Fixed

**1. The headline claimed an effect the data does not support.** The paper said the operating
system gap was 2.6 points. At 76 of 76 against 74 of 76 the two sided Fisher exact p is
**0.497**, and the Wilson intervals, 95.2 to 100.0 against 90.9 to 99.3, overlap almost
entirely. There is no evidence of a platform effect in that metric.

This was the most serious defect in the paper and the fix improved the argument rather than
weakening it. The claim is now that **the survival metric cannot see a platform effect at all**,
while a second instrument, looking elsewhere, finds six. The paper's thesis was always that the
operating system's contribution is invisible to conventional agent metrics. It is now literally
true, with a p value.

**2. No confidence intervals anywhere.** Four percentages were quoted to one decimal place with
no indication of precision, which is the clearest signal of an unserious quantitative paper.
Wilson intervals are now in the results table. They also make the shell result stronger:
90.9 to 99.3 against 17.7 to 37.2 does not need a test.

**3. A confounded finding was attributed to the operating system.** Divergence six, the `grep`
end anchor against CRLF, was presented as a platform behaviour. A follow up test showed the file
is byte identical on both arms, five bytes, `b'abc\r\n'`, so the difference is entirely inside
`grep`. **And the arms run different greps: 3.0 under Git Bash, 3.11 on Debian.** Version and
MSYS2 text handling cannot be separated without pinning, which was not attempted.

Divergence six is withdrawn from the platform claim and relabelled a tool difference of unknown
cause. `bash` 5.2.37 and `sed` 4.9 are identical on both arms, so divergences one to five are
not version confounded, and that is now stated rather than assumed.

**4. A contaminated denominator was used without offering the clean one.** Nine of the 76 probes
test tool presence with `> /dev/null`, invalid in PowerShell, so they measure redirection. The
paper disclosed this and still quoted 26.3%. Both figures are now given: excluding those nine,
PowerShell 5.1 is 29.9% and 7.6.6 is 49.3%. The contaminated pair stays in the body because it
is the conservative direction for the Git Bash claim, and the reader can now use either.

**5. A directed sample was formatted to look like a rate.** "6 of 24" invites a reading of 25%
of commands diverging. The 24 probes were chosen by looking where divergence was suspected, so
the denominator describes the author's attention, not the environment. Restated as counts, with
an explicit refusal to offer a rate, and with the note that **the zero does not depend on the
sampling**: however the probes were chosen, none produced a non-zero exit.

**6. Two ledger rows were mistyped.** G027 and G042 were marked `external-cited` while their
sources are the author's own observations on this machine. Retyped `own-data-derived`. The
ledger is now 70 claims: 66 own-data-derived, 4 assumption, **0 external-cited**, which is
itself worth noticing. This paper cites no outside source for any load bearing claim, and that
is a weakness listed below rather than a virtue.

---

## Disclosed, not fixed

**7. Toolchain and locale were never captured by the instruments.** They were collected by hand
after the runs, which is how the `grep` confound was found. Locale differs on every arm and none
matches another: Debian `en_US.UTF-8`, Git Bash `LANG` unset with `LC_CTYPE=C.UTF-8`, PowerShell
code page 437. No result is known to be caused by this and none is known not to be. The probes
should record their own environment; they now do not.

**8. One pass per arm, no repetition, no measured variance.** The Wilson intervals are binomial
on 76 commands, which is a statement about sampling commands, not about run to run stability.
Repeatability is asserted.

**9. One coder, no second reader, no inter-rater figure.** Paper 01 of this programme reported
0.93 after a failed first round. This has nothing equivalent. The mitigation is that the
classifications are mechanical and the raw bytes are published.

**10. No external sources at all.** Every load bearing claim is own-data-derived. The programme
standard is that load bearing claims target three of five source classes; this paper reaches
one. It has no vendor documentation on shell behaviour, no published work on POSIX portability,
nothing from the tool's own release notes. That is the single largest gap between this and the
other papers in the programme, and closing it needs reading rather than measuring.

---

## What would make it materially better, in order

1. **Run the model.** Everything here is the floor. Whether an agent falls into the six traps is
   the question a reader actually has.
2. **A native idiom corpus per shell.** Without it, measuring PowerShell on POSIX is close to
   tautological, and the shell finding is weaker than it looks.
3. **Cite something.** One source class is not a triangulated claim.
4. **Pin the toolchain and record it from inside the instruments.**
5. **A second coder on the silent probe**, or at minimum a second machine per platform.
6. **WSL2**, which removes the hardware question entirely.

Items 2 and 3 are the difference between an interesting measurement and a paper. Item 1 is the
difference between this and the study that was planned.
