# Paper 05 progress

**Status: assembled, 10,179 words, 39 claims, all cited. Not published to the live
site.**

## Finished on 22 August 2026

| step | outcome |
|---|---|
| house rules | 57 dashes removed across eleven sections |
| ledger | 39 claims written from the paper's own claim schedule, D001 to D039, every one discharged by a section |
| ledger section | generated from the CSV into section 12 |
| assembly | twelve sections, house rules enforced, `draft/paper.md` written |

## What the paper says

Seven versions of a Windows automation agent, 270 task-executions, one model held
constant. Two of six transitions made the agent worse and both are in the body.
The best-scoring version is 25 per cent worse value than the best version, and the
paper says so.

The most useful section is the one about a chart that would have been wrong: read
from the run-level totals the last version costs 3.3 times the one before it, and
that is an artefact of a protocol change to three repetitions per task which
nothing in the output announced.

The negative control, added after the fact, found that an agent doing nothing
scores 22.8 per cent and one emitting a plausible paragraph scores 36.8. Every
comparison survives, because every version sat on the same scale; the absolute
percentages do not.

## Open

- **Publishing.** Nothing has been uploaded.
- **The transcripts are unread.** Every run wrote one, all are published, none was
  coded. The paper names this as its largest piece of unfinished work.
