# Calibration: L-bash arm

Run 28 September 2026 on the Linux system, at load 0.02 one minute, under `nice -n 15` and
`ionice -c3`. All four bot services verified active before and after. Raw result in
`L-bash.json`.

| Measure | Value |
|---|---|
| CPU, fixed loop | 35.1 ms |
| Bulk write, 64 MB | 1,608 MB/s |
| Small files, create and close | 86,836 files/s |
| Stat | 852,152 /s |
| Shell spawn, `/bin/sh -c` | **0.765 ms median**, 1.06 ms max |
| TCP to the API host | 5.5 ms, TLS handshake 9.4 ms |

**Machine:** Debian 13, kernel 6.12.95, AMD Ryzen 3 4300U, 4 cores, 15 GB RAM, M.2 SSD.

The shell spawn figure is the one to watch. It is the unit of a tool call, and at 0.765 ms
it is effectively free on this arm. Whether that holds on Windows is close to the whole
question the paper asks.

## A correction, on the first run of the first instrument

The first version reported **81,000 fsync'd file creations per second**. That is not a
believable number for any storage device, and asking whether it was believable took about
two seconds, which is the check the programme's retrospective says to apply before asking
whether a result reproduces. It reproduced perfectly. It was also wrong.

A direct test settled it: 300 files with `fsync` and 300 files without take the same time,
**ratio 0.99**. The `fsync` is not reaching the device on this machine, so the disk measures
were timing the page cache and the drive's volatile cache, not durability.

The danger was not the wrong number. It was that **this is not uniform across arms.** Had
the instrument gone on to the Windows arms unchanged, an arm whose `fsync` is real would
have been compared against an arm whose `fsync` is a no-op, and the difference would have
appeared in the results as a large filesystem gap between operating systems. A fabricated
headline, produced by the calibration instrument that exists to prevent fabricated headlines.

**What changed.** The disk measures no longer call `fsync` at all, on any arm, so they
measure cache throughput and are labelled as such. A new `fsync_honoured` diagnostic records
the with and without ratio per arm, so a reader can see whether durability was measurable
there. It is reported, not corrected: whether a machine's disk honours a flush is a property
of the arm, and hiding it would be the same error in the other direction.

This is the third instrument in this programme to produce a confident falsehood on an early
run, and like the others it was caught by looking at the output rather than by reading the
code, which looked correct throughout.

---

## Results: three arms calibrated

Best of five repeats per Windows arm, best of three on Linux, taken as the minimum elapsed time because that is the sample least disturbed by everything else on the machine. Linux figures use `bash`, not `dash`, so the shell comparison is like for like. Spread is the ratio of worst to best across repeats.

| Measure | L-bash | W-gitbash | W-ps51 |
|---|---|---|---|
| CPU, fixed loop | 35.1 ms | 36.2 ms | 35.6 ms |
| **Shell spawn** | **1.4 ms** | **46.3 ms** | **223.6 ms** |
| Small files created/s | 86,550 | 2,137 | 2,108 |
| Stat/s | 862,813 | 98,551 | 109,745 |
| Run to run spread | 1.01x to 1.10x | 1.11x to 1.59x | 1.07x to 1.74x |
| fsync honoured | no | yes | yes |

### What the numbers say, and what they do not

**The two machines have the same core speed.** 35.1 ms against 36.2 and 35.6 on a fixed loop, inside 3%. That is the single most useful line in the table, because it removes the hardware objection from every ratio below it. Neither is a faster computer. The difference is not the hardware.

**Spawning a shell is 33 times slower under Git Bash and 160 times slower under PowerShell 5.1.** This is the unit of a tool call. An agent that issues forty commands to finish a task spends about 60 ms of its life on process creation on Linux and about 9 seconds on PowerShell, before any of those commands does anything. Shell spawn is also the most stable measure in the set, varying by 7 to 11% across repeats where the others vary by up to 74%, so it is the number to build on.

**Creating small files is roughly 41 times slower on Windows, and stat about 9 times.** Agents read and write many small files, so this is not an academic measure. Both Windows arms agree closely with each other, which is the instrument validating itself: they share a machine, so they should agree on everything except the shell, and they do.

**Windows is also less repeatable.** Linux repeats land within 1 to 10% of each other. Windows repeats spread up to 1.74x on the same machine, minutes apart, on identical work. Whatever causes that (file scanning is the first suspect, and the plan already has it as a recorded sub-condition) is noise the agent lives inside. **A platform whose timings move by 74% between identical runs needs more repetitions to say anything, and that is a cost the study must pay rather than wish away.**

**What this is not.** It is a hardware and environment calibration, not the study. It says the environment is slower and noisier, and it says nothing yet about whether the model does worse work in it. Slow is not wrong. The paper only has a finding when the task suite shows failed commands and corrective reissues, and that run has not happened.

## Blocked arms

**W-wsl2 is not available: WSL is not installed on the Windows system.** This is the arm the plan calls the single most informative contrast, because it is the only one that runs a Linux kernel on the same hardware and so removes the machine from the comparison entirely. Without it, every Linux against Windows ratio above rests on the CPU loop agreeing to within 3%, which is good evidence and is not the same thing as no hardware difference at all. Installing it needs virtualisation features enabled and probably a reboot, so it needs Nikolaos to say go.

**W-ps7 is not available: PowerShell 7 is not installed**, only 5.1. It is a sensitivity arm rather than a main one, and PowerShell 7 starts faster than 5.1, so its absence currently makes the Windows shell figures look worse than a modern Windows setup would. Worth installing before the paper quotes a PowerShell number.
