Everything behind the paper: the data, the code that produced every table, and the ledger of every claim considered, including the ones that failed checking.
| file | size | what it is |
|---|---|---|
| CORRECTION-2026-08-21.md | 3.1 KB | a note published with the paper |
| LABELLING-RULE.md | 4.3 KB | a note published with the paper |
| PLAN.md | 6.9 KB | a note published with the paper |
| PROGRESS.md | 2.2 KB | what was done, what is open, and what was found wrong |
| analysis-ablation.md | 6.3 KB | a note published with the paper |
| analysis-axium-versus.md | 9.2 KB | a note published with the paper |
| analysis-brownfield.md | 6.6 KB | a note published with the paper |
| analysis-classifier.md | 7.5 KB | a note published with the paper |
| analysis-shopkit.md | 5.9 KB | a note published with the paper |
| assemble.py | 3.0 KB | the code that produced or checked the numbers |
| axium-bench.jsonl | 141.9 KB | one record per line, as the run produced it |
| axium-benchmarkability.md | 4.4 KB | a note published with the paper |
| bench-reps.jsonl | 63.9 KB | one record per line, as the run produced it |
| bench-results.jsonl | 3.9 KB | one record per line, as the run produced it |
| brownfield-deduped.csv | 8.8 KB | a table, recomputable by anyone holding this file |
| claim-ledger.csv | 11.5 KB | every claim in the paper, its status and its source |
| comparison-feasibility.md | 7.3 KB | a note published with the paper |
| existing-evals.csv | 151.2 KB | a table, recomputable by anyone holding this file |
| field-dictionary.md | 6.7 KB | a note published with the paper |
| fix_graders.py | 5.8 KB | the code that produced or checked the numbers |
| grader_probe.py | 2.8 KB | the code that produced or checked the numbers |
| hermes_adapter.py | 8.2 KB | the code that produced or checked the numbers |
| licence.txt | 0.5 KB | the licence for this package |
| nul_probe.py | 2.9 KB | the code that produced or checked the numbers |
| parity_notes.md | 2.9 KB | a note published with the paper |
| prompts.csv | 8.9 KB | a table, recomputable by anyone holding this file |
| results.csv | 13.8 KB | a table, recomputable by anyone holding this file |
| run_hermes.py | 3.7 KB | the code that produced or checked the numbers |
| shopkit-bench-README.md | 4.7 KB | a note published with the paper |
| source-log.csv | 2.8 KB | every source consulted, with its retrieval date |
| technique-inventory.md | 11.4 KB | a note published with the paper |
| threeway.py | 4.6 KB | the code that produced or checked the numbers |
| v4_toolargs.json | 6.8 KB | structured output kept exactly as recorded |
| versus-axium-20260807.jsonl | 98.2 KB | one record per line, as the run produced it |
| versus-axium-fixedgrader.jsonl | 98.2 KB | one record per line, as the run produced it |
| versus-axium.jsonl | 99.7 KB | one record per line, as the run produced it |
| versus-hermes.jsonl | 78.0 KB | one record per line, as the run produced it |
| versus-orange-20260807.jsonl | 66.1 KB | one record per line, as the run produced it |
| versus-orange-baseline-20260807.jsonl | 56.0 KB | one record per line, as the run produced it |
| versus-orange-fixedgrader.jsonl | 56.0 KB | one record per line, as the run produced it |
| versus-orange.jsonl | 57.4 KB | one record per line, as the run produced it |
| versus_head2head.py | 4.3 KB | the code that produced or checked the numbers |
Every figure in the paper can be recomputed from these files. If one of them disagrees with the paper, the file is right and the paper is wrong, and it will be corrected in the change log.