Both corpora, the classification of all sixty companies with a reason on every row, and the rules that governed it, written before any company was scored.
| file | size | what it is |
|---|---|---|
| coding-rules.md | 5.8 KB | the seven product shapes, the classification rule, and the threshold that would have falsified the paper. Written before any company was scored and unamended since |
| failures-coded.csv | 6.0 KB | the sixty sampled failures, each with its description, its classification and the reason for it. The eight that could not be classified are marked rather than guessed at |
| unicorns.json | 223.7 KB | 1,404 companies valued at a billion dollars or more, with valuation, entry date, country and industry, as retrieved |
| failures.json | 177.1 KB | 410 startup failure post-mortems parsed to structured records, including the 144 whose entries never say what the product did |
| claim-ledger.csv | 7.2 KB | every claim in the paper, what kind of claim it is, what it rests on, and the two that failed checking before publication |
| source-log.csv | 1.9 KB | every source consulted, with its class, its retrieval date and what it was used for |
The two files at the top are the ones to open if you want to disagree with the paper. The rules say what a classification means and what result would have killed the argument. The coded file says how every judgement went and why.
Re-classify the sixty and get a materially different number, and the paper is wrong. That is the intended use.
Every entry in the failure compilation carrying a description of what the product did was in frame, which is 266 of the 410. Sixty were drawn from those with a fixed random seed of 20260901, so the same sixty come out of every run and the sample could not be adjusted after the result was known.
The winners were not classified for product shape. The billion-dollar list carries no product descriptions, and classifying fourteen hundred companies from their names would have imported exactly the outcome knowledge the blind rule exists to keep out. The paper therefore compares against a ceiling of 73.9 per cent rather than a measured winner rate, and says so in its limitations.
Text and data under CC BY 4.0. The two corpora are parsed reproductions of publicly published listings, retrieved 1 September 2026 and preserved here so the paper's figures remain checkable after the sources change.