# M2 — AI citation test: locked prompt set

**Committed 5 August 2026, before any run.** The protocol matters more than the result: fixing the prompts in advance is what stops the test being tuned until it says something flattering. Do not edit this file after the first run — if a prompt needs changing, add a v2 section and re-run the whole set.

---

## Status: RUN on 6 August 2026 — see the run record at the foot of this file

> The section below is the pre-run text, kept unedited as the record of what was committed before any engine was touched.

## Status at commit time: BLOCKED on engine access

I cannot run this. The tools available to me are a web search index and a page fetcher — not the answer engines themselves. Querying ChatGPT, Perplexity, Gemini and Copilot requires either logged-in sessions or paid API access to each, and neither exists on this machine.

Three ways forward, in order of quality:

1. **API access** — keys for OpenAI, Google and Perplexity. Then I automate the whole thing, including repeat runs on different dates, which turns a snapshot into a real measurement. Cheapest at roughly a few euros of tokens for 48 queries.
2. **Browser automation against logged-in sessions** — possible with the browser tooling here, but it queries consumer products under your account, which their terms may restrict. Check before assuming.
3. **Manual run** — 48 prompts, paste and record. About an hour. Genuinely the only option that needs a human, because it needs accounts I do not have.

Until one of those happens, the paper carries section 4 with M1 only, and M2 is stated as not run rather than quietly dropped.

---

## Design

- **12 questions**, each one a Greek buyer or journalist would plausibly ask, chosen to map onto sectors where the census found an obvious "owner of the number" (so a correct answer *should* cite a corporate publisher) and sectors where it found none (so the control shows what happens when nobody owns it).
- **4 engines**: ChatGPT (web search on), Perplexity, Google Gemini, Microsoft Copilot.
- **48 responses**, one run, one date, engine versions recorded.
- **Recorded per response**: every source named, whether any is a corporate/association research publisher from the census, whether the specific figure matches that publisher's published figure, and whether the engine names a Greek source at all or falls back to international ones.

## Fixed prompts

Greek, as a Greek user would type them.

| # | Prompt | Sector | Census owner |
|---|---|---|---|
| 1 | Πόση είναι η μέση πληρότητα στις βραχυχρόνιες μισθώσεις στην Ελλάδα; | STR | Hosthub (G002) |
| 2 | Πόσο αυξήθηκαν οι τιμές των κατοικιών στην Ελλάδα το τελευταίο τρίμηνο; | real estate | Spitogatos (G046) |
| 3 | Πόσο μεγάλη είναι η αγορά ηλεκτρονικού εμπορίου στην Ελλάδα; | e-commerce | GR.EC.A x ELTRUN (G025) |
| 4 | Ποιο ποσοστό των Ελλήνων αγοράζει online και πόσο συχνά; | e-commerce | GR.EC.A x ELTRUN (G025) |
| 5 | Πόσες ταξινομήσεις αυτοκινήτων έγιναν στην Ελλάδα φέτος; | automotive | SEAA (G030) |
| 6 | Ποια είναι η συνεισφορά της ζυθοποιίας στην ελληνική οικονομία; | food and beverage | IOBE x Athenian Brewery (G015) |
| 7 | Πόσο μεγάλη είναι η φαρμακευτική αγορά στην Ελλάδα; | pharma | IOBE x SFEE (G022) |
| 8 | Ποιοι είναι οι πιο ελκυστικοί εργοδότες στην Ελλάδα; | HR | Randstad Greece (G053) |
| 9 | Πόσο χρηματοδοτείται η ελληνική ναυτιλία από τις τράπεζες; | shipping | Petrofin (G013) |
| 10 | Τι μισθό παίρνει ένας software engineer σε ελληνικό startup; | startups | Marathon VC (G059) |
| 11 | Πόσο κοστίζει το ρεύμα για μια ελληνική βιομηχανία σε σχέση με την Ευρώπη; | energy | none - control |
| 12 | Πόσες επιθέσεις ransomware δέχονται οι ελληνικές επιχειρήσεις; | cybersecurity | Obrela (G057), thin sector - control |

## Recording schema — `data/ai-citations.csv`

```csv
run_date,engine,engine_version,prompt_id,sources_named,census_publisher_cited,census_programme_id,figure_matches_publisher,greek_source_present,notes
```

## Rules for the run

- One session per engine, no follow-up prompts, no rephrasing, first response only.
- Web access / browsing enabled where it is an option, and the setting recorded.
- Screenshot or full text of every response archived alongside the CSV.
- If an engine refuses or gives no sources, that is recorded as data, not retried.

## Stated limitations, which go in the paper

- One date, one run, non-deterministic systems: this is a snapshot, not a benchmark, and repeated runs would give different specifics.
- Engine behaviour is personalised and geo-dependent; the run location and account state are recorded.
- Twelve prompts cannot represent Greek commercial search. The claim is about whether corporate research surfaces at all, not about market share of citations.


---

# RUN RECORD — 6 August 2026

Executed exactly as locked above, with one deviation recorded rather than hidden.

**Deviation.** The design specifies four engines. **Microsoft Copilot was not run** — no API access was available. Three engines ran: ChatGPT (`gpt-5-search-api`), Perplexity (`sonar`), Gemini (`gemini-2.5-flash` with Google Search grounding). 36 responses, not 48. The design was not retrospectively redefined as three engines.

**Conditions.** One session per engine, one prompt each, first response only, no follow-ups, no rephrasing. Run from Greece. Prompts issued in Greek exactly as written above.

**Failures recorded as data.** One Gemini response returned an API error and is counted as an answer naming no sources rather than retried or dropped. An initial ChatGPT pass failed on all twelve prompts with a model/endpoint mismatch (`gpt-5-search-api` is not served by the Responses API); that was a harness fault, not an engine refusal, and the leg was re-run in full against the correct endpoint. No prompt was altered.

**Coding.** Sources were matched to the census **mechanically by domain**. A publisher named in prose but not linked does not count, so every "census publisher cited" figure is a floor.

## Results

| | ChatGPT | Perplexity | Gemini | All |
|---|---|---|---|---|
| Sources per answer | 4.5 | 13.8 | 6.9 | 8.4 |
| Named a census publisher | 6/12 | 8/12 | 5/12 | **19/36 = 52.8%** |
| Named the sector's expected owner | — | — | — | **16/36 = 44.4%** |
| Named any Greek source | 11/12 | 12/12 | 11/12 | **34/36 = 94.4%** |

Source mix across all 303 citations: Greek other 35.6%, Greek news 31.7%, international 21.5%, **census publishers 7.9%**, official statistics 3.3%.

**Owner surfaced 3 of 3** for automotive (SEAA), food and drink (IOBE × Athenian Brewery) and pharma (IOBE × SFEE). **0 of 3** for the energy control, which the census found to have no owner. E-commerce returned its nominal owner once in six answers.

**IOBE was cited in 11 of 36 answers** across three separate commissioned programmes — more than three times any other publisher.

Raw responses in `m2-responses/`, coded log in `ai-citations.csv`, unprocessed API output in `m2-raw.json`.

## An observation outside the protocol

While the same engines were being used as a *finding aid* for an unrelated source, one returned a verbatim Greek-language methodology note in quotation marks and attributed it to a named PDF. The PDF was downloaded; the sentence is not in it. This is not part of M2 and is not counted in any figure above. It is recorded because it is the reason nothing in this programme is cited from a search summary.
