v1.0 · August 2026

Who publishes
the numbers

When a Greek business looks up a number about its own market, where does that number come from, and now that a machine chooses which one to repeat, whose gets repeated?

Download the paper
01

Why I counted

I run businesses in Greece, and almost every decision I make rests on a number somebody else published. What the market is worth. How fast it is growing. What share of it moved online. I have quoted those numbers in pitches, priced against them, and planned a year around them, without once being able to say who produced them, why, or whether anybody had ever checked.

So I set out to answer a question that turns out to have no published answer: when a Greek business looks up a number about its own market, where does that number actually come from?

The short answer is that it comes from a company with something to sell. In Greek short-term rentals the definitive figure is published by a property-management software company. In residential property prices it is a listings platform. In ship finance it is a small consultancy that has kept an index running since 2001. None of them is a statistical agency, all of them have a commercial interest in the subject, and all of them give the number away for free.

That is not a scandal. It is an arrangement, and it is the one the Greek economy actually runs on. But an arrangement nobody has described cannot be judged, and I wanted to judge it before trusting it any further.

Why this is a 2026 question and not a 2019 one

Until recently, a published number reached a decision-maker through a journalist. Somebody chose it, attributed it, and put it in a sentence a reader could trace. That path is now the minority path.

A Greek business owner asking what the e-commerce market is worth increasingly asks a model, and the model returns one number with a citation attached. The editorial layer has been replaced by a retrieval layer, and the retrieval layer has different tastes: it does not care about a press relationship, it cannot be pitched, and it will happily repeat a vendor's marketing figure with the confidence of a national statistic.

Which means the question “who publishes the numbers” has quietly become “whose numbers does the machine repeat”, and I could not find anyone who had measured that for Greece. So this paper does both: it censuses who publishes, and then it asks three answer engines thirty-six questions about the Greek economy and codes every source they cite.

What this paper is the foundation for

This is the first of seven, and it exists to establish the standard the other six are held to. If the finding is that Greek business research is dominated by commissioned surveys with undisclosed methods, then a paper claiming to do better has an obligation to show its workings, publish its data, and record what failed verification. Everything the other six papers do, the claim ledgers, the published instruments, the retractions in the body, is a response to what this census found.

It is also, deliberately, a test of whether the category is open. It is: only a quarter of Greek publication programmes are built on data the publisher already owns, against a considerably higher share internationally. That gap is the reason the other six papers were worth writing at all.

02

What this found

Somebody publishes the number your industry quotes. In Greek short-term rentals it is a property-management software company. In residential property prices it is a listings platform. In ship finance it is a small consultancy that has kept an index running since 2001. None of them is a statistical agency, and all of them give the number away.

This is the first census of that behaviour in Greece that I am aware of: 68 publication programmes from 64 publishers across 15 sectors, assembled through 36 documented searches, coded on fourteen variables1,2, and checked for coding reliability three times, once by an independent coder.

Finding 01

Greek publishers under-use the data they already own

Only 18 of 68 programmes are built on the publisher's own operational data. Against a comparable international corpus on a matched schema, Greece runs 31.6% against 42.7%8: an eleven-point deficit. Commissioned surveys are equally dominant in both.

Finding 02

Almost nobody gates

Two programmes out of 68 sit behind an email form, and both belong to software companies selling outside Greece. Internationally the figure is 13.4%, more than four times higher.

Finding 03

It is not CSR

Seven of 68 programmes are framed as contribution, footprint or responsibility: the one coding column that survived independent replication at 89.7% agreement. Category ownership, at 23 programmes, is the most common motive coded, though motive itself proved the corpus’s least reproducible column.

Finding 04

Fewer than half show their workings

Thirty of 68 disclose method fully. Six disclose nothing at all: asking readers to take a number entirely on trust.

Finding 05

The press picks up numbers, not documents

Across 25 programmes tracked, deep practical whitepapers drew zero Greek media coverage. Own-data releases averaged 4.6 outlets; free summaries over a paid dataset averaged 6.0.

Finding 06

English costs you Greece

Programmes published only in English averaged 0.8 Greek outlets. Greek or bilingual programmes averaged 4.4 to 4.5: more than five times as many.

This is a paper about corporate research, published by somebody who produces corporate research, and a reader is entitled to hold that against it. My defence is in section 3, where I print the coding failure that nearly put a wrong number on the cover.

03

About this research

I censused organisations that are Greek or substantively Greek-operating, that publish original research, their own data, their own survey, or a study they commissioned, and that make it free in whole or as a substantive teaser. Translated global editions, news roundups and paid-only research were excluded.

The corpus was built by systematic search across fifteen sectors in Greek and English, by walking the member lists of six business associations, by sweeping Greek business-press archives for “έρευνα της [εταιρείας]” patterns, and by snowballing from co-publishers and named institutes. Every query is logged with its date and its yield, including the eleven that returned nothing.

State bodies are excluded throughout. ELSTAT, the Bank of Greece, EETT and the regulators publish more free research than any company here, but statutory publication is not the behaviour being studied. Business associations funded by their members are included, because that is the same behaviour with the cost pooled.

The unit is the publication programme, a publisher plus a recurring title, not the organisation and not the individual edition. One consultancy runs three programmes with different formats and audiences; collapsing them to one row would lose the variation this paper is about.

Motive is inferred from four observable signals, never asserted: who authored the artefact, whether it is gated, what its headline metric measures, and who it addresses. The decision rule is published so any row can be re-derived or disputed.

Figures are descriptive. No claim of statistical significance is made anywhere: the corpus is what systematic search surfaces, not a random sample of Greek companies. Every count is a floor and is written as “at least”.

04

What they publish

Six formats appear in the corpus.

Commissioned survey
26
Own operational data
18
Desk / analytical study
11
Teaser over paid data
6
Footprint study
4
Deep whitepaper
3

Only about one programme in four is built on data the publisher already owns. Fifty-four per cent are surveys of respondents or analysis of somebody else's figures.

To test whether that is a Greek trait or simply what corporate research looks like everywhere, I coded the same variables across a comparison corpus of 82 European and American publishers, matching the schemas by setting the Greek desk-analysis rows aside.

Format mix, Greece against an international comparison corpus
FormatGreece (n=57)International (n=82)
Own operational data31.6%42.7%
Commissioned survey45.6%43.9%
Teaser over a paid dataset10.5%4.9%
Institute-authored footprint study7.0%1.2%
Deep practical whitepaper5.3%7.3%
Email-gated (all rows)2.9%13.4%

The result is more specific than the intuition I started with. Commissioned surveys are equally dominant in both corpora, so buying a survey is not a Greek habit; it is the global default. What is distinctively Greek is the eleven-point deficit in own-data research, the four-and-a-half-fold gap in gating, and an unusually heavy reliance on institute-authored footprint studies: almost six times more common here.

Greek companies are sitting on datasets and commissioning surveys about the same subjects.

The comparison corpus is a convenience sample assembled by a looser protocol than this census, coded by the same person on an earlier version of the schema. It supports a direction, not a precise gap.

Fewer than half show their workings

“Fully disclosed” is not a demanding bar. It requires stating what the data is, what period it covers, and how it was collected. Thirty of 68 clear it. Thirty-two disclose partially. Six disclose nothing at all.

A finding about this paper

While coding this corpus I ran a reliability check, re-coding a fifth of it blind. It failed, at 0.64 against a 0.80 threshold.

The reason was a missing code. Bank and institute desk analyses, original work on third-party data with no new collection, had nowhere to go in the schema and had been defaulting into the commissioned-survey category. Eleven programmes sat in the wrong bucket, and the resulting headline, “36 of 70 programmes are commissioned surveys”, was wrong by ten.

I added a sixth format, recoded the corpus, drew a fresh sample and re-checked, reaching 0.93. The corrected figure is 26 of 68.

I am printing this because a paper whose subject is the transparency of other people's research has no standing to conceal a defect in its own, and because it carries an obvious implication for the six programmes that disclose no method at all. Nobody catches what nobody can see.

05

Why they do it, and the CSR question

Primary motive, inferred from four observable signals
MotiveProgrammesShare
Category ownership2333.8%
Policy1319.1%
Demand generation1319.1%
Teaser to paid68.8%
PR amplification57.4%
CSR marketing45.9%
Licence to operate34.4%
Employer brand11.5%

The single most common reason a Greek company publishes research is to own the recurring number in its category. A third of the corpus. Not leads, not reputation in the abstract, not social contribution: becoming the organisation that gets called when a journalist needs the figure.

So: is it CSR?

No. Seven programmes of 68, 10.3%, are framed as contribution, footprint or responsibility. Licence to operate is the primary motive for three programmes, 4.4% of the corpus.

And those three have an identical shape. Each is an institute-authored footprint study funded by a manufacturer or an industry association. The company pays, an independent institute writes and signs, the output is measured in jobs and multipliers and share of GDP, nothing is gated, and the audience is the state.

Read plainly, these are arguments addressed to government: this sector is too economically load-bearing to tax, restrict or ignore. That is legitimacy-seeking, and it is a perfectly respectable thing for an industry to do. It is not philanthropy, and treating it as philanthropy misreads both the document and the deadline it was written against.

Research motivated by licence to operate defends the publisher's right to operate. Research motivated commercially defends its right to be considered.

Sixty-one of the 68 programmes are the second kind.

There is a second, quieter pattern. Four programmes are coded CSR-marketing: research genuinely about a social or environmental question that also happens to promote the publisher's own programme in that area. A beverage company surveying whether consumers will pay more for sustainable hospitality, having a sustainable-hospitality scheme to sell. A telecoms operator measuring the digital maturity of the small businesses it sells digital services to. These are not cynical, and they are not disinterested either. The honest description is that the question is real and the choice of question is commercial.

06

Does it work in Greece?

I measured one thing: how far a publication travels in the Greek press. For 25 programmes stratified across the formats, I ran a single dated search and counted the distinct Greek outlets carrying the latest edition.4

This is a floor, and a rough one. It counts online press only, misses print and broadcast entirely, and in seven of 25 cases it hit the search-result ceiling, so the strongest performers are undercounted. It separates low pickup from high pickup, not high from higher.

Overall: median 4 outlets, mean 3.7.

Teaser over paid data
6.0
Footprint study
5.3
Own operational data
4.6
Commissioned survey
2.7
Deep whitepaper
0.0

Deep whitepapers drew no Greek press coverage at all. Both cases in the sample, zero. That includes a well-made 47-page document built on more than a million records, whose publisher's statistics release in the same year drew four outlets. Same company, same data asset, same year: the number travelled and the document did not.

The Greek press covers findings, not artefacts. A whitepaper is bought with attention from the reader who already has the problem; it is not a press play, and expecting it to be one is how a good document gets mistaken for a failure.

The most-used format is among the weakest performers. Commissioned surveys are 38% of the corpus and average 2.7 outlets. The ranges overlap and the samples are small, so this is a tendency rather than a rule, but the tendency runs against the corpus's own investment pattern.

English costs you Greece

Media pickup by publication language
LanguageProgrammesMean Greek outlets
Greek and English104.5
Greek only104.4
English only50.8

Roughly a five-fold difference. Publishing bilingually costs a translation and appears to lose nothing; publishing only in English forfeits the Greek press almost entirely. For a company whose customers are Greek, this is the single most actionable line in the paper.

Giving it away outperforms gating it

Programmes coded demand-generation averaged zero Greek outlets across three cases. Programmes coded teaser-to-paid averaged 6.0 across three.

Both models withhold something. The difference is what. The teaser withholds the dataset and publishes the findings, so there is something to report. The gate withholds the findings, so there is not. In a market this size, where the same fifteen or twenty business outlets carry everything, that appears to be the difference between existing and not.

Repetition compounds, modestly

Pickup against programme age
Editions to dateProgrammesMean outlets
1-693.0
7-2063.8
21+104.3

Monotonic and in the expected direction. The gradient is gentle and edition counts are mostly estimated, so this is a pattern rather than a coefficient. But it matches what the long-running publishers look like from outside: they are not covered because any single edition is remarkable, they are covered because they are the series of record.

And the machines that now answer the question

Press pickup measures whether a journalist repeats your number. That is a proxy for something which now matters more and can be measured directly: whether a machine repeats it. The twelve prompts for this test were fixed and published on 5 August 2026, before any engine was touched, so the protocol could not be tuned to a flattering result.

On 6 August 2026 I put those twelve questions, in Greek, as a Greek user would type them, to three answer engines with web access, one session each, first response only, no follow-ups. Thirty-six answers naming 303 sources.

5,6
M2: answer-engine citation test, 6 August 2026. One run, twelve fixed prompts per engine.
EngineSources per answerNamed a census publisherNamed any Greek source
Perplexity sonar13.88 of 1212 of 12
ChatGPT gpt-5-search-api4.56 of 1211 of 12
Gemini 2.5-flash, Google Search6.95 of 1211 of 12
All three8.419 of 36: 52.8%34 of 36: 94.4%

Across all 36 answers, 52.8% named at least one publisher from the census, and 44.4% named the specific publisher that owns that sector's recurring number. Almost every answer, 94.4%, named a Greek source, so this is not engines falling back on international material for want of anything local.

The share of citations is the sobering figure. Of the 303 sources named, only 24, 7.9%, were census publishers.

Greek websites, other
35.6%
Greek news media
31.7%
International, other
21.5%
Census publishers
7.9%
Official statistics
3.3%

Source mix across all 303 citations, classified by domain against the census. Full log in ai-citations.csv.

The engines mostly cite the press reporting the research rather than the research itself. A corporate publisher's number reaches the answer through a news site that quoted it, which means the publisher gets the number into circulation and the news site gets the citation.

Where a company owns the recurring number, every engine named it. Where nobody owns it, no engine named anyone.

The sector pattern is the actual finding. The prompts were chosen in advance to cover sectors where the census found an obvious owner of the number, plus one control sector where it found none.

Whether the sector's census owner was named, by prompt. Three engines, so three chances per row.
QuestionSectorOwner named
How many vehicles were registered in Greece this year?Automotive3 of 3
What does brewing contribute to the Greek economy?Food and drink3 of 3
How big is the Greek pharmaceutical market?Pharma3 of 3
What is average occupancy in Greek short-term rentals?Short-term rental2 of 3
How much did Greek house prices rise last quarter?Real estate1 of 3
How is Greek shipping financed by the banks?Shipping1 of 3
What does a software engineer earn at a Greek startup?Startups1 of 3
Who are the most attractive employers in Greece?HR1 of 3
How big is the Greek e-commerce market?E-commerce1 of 3
What share of Greeks buy online, and how often?E-commerce0 of 3
How many ransomware attacks hit Greek businesses?Cybersecurity0 of 3
What does industrial electricity cost in Greece versus Europe?Energy: control, no owner0 of 3

Three sectors returned their owner in every answer from every engine. The control returned no corporate publisher at all, from any engine, which is what the design predicted and is the cleanest evidence in this paper that ownership of a number is a real and machine-visible property, not a story publishers tell about themselves.

E-commerce is again the largest miss. It is the sector with the most Greek publishers and the loudest trade press, and across six answers about market size and shopper behaviour its nominal owner surfaced once.

One organisation dominates the citations. IOBE, the Foundation for Economic and Industrial Research, was named in 11 of 36 answers, through three separate commissioned programmes: brewing, food and drink, and pharmaceuticals. No other publisher appears more than three times. The most citable object a Greek company can pay for is an independent institute's name on a study of its own sector's economic contribution. That is the same format that led press pickup at 5.3 outlets, and it is the format this corpus produces least.

What this test is and is not. One date, one run, three engines, twelve prompts. Answer engines are non-deterministic and personalised; a re-run would give different specifics. This is a snapshot, not a benchmark, and the claim it supports is narrow: that corporate research surfaces in generated answers at all, and surfaces far more reliably where a publisher owns the recurring number.

The locked protocol specified four engines. Microsoft Copilot was not run, no API access, and that is recorded rather than quietly redefining the design as three. One Gemini response returned an API error and is counted as an answer naming no sources, not dropped. Sources were matched to the census mechanically by domain, so a publisher named in prose but not linked is not counted; the 52.8% is therefore a floor. Full log, prompts and every raw response are published.

A separate observation, outside this protocol. While using the same engines as a finding aid for an unrelated source, one returned a verbatim Greek-language “methodology note” in quotation marks and attributed it to a specific PDF. I downloaded the PDF. The quoted sentence is not in it. That cost nothing here because nothing entered this paper without being opened at source, but it is the reason nothing does.

An earlier draft carried two external statistics about why published research works commercially. Both were cut at citation audit: one could not be opened at source, and the other, the widely-quoted claim that only 5% of buyers are in-market at any moment, turned out to originate not with the report I had attributed it to but with separate research quoted inside it. The claim may well be sound. My attribution was not, and a paper that spends a section on other people's method disclosure does not get to publish a citation it has not opened.

07

How Greece differs from everyone else

The Greek census has a companion: a lighter scan of 119 programmes across Greece, Europe and America, coded on a matched schema. It is not the same quality of instrument and is used here only for shape. On that basis three differences stand out, and one of them explains most of the others.

What this corpus is, and what it is not. The 119-row comparison set was assembled by systematic search across three regions with a deliberately lighter verification standard than the Greek census: 64 of the 119 rows are confirmed and 55 are unverified. It carries no reliability check and no second coder.

It is therefore used for one purpose only: comparing the distribution of formats and access models between regions, where a consistent error would have to be regionally biased to matter. No count from it appears as a headline in this paper, and no Greek claim rests on it. Where the Greek census and this corpus disagree, the census wins.

America publishes its own data. Greece commissions surveys.

Format distribution across the 119-programme comparison corpus. Percentages are of each region's own total: Greece 37, Europe 36, America 46.8
FormatGreeceEuropeAmerica
Platform benchmark: the publisher's own operational data19%19%48%
Commissioned survey43%47%30%
Research programme8%3%9%
Benchmark, other8%11%4%
Deep white paper5%17%0%
Free teaser over paid data8%3%7%
Footprint study and advocacy8%0%2%

Nearly half of the American programmes are built on the publisher's own platform data, against a fifth in Greece and a fifth in Europe. The mirror image is the commissioned survey: 43% of Greek programmes and 47% of European ones, against 30% American.

This is the same finding the Greek census produced from a better instrument, 18 of 68 programmes on own operational data, arriving from a different direction. Greek publishers buy their evidence. American publishers extract theirs from the systems they already run.

The economics of that difference are not subtle. A commissioned survey is a recurring external cost with a fieldwork lead time, and it produces a number anybody else can also commission. Own operational data has no fieldwork cost, no lead time, and cannot be replicated by a competitor at any price. The Greek corpus is concentrated in the expensive, replicable format and thin in the cheap, defensible one.

The format Greek publishers use most is the one they cannot own. The format they use least is the one nobody can take from them.

The deep white paper is a European habit

The document type that gives this genre its name barely exists outside Europe in this sample: 17% of European programmes, 5% of Greek, and none at all of the 46 American ones.

Set that beside the Greek media result reported earlier, deep practical whitepapers drew zero Greek press coverage across both cases in the amplification sample, and a consistent picture appears. The long teaching document is a real format with a real audience, and that audience is a reader with a problem rather than a journalist with a deadline. American corporate publishing, which is the most commercially optimised of the three, has largely stopped producing it.

That is not an argument against writing one. It is an argument against expecting it to travel. A whitepaper is bought with the attention of somebody already looking for the answer; it is not a distribution strategy, and the American corpus suggests the market worked that out some time ago.

Gating is an American practice

America: email-gated
17%
Europe: email-gated
8%
Greece: email-gated
5%

Share of each region's programmes placing the full document behind an email form. Comparison corpus, n=119.

American publishers gate more than three times as often as Greek ones. The Greek census puts the domestic figure lower still, at two programmes in 68, and both of those belong to software companies selling outside Greece.

Two readings are available and the data cannot separate them. Either Greek publishers have understood that gating suppresses the reach that makes the exercise worthwhile, or they have not built the lead-capture machinery that makes gating worth doing. The Greek amplification result, ungated programmes reaching more Greek outlets than gated ones, is consistent with the first, and the near-total absence of marketing-technology publishers from the Greek corpus is consistent with the second.

The sectors that do not exist in Greece

The comparison corpus carries 31 sector labels that appear in Europe or America and not once in Greece. They are not random. They cluster into three groups, and the clustering is the point.

  1. Infrastructure and network operators: internet infrastructure, telecom security, network traffic. Companies whose product generates a global dataset as a by-product, and who publish it because it is free to produce and impossible to contest.
  2. Payments and insurance: payments/retail, payments/e-commerce, insurance, reinsurance. Sectors that sit on transaction and loss data and, abroad, publish from it routinely.
  3. Developer and marketing technology: developer tools, marketing tech, B2B marketing, media/advertising. The sectors that invented gated research, absent from the Greek corpus almost entirely.

The Greek census found publishers in Greek payments and Greek banking, so the second group is not empty in reality: it is thin. The first and third are close to genuinely absent, and both are sectors where the publishable asset is a by-product of operations rather than a purchase.

Which returns the argument to where the census left it. The gap between Greek and international corporate research is not mainly a gap in effort, budget or sophistication. It is a gap in whether the publisher looks at the data they already hold and recognises it as a publication. Every sector in the first group abroad is a sector whose Greek equivalent runs the same systems and generates the same logs.

08

The white space

If a sector has no owner of its recurring number, the position is open.

Single-publisher sectors in this corpus: payments and energy. One programme each. Any credible second entrant with real data would be competing against a near-vacuum.

Sectors searched that returned nothing at all: industrial manufacturing footprint studies, courier and last-mile logistics, consumer electronics retail, gambling and gaming, and telecoms beyond a single university-partnered study. In courier and last-mile the absence is conspicuous, because the parcel operators sit on exactly the kind of operational data that produces a category-owning index, and the Greek e-commerce press covers the topic constantly using figures scraped from filed accounts and regulator reports.

If you are deciding whether to publish

  1. Check whether your category is taken. If somebody has run the reference number for a decade, entering means a different angle or a better dataset: not the same survey with your logo on it.
  2. Prefer the data you own. It is cheaper, faster, cannot be copied, and in this corpus it travels roughly seventy per cent further than a commissioned survey.
  3. Publish in Greek, or bilingually, unless your customers genuinely are not here.
  4. Separate the news from the document. Publish the finding openly and let the depth sit behind it. Treating a whitepaper as a press release produces a document nobody reports.
  5. Budget for the third edition. One edition is a blog post. The organisations that own their categories here have been at it for twelve to twenty-four years.
One edition is a blog post. Three is an index.
09

Limitations

LimitationWhat I did about it
The corpus is what systematic search surfaces. Publishers with no web presence are invisible to itSearch protocol and all 36 queries published, including the 11 that found nothing2. Every count written as “at least”
Motive is inferred from artefacts, not stated intentDecision rule published; CSR framing coded separately from motive so the CSR answer can be re-derived without trusting the motive column
Media pickup counts online Greek press only, one dated search per programme, censored in 7 of 25 casesReported as a floor throughout, and used only to compare formats, never to score individual publishers
Several sub-findings rest on two or three cases. The zero-pickup result for whitepapers is n=2Sample sizes printed in the sentence, not hidden in a note
Edition counts are mostly estimated from cadence rather than counted from archivesFlagged per row; no headline claim rests on an estimated count
The motive column failed inter-rater reliability at κ 0.34. Two coders working from the published rule agreed on why a programme was published just over half the timeMeasured on all 68 rows and printed in full in Appendix B, including the two codes the rule had failed to define. Motive figures are labelled indicative. The corpus was deliberately not recoded afterwards. The CSR answer rests on csr_framed, agreement 89.7%, which was coded independently for exactly this reason
The independent second coder was a language model, not a trained human researcherStated. It tests whether the published rule reproduces a decision, which is what a published rule is for; it is not a substitute for a human second coder and the result is reported as the weaker instrument it is
The answer-engine test is one run on one date against three of the four engines specifiedPrompts locked and published before the run; Copilot recorded as not run rather than the design being redefined; all 36 raw responses published
I publish research and would sell research servicesMethod, corpus and coding rules published in full; my own coding failure printed in section 3

The materially weakening limitation is the first one. This is a floor, not a total. Treat every count as the minimum true value, and the sectors reported as empty as sectors where I found nothing: not sectors where nothing exists.

10

Sector by sector: who owns the number

The census answers a question no Greek company can currently answer about its own market: in my sector, has somebody already become the source everyone quotes?

The table below lists every sector in the corpus, how many publication programmes it contains, and who, if anyone, publishes research built on data they already own. That last column is the one that matters, because own-data programmes are the ones competitors cannot replicate.

All 15 sectors, by programme count
SectorProgrammesPublishersWho publishes own-data research
Tourism and hospitality1210Athens International Airport, Hosthub, INSETE
E-commerce and retail99Skroutz, Wolt Greece
Banking, finance and insurance66EAEE
Professional services55nobody
Real estate55Cushman and Wakefield Proprius, Danos / BNP Paribas Real Estate, Spitogatos
HR and labour55nobody
Cross-sector bodies55nobody
Food and beverage44nobody
Automotive and mobility44Ayvens Greece, Car.gr, SEAA
Startups and venture capital43Marathon Venture Capital
Shipping33Moore Greece, Petrofin Research
Pharma and health22nobody
Technology and cyber22Obrela Security Industries
Payments and fintech11Nexi Greece
Energy11nobody

Six sectors have nobody

22 of the 68 programmes sit in sectors where not one publisher builds on data they already hold. Professional services, HR and labour, the cross-sector business bodies, food and beverage, pharma and health, and energy: every one of those programmes is a survey, a desk analysis, a footprint study or a paid-data teaser.

That is not a criticism of those programmes. A commissioned survey is a legitimate instrument and some of the corpus's best work is survey-based. It is an observation about competitive position: in six sectors, the recurring number is up for the taking by whoever holds a dataset and decides to publish it.

Shipping is the smallest sector and the best run

Shipping contains three programmes. It also contains the two longest evidenced runs in the entire corpus, a ship-finance index running since 2001, and a ferry-market report in its 24th edition, and two of its three programmes are built on data the publisher compiles itself.

Both are small organisations. Neither is a statistical agency, a bank or a consultancy of any scale. They simply started early, published every year without exception, and became the reference. If the census contains a model worth copying, it is this one, and it did not require size.

E-commerce is the largest miss

E-commerce and retail is the second-densest sector at nine programmes. Only two are built on the publisher's own data.

This is the sector where every participant, marketplace, courier, payment processor, platform, merchant, holds transaction-level data by default. It generates more usable proprietary data per operator than almost any other sector in the corpus, and it is publishing surveys about itself instead.

Where the data is the product, the pattern reverses

Real estate and automotive have the corpus's highest own-data ratios: three of five and three of four. Both are sectors where the publisher's business is a dataset: listings, registrations, transactions, fleet costs. When the data is already the product, publishing an index is a small step rather than a new capability.

The lesson generalises past those two sectors. The organisations publishing own-data research in this corpus are rarely the ones with the biggest research budgets. They are the ones for whom the dataset already existed and someone decided to point it outward.

11

The six formats as a playbook

Each format costs something different to produce and returns something different. The pickup column is the median distinct Greek outlets covering a release, from the sample of 25 programmes measured in section 5.

What each format costs and what it returns
FormatIn corpusMean pickupCostsBest for
Teaser over paid data66.0Nothing extra: the dataset is already soldData businesses. Highest pickup in the corpus
Footprint study45.3Highest. An independent institute, commissionedLicence to operate. Aimed at government, not buyers
Own operational data184.6Analyst time. No fieldworkCategory ownership. Cannot be replicated
Commissioned survey262.7Panel fieldwork, recurringMarkets where you hold no data of your own
Desk analysis11n/aAnalyst time onlyInstitutions with standing economics teams
Deep whitepaper30.0High. Long-form writing and designThe reader who already has the problem. Not a press play

Read the first and last rows together. The format with the highest media pickup costs almost nothing extra to produce, because the dataset is already being sold. The format with the lowest, zero, in both sampled cases, is the most expensive to write.

The cheapest format travelled furthest. The most expensive did not travel at all.

That is not an argument against deep whitepapers. It is an argument against expecting one to do a press release's job. The two formats answer different questions: one earns coverage, the other earns a meeting.

The corpus's own investment pattern runs against this. Commissioned surveys are the most common format at 26 of 68, and they sit fourth of five on measured pickup.

A

Appendix A: the census

All 68 included publication programmes. Two further programmes were coded and excluded under the published inclusion rules and are retained in the dataset with their exclusion reason.

The full corpus
PublisherSectorProgrammeFormatCadenceAccess
HosthubTourism and hospitalityThe Top 100 Issues Hosts Have to Handledeep whitepaperone-offemail-gated
HosthubTourism and hospitalityAnnual Greek STR statisticsown dataannualungated
HosthubTourism and hospitalitySTR share of Greek GDP researchown dataone-offungated
Grant Thornton GreeceTourism and hospitalityHotels in Greece / hospitality studiescommissioned surveyannualungated
Deloitte GreeceTourism and hospitalityThe Greek Hospitality Reimagineddeep whitepaperone-offungated
Deloitte Greece x Hellenic Hoteliers FederationTourism and hospitalityThe Future of the Greek Hotelcommissioned surveyone-offungated
GBR ConsultingTourism and hospitalityHospitality Newsletterteaser over paidquarterlyungated-teaser
ITEP x Hellenic Chamber of HotelsTourism and hospitalityAnnual Survey of the Hotel Sectorcommissioned surveyannualungated
INSETETourism and hospitalityGreek tourism annual report and regional studiesown dataannualungated
INSETE x Deloitte-RemacoTourism and hospitalityGreek Tourism 2030 Action Plansdesk analysisone-offungated
SETE / Marketing GreeceTourism and hospitalityDestination and inbound-visitor researchdesk analysisperiodicungated
XRTC Business ConsultantsShippingAnnual Report on the Greek Ferry Marketcommissioned surveyannualungated
Petrofin ResearchShippingPetrofin Bank Research and Index of Greek Ship Financeown dataannualungated
Moore GreeceShippingMoore Maritime Indexown dataannualungated
IOBE x Athenian BreweryFood and beverageThe Brewing Sector in Greecefootprint studyannualungated
IOBE x Coca-Cola HBCFood and beverageSocio-economic footprint incl. HORECAfootprint studyperiodicungated
Coca-Cola Greece x AUEBFood and beverageSustainability in HORECA - what consumers wantcommissioned surveyone-offungated
IOBE x SEVTFood and beverageThe food and beverage industry in Greecefootprint studyannualungated
IELKAE-commerce and retailRetail and consumer research programmecommissioned surveycontinuousungated
NielsenIQ GreeceE-commerce and retailFMCG market data releasesteaser over paidperiodicungated-teaser
Circana GreeceE-commerce and retailConsumer trends and FMCG releasesteaser over paidperiodicungated-teaser
IOBE x SFEEPharma and healthThe pharmaceutical market in Greece - Facts and Figuresfootprint studyannualungated
IQVIA GreecePharma and healthePharmacy audit insightsteaser over paidperiodicungated-teaser
Convert GroupE-commerce and retailGreek ePharmacy / eGrocery reportsteaser over paidannual+quarterlyteaser-plus-paid
GR.EC.A x ELTRUN AUEBE-commerce and retailAnnual Greek E-Commerce Surveycommissioned surveyannualungated
GR.EC.A x TruberriesE-commerce and retailE-shop and courier B2B market researchcommissioned surveyone-offungated
GR.EC.A x SySPaL University of the AegeanE-commerce and retailGreek logistics sector studycommissioned surveyone-offungated
SkroutzE-commerce and retailAnnual purchasing behaviour and merchant reviewown dataannualungated
Nexi GreecePayments and fintechEcommerce Report Greeceown dataannualungated
SEAAAutomotive and mobilityVehicle registration statisticsown datamonthlyungated
Car.grAutomotive and mobilityAnnual search and market infographicsown dataannualungated
Ayvens GreeceAutomotive and mobilityCar Cost Index - Greek cutown dataannualungated
INTERAMERICAN and Anytime x NTUAAutomotive and mobilityRoad safety and driving behaviour researchcommissioned surveyperiodicungated
EurobankBanking, finance and insuranceEconomic Analysis and Researchdesk analysisweekly+studiesungated
Alpha BankBanking, finance and insuranceEconomic Research bulletinsdesk analysisweeklyungated
National Bank of GreeceBanking, finance and insuranceSectoral studies and economic analysesdesk analysisperiodicungated
Piraeus BankBanking, finance and insuranceEpi Gis agricultural economy periodicaldesk analysisperiodicungated
Piraeus Financial HoldingsBanking, finance and insuranceSectoral studies / economic analysisdesk analysisperiodicungated
PwC GreeceProfessional servicesGreek CEO Survey and publicationscommissioned surveyannualungated
EY GreeceProfessional servicesEY Attractiveness Survey Greececommissioned surveyannualungated
Deloitte GreeceProfessional servicesGreek CFO Surveycommissioned surveyannualungated
KPMG GreeceProfessional servicesSector surveys - real estate, transport and logisticscommissioned surveyperiodicungated
ICAP CRIFProfessional servicesGreek sector studiesteaser over paidcontinuousteaser-plus-paid
EAEEBanking, finance and insuranceAnnual statistical report of the insurance marketown dataannualungated
SpitogatosReal estateSpitogatos Property Indexown dataquarterlyungated
Danos / BNP Paribas Real EstateReal estateGreek market reportsown dataquarterlyungated
Cushman and Wakefield PropriusReal estateGreece Marketbeat snapshotsown dataquarterlyungated
Algean PropertyReal estateResearch and yields reportsdesk analysisannualungated
American-Hellenic Chamber of CommerceReal estateProperty Market Outlook for Greecedesk analysisbiannualungated
WorkableHR and labourFree HR reports, ebooks and Hiring Pulsedeep whitepapermonthly+annualemail-gated
Epignosis / TalentLMSHR and labourAnnual L&D Benchmark Reportcommissioned surveyannualungated
Randstad GreeceHR and labourEmployer Brand Research Greececommissioned surveyannualungated
ManpowerGroup GreeceHR and labourEmployment Outlook Surveycommissioned surveyquarterlyungated
kariera.gr x AUEBHR and labourPeople Pay and AI researchcommissioned surveyannualungated
Obrela Security IndustriesTechnology and cyberDigital Universe Reportown dataannualungated
Endeavor GreeceStartups and venture capitalGreek tech ecosystem reports and Year in Reviewdesk analysisannualungated
Marathon Venture CapitalStartups and venture capitalGreek Startup Compensation Reportcommissioned surveyannualungated
Marathon Venture CapitalStartups and venture capitalGreek startup funding rounds and exitsown dataannualungated
Found.ation x EIT DigitalStartups and venture capitalStartups in Greece annual reportcommissioned surveyannualungated
IENEEnergyThe Greek Energy Sector annual reportdesk analysisannualungated
SEVCross-sector bodiesO Sfygmos tou Epicheirein annual business surveycommissioned surveyannualungated
IME GSEVEECross-sector bodiesBiannual SME economic climate surveycommissioned surveybiannualungated
diaNEOsisCross-sector bodiesResearch programmecommissioned surveycontinuousungated
Focus BariCross-sector bodiesPublic opinion and media research releasescommissioned surveyperiodicungated
Metron AnalysisCross-sector bodiesMetron Forum pollscommissioned surveymonthlyungated
Athens International AirportTourism and hospitalityPassenger traffic statisticsown datamonthlyungated
Wolt GreeceE-commerce and retailConsumer Reportown dataannualungated
COSMOTE x ELTRUN AUEBTechnology and cyberDigital Readiness of Greek SMEscommissioned surveybiennialungated
C

Appendix C: every search, including the empty ones

All 36 queries run to build the corpus, with what each returned. 11 returned nothing. They are printed because a census that shows only its successful searches is not reproducible, and because the empty rows are where the white space in section 6 comes from.

Search log: highlighted rows found nothing
LangQueryFound
ENPetrofin Research Greek shipping annual report free bank research index1
ENMarathon VC Found.ation Greek startup ecosystem annual report free download3
ELΣΦΕΕ ΙΟΒΕ μελέτη φαρμακευτική αγορά Ελλάδα δωρεάν έκθεση1
ELΤράπεζα Πειραιώς κλαδικές μελέτες αγροτικός τομέας δωρεάν pdf2
ELΙΕΝΕ ΔΕΠΑ ΔΕΗ ετήσια έκθεση ενέργεια Ελλάδα μελέτη δωρεάν1
ELΕΑΕΕ ασφαλιστική αγορά Ελλάδα ετήσια έκθεση στατιστικά δωρεάν1
ELAlgean Property Prosperty Colliers Greece market report ακίνητα έρευνα δωρεάν1
ELKariera.gr Skywalker έρευνα αγορά εργασίας Ελλάδα report δωρεάν μισθοί2
ELCOSMOTE Nova Vodafone Ελλάδα έρευνα ψηφιακές συνήθειες report δωρεάν μελέτη0
ELLinkwise Generation Y ελληνικό digital marketing report έρευνα δωρεάν affiliate0
ELΕΛΤΑ Courier Speedex Geniki Tachydromiki έρευνα ecommerce delivery Ελλάδα report0
ELNielsen Circana Ελλάδα λιανική έρευνα καταναλωτές δωρεάν report FMCG2
ENTalentLMS Epignosis research report free download training survey1
ENMoore Greece XRTC shipping survey Greek ship finance annual report free2
ELΣΕΒ έρευνα μελέτη επιχειρήσεις Ελλάδα δωρεάν special report1
ELΕθνική Τράπεζα κλαδική μελέτη οικονομική ανάλυση δωρεάν pdf sectoral report1
ELΓΣΕΒΕΕ εξαμηνιαία έρευνα μικρομεσαίες επιχειρήσεις Ελλάδα δωρεάν αποτελέσματα1
ELdiaNEOsis έρευνες μελέτες Ελλάδα δωρεάν χρηματοδότηση επιχειρήσεις1
ELΤΙΤΑΝ Μυτιληναίος Metlen ΙΟΒΕ μελέτη κοινωνικοοικονομικό αποτύπωμα δωρεάν0
ELPublic Πλαίσιο Κωτσόβολος έρευνα καταναλωτών τεχνολογία Ελλάδα report δωρεάν0
ELFocus Bari MRB Metron Analysis Kapa Research δωρεάν έρευνα αποτελέσματα Ελλάδα2
ELIQVIA Ελλάδα φαρμακείο αγορά έρευνα report δωρεάν στοιχεία1
ENColliers Greece Cushman Wakefield Proprius Athens market report free download real estate1
ELΣΕΒΤ τρόφιμα ποτά ελληνική βιομηχανία ετήσια έκθεση κλάδου δωρεάν1
ELΟΠΑΠ ΙΟΒΕ μελέτη τυχερά παιχνίδια Ελλάδα αποτύπωμα δωρεάν έκθεση0
ELViva.com έρευνα πληρωμές μικρές επιχειρήσεις Ελλάδα report δωρεάν0
ELEnterprise Greece Athens Exchange ΕΒΕΑ έρευνα επενδύσεις Ελλάδα έκθεση δωρεάν0
ELOdyssey Cybersecurity Encode Ελλάδα threat report δωρεάν λήψη0
ELMarketing Greece ΣΕΤΕ έρευνα τουριστικό brand Ελλάδα μελέτη δωρεάν1
ELDeloitte Greece CFO Survey EY Greece CEO Outlook ελληνικές επιχειρήσεις έρευνα δωρεάν2
ELΕλληνο-Αμερικανικό Εμπορικό Επιμελητήριο έρευνα μελέτη οικονομία δωρεάν έκδοση1
ELΔιεθνής Αερολιμένας Αθηνών ΟΛΠ ετήσια έκθεση κίνησης στατιστικά δωρεάν δεδομένα1
ELHELLENiQ ENERGY Motor Oil ΔΕΗ ΙΟΒΕ μελέτη αποτύπωμα ενέργεια Ελλάδα έκθεση0
ELWorldline Cardlink Ελλάδα δεδομένα συναλλαγών έρευνα report δωρεάν καταναλωτική δαπάνη0
ELefood Wolt Ελλάδα ετήσια στοιχεία παραγγελίες report δεδομένα δημοσίευση1
ELQuest Group Info Quest COSMOTE έρευνα ψηφιακή ωριμότητα ελληνικές επιχειρήσεις report1
B

Appendix B: how everything was coded

Every row in Appendix A was coded on fourteen variables. The rules below were written when each ambiguity was first hit, not reconstructed afterwards, so any row can be re-derived or disputed.

The unit

The unit is the publication programme, a publisher plus a recurring title, not the organisation and not the individual edition. Cadence, format, gating and method disclosure are properties of a programme, not of a company. One consultancy in this corpus runs three programmes with different formats and audiences; collapsing them to a single row would lose exactly the variation the paper is about. Both counts are reported: 68 programmes across 64 publishers.

Inclusion

All three had to hold: the publisher is Greek or substantively Greek-operating; the research is original: own data, own survey, or a commissioned study, not a translated global edition; and it is free in whole or as a substantive teaser, where a teaser must carry a finding rather than a table of contents.

Exclusions, and why

CaseDecisionReason
State bodies and regulatorsExcludedStatutory publication is not the behaviour being studied
Universities publishing aloneExcludedIncluded only as a co-signer on company- or association-funded work
Business associationsIncludedCompany-funded bodies publishing on members' behalf: the same behaviour with the cost pooled
Free software tools with dashboardsExcludedA tool is not a publication
Figures reaching the press only via filed accountsExcludedJournalism about a company is not publication by it

The six formats

A own operational data · B deep practical whitepaper · C commissioned survey · D institute-authored footprint study · E free teaser over a paid dataset · F desk or analytical study.

Where a programme could be two, the data basis decides: own operational data is A, new primary collection is C, neither is F. If the full dataset is sold, E overrides everything. Code D requires an independent author: a self-authored footprint claim is A or F with a policy motive, not D.

The motive decision rule

Motive is inferred from four observable signals, applied in order, and never asserted:

SignalReading
AuthorIndependent institute or university on the cover → licence to operate or policy. Self-authored → commercial
GateEmail-gated → demand generation, always
Headline metricJobs, GDP share, multipliers → licence to operate. Behaviour, prices, indices → category ownership. Practical how-to → demand generation
AddresseeGovernment, media, society → policy. Buyers → demand generation or category ownership

Ties resolve by addressee. Where the publisher is an association and the addressee is the state, the code is policy even where the work also serves members commercially. CSR framing is coded separately from motive, so a reader can re-derive the CSR answer without trusting the motive column.

The full code list. Eight codes are in use. An earlier version of this appendix defined only six, which the round-three reliability check caught:

CodeDefinition
category-ownershipOwning the recurring number in a category
demand-genGenerating or qualifying commercial demand
policyAddressed to the state or to public debate
teaser-to-paidFree findings existing to sell the dataset
pr-amplificationA press release rather than a document; no deeper purpose visible
licence-to-operateDefending the publisher’s right to operate
csr-marketingContribution or responsibility framing used as brand marketing rather than to defend a licence. Four programmes
employer-brandAddressed to candidates rather than buyers or the state. One programme

csr_framed takes three values, not two: yes, no, and partial where the artefact carries contribution framing in part of the document only. Two programmes are partial.

Reliability, and the round that failed

Coding reliability was checked three times, twice by re-coding a systematic 20% sample blind against the same source artefacts, and once by an independent second coder against the whole corpus.

Round one failed at 0.64 against a 0.80 threshold. It produced three rule corrections and one schema change: desk analyses had no code and were defaulting into commissioned surveys, putting eleven programmes in the wrong bucket. Round two, on a fresh sample drawn with a different start after recoding, reached 0.93.

Both of those rounds were intra-rater: one coder re-checking their own work at a distance. That catches inconsistent application of a rule. It does not catch a rule that is wrong in a way one person repeats consistently.

Round three: independent coding, and what it broke

On 6 August 2026 the whole corpus was coded a second time by an independent agent working only from the published rules and the same visible facts about each programme, blind to the original codes. Not a sample: all 68 programmes. The second coder was a large language model, which is a weaker instrument than a trained human second coder and is stated as such below. What it tests precisely is whether the published rule, applied by somebody who was not there when it was written, reproduces the decision.

7
Inter-rater agreement, n=68, all three judgement columns. Threshold for this project is κ 0.80.
ColumnRaw agreementCohen’s κVerdict
format85.3%0.80Meets the threshold
csr_framed89.7%0.34High agreement, unstable κ: see below
motive_primary51.5%0.34Fails

The format schema holds. The motive column does not. Two coders working from the same published rule agreed on the format of a programme 85% of the time, and on why it was published just over half the time. That is the single most important thing this paper learned about itself, and it lands on the column that answers its own headline question.

Three things caused it, and only two are fixable.

One: the published rule was incomplete. The corpus uses eight motive codes. The decision rule printed above listed six: csr-marketing and employer-brand appear in the results table but were never defined, so five programmes could not possibly have been matched. The same defect applied to csr_framed, which the rule described as yes/no while the data carries a third value, partial, in two rows. Both are now documented above. Excluding those five rows moves κ to 0.38: it was not the main cause.

Two: the gate signal was read as an equivalence. Eleven of the 34 motive disagreements are the same disagreement: programmes I coded demand-generation, the second coder coded category-ownership, and all eleven are ungated. The rule says email-gated implies demand-generation. It does not say the reverse, but it does not say what ungated implies either, so the decision falls through to signals three and four, and for a professional-services firm's free sector report, “practical how-to” and “benchmark” are genuinely both true. The rule under-determines the answer for a whole class of publisher.

Three: the column may not be reliably codeable from artefacts at all. Motive is inferred, and this result is the measurement of how much that inference costs.

What follows from it. The motive distribution stays in the paper, because deleting an inconvenient result after measuring it is worse than publishing it with its reliability attached. But it is now labelled: every motive figure in this paper is a single-rater judgement with κ 0.34, and should be read as indicative, not measured. The corpus was not recoded to chase a better number: recoding after seeing the disagreements is how a reliability check becomes theatre.

The CSR answer survives this intact, and that is by design rather than luck. csr_framed was coded independently of motive from the beginning precisely so the headline question would not depend on the softest column. Its raw agreement is 89.7%. Its κ is low for a well-understood reason: 59 of 68 rows carry the same value, so chance agreement is already 84.5% and κ becomes unstable: the arithmetic penalises a skewed column even when the coders agree. On a question with a lopsided true answer, raw agreement is the more honest statistic, and it is reported alongside.

Second coder: DeepSeek deepseek-chat, temperature 0, one pass, no retries on disagreement. Prompt carried the published rules verbatim plus nine observable fields per programme; the original codes were withheld. Per-row output in inter-rater-check.csv.

Link integrity

All 70 coded rows carry a source URL. At the most recent check, 66 resolved live, four returned 403 to an automated client and were reached by other means, and none were broken. Two URLs found broken in an earlier check were corrected rather than dropped.

12

References and instruments

This paper rests almost entirely on primary research conducted for it. That makes the instruments themselves the references, and all of them are published in full at the URLs below. The two external statistics an earlier draft relied on were both cut at citation audit, and they are listed here rather than deleted from the record.

Instruments and datasets, all published with this paper

  1. Broikos, N. (2026). Census of Greek corporate research publishing. Dataset. 70 publication programmes from 65 publishers, coded on 20 variables; 68 included and 2 excluded with their exclusion reason recorded. Coded 5-6 August 2026. Every row carries the source URL of the artefact coded. https://broikos.gr/research/data/paper-01/census.csv This file is the answer to any question about an individual publisher in this paper. The 68 organisations analysed are not listed separately in these references because each is a row in this dataset with its own URL.
  2. Broikos, N. (2026). Search log. Dataset. All 36 documented searches in Greek and English, including the 11 that returned nothing. https://broikos.gr/research/data/paper-01/search-log.csv Published so that the corpus boundary can be audited. Every count in this paper is a floor, and this file is the evidence for how the floor was reached.
  3. Broikos, N. (2026). Coding rules. Full rule set, with each ambiguity recorded at the point it was first encountered rather than reconstructed afterwards. Includes the eight motive codes and the three values of the CSR-framing variable. https://broikos.gr/research/data/paper-01/coding-rules.md
  4. Broikos, N. (2026). Media amplification measurement (M1). Dataset. 25 programmes stratified across formats; one dated search each; distinct Greek online outlets counted. Measured 5-6 August 2026. https://broikos.gr/research/data/paper-01/amplification.csv Reported throughout as a floor. Censored at the search-result ceiling in 7 of 25 cases; online press only.
  5. Broikos, N. (2026). Answer-engine citation test (M2): locked prompt set. Twelve prompts in Greek, committed and dated 5 August 2026, before any engine was queried. https://broikos.gr/research/data/paper-01/prompts-m2.md Pre-registration of the protocol. The file also carries the run record and the deviation from design: three engines were run, not the four specified, because no Microsoft Copilot API access was available.
  6. Broikos, N. (2026). Answer-engine citation log (M2). Dataset. 36 responses, 303 source citations, coded by domain against the census. Run 6 August 2026. https://broikos.gr/research/data/paper-01/ai-citations.csv Engines and versions as queried: OpenAI gpt-5-search-api via the chat-completions endpoint; Perplexity sonar; Google gemini-2.5-flash with Google Search grounding. One session per engine, first response only, no follow-ups, run from Greece. One Gemini response returned an API error and is counted as an answer naming no sources.
  7. Broikos, N. (2026). Coding reliability records. Three rounds: two intra-rater test-retest on systematic 20% samples, and one inter-rater round across all 68 rows. https://broikos.gr/research/data/paper-01/reliability-check.csv https://broikos.gr/research/data/paper-01/reliability-check-2.csv https://broikos.gr/research/data/paper-01/inter-rater-check.csv Round 1 failed at 0.64 and is published in full, including the three rule corrections and the schema change it forced. Round 3 used DeepSeek deepseek-chat at temperature 0 as an independent second coder, blind to the original codes, and is published row by row including every disagreement. It failed on the motive column at κ 0.34.
  8. Broikos, N. (2026). International comparison corpus. Dataset. 119 publication programmes across Greece, Europe and America on a matched schema. https://broikos.gr/research/data/paper-01/international-corpus-119.csv Lower evidential standard than the census, and used only for distributional comparison. 64 of the 119 rows are confirmed and 55 unverified; no reliability check and no second coder. Stated as such wherever it is used.
  9. Broikos, N. (2026). Claim ledger and source log. The verification status of every figure considered for this paper. https://broikos.gr/research/data/paper-01/claims-claim-ledger.csv https://broikos.gr/research/data/paper-01/source-source-log.csv

Sources considered and cut

Both external statistics an earlier draft used to argue that published research works commercially were removed at citation audit. They are recorded here because a paper that grades other people’s sourcing owes the reader its own casualties.

  1. Edelman and LinkedIn. 2025 B2B Thought Leadership Impact Report. https://www.edelman.com/expertise/Business-Marketing/2025-b2b-thought-leadership-report CUT. Could not be opened at source within the verification window. The finding may well be sound; the citation was not verifiable, so it was removed rather than carried on the strength of secondary reporting.
  2. Demand Gen Report. Content Preferences Benchmark Survey. https://www.demandgenreport.com/resources/content-preferences-survey/ CUT: misattribution. The widely-quoted claim that only 5% of buyers are in-market at any moment was traced and found not to originate with the report it had been attributed to, but with separate research quoted inside it. This is failure mode F2 in the citation taxonomy used for this project: attributing a finding to the document that repeated it rather than to the one that produced it.

Coded exemplars, with their URLs printed

The census codes 68 publication programmes from their published artefacts. Naming all 68 in the text would make this a directory rather than a paper, and the full coded dataset is published separately. These seven are the exemplars the text discusses by name, listed here with live URLs so a reader can check the coding against the artefact rather than taking it on trust. All seven were re-checked on 7 August 2026.

  1. Hosthub. The Top 100 Issues Hosts Have to Handle (Greek edition). Own-data format exemplar; more than a million guest messages; email-gated. https://www.hosthub.com/el/blog/whitepaper-ta-100-korifaia-zitimata/
  2. Hosthub. Short-term rentals in Greece. Second programme from the same publisher, used for the edition-count coding. https://www.hosthub.com/blog/str-in-greece/
  3. Convert Group. Greek ePharmacy report, 2024. Own-data exemplar built on transaction panel data. https://convertgroup.com/reports_posts/greek-epharmacy-2024/
  4. Randstad Greece. Press releases and research index. Commissioned-survey exemplar with a continuous edition series. https://www.randstad.gr/h-randstad/deltia-typou/
  5. Obrela Security Industries. Reports index. Own-operational-data exemplar in a technical sector. https://www.obrela.com/resources/reports
  6. Workable. Free ebooks and reports. International comparator for the format coding. https://resources.workable.com/free-ebooks-and-reports
  7. Broikos, N. Prior whitepaper research dataset. The 119-programme lighter scan used for the international comparison in section 06. https://broikos.gr/research/data/paper-01/

A note on what is not cited here

There is no literature review in this paper and no academic citation apparatus, and that is a deliberate limitation rather than an oversight. The paper makes an empirical claim about a population of documents in one country and defends it with a census, not with a theoretical framework. A reader looking for the signalling-theory or content-marketing literature that would situate this work will not find it, and should treat the absence as a boundary on what the paper claims.

13

Citation, licence and interests

How to cite this paper

Broikos, N. (2026) Who publishes the numbers: a census of Greek corporate research publishing. Version 1.0, 6 August 2026. Athens. Available at: https://broikos.gr/research/who-publishes-the-numbers.pdf (Accessed: DD Month YYYY).

Version

Version 1.0, published 6 August 2026. This is the first public version. Substantive changes will increment the version number and be listed in the change log published with the data.

Data availability

Everything underlying this paper is published at https://broikos.gr/research/data/paper-01/ under the same licence as the text. This is the full working apparatus, not a summary extract:

census.csv: all 70 coded programmes with 20 variables each, including the two excluded rows and their exclusion reason, and the source URL for every row. search-log.csv: all 36 documented searches including the 11 that found nothing. coding-rules.md: the full rule set with every ambiguity recorded at the point it was hit. amplification.csv: the media pickup measurement. reliability-check.csv and reliability-check-2.csv: the two intra-rater rounds, including the round that failed at 0.64. inter-rater-check.csv: the round-three independent coding, row by row, including every disagreement. ai-citations.csv and prompts-m2.md: the answer-engine test: the prompts as locked before the run, and every source named in all 36 responses. claim-ledger.csv and source-log.csv.

The paper is designed so that a reader who disagrees with a coding decision can find the row, read the rule, and re-derive it. That is the point of publishing the failed reliability round and the disagreement log rather than only the final figures.

No confidential data was used. Every coded programme is a publicly available document.

Declaration of interests

The author produces corporate research and would sell research services. This paper argues that Greek companies under-use a capability the author sells, which is a conflict of interest of the most direct kind and is stated here as plainly as it can be.

Three specific disclosures. The author's own e-commerce businesses are not in the census and were not eligible for it. No organisation coded here was contacted, consulted or given sight of the paper before publication, so no coding decision could have been influenced by a commercial relationship. And the author's own coding failed its first reliability check at 0.64 and its motive column failed the independent check at κ 0.34; both are printed in full rather than repaired quietly, which is the only real defence available against the interest declared above.

Licence

This paper and the datasets published with it are released under a Creative Commons Attribution 4.0 International licence (CC BY 4.0). You may copy, redistribute, quote, chart and build on this material, including commercially, provided you credit the source. Licence text: https://creativecommons.org/licenses/by/4.0/, and served beside this paper at https://broikos.gr/research/LICENSE-CC-BY-4.0.txt. What the grant covers and what it withholds is itemised in https://broikos.gr/research/LICENSING.md. Analysis code is MIT: https://broikos.gr/research/LICENSE-MIT.txt

Funding and independence

Unfunded. No sponsor, client, trade body or commissioning party paid for, commissioned, reviewed or approved this research. No organisation named in it was given sight of it before publication. This matters because a substantial share of the sources graded here were commissioned, and the paper argues that a commissioning interest should be stated wherever it exists.

Corrections

If a figure here is wrong, or an organisation is described incorrectly, write and it will be corrected with the date and the reason recorded in a public change log rather than silently edited. That log is published at https://broikos.gr/research/corrections.md and already records every figure cut, corrected or downgraded during production. Every correction made during production is already recorded in the claim ledger published with this paper, including the figures that were cut.

Contact: https://broikos.gr/contact · https://broikos.gr

About the author

Nikolaos Broikos operates e-commerce businesses in Greece, works in web development and digital strategy, and builds agent harnesses. He writes from Athens.

This programme is the record of that work rather than a commentary on it. The e-commerce operations supply the transaction data behind the landed-cost model in What a Greek online order actually costs; the harnesses are the instruments measured in What actually makes an agent harness work. Where the author’s own systems are the subject, that is stated in the first paragraph of the paper concerned as well as in its declaration of interests.

No academic affiliation, no institutional backing, no funding, and no client commissioned any of this. The papers therefore ask to be judged on their published instruments, data and corrections rather than on credentials: every dataset, every claim ledger including the claims that failed, and every retraction is published alongside the text, so a reader who distrusts the author can check the work instead.


The other papers in this programme

Independent, unfunded and published free under CC BY 4.0, each with its underlying data. Read separately; they share a method, not an argument.

  1. What a Greek online order actually costs https://broikos.gr/research/what-a-greek-online-order-costs.pdf What the €36.1bn e-commerce figure counts and what it does not, the convergence of Greek and European online buying, and a landed-cost model built from published tariffs rather than quoted rates.
  2. The adoption gap: what businesses say about AI and what the statistics measure https://broikos.gr/research/the-adoption-gap.pdf Where Greek firms actually sit on AI adoption once the size classes are separated, why the skills-gap explanation does not survive the spending data, and what Greek firms bought instead.
  3. What actually makes an agent harness work https://broikos.gr/research/agent-harness.pdf Techniques, prompt engineering and measured results from three agent harnesses, two built by the author and one not, across 45 graded sessions on a single model, including three faults found in the measuring apparatus itself.
  4. The Greek e-shop technical audit https://broikos.gr/research/greek-eshop-audit.pdf 900 Greek domains measured on one day: what Greek online retail actually sends to a browser. 36.8% send no security headers, 11.1% fail a validating TLS handshake, and 5 shops out of 505 pass five elementary checks.
  5. The Greek SME digital bill https://broikos.gr/research/greek-sme-digital-bill.pdf What it costs per year to sell online in Greece, priced entirely from published pages: EUR 157.50 at the floor. Of 25 cost lines, 7 have no published price at all, and they are disproportionately the mandatory ones.
  6. What an agent actually costs https://broikos.gr/research/what-an-agent-costs.pdf Seven versions of one production agent on the same benchmark and the same model, including the two versions that got worse, the failure taxonomy, and the chart that would have shown a cost explosion that never happened.