September 2026 measurement · compared with —

Can AIs read small-business websites? We measured 3 054 sites across 12 countries.

Nobody was publishing the figure for France. We measure it every month, on a sample of shops, tradespeople, restaurants and practices drawn at random across fifteen départements. The first measurement contradicted our own headline — we had written that “many sites” had shut themselves off from AI without knowing. That is false, and we corrected it the same day. Here is what is true, and how it moves.

1,1 %
of sites block a bot that answers inside the AIs
71,2 %
show no price written out in full
74,9 %
show no update date at all
1 200
sites read, none unreachable

The sample: small businesses, not media

The addresses come from OpenStreetMap, which records the website of hundreds of thousands of French shops, tradespeople, offices and restaurants. Fifteen départements, from Paris to the Cantal, so as not to measure a single metropolis. Above all: no popularity ranking. A top 1,000 is full of media outlets, and the press blocks heavily — 49% of US and UK news sites shut the door on OAI-SearchBot, against 0.2% of ordinary business sites. Inferring anything about a bakery from that would be a reasoning error. Random draw with a fixed seed after de-duplication by domain, so the measurement can be replayed identically. Facebook pages, booking platforms and directories are excluded: those are somebody else's sites.

The results, line by line

What was foundSitesShareChange
At least one search bot blocked — the site cannot be quoted13 / 1 2001,1 %
of which OAI-SearchBot (ChatGPT)11 / 1 2000,9 %
of which Claude-SearchBot7 / 1 2000,6 %
of which PerplexityBot11 / 1 2000,9 %
GPTBot blocked — training only, costs no citation42 / 1 2003,5 %
robots.txt unreadable: the server refuses or does not answer239 / 1 20019,9 %
noindex or nosnippet tag on the home page11 / 8961,2 %
Content that probably depends on JavaScript41 / 8964,6 %
No price written out in full638 / 89671,2 %
No readable update date671 / 89674,9 %

Two denominators, and they must not be mixed: what is read in robots.txt covers every site in the sample, what is read in the page covers only the pages actually read. Each row therefore shows its own. The “Change” column is in percentage points since the previous measurement; it stays empty in the first month, for want of a comparison.

And elsewhere? The same measurement, country by country

The same diagnostic, on an OpenStreetMap sample specific to each country. Each row has its own month and its own sample: nothing is mixed, only like is compared with like.

CountryMeasuredSites readAt least one search bot blocked — the site cannot be quotedNo price written out in fullNo readable update date
United Arab EmiratesSeptember 20261981,5 %93,6 %79,8 %
BelgiumSeptember 20261981,0 %71,2 %80,1 %
SwitzerlandSeptember 20261910,5 %94,6 %81,3 %
EgyptSeptember 20261771,7 %86,5 %85,4 %
JordanSeptember 20261171,7 %91,5 %85,4 %
KuwaitSeptember 20261420,7 %94,1 %87,1 %
LebanonSeptember 20261421,4 %75,6 %77,9 %
MoroccoSeptember 20262150,0 %82,1 %73,7 %
QatarSeptember 20261372,2 %90,1 %84,6 %
Saudi ArabiaSeptember 20261651,2 %93,8 %87,5 %
TunisiaSeptember 20261721,7 %84,9 %66,7 %

What it means: blocking is rare, missing information is not

About one site in a hundred has removed itself from AI answers. That is rare — but it is total, it lasts as long as the line stays in the file, and nobody is ever notified: it is an argument about severity, not frequency. The gesture itself is common: 3.5% block GPTBot, which only trains models and costs no citation. Many sites therefore believe they have protected themselves while closing nothing that matters. What is genuinely massive lies elsewhere: seven sites in ten do not write their price, three in four show no date. Yet the SIGIR 2026 study, over 252,000 matched trials, ranks explicit price among the four decisive factors for citation, and Ahrefs, over 17 million citations, shows AIs quote fresher content than organic search does.

What it does not mean

That writing a price will get your client quoted. Nobody has published a before/after study with a control group showing that, and we will not promise it while that remains true. What these figures establish is a frequency: how many sites carry a verifiable obstacle, documented by the providers themselves or measured with a control group. Removing a verifiable obstacle is useful work; selling it as a guarantee of citation would be a lie.

The limits, spelled out

1 200 sites are not a census: the intervals above state the margin. The sample only contains businesses visible enough to have been mapped along with their website — the smallest are absent. One site in five did not let its robots.txt be read, usually because its host refuses any unknown bot; those are counted as “unknown”, never as “open”. Only the home page was read, not product pages. And it is a snapshot taken at the time of measurement: a robots.txt changes in one line.

Redo the measurement

The diagnostic behind these figures is exactly the one in the free tool: two HTTP requests per site, no model call at all, so it costs nobody anything. You can run it on any address, five times a day without an account, without limit once signed in. The sample data comes from OpenStreetMap's Overpass API, freely queryable: the method described here is enough to redo the measurement and contradict us, which is the only serious way to validate it.

And your client's site — where does it stand?

The same diagnostic, on any address. Free, no account, ten seconds.

Check a site →

← Back to home