AI visibility

Which websites block AI crawlers, and why most of them never decided to

robots.txt and llms.txt look like a clean buying signal for anyone selling SEO, content or AI visibility. Across 151.6 million live websites, most of those rules were written by a platform default. Here is how to tell a decision from a template.

151.6M websites7 sourcesChecked September 2026No survey data

Every go-to-market team selling into content, SEO or AI search has had the same idea this year.

If a website blocks GPTBot, it cares about AI. If it publishes an llms.txt, it wants to be found by AI. Either way it looks like a warm account.

It is a good idea with a catch. StackScan fetched the robots.txt and llms.txt of every live website on the internet, 151.6 million of them, for its AI crawler statistics and its study of 17.9 million llms.txt files. The numbers show that most of these files were written by software, not by the people who run the sites.

Why this is a go-to-market question

A rule somebody chose is intent. A rule a platform wrote for them is noise.

Mixing the two is how a promising segment turns into a list of small hosted sites whose owners have never heard of GPTBot.

The blocks

Blocking GPTBot is common, and mostly one vendor's default

6,250,586 websites block GPTBot. That is 6.88% of the 90.8 million sites with a robots.txt, 59.4% of the sites that mention GPTBot at all, and 4.1% of every live website. All three figures are correct, and published reports routinely mix them up.

The bigger finding is where the blocks sit. Sites on Cloudflare's addresses block GPTBot at 23.0% and every other site at 1.5%, so 84.0% of all GPTBot blocks are behind Cloudflare.

That concentration has a start date. On 1 July 2025 Cloudflare began blocking AI crawlers by default on new websites, with the Associated Press, Condé Nast, The Atlantic, Gannett and other publishers backing the change, as TechCrunch and Press Gazette reported. A year later it announced that crawlers mixing search, agent use and training will be blocked by default on ad-supported pages from 15 September 2026, for new customers, new sites and every free plan (TechCrunch).

The block counts for the eight most blocked AI crawlers sit within roughly half a million sites of one another, which is what a single template looks like in the data. And 96.0% of the sites using Cloudflare's Content-Signals block also use Cloudflare nameservers.

What to do with it

Treat a GPTBot block on Cloudflare nameservers as a default until proven otherwise.

Treat one on any other network, or one sitting next to rules the owner clearly wrote, as a decision worth a conversation.

Remember too that robots.txt is a request, not a wall. When Patreon switched from robots.txt to active blocking in July 2026, weekly attempts by individual AI training crawlers went from thousands to zero, which suggests they had been ignoring the file (TechCrunch).

The allows

The deliberate choices are more useful than the blocks

Of the sites that mention GPTBot by name, 4.3 million let it in and 6.3 million shut it out. Roughly 40% of explicit GPTBot rules are allows rather than blocks.

A larger group, 5.9 million sites, keeps OpenAI's training crawler out while letting its search crawler and user-triggered fetches in. The opposite setup is almost unheard of, at 13,850 sites. These are sites that want to be found in AI search without handing over their content for training, which is exactly the audience for AI visibility and citation tracking tools.

A small group, 231 thousand sites, actively opts in to training with ai-train=yes in a Content-Signals block.

The economics explain the mood. Cloudflare's June 2025 figures, reported by TechCrunch, counted 14 Google crawls for every visitor Google sent back to a site, against 1,700 for OpenAI and 73,000 for Anthropic. GPTBot's share of AI-only crawling rose from 5% in May 2024 to 30% a year later, according to Cloudflare figures reported by Press Gazette.

The biggest sites mostly let it in. Among the top 1,000 with a robots.txt, 87.0% allow GPTBot, and the top 10,000 are almost identical at 86.9%.

llms.txt

llms.txt is everywhere, and platforms wrote most of it

Of the 151,649,745 live websites on the internet, 17,956,268 serve an llms.txt, 11.8% of the web. That is far above earlier estimates, which sampled large publishers.

The reason is who wrote them. Wix generated the file for 3.5 million websites and GoDaddy for 3.6 million, two fifths of the total between them. Shopify's storefront template accounts for another 2.2 million, and three WordPress SEO plugins wrote 1.8 million more.

One file in five, 19.9% or 3,571,018 sites, is empty or belongs to a site that is closed or not launched yet.

Adoption falls as companies grow: 26.3% of sole traders serve an llms.txt, but only 14.9% of firms with more than ten thousand employees.

Whether any model reads the file is another matter. In June 2026 Google's John Mueller called its value 'purely speculative for now' and noted that none of the AI systems use it, as Search Engine Journal reported. For a seller, that makes an llms.txt a sign of what the site owner hopes for rather than proof of AI traffic.

The filter that works

Keep files that are not a platform template and that contain links. The median llms.txt has 5 links, and 1.1 million files over 200 bytes have none at all.

Segments

Four segments worth building

  1. Wants AI search, not AI training
    Blocks the training crawler but allows search and user fetches: 5.9 million sites. The natural buyers of AI visibility and citation tracking.
  2. Hand-written llms.txt with real links
    Leave out the Wix, GoDaddy and Shopify templates and the closed sites, and what remains was written by someone who cares how models read them. A good fit for documentation and content tooling.
  3. Rules nobody maintains
    54,110 sites still name Anthropic's old anthropic-ai and Claude-Web user agents in a block while never mentioning the current ClaudeBot, so the crawler they meant to stop walks straight in. chatgpt.com is one of them. A clean opener for technical SEO audits.
  4. Mixed signals
    548.4 thousand websites publish an llms.txt while blocking every unlisted crawler in robots.txt. Someone is inviting the models in and locking the door. That is a conversation about AI policy, not a product demo.
Questions

Things people ask

How many websites block GPTBot?

6,250,586 websites, which is 6.88% of the 90.8 million sites that publish a robots.txt and 4.1% of every live website. Among the sites that name GPTBot at all, 59.4% block it.

Is llms.txt an official standard?

No. It was proposed in 2024 as a Markdown file that tells a language model what a site is and where its useful content lives. Adoption is high mainly because hosting platforms now generate it by default.

Does blocking GPTBot stop ChatGPT from citing a site?

Not by itself. OpenAI runs separate agents for training, for search and for fetches a user asks for, and 5.9 million websites block the training crawler while allowing the other two.

Method

Where the numbers come from

Every figure comes from two StackScan studies published on 10 September 2026: AI Crawler Statistics 2026 and llms.txt Statistics 2026.

Both measured all 151,649,745 live websites on the internet. Blocking shares are of the 90.8 million sites that publish a robots.txt unless a line says otherwise, and llms.txt shares are of every live website.

Context on Cloudflare's policy, crawl-to-referral ratios, Patreon and llms.txt comes from TechCrunch, Press Gazette and Search Engine Journal, each linked where it is used.

Related on iNetZeal

Build the list with our 17 lead generation tools by cost per lead, size the category with 15 competitor analysis tools, or see 27 RevOps statistics on how the buying side is changing.