robots.txt and llms.txt look like a clean buying signal for anyone selling SEO, content or AI visibility. Across 151.6 million live websites, most of those rules were written by a platform default. Here is how to tell a decision from a template.
Every go-to-market team selling into content, SEO or AI search has had the same idea this year.
If a website blocks GPTBot, it cares about AI. If it publishes an llms.txt, it wants to be found by AI. Either way it looks like a warm account.
It is a good idea with a catch. StackScan fetched the robots.txt and llms.txt of every live website on the internet, 151.6 million of them, for its AI crawler statistics and its study of 17.9 million llms.txt files. The numbers show that most of these files were written by software, not by the people who run the sites.
A rule somebody chose is intent. A rule a platform wrote for them is noise.
Mixing the two is how a promising segment turns into a list of small hosted sites whose owners have never heard of GPTBot.
6,250,586 websites block GPTBot. That is 6.88% of the 90.8 million sites with a robots.txt, 59.4% of the sites that mention GPTBot at all, and 4.1% of every live website. All three figures are correct, and published reports routinely mix them up.
The bigger finding is where the blocks sit. Sites on Cloudflare's addresses block GPTBot at 23.0% and every other site at 1.5%, so 84.0% of all GPTBot blocks are behind Cloudflare.
That concentration has a start date. On 1 July 2025 Cloudflare began blocking AI crawlers by default on new websites, with the Associated Press, Condé Nast, The Atlantic, Gannett and other publishers backing the change, as TechCrunch and Press Gazette reported. A year later it announced that crawlers mixing search, agent use and training will be blocked by default on ad-supported pages from 15 September 2026, for new customers, new sites and every free plan (TechCrunch).
The block counts for the eight most blocked AI crawlers sit within roughly half a million sites of one another, which is what a single template looks like in the data. And 96.0% of the sites using Cloudflare's Content-Signals block also use Cloudflare nameservers.
Treat a GPTBot block on Cloudflare nameservers as a default until proven otherwise.
Treat one on any other network, or one sitting next to rules the owner clearly wrote, as a decision worth a conversation.
Remember too that robots.txt is a request, not a wall. When Patreon switched from robots.txt to active blocking in July 2026, weekly attempts by individual AI training crawlers went from thousands to zero, which suggests they had been ignoring the file (TechCrunch).
Of the sites that mention GPTBot by name, 4.3 million let it in and 6.3 million shut it out. Roughly 40% of explicit GPTBot rules are allows rather than blocks.
A larger group, 5.9 million sites, keeps OpenAI's training crawler out while letting its search crawler and user-triggered fetches in. The opposite setup is almost unheard of, at 13,850 sites. These are sites that want to be found in AI search without handing over their content for training, which is exactly the audience for AI visibility and citation tracking tools.
A small group, 231 thousand sites, actively opts in to training with ai-train=yes in a Content-Signals block.
The economics explain the mood. Cloudflare's June 2025 figures, reported by TechCrunch, counted 14 Google crawls for every visitor Google sent back to a site, against 1,700 for OpenAI and 73,000 for Anthropic. GPTBot's share of AI-only crawling rose from 5% in May 2024 to 30% a year later, according to Cloudflare figures reported by Press Gazette.
The biggest sites mostly let it in. Among the top 1,000 with a robots.txt, 87.0% allow GPTBot, and the top 10,000 are almost identical at 86.9%.
Of the 151,649,745 live websites on the internet, 17,956,268 serve an llms.txt, 11.8% of the web. That is far above earlier estimates, which sampled large publishers.
The reason is who wrote them. Wix generated the file for 3.5 million websites and GoDaddy for 3.6 million, two fifths of the total between them. Shopify's storefront template accounts for another 2.2 million, and three WordPress SEO plugins wrote 1.8 million more.
One file in five, 19.9% or 3,571,018 sites, is empty or belongs to a site that is closed or not launched yet.
Adoption falls as companies grow: 26.3% of sole traders serve an llms.txt, but only 14.9% of firms with more than ten thousand employees.
Whether any model reads the file is another matter. In June 2026 Google's John Mueller called its value 'purely speculative for now' and noted that none of the AI systems use it, as Search Engine Journal reported. For a seller, that makes an llms.txt a sign of what the site owner hopes for rather than proof of AI traffic.
Keep files that are not a platform template and that contain links. The median llms.txt has 5 links, and 1.1 million files over 200 bytes have none at all.
6,250,586 websites, which is 6.88% of the 90.8 million sites that publish a robots.txt and 4.1% of every live website. Among the sites that name GPTBot at all, 59.4% block it.
No. It was proposed in 2024 as a Markdown file that tells a language model what a site is and where its useful content lives. Adoption is high mainly because hosting platforms now generate it by default.
Not by itself. OpenAI runs separate agents for training, for search and for fetches a user asks for, and 5.9 million websites block the training crawler while allowing the other two.
Every figure comes from two StackScan studies published on 10 September 2026: AI Crawler Statistics 2026 and llms.txt Statistics 2026.
Both measured all 151,649,745 live websites on the internet. Blocking shares are of the 90.8 million sites that publish a robots.txt unless a line says otherwise, and llms.txt shares are of every live website.
Context on Cloudflare's policy, crawl-to-referral ratios, Patreon and llms.txt comes from TechCrunch, Press Gazette and Search Engine Journal, each linked where it is used.
Build the list with our 17 lead generation tools by cost per lead, size the category with 15 competitor analysis tools, or see 27 RevOps statistics on how the buying side is changing.