This summer, US Tech Automations released a census of the top 100,000 domains. They found that 19.1% of the ones with a working robots.txt file block at least 1 of 21 AI crawlers. It’s a good headline, but it’s worth diving into the details.
A robots.txt file tells search engines and web bots which website pages they are allowed to visit. Plenty of these would have been set up years ago via a CMS or an old SEO plugin, so blocking the latest AI bots wouldn’t have been on their mind at the time.
There’s also an important distinction worth factoring in. AI bots broadly split into two camps – ones that scan your site to train a model, and the ones that fetch your live webpage when someone is asking about you in AI-search. GPTBot, OpenAI's training crawler, is the single most-blocked in the whole census at 15.4%. It’s not the same as blocking ChatGPT, because OpenAI runs separate bots for live search and for fetching a page when a user asks about you directly.
Blocking GPTBot alone is a reasonable call for a lot of businesses. It stops your content from training a model when you're not being paid for it, and it has limited impact on your visibility in AI answers today, since that runs through a different type of bot. The risk is in how this gets implemented. An old robots.txt file is one way to block more than a brand means to, while a blanket "block AI bots" toggle in a plugin or a CDN setting can equally block the crawlers you actually need to stay open alongside the ones you want to block.
This is about to get a lot more relevant, as Cloudflare is changing its own defaults on 15th September. From that date, both training and agent crawlers get blocked by default on any page carrying ads, while search remains permitted, and it applies to every new customer plus any existing free-plan site that hasn't changed its settings. Cloudflare hasn't published how many sites it will affect, but given W3techs has identified that they sit in front of 25% of all websites, this isn't a niche change. Agent crawlers, the ones that act on a live user's behalf rather than bulk-scraping for training data, sit much closer to the tier that actually affects whether you show up when someone asks about you right now.
If you'd like to find out whether your website is blocking the crawlers that actually matter, head on over to our free Technical GEO Checker tool, which has just launched. It checks 35 markers across 7 technical areas for the exact things that decide whether ChatGPT, Claude, Perplexity and Copilot can reach, read, and cite your content effectively.
Take the challenge now and see how your site performs. Better still, take the results directly to your web team and have them correct any flagged areas to ensure every pound / dollar spent on your content has the best chance of being picked up by AI-crawlers to aid your AI-search visibility.
