Is Cloudflare blocking AI crawlers?
Quite possibly, and your robots.txt won’t tell you. The file can say yes while Cloudflare says no at the door, and only one of those is visible from the front of your site. Registry verified 19 September 2026.
What Cloudflare blocks now
Cloudflare sorts AI bots into three kinds. Search crawlers index your content to answer questions about it later. Agent traffic fetches a page in real time for a person, such as a chat assistant opening a link. Training crawlers take content to train or fine-tune a model, and mixed-purpose crawlers count as training.
- Since September 15, 2026, Cloudflare blocks training and agent bots on ad-supported pages, for new domains and for every site on its free plan. Search stays allowed.
- The same release replaced Managed robots.txt with Bot Preference Sync, which writes your robots.txt from your Search, Training and Agent settings. While it is on, editing the file on your own server changes nothing.
- It also brought mixed-use crawlers inside the scope of Cloudflare’s AI blocking for the first time, including Googlebot, Bingbot and Applebot. A switch flipped to keep AI out can now take ordinary search with it.
That was not the first default. Since July 2025, new sites on Cloudflare have been set to block AI crawlers unless their owner turned it off, so an older site may carry a block nobody remembers enabling.
None of this is a bad idea on its own. The trouble is that plenty of people never chose it, and most of them have no idea it is on.
Why every robots.txt checker says you’re fine
Because it read the file, and the file is fine. The block happens earlier. Cloudflare answers the crawler at the edge, before the request ever reaches the server that would have followed your robots.txt. A tool that only parses the file can’t see that layer, however good it is.
How to check, from quickest to most certain
1. Look at your robots.txt
Open yoursite.com/robots.txt. If it has Content-Signal lines and a long comment quoting Article 4 of the EU copyright directive, Cloudflare probably wrote it, not you. That combination is the fingerprint our scanner looks for.
2. Ask as the crawler
Fetch your home page twice, once as a browser and once wearing a crawler’s name:
curl -sI https://yoursite.com
curl -sI -A "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot" https://yoursite.comIf the first answers 200 and the second gets a 403, a page titled “Just a moment...”, or a cf-mitigated: challenge header, something at the edge treats that user agent differently. An HTTP 402 with a crawler-price header is its own thing: that’s Pay-Per-Crawl, and the crawler is being asked to pay rather than turned away.
One honest caveat. Your curl isn’t really OAI-SearchBot. It carries the name without coming from OpenAI’s published IP ranges, and Cloudflare can tell. A 403 to the imposter proves a user-agent rule exists. It doesn’t prove the real crawler gets the same treatment, and a 200 doesn’t prove it gets in either. Only the dashboard knows that.
3. Look in the dashboard
If you can log in to Cloudflare for the domain, open AI Crawl Control. Check which crawlers are set to block, and whether Bot Preference Sync is writing your robots.txt. This is the only place the full answer lives.
What to change
Decide by what each crawler does, not by which company runs it. These are the ones that put you in AI answers, and they are worth allowing even if you block everything else:
| OAI-SearchBot | Your site disappears from ChatGPT search answers and citations. The most damaging accidental block we see. |
|---|---|
| Claude-SearchBot | Reduced visibility in Claude's search results. |
| PerplexityBot | Your site will not be surfaced or linked in Perplexity answers. |
The training crawlers are the ones you can refuse without losing traffic. Which AI crawlers to allow walks through the choice by kind of site. Whatever you pick, check afterwards that Googlebot and Bingbot didn’t get caught in the new mixed-use scope.
What we can’t tell you
What your dashboard says, or how Cloudflare treats a crawler arriving from its verified IPs. From outside, we can show whether the response changes when the user agent does, for every AI crawler at once, which is the half nobody else checks.
The scan does that in about ten seconds, with no account.
Crawlers in this guide
- OAI-SearchBot: Search & citations, run by OpenAI
- Claude-SearchBot: Search & citations, run by Anthropic
- PerplexityBot: Search & citations, run by Perplexity
- Googlebot: Reference, run by Google
- Bingbot: Reference, run by Microsoft
- Applebot: Search & training, run by Apple
Where to go next
- Why your site is not showing up in ChatGPT
The common failure modes, in order, with a check you can run for each one. - AI SEO setup: the five things worth checking
Five layers you can check from outside, each capable of hiding a site. We go through them in the order they tend to break. - Which AI crawlers should you allow?
The tokens have four jobs. Which ones you want depends on what you are trying to achieve, and the costly mistake is blocking the citation crawler when you meant to block the training one.