Is Cloudflare blocking AI crawlers?

Quite possibly, and your robots.txt won’t tell you. The file can say yes while Cloudflare says no at the door, and only one of those is visible from the front of your site. Registry verified 19 September 2026.

What Cloudflare blocks now

Cloudflare sorts AI bots into three kinds. Search crawlers index your content to answer questions about it later. Agent traffic fetches a page in real time for a person, such as a chat assistant opening a link. Training crawlers take content to train or fine-tune a model, and mixed-purpose crawlers count as training.

That was not the first default. Since July 2025, new sites on Cloudflare have been set to block AI crawlers unless their owner turned it off, so an older site may carry a block nobody remembers enabling.

None of this is a bad idea on its own. The trouble is that plenty of people never chose it, and most of them have no idea it is on.

Why every robots.txt checker says you’re fine

Because it read the file, and the file is fine. The block happens earlier. Cloudflare answers the crawler at the edge, before the request ever reaches the server that would have followed your robots.txt. A tool that only parses the file can’t see that layer, however good it is.

How to check, from quickest to most certain

1. Look at your robots.txt

Open yoursite.com/robots.txt. If it has Content-Signal lines and a long comment quoting Article 4 of the EU copyright directive, Cloudflare probably wrote it, not you. That combination is the fingerprint our scanner looks for.

2. Ask as the crawler

Fetch your home page twice, once as a browser and once wearing a crawler’s name:

curl -sI https://yoursite.com
curl -sI -A "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot" https://yoursite.com

If the first answers 200 and the second gets a 403, a page titled “Just a moment...”, or a cf-mitigated: challenge header, something at the edge treats that user agent differently. An HTTP 402 with a crawler-price header is its own thing: that’s Pay-Per-Crawl, and the crawler is being asked to pay rather than turned away.

One honest caveat. Your curl isn’t really OAI-SearchBot. It carries the name without coming from OpenAI’s published IP ranges, and Cloudflare can tell. A 403 to the imposter proves a user-agent rule exists. It doesn’t prove the real crawler gets the same treatment, and a 200 doesn’t prove it gets in either. Only the dashboard knows that.

3. Look in the dashboard

If you can log in to Cloudflare for the domain, open AI Crawl Control. Check which crawlers are set to block, and whether Bot Preference Sync is writing your robots.txt. This is the only place the full answer lives.

What to change

Decide by what each crawler does, not by which company runs it. These are the ones that put you in AI answers, and they are worth allowing even if you block everything else:

OAI-SearchBotYour site disappears from ChatGPT search answers and citations. The most damaging accidental block we see.
Claude-SearchBotReduced visibility in Claude's search results.
PerplexityBotYour site will not be surfaced or linked in Perplexity answers.

The training crawlers are the ones you can refuse without losing traffic. Which AI crawlers to allow walks through the choice by kind of site. Whatever you pick, check afterwards that Googlebot and Bingbot didn’t get caught in the new mixed-use scope.

What we can’t tell you

What your dashboard says, or how Cloudflare treats a crawler arriving from its verified IPs. From outside, we can show whether the response changes when the user agent does, for every AI crawler at once, which is the half nobody else checks.

The scan does that in about ten seconds, with no account.

Crawlers in this guide

Where to go next

All AI crawler guides