CCBot

Common Crawl's crawler — the open corpus a great many AI labs train on. It is the highest-leverage single training opt-out there is.

OperatorCommon Crawl
What it doesOpen training corpus
Obeys robots.txtYes, per its operator's docs
GovernsAI model training
Has a crawlerYes
Last verified13 August 2026

What blocking it costs you

Excluded from future Common Crawl snapshots — the open corpus many AI labs train on. Indirectly reduces presence in many models' training data (not retroactive).

What to know

One block here reaches downstream to every lab that trains on Common Crawl snapshots, which is a long list. No other single robots.txt rule covers as much ground.

Two limits are worth being clear about. It is not retroactive: existing snapshots already contain what they contain, and blocking CCBot today affects future ones only. And Common Crawl itself warns about spoofers sending its user agent, which is a good reminder that a declared user agent is a claim rather than an identity.

The robots.txt rules

To block CCBot:

User-agent: CCBot
Disallow: /

To allow it explicitly — worth doing when your file also contains a blanket User-agent: * disallow, since the most specific matching group wins and a token with no group of its own falls back to the wildcard:

User-agent: CCBot
Allow: /

If neither rule is present, CCBot is allowed. That is the default, and it is why an accidental block is nearly always something that was added rather than something that was forgotten — usually a snippet copied from a blog post that listed more tokens than the author intended to block.

Is it blocked on your site?

Your robots.txt is only the half you can read. A CDN or bot-management rule can turn CCBot away while your file still says it is welcome, and that mismatch is invisible from the front of the site. We check both layers separately and tell you which one is doing what.

We request two things: your robots.txt and your homepage. No other pages, ever.

Official sources

← All AI crawlers