CCBot
Common Crawl's crawler — the open corpus a great many AI labs train on. It is the highest-leverage single training opt-out there is.
| Operator | Common Crawl |
|---|---|
| What it does | Open training corpus |
| Obeys robots.txt | Yes, per its operator's docs |
| Governs | AI model training |
| Has a crawler | Yes |
| Last verified | 13 August 2026 |
What blocking it costs you
What to know
One block here reaches downstream to every lab that trains on Common Crawl snapshots, which is a long list. No other single robots.txt rule covers as much ground.
Two limits are worth being clear about. It is not retroactive: existing snapshots already contain what they contain, and blocking CCBot today affects future ones only. And Common Crawl itself warns about spoofers sending its user agent, which is a good reminder that a declared user agent is a claim rather than an identity.
The robots.txt rules
To block CCBot:
User-agent: CCBot
Disallow: /To allow it explicitly — worth doing when your file also contains a blanket User-agent: * disallow, since the most specific matching group wins and a token with no group of its own falls back to the wildcard:
User-agent: CCBot
Allow: /If neither rule is present, CCBot is allowed. That is the default, and it is why an accidental block is nearly always something that was added rather than something that was forgotten — usually a snippet copied from a blog post that listed more tokens than the author intended to block.
Is it blocked on your site?
Your robots.txt is only the half you can read. A CDN or bot-management rule can turn CCBot away while your file still says it is welcome, and that mismatch is invisible from the front of the site. We check both layers separately and tell you which one is doing what.
We request two things: your robots.txt and your homepage. No other pages, ever.
Official sources
- Common Crawl’s crawler documentation
- Published IP ranges — the only way to verify a request claiming this user agent really came from Common Crawl