Which AI crawlers should you allow?
The 26 crawler tokens in our registry do four different jobs. The costly mistake is blocking the crawler that puts you in answers when you meant to block the one used for training. Verified 19 September 2026.
The short answer
If you have never touched these settings, you are usually fine. With no matching rule, a crawler is allowed. Accidental blocks are more often something added, such as a copied snippet that listed more tokens than intended.
If you want to appear in AI answers, allow the search and citation crawlers. If you do not want your content used for training, block the training crawlers. They are separate choices.
The four jobs
Search & citations
These fetch your pages so an assistant can quote you and link you when somebody asks a question you have answered. This is the group that decides whether you are in the answer.
| OAI-SearchBot | OpenAI. Your site disappears from ChatGPT search answers and citations. The most damaging accidental block we see. |
|---|---|
| Claude-SearchBot | Anthropic. Reduced visibility in Claude's search results. |
| PerplexityBot | Perplexity. Your site will not be surfaced or linked in Perplexity answers. |
| meta-webindexer | Meta. Your site will not be cited or linked in Meta AI answers. |
| MistralAI-Index | Mistral. Out of Le Chat's search results and citations. |
| DuckAssistBot | DuckDuckGo. Out of DuckAssist AI answers. |
Model training
These collect text to train future models. Blocking them is a policy choice with no search cost attached, which makes it the one group you can refuse without losing traffic.
| GPTBot | OpenAI. Excluded from future OpenAI model training. Does NOT affect ChatGPT search citations (that's OAI-SearchBot). |
|---|---|
| ClaudeBot | Anthropic. Excluded from future Anthropic model training. Does not affect Claude's search visibility (that's Claude-SearchBot). |
| Google-Extended | Google. Blocks Gemini model training and grounding (content fed from the Search index to Gemini at prompt time). DOES NOT remove you from AI Overviews or AI Mode, and does not affect Search ranking. If you blocked this to hide from AI Overviews, it isn't doing what you think. |
| Applebot-Extended | Apple. Content excluded from Apple foundation-model training. Siri/Spotlight presence unaffected. |
| meta-externalagent | Meta. Content excluded from Meta AI training and direct indexing. |
| Bytespider | ByteDance. Best-effort only. ByteDance publishes no official crawler documentation or IP ranges, and Bytespider is widely reported to ignore robots.txt. |
| MistralAI-Training | Mistral. Out of Mistral model training. |
Fetches when a user asks
These run when a person pastes your URL or asks about you by name. One request, one page, prompted by somebody who has already found you and wants to know more.
| ChatGPT-User | OpenAI. Signals opt-out of live page fetches when ChatGPT users ask about your site, but OpenAI notes robots.txt rules 'may not apply' to user-initiated fetches. |
|---|---|
| Claude-User | Anthropic. Claude cannot fetch your pages when users ask about your site or product. |
| Perplexity-User | Perplexity. Advisory only. Perplexity's own docs state this user-triggered fetcher 'generally ignores robots.txt rules'. |
| meta-externalfetcher | Meta. Advisory only. Meta's own documentation says this fetcher may bypass robots.txt, because the fetch is something a user asked for. |
| MistralAI-User | Mistral. Le Chat cannot open your pages when somebody asks it about them. |
Open training corpus
An open archive that many labs train from, rather than one company's crawler. Blocking it reaches more models than any other single rule, and it is not retroactive.
| CCBot | Common Crawl. Excluded from future Common Crawl snapshots, the open corpus many AI labs train on. Indirectly reduces presence in many models' training data (not retroactive). |
|---|
The mistake that costs the most
Many operators separate training from search and citations. Blocking the training token is a policy choice. Blocking the search token can remove you from answers.
| GPTBot | Excluded from future OpenAI model training. Does NOT affect ChatGPT search citations (that's OAI-SearchBot). |
|---|---|
| OAI-SearchBot | Your site disappears from ChatGPT search answers and citations. The most damaging accidental block we see. |
| ClaudeBot | Excluded from future Anthropic model training. Does not affect Claude's search visibility (that's Claude-SearchBot). |
| Claude-SearchBot | Reduced visibility in Claude's search results. |
| Google-Extended | Blocks Gemini model training and grounding (content fed from the Search index to Gemini at prompt time). DOES NOT remove you from AI Overviews or AI Mode, and does not affect Search ranking. If you blocked this to hide from AI Overviews, it isn't doing what you think. |
The pattern to take from that table: a rule written to keep your writing out of a training set should name the training token only. A blanket rule aimed at one company catches both of its crawlers, and only one of them was the problem.
By what your site is for
The registry tells you what each token does. The site-type recommendations below are our practical guidance, not facts claimed by the operators.
You sell to people nearby
Somebody asks an assistant who does this near them, and you want to be the name it gives. Nothing about your pages is worth training on, so the training question barely applies to you.
Worth allowing deliberately: OAI-SearchBot, Claude-SearchBot, PerplexityBot, meta-webindexer.
There is nothing here worth blocking. Every token that could hide you is a citation crawler, and the training crawlers are not reading your opening hours for a model.
You publish writing
You would like the citation when an assistant answers a question you already answered at length. This is also the one profile where refusing to be trained on is a real position.
Worth allowing deliberately: OAI-SearchBot, Claude-SearchBot, PerplexityBot, DuckAssistBot.
Worth considering a block on: CCBot, GPTBot, ClaudeBot.
Blocking the training crawlers costs you no citations and no search traffic, because a different token does that job. CCBot is the highest-leverage single rule, since many labs train from that one archive rather than crawling you themselves.
You sell software or a service
Somebody asks an assistant how to do a thing your product does, or asks about your product by name. Your docs are the pages that answer both, and they are usually the ones a site renders client-side.
Worth allowing deliberately: OAI-SearchBot, Claude-SearchBot, PerplexityBot, ChatGPT-User, Claude-User.
The user fetchers matter more here than anywhere else. When somebody evaluating you pastes your pricing page into an assistant, that is the request being made, and blocking it means the assistant answers about you from whatever it remembers instead.
You sell things online
Product pages that can be found, and links that look like something when a customer shares one. The token that bites here is not an AI crawler at all.
Worth allowing deliberately: OAI-SearchBot, PerplexityBot, meta-webindexer, facebookexternalhit.
facebookexternalhit is the one to be careful with. It builds the preview card when somebody shares you on Facebook, Instagram, WhatsApp or Messenger, and a blanket block on meta tokens takes it out. Every shared link then degrades to a bare URL, which is usually noticed weeks later as posts quietly underperforming.
The ones that are not a choice
Three things on this list look like decisions and are not. Blocking Googlebot removes you from Google Search entirely, including AI Overviews and AI Mode, because those are built on ordinary Search crawling rather than on a separate AI crawler. It is almost never what somebody setting out to keep AI off their site intended.
Applebot has a quirk worth knowing: if your file names Googlebot and says nothing about Applebot, Applebot follows your Googlebot rules. A file that carefully allows Google is allowing Apple too, and one that disallows Googlebot has just disallowed Apple without saying so.
And no rule reaches xAI (Grok) at all. Grok fetches with ordinary browser user agents, so there is no token to address and a wildcard disallow aimed at it would turn away every crawler you want. A rule naming it reads as protection and does nothing.
Writing the file
A crawler obeys exactly one group, the most specific one that names it, and ignores every other group in the file. It does not merge them. So a token with its own group inherits nothing from User-agent: *, and a token without one inherits everything.
The generator writes this out with the cost of each choice shown before you paste it, and the checker reads a file you already have and says which of these 26 tokens it blocks. Neither asks for a domain or an account.
Where to go next
- GPTBot vs OAI-SearchBot: which to block?
OpenAI's two crawlers do opposite jobs. One costs you nothing to block and the other takes you out of ChatGPT's answers. - ClaudeBot vs Claude-SearchBot: which to block?
Anthropic's crawlers split the same way OpenAI's do, and the legacy tokens in most files do nothing at all. - Does Google-Extended block AI Overviews?
A common robots.txt mix-up, and the controls that actually decide what AI Overviews can quote.