What does your robots.txt actually say to AI crawlers?

Paste your robots.txt and we’ll show you what it says to all 26 crawler tokens: what is blocked, which rule caused it, and what that means. Nothing is uploaded and no domain is needed.

nothing is sent anywhere

Open yoursite.com/robots.txt, copy the file, and paste it here.

What this file does

3 of 26 crawler tokens are blocked by this file. What each one costs you is below.

  • Not what it looks likeYou will not be cited in ChatGPT search answersOAI-SearchBot is the citation crawler, not the training one. If the intent was to opt out of training, block GPTBot instead and leave this allowed.
  • One sitemap is declared. A scan reads it to measure how much of your published content a rule actually excludes, rather than guessing from the path name.

Every crawler, and what your file says to it

ChatGPT Search

Blocked

Your site will not appear in or be cited by ChatGPT search answers.

  • OAI-SearchBotblocked in robots.txtMatched Disallow: / under User-agent: oai-searchbotYour site disappears from ChatGPT search answers and citations. The most damaging accidental block we see.
  • ChatGPT-Userallowed by this fileMatched Allow: / under User-agent: *

Claude

No rule blocks it

Claude can access your site.

  • Claude-SearchBotallowed by this fileMatched Allow: / under User-agent: *
  • Claude-Userallowed by this fileMatched Allow: / under User-agent: *

Perplexity

No rule blocks it

Perplexity can access your site.

  • PerplexityBotallowed by this fileMatched Allow: / under User-agent: *
  • Perplexity-Userallowed by this fileMatched Allow: / under User-agent: *

Meta AI

No rule blocks it

Meta AI can access your site.

  • meta-webindexerallowed by this fileMatched Allow: / under User-agent: *
  • meta-externalfetcherallowed by this fileMatched Allow: / under User-agent: *

Google AI (AI Overviews / Gemini)

No rule blocks it

No Google-Extended restriction declared. AI Overviews appearance is governed by normal Google indexing plus snippet controls (nosnippet / max-snippet / noindex).

  • Google-Extendedallowed by this fileMatched Allow: / under User-agent: *
  • Googlebotallowed by this fileMatched Allow: / under User-agent: *

AI model training

Deliberate

Content excluded from the blocked operators' future model training. This may be intentional, and it does not affect AI search visibility by itself.

  • GPTBotblocked in robots.txtMatched Disallow: / under User-agent: gptbotExcluded from future OpenAI model training. Does NOT affect ChatGPT search citations (that's OAI-SearchBot).
  • ClaudeBotblocked in robots.txtMatched Disallow: / under User-agent: claudebotExcluded from future Anthropic model training. Does not affect Claude's search visibility (that's Claude-SearchBot).
  • Applebot-Extendedallowed by this fileMatched Allow: / under User-agent: *
  • meta-externalagentallowed by this fileMatched Allow: / under User-agent: *
  • Amazonbotallowed by this fileMatched Allow: / under User-agent: *
  • Bytespiderallowed by this fileMatched Allow: / under User-agent: *
  • CCBotallowed by this fileMatched Allow: / under User-agent: *
  • MistralAI-Trainingallowed by this fileMatched Allow: / under User-agent: *

Other AI assistants

No rule blocks it

Accessible to these assistants.

  • OAI-AdsBotallowed by this fileMatched Allow: / under User-agent: *
  • Applebotallowed by this fileMatched Allow: / under User-agent: *
  • meta-externaladsallowed by this fileMatched Allow: / under User-agent: *
  • facebookexternalhitallowed by this fileMatched Allow: / under User-agent: *
  • MistralAI-Indexallowed by this fileMatched Allow: / under User-agent: *
  • MistralAI-Userallowed by this fileMatched Allow: / under User-agent: *
  • DuckAssistBotallowed by this fileMatched Allow: / under User-agent: *
  • Bingbotallowed by this fileMatched Allow: / under User-agent: *

Grok (xAI), no robots.txt control

No rule reaches it

Grok (xAI) publishes no official crawler identity and has been observed fetching with standard browser user agents from datacenter IPs. No robots.txt rule can control it. We tell you this because pretending otherwise would be dishonest.

    This is the half you can read. A robots.txt says what you intend. It cannot show a CDN or bot-management rule turning a crawler away while the file still says it is welcome, and that mismatch is invisible from the front of a site. A scan sends each crawler’s own user agent and compares what comes back.

    We request your robots.txt and homepage. We read your sitemap only when robots.txt hides pages from a crawler and we need to measure how much is affected. No other page.

    Want the file instead of a verdict on one? Build a robots.txt from the same registry.

    Why the answer is not just “allowed” or “disallowed”

    A robots.txt checker can tell you whether a rule matches. The harder part is knowing what each token actually controls. Blocking GPTBot does not remove you from ChatGPT search, while blocking OAI-SearchBot does.

    Each verdict includes the consequence, using the same sourced, dated registry as the scanner. The goal is to tell you not just what the file says, but why it matters.

    A verdict about your homepage is not a verdict about your site

    A homepage-only check can give a false sense of safety. Your file might allow / but block /blog/, leaving the homepage open while your articles are invisible to a crawler.

    Every rule in the group governing a crawler is evaluated here, not only the one that matches /, and an Allow: sitting beside a Disallow: is counted properly, so a rule that carves its own exception back out is not reported as a block. Try the Homepage fine, content not example above.

    What a file cannot tell you

    robots.txt tells us what you intend. It does not tell us whether a CDN, WAF or bot-management rule actually lets the request through. This page only evaluates the file. The full scan checks the request too.

    Two more things live outside the file. Grok publishes no crawler token at all, so no rule you write reaches it. And Copilot is controlled by meta tags rather than by user agent, so a robots.txt-only answer about it is necessarily incomplete.

    Where these answers come from

    The registry covers 26 tokens across 13 operators. Each entry has a source and a verification date, and the registry was last reviewed 19 September 2026. See the crawler reference for the full list.

    Want to change what your file says?

    The robots.txt generator writes a file from the crawlers you choose to allow, shows what each choice costs, and will not let a preset block Googlebot. Paste the result back here to check it.

    ← DeployRadar