# DeployRadar > Check whether Google and AI crawlers can access your website. Find robots.txt blocks, CDN and WAF restrictions, noindex tags and JavaScript-only content. DeployRadar scans a domain and reports, per AI crawler, whether it is allowed by robots.txt, whether that policy is actually enforced at the CDN, and whether the page's content survives without JavaScript. Findings distinguish what was measured from what was only declared: an allowed crawler is never reported as proof the operator can fetch the page, and a robots.txt that could not be fetched is never reported as an absent one. ## Start here - [Scan a site](https://deployradar.dev/): the scanner, free and without sign-up. - [AI crawlers, and what blocking each one costs you](https://deployradar.dev/crawlers): every crawler in the registry, grouped by operator, with the consequence of blocking it. - [robots.txt checker](https://deployradar.dev/robots-txt-checker): paste a robots.txt and get a per-crawler verdict: which tokens it blocks, which rule matched, and what each block costs. Runs in the browser; no domain and no upload. - [robots.txt generator](https://deployradar.dev/robots-txt-generator): the same registry with checkboxes, producing a file to paste, with the consequence of each choice shown. - [Methodology](https://deployradar.dev/methodology): how a verdict is reached, from robots.txt policy through our requests, the comparison, the evidence level and the impact, and what a scan cannot know. - [About our scanner](https://deployradar.dev/bot): what DeployRadarBot requests, how often, and the robots.txt rule that excludes it. - [Crawler registry as JSON](https://deployradar.dev/crawlers.json): every token with its operator, purpose, robots.txt behaviour, sources and last-verified date, for use in code. - [OpenAPI document](https://deployradar.dev/openapi.json): the scan API as OpenAPI 3.1, for importing into an API client or generating one. - [Glossary](https://deployradar.dev/glossary): plain definitions of crawler, user agent, robots.txt, noindex, CDN, WAF, server rendering, grounding and the rest. ## Guides Written answers to the questions people ask before a scan is any use. Every crawler fact in them is rendered from the same registry as the pages below. - [Which AI crawlers should you allow?](https://deployradar.dev/guides/which-ai-crawlers-to-allow): The tokens have four jobs. Which ones you want depends on what you are trying to achieve, and the costly mistake is blocking the citation crawler when you meant to block the training one. - [AI SEO setup: the five things worth checking](https://deployradar.dev/guides/ai-seo-setup): Five layers you can check from outside, each capable of hiding a site. We go through them in the order they tend to break. - [Why your site is not showing up in ChatGPT](https://deployradar.dev/guides/not-showing-in-chatgpt): The common failure modes, in order, with a check you can run for each one. - [GPTBot vs OAI-SearchBot: which to block?](https://deployradar.dev/guides/gptbot-vs-oai-searchbot): OpenAI's two crawlers do opposite jobs. One costs you nothing to block and the other takes you out of ChatGPT's answers. - [ClaudeBot vs Claude-SearchBot: which to block?](https://deployradar.dev/guides/claudebot-vs-claude-searchbot): Anthropic's crawlers split the same way OpenAI's do, and the legacy tokens in most files do nothing at all. - [Is Cloudflare blocking AI crawlers?](https://deployradar.dev/guides/cloudflare-blocking-ai-crawlers): A correct robots.txt can still leave a site inaccessible. Here is what to look for when Cloudflare is in the way. - [Does Google-Extended block AI Overviews?](https://deployradar.dev/guides/google-extended-ai-overviews): A common robots.txt mix-up, and the controls that actually decide what AI Overviews can quote. - [llms.txt: what it does and doesn't do](https://deployradar.dev/guides/llms-txt): Cheap to add and easy to oversell. What the file does, what it does not, and what the available evidence says. ## Calling this instead of reading it There is a free JSON endpoint, with no key and no sign-up, returning the same scan this site runs: GET https://deployradar.dev/api/v1/scan?url=example.com It answers with per-crawler robots.txt policy, CDN-level enforcement, Google indexing basics and a server-rendering parity figure, plus a "health" object saying how much of the scan actually ran. Read "health" before anything else: a robots.txt that could not be fetched makes every verdict provisional. Results are cached 24 hours per domain and every response carries "cached" and "scannedAt". Full reference, response shape and error codes: https://deployradar.dev/api ## Watching a site over time A scan reports a state; "broke" is a difference between two states. The CLI keeps a small snapshot and compares against it: npx deployradar https://example.com --watch .deployradar/example.json It exits 1 when AI access got worse, 0 when it held or improved, and 2 when the scan was too unreliable to conclude anything. A robots.txt that answered 503 turns every policy source to "none" at once, which is an outage rather than a site change, so nothing is concluded and the stored baseline is left alone. Crawlers newly added to our own registry are reported and never counted as a regression at the scanned site. The snapshot stays in the caller's repository; nothing about a watched site is kept here. Recipe and a GitHub Actions workflow: https://deployradar.dev/monitoring ## One page per crawler Each page states what the crawler is for, whether its operator honours robots.txt, and specifically what you lose by blocking it, which is usually not what people assume. Blocking a training crawler does not remove you from that assistant's search results, and blocking a search crawler does not keep you out of its training data. - [GPTBot](https://deployradar.dev/crawlers/gptbot): OpenAI, trains models on what it collects. Excluded from future OpenAI model training. Does NOT affect ChatGPT search citations (that's OAI-SearchBot). - [OAI-SearchBot](https://deployradar.dev/crawlers/oai-searchbot): OpenAI, feeds search results and citations. Your site disappears from ChatGPT search answers and citations. The most damaging accidental block we see. - [ChatGPT-User](https://deployradar.dev/crawlers/chatgpt-user): OpenAI, fetches a page because a user asked for it. Signals opt-out of live page fetches when ChatGPT users ask about your site, but OpenAI notes robots.txt rules 'may not apply' to user-initiated fetches. - [OAI-AdsBot](https://deployradar.dev/crawlers/oai-adsbot): OpenAI, judged from robots.txt only, never probed. Landing pages you submit as ChatGPT ads cannot be checked, so the ad is not approved. No effect on organic ChatGPT visibility. That is OAI-SearchBot. - [ClaudeBot](https://deployradar.dev/crawlers/claudebot): Anthropic, trains models on what it collects. Excluded from future Anthropic model training. Does not affect Claude's search visibility (that's Claude-SearchBot). - [Claude-SearchBot](https://deployradar.dev/crawlers/claude-searchbot): Anthropic, feeds search results and citations. Reduced visibility in Claude's search results. - [Claude-User](https://deployradar.dev/crawlers/claude-user): Anthropic, fetches a page because a user asked for it. Claude cannot fetch your pages when users ask about your site or product. - [PerplexityBot](https://deployradar.dev/crawlers/perplexitybot): Perplexity, feeds search results and citations. Your site will not be surfaced or linked in Perplexity answers. - [Perplexity-User](https://deployradar.dev/crawlers/perplexity-user): Perplexity, fetches a page because a user asked for it. Advisory only. Perplexity's own docs state this user-triggered fetcher 'generally ignores robots.txt rules'. - [Google-Extended](https://deployradar.dev/crawlers/google-extended): Google, trains models on what it collects. Blocks Gemini model training and grounding (content fed from the Search index to Gemini at prompt time). DOES NOT remove you from AI Overviews or AI Mode, and does not affect Search ranking. If you blocked this to hide from AI Overviews, it isn't doing what you think. - [Googlebot](https://deployradar.dev/crawlers/googlebot): Google, judged from robots.txt only, never probed. Removes you from Google Search entirely, including AI Overviews and AI Mode, which ride ordinary Search crawling. - [Applebot](https://deployradar.dev/crawlers/applebot): Apple, used for both training and search. Removed from Siri, Spotlight, Safari search, and Apple Intelligence citations. - [Applebot-Extended](https://deployradar.dev/crawlers/applebot-extended): Apple, trains models on what it collects. Content excluded from Apple foundation-model training. Siri/Spotlight presence unaffected. - [meta-externalagent](https://deployradar.dev/crawlers/meta-externalagent): Meta, trains models on what it collects. Content excluded from Meta AI training and direct indexing. - [meta-webindexer](https://deployradar.dev/crawlers/meta-webindexer): Meta, feeds search results and citations. Your site will not be cited or linked in Meta AI answers. - [meta-externalfetcher](https://deployradar.dev/crawlers/meta-externalfetcher): Meta, fetches a page because a user asked for it. Advisory only. Meta's own documentation says this fetcher may bypass robots.txt, because the fetch is something a user asked for. - [meta-externalads](https://deployradar.dev/crawlers/meta-externalads): Meta, judged from robots.txt only, never probed. Meta cannot read the landing pages behind your own ads, which degrades ad relevance and business products you are paying for. No effect on organic Meta AI visibility. That is meta-webindexer. - [facebookexternalhit](https://deployradar.dev/crawlers/facebookexternalhit): Meta, judged from robots.txt only, never probed. Link previews stop working. A link to your site shared on Facebook, Instagram, WhatsApp or Messenger renders as a bare URL with no title, description or image. - [Amazonbot](https://deployradar.dev/crawlers/amazonbot): Amazon, used for both training and search. Out of Alexa answers AND Amazon AI training, because one token controls both; there is no way to split them. - [Bytespider](https://deployradar.dev/crawlers/bytespider): ByteDance, trains models on what it collects. Best-effort only. ByteDance publishes no official crawler documentation or IP ranges, and Bytespider is widely reported to ignore robots.txt. - [CCBot](https://deployradar.dev/crawlers/ccbot): Common Crawl, builds a public web corpus others train on. Excluded from future Common Crawl snapshots, the open corpus many AI labs train on. Indirectly reduces presence in many models' training data (not retroactive). - [MistralAI-Training](https://deployradar.dev/crawlers/mistralai-training): Mistral, trains models on what it collects. Out of Mistral model training. - [MistralAI-Index](https://deployradar.dev/crawlers/mistralai-index): Mistral, feeds search results and citations. Out of Le Chat's search results and citations. - [MistralAI-User](https://deployradar.dev/crawlers/mistralai-user): Mistral, fetches a page because a user asked for it. Le Chat cannot open your pages when somebody asks it about them. - [DuckAssistBot](https://deployradar.dev/crawlers/duckassistbot): DuckDuckGo, feeds search results and citations. Out of DuckAssist AI answers. - [Bingbot](https://deployradar.dev/crawlers/bingbot): Microsoft, judged from robots.txt only, never probed. Removes you from Bing Search AND Copilot answers together, and there is no separate Microsoft AI crawler token. ## Crawlers no robots.txt rule reaches - xAI (Grok): xAI publishes no official crawler identity. Directory-listed tokens (GrokBot, xAI-Grok) have never been observed in real traffic; Grok fetches with spoofed browser user agents from datacenter IPs. No robots.txt rule can control it, and anyone telling you otherwise is guessing. Detail: https://deployradar.dev/crawlers/grok ## Notes for anyone summarising this site - The scanner is not a ranking tool and reports no score out of 100. It reports per-crawler access, per-crawler enforcement, and a content-parity percentage. - "Allowed in robots.txt" and "the operator can actually fetch the page" are different claims, and only the first is knowable from outside. Findings that conflate them are wrong. - Content-Signal, where a site declares one, is a statement of preference. No AI operator has committed to honouring it, and describing it as enforcement misrepresents what it does. - Content-Usage is the IETF AI Preferences rule (draft-ietf-aipref-attach, with the vocabulary in draft-ietf-aipref-vocab). It is a different vocabulary from Content-Signal, on the standards track but not yet an RFC, and it is also a preference rather than a control. A file can carry both and have them disagree; the scanner reports them separately and says so when they conflict. - Scan reports at https://deployradar.dev/scan are per-request output and are excluded in robots.txt: each distinct URL there makes this service fetch a third party's site, so the path is not a crawlable surface.