robots.txt generator for AI crawlers
Choose which crawlers you want to allow, then copy a ready-to-use robots.txt. Each choice comes with its consequence, using the same sourced registry as our scanner. 26 tokens across 13 operators, verified 19 September 2026.
What do you want your site to allow?
Who gets in
How your content may be used
This is optional and separate from robots.txt. Content-Usage is an IETF AI Preferences draft, not yet an RFC. It expresses a preference, not an enforcement rule. A category you leave undeclared means no preference, not a refusal.
Your robots.txt8 of 26 blocked
# robots.txt for AI crawlers
# Generated by https://deployradar.dev/robots-txt-generator
# Crawler registry 2026-09-19. Every rule below is sourced.
#
# Paste this alongside your existing rules. Each token gets
# its own group because the most specific match wins: a token
# with no group of its own falls back to your `User-agent: *`
# block and inherits whatever it says.
# ── ChatGPT search ──
# These decide whether ChatGPT can cite you. Not the same as
# training.
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
# ── Claude ──
# Whether Claude can reach your pages when it searches or a user
# asks.
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /
# ── Perplexity ──
# Whether Perplexity answers can surface and link to you.
User-agent: PerplexityBot
Allow: /
User-agent: Perplexity-User
Allow: /
# ── Google ──
# Googlebot is ordinary Search. Google-Extended is Gemini
# training and grounding. Blocking it does NOT remove you from
# AI Overviews.
User-agent: Google-Extended
Disallow: /
User-agent: Googlebot
Allow: /
# ── Other AI assistants ──
# Siri and Apple Intelligence, DuckAssist, Le Chat, Meta AI,
# Alexa.
User-agent: OAI-AdsBot
Allow: /
User-agent: Applebot
Allow: /
User-agent: meta-externalads
Allow: /
User-agent: facebookexternalhit
Allow: /
User-agent: MistralAI-Index
Allow: /
User-agent: MistralAI-User
Allow: /
User-agent: DuckAssistBot
Allow: /
User-agent: Bingbot
Allow: /
# ── AI model training ──
# Blocking these is a policy choice about your content being
# trained on. On its own it does not affect whether AI search
# can cite you.
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: meta-externalagent
Disallow: /
User-agent: Amazonbot
Allow: /
User-agent: Bytespider
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: MistralAI-Training
Disallow: /
# ── Other crawlers ──
User-agent: meta-webindexer
Allow: /
User-agent: meta-externalfetcher
Allow: /
# ── Not reachable from this file ──
# xAI (Grok) publishes no crawler token, so no rule here can
# control it. Said plainly rather than left out.
What this actually does
- Not what it looks likeThis does not remove you from AI OverviewsGoogle-Extended controls Gemini training and grounding. AI Overviews and AI Mode are built on ordinary Google Search crawling, so they are governed by Googlebot plus your snippet controls, meaning nosnippet, max-snippet and noindex, not by this rule.
- Worth knowingGPTBot blocked, ChatGPT search still allowedThis is usually exactly what people mean: your content stays out of OpenAI's training data, and ChatGPT search can still find and cite you. Flagged only so you know it is deliberate rather than half-finished.
- Worth knowingBytespider is widely reported not to obey robots.txtKeep the rule. It states your policy, and it is what a well-behaved fetcher reads. Just do not treat it as enforcement. Blocking at the CDN is the layer that actually stops a crawler that has decided to ignore you.
Three things people get wrong here
These are common mistakes caused by a robots.txt rule meaning something different from what its author intended.
- Blocking GPTBot to disappear from ChatGPT. It does not. GPTBot is the training crawler; OAI-SearchBot is the one that fetches pages to cite in ChatGPT search answers. Block GPTBot and ChatGPT can still find and link you, which is what most people wanted in the first place.
- Blocking Google-Extended to get out of AI Overviews. Also no. Google-Extended governs Gemini training and grounding. AI Overviews and AI Mode are built on ordinary Google Search crawling, so they answer to Googlebot and to your snippet controls, meaning
nosnippet,max-snippetandnoindex, not to this rule at all. - Catching Googlebot in the net. Someone sets out to keep AI off the site, ticks everything that looks like a robot, and takes the site out of Google Search. This generator will not let a preset do that, and it says so loudly if you do it by hand.
Where to put the file
robots.txt lives at the root of each origin, at https://yoursite.com/robots.txt. Not in a subdirectory, and a file on www.yoursite.com does not govern yoursite.com; they are separate origins and each needs its own.
If you already have one, paste these groups alongside your existing rules rather than replacing the file. Keep one group per user-agent: under the Robots Exclusion Protocol the most specific matching group wins outright, and a token with no group of its own falls back to your User-agent: * block and inherits whatever that says. That fallback is why the generator writes an explicit Allow: / even for crawlers that would have been allowed anyway. Without it, a wildcard Disallow you forgot about quietly overrides the intent.
A robots.txt is a request, not a wall
This page describes what a well-behaved crawler should do with your file. Two important things remain outside robots.txt:
- Whether the crawler obeys it. Some are widely reported not to. The rule is still worth writing, because it states your policy and it is what an honest fetcher reads, but it is not enforcement. Blocking at the CDN is the layer that actually stops someone who has decided to ignore you.
- Whether something else is already blocking them. This is the failure we built the scanner for. A CDN or bot-management rule can turn a crawler away while your robots.txt still says it is welcome, and nothing about the front of your site shows it. Sites lose AI visibility this way without anyone changing a line of the file.
Check what your site does today
Before changing the file, check what it currently tells all 26 tokens. Paste it into the checker, with no domain or upload required.
The one thing neither page can read off a file is whether your CDN agrees with it. That takes a scan: about ten seconds, and no sign-up.
We request your robots.txt and homepage. We read your sitemap only when robots.txt hides pages from a crawler and we need to measure how much is affected. No other page.