Why your site is not showing up in ChatGPT
In rough order of how often each one turns out to be the cause. The first two account for most of it, and both are invisible from the front of a site, which is why they survive so long. Registry verified 19 September 2026.
First, the thing to know about OpenAI’s crawlers
OpenAI runs several, and they are not interchangeable. This is the single most common reason a site disappears from ChatGPT answers on purpose without anybody meaning to.
| OAI-SearchBot | Your site disappears from ChatGPT search answers and citations. The most damaging accidental block we see. |
|---|---|
| GPTBot | Excluded from future OpenAI model training. Does NOT affect ChatGPT search citations (that's OAI-SearchBot). |
| ChatGPT-User | Signals opt-out of live page fetches when ChatGPT users ask about your site, but OpenAI notes robots.txt rules 'may not apply' to user-initiated fetches. |
1. You blocked the citation crawler while meaning to block the training one
The usual route is a snippet. Somebody copies a “block AI bots” block from a blog post or a forum answer, and it lists more tokens than the author meant to block. Keeping your writing out of a training set is a coherent decision. OAI-SearchBot is not the token that does that, and blocking it is what removes you from the answer.
Worth checking even if you are sure: a rule written under User-agent: * reaches every crawler that has no group of its own. A crawler obeys exactly one group, the most specific that names it, and ignores the rest.
2. Your robots.txt is fine and something else refuses the request
This is the one that wastes the most time, because every robots.txt checker will tell you the file is correct. It is. The request is being turned away before it reaches the server that would have obeyed it, by bot management at a CDN, a WAF rule, or a firewall setting.
Cloudflare turned AI-crawler blocking on by default for new sites, and as of the September 2026 change it classes some mixed-use crawlers, including Googlebot and Bingbot, inside the scope of those settings. A switch flipped to keep ChatGPT out can reach further than intended, and nothing about your robots.txt reflects any of it.
The only way to see this layer is to make the request. Fetch the page as an ordinary browser, then again carrying each AI crawler’s user agent, and compare the two responses from the same client in the same run.
3. Your content only exists after JavaScript runs
Googlebot renders JavaScript. Most AI crawlers do not, so what they receive is the shell: a nav, a footer, and a spinner where the article was. The page is not blocked and not broken, and there is nothing in it to quote.
The measurement that matters is what share of the rendered text is absent from the initial HTML response, and whether the title is part of it. A crawler with no title for a page has nothing to show as the result, which is worse than missing body text.
4. A noindex nobody remembers adding
Usually a staging environment that shipped its settings, or a template copied from a page that was meant to be hidden. It can be a meta tag in the page or an X-Robots-Tag header from the server, and those are configured in different systems, so finding one does not mean there is not another.
5. Nothing is blocking you
Sometimes this is the answer, and a guide that cannot say so is not worth much. If the crawlers can reach you, your content is in the HTML, and you are indexed, then you are eligible and simply have not been the best answer to that question yet. That is a content and authority problem rather than an access one, and no robots.txt edit will move it.
The distinction is worth establishing before you spend a month on the wrong half. Being blocked and being unconvincing look identical from where you are standing, and only one of them is fixable in ten minutes.
What we cannot tell you
Whether ChatGPT will cite you tomorrow. Nobody outside OpenAI can measure that, and a tool that claims to is guessing. What is checkable from outside is whether each crawler is allowed to read you, whether anything in front of your server disagrees with your robots.txt, and whether there is text in the response for them to read at all.
The scan answers those three for one domain in about ten seconds, with no account. If you would rather not hand over a domain, paste your robots.txt instead and it will tell you which tokens it blocks and which rule did it.
Crawlers in this guide
- OAI-SearchBot: Search & citations, run by OpenAI
- GPTBot: Model training, run by OpenAI
- ChatGPT-User: Fetches when a user asks, run by OpenAI
Where to go next
- GPTBot vs OAI-SearchBot: which to block?
OpenAI's two crawlers do opposite jobs. One costs you nothing to block and the other takes you out of ChatGPT's answers. - Is Cloudflare blocking AI crawlers?
A correct robots.txt can still leave a site inaccessible. Here is what to look for when Cloudflare is in the way. - AI SEO setup: the five things worth checking
Five layers you can check from outside, each capable of hiding a site. We go through them in the order they tend to break.