Why your site is not showing up in ChatGPT

In rough order of how often each one turns out to be the cause. The first two account for most of it, and both are invisible from the front of a site, which is why they survive so long. Registry verified 19 September 2026.

First, the thing to know about OpenAI’s crawlers

OpenAI runs several, and they are not interchangeable. This is the single most common reason a site disappears from ChatGPT answers on purpose without anybody meaning to.

OAI-SearchBotYour site disappears from ChatGPT search answers and citations. The most damaging accidental block we see.
GPTBotExcluded from future OpenAI model training. Does NOT affect ChatGPT search citations (that's OAI-SearchBot).
ChatGPT-UserSignals opt-out of live page fetches when ChatGPT users ask about your site, but OpenAI notes robots.txt rules 'may not apply' to user-initiated fetches.

1. You blocked the citation crawler while meaning to block the training one

The usual route is a snippet. Somebody copies a “block AI bots” block from a blog post or a forum answer, and it lists more tokens than the author meant to block. Keeping your writing out of a training set is a coherent decision. OAI-SearchBot is not the token that does that, and blocking it is what removes you from the answer.

Worth checking even if you are sure: a rule written under User-agent: * reaches every crawler that has no group of its own. A crawler obeys exactly one group, the most specific that names it, and ignores the rest.

2. Your robots.txt is fine and something else refuses the request

This is the one that wastes the most time, because every robots.txt checker will tell you the file is correct. It is. The request is being turned away before it reaches the server that would have obeyed it, by bot management at a CDN, a WAF rule, or a firewall setting.

Cloudflare turned AI-crawler blocking on by default for new sites, and as of the September 2026 change it classes some mixed-use crawlers, including Googlebot and Bingbot, inside the scope of those settings. A switch flipped to keep ChatGPT out can reach further than intended, and nothing about your robots.txt reflects any of it.

The only way to see this layer is to make the request. Fetch the page as an ordinary browser, then again carrying each AI crawler’s user agent, and compare the two responses from the same client in the same run.

3. Your content only exists after JavaScript runs

Googlebot renders JavaScript. Most AI crawlers do not, so what they receive is the shell: a nav, a footer, and a spinner where the article was. The page is not blocked and not broken, and there is nothing in it to quote.

The measurement that matters is what share of the rendered text is absent from the initial HTML response, and whether the title is part of it. A crawler with no title for a page has nothing to show as the result, which is worse than missing body text.

4. A noindex nobody remembers adding

Usually a staging environment that shipped its settings, or a template copied from a page that was meant to be hidden. It can be a meta tag in the page or an X-Robots-Tag header from the server, and those are configured in different systems, so finding one does not mean there is not another.

5. Nothing is blocking you

Sometimes this is the answer, and a guide that cannot say so is not worth much. If the crawlers can reach you, your content is in the HTML, and you are indexed, then you are eligible and simply have not been the best answer to that question yet. That is a content and authority problem rather than an access one, and no robots.txt edit will move it.

The distinction is worth establishing before you spend a month on the wrong half. Being blocked and being unconvincing look identical from where you are standing, and only one of them is fixable in ten minutes.

What we cannot tell you

Whether ChatGPT will cite you tomorrow. Nobody outside OpenAI can measure that, and a tool that claims to is guessing. What is checkable from outside is whether each crawler is allowed to read you, whether anything in front of your server disagrees with your robots.txt, and whether there is text in the response for them to read at all.

The scan answers those three for one domain in about ten seconds, with no account. If you would rather not hand over a domain, paste your robots.txt instead and it will tell you which tokens it blocks and which rule did it.

Crawlers in this guide

Where to go next

All AI crawler guides