AI SEO setup: the five things worth checking
A lot of SEO advice depends on data only the site owner or search engine can see. These five checks are different: you can verify them from outside, and each can quietly remove a page from search or AI answers.
Why these five
They have one thing in common: when they break, the site still looks fine to a person. There is usually no warning, so the problem can sit unnoticed for weeks.
The order matters. There is little point tuning a description if a crawler cannot reach the page in the first place.
01. robots.txt says what you think it says
What breaks it. A copied snippet, a CMS default, or a rule meant for one crawler placed under User-agent: *. A crawler uses the most specific group that names it; otherwise it falls back to the wildcard group.
Why you would not notice. Nothing warns you. The file is syntactically valid either way, and the crawler simply stops arriving.
How to check. Read the file as each crawler reads it rather than as English. Which token is blocked, and which rule did it.
02. Nothing turns the crawler away at the edge
What breaks it. A CDN, WAF or firewall rule that treats crawler traffic differently. Cloudflare has switched AI-crawler blocking on by default for new sites, and its controls can now affect mixed-use crawlers that also do ordinary search.
Why you would not notice. robots.txt cannot see this layer. The file can look completely open while the request is stopped at the edge.
How to check. Fetch the page as an ordinary browser and again as each AI crawler, from the same client in the same run, then compare. A difference between those two responses is the whole finding.
03. No stray noindex
What breaks it. A meta tag or X-Robots-Tag left behind by staging, a plugin or a hidden-page template.
Why you would not notice. The page looks normal to visitors, but search engines can still be told not to index it.
How to check. Look for the directive in both places, because they are set in different systems: the meta tag lives in your page template, the header in your server or CDN config.
04. The content is in the HTML, not only in JavaScript
What breaks it. Client-side rendering. Googlebot can render JavaScript, but many AI crawlers do not, so they may receive only the page shell.
Why you would not notice. Every human sees the finished page. The difference only appears for clients that do not run JavaScript.
How to check. Compare the server-rendered HTML against a real browser render and measure the difference. The number worth knowing is what share of the rendered text is missing from the initial response, and whether the title is among it.
05. The basics Google still reads
What breaks it. A missing title, a wrong canonical, a missing or weak description, or invalid structured data.
Why you would not notice. The page still loads. These issues change how search engines understand or display it.
How to check. A canonical has to be absolute and name the page it sits on. Invalid structured data is ignored by every consumer, which makes it a plain syntax error rather than a judgement call.
Which crawlers this is for
All of it applies to Google. Layers 1, 2 and 4 are the ones that decide whether an AI assistant can read you, and they are answered by different tokens for different jobs. Blocking GPTBot keeps you out of a training set and costs you nothing in search, while blocking OAI-SearchBot takes you out of the answer. Which to allow is its own guide.
What this guide will not tell you
This guide will not tell you how to rank. Nobody outside Google can predict that reliably. What we can check is whether Google and AI crawlers can read your site, which is a much more concrete question.
Some things cannot be verified from outside, such as how a site treats an operator’s verified IP ranges or whether a declared User-Agent is genuinely from that operator. When we cannot know, we say so.
Running the scan checks all five layers against one domain in about ten seconds, against a registry of 26 tokens verified 19 September 2026. It asks for no account.
Where to go next
- Which AI crawlers should you allow?
The tokens have four jobs. Which ones you want depends on what you are trying to achieve, and the costly mistake is blocking the citation crawler when you meant to block the training one. - Is Cloudflare blocking AI crawlers?
A correct robots.txt can still leave a site inaccessible. Here is what to look for when Cloudflare is in the way. - llms.txt: what it does and doesn't do
Cheap to add and easy to oversell. What the file does, what it does not, and what the available evidence says.