How DeployRadar decides what an AI crawler can read

Every finding in a report is the end of the same chain: what your robots.txt says, what our requests received, how the two compare, how sure we can be, and what a block on that crawler changes. This page walks through each link, and through what the chain cannot establish from outside your site.

First: was the scan itself reliable?

Before any crawler verdict, a scan records its own health: whether robots.txt answered, whether an ordinary browser could load the page, and how many crawler requests completed. From those it sets a confidence of high, medium or low.

A robots.txt that answered 503 is the case this exists for. Crawlers treat that as a temporary block on everything, so every policy below it describes the file and not what crawlers are applying right now. The report shows that as a banner above the results, and the API puts health first for the same reason: a verdict under low confidence is provisional, whatever colour it is.

1. Policy: what your robots.txt declares

We read robots.txt the way RFC 9309 says a crawler must, for each of 26 crawler tokens across 13 operators. For each token the result is allow, disallow, or no rule, together with the rule that decided it and whether that rule named the crawler or was a blanket User-agent: *. A file that allows your home page but disallows /blog/ is reported as the partial block it is, not as clear.

Policy is a statement of intent. It says nothing about whether a request is actually let through, which is why it is only the first link.

2. Request: what our requests received

For 5 crawlers, GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and meta-externalagent, we request your home page under that crawler’s user agent. In the same scan we request it as an ordinary browser, and under our own bot name as a control. A refusal is asked once more, so a one-off error is not reported as a block. The scanner page lists every request a scan makes and how to block it.

The other tokens are not requested, for one of three reasons. Some are robots.txt tokens that no crawler sends, such as Google-Extended. Googlebot and Bingbot are never impersonated, because a request claiming to be them trips the fake-bot rules that protect sites from exactly that, and the answer would describe the rule rather than the crawler. The rest sit outside the small set we request, since every request costs the site something. For all of them, the verdict rests on robots.txt alone and says so.

3. Enforcement: how the answers compare

The crawler’s answer is read against the browser’s. Each requested crawler ends in one of these outcomes:

No user-agent block detectedThe crawler’s request reached the page. Never reported as “allowed”: our request comes from our address, not the operator’s published ranges, and an edge that checks addresses may treat the real crawler differently.
Blocked by user agentThe browser got the page and the crawler was refused, and asking once more did not get through either. A rule somewhere matches that user agent.
Possibly blocked at the edgeThe crawler was refused behind a bot-management CDN while robots.txt allows it. Our request is itself an unverified bot, so this may be the CDN turning away impostors rather than the real crawler. The control request narrows it down where it can.
Pay-per-crawlThe site answered 402 with a crawl price, which is a definite signature of a paid crawling arrangement.
InconclusiveA server error, a timeout, a refusal that did not repeat, or a browser that could not load the page either. Neither a block nor access, and reported as neither.
Not testedNo request was sent as this crawler, for one of the reasons above.

4. Evidence: how each finding is established

Every finding carries one of four evidence levels, and the report prints it beside the verdict:

VerifiedWe sent the request and saw the answer ourselves.
DeclaredIt rests on what robots.txt or the operator’s documentation says.
InferredA response pointed at a cause we could not attribute for certain.
UnknownNothing that can be checked from outside the site decides it.

5. Impact: what a block on that crawler changes

A block matters according to what the crawler is for, and that comes from the operator’s own description of it. For each crawler we state whether a block affects AI search results, citations in AI answers, model training, and live fetches for a person asking about your page. A block on GPTBot changes training and nothing else. A block on OAI-SearchBot changes ChatGPT search and citations. Where an operator runs one crawler for several purposes without saying which, the impact is reported as unknown rather than guessed.

We show the consequence. Which crawlers to allow is your decision, and the report does not score it.

JavaScript: what a non-rendering crawler receives

Most AI crawlers do not run JavaScript. We load the page in a real browser and compare its text with the text in the initial HTML response, and report the share that is present without JavaScript. A page whose content only appears after scripts run is, to those crawlers, mostly empty.

What a scan cannot know

How the crawler registry is kept

Every crawler fact above comes from one registry. Each entry is sourced from the operator’s documentation where one exists and carries the date it was last checked against that source. Entries are re-checked every 91 days, and a scheduled job opens an issue when one is due. The registry was last reviewed 19 September 2026. Read it as a page at the crawler reference, or as data at /crawlers.json.

The terms used on this page, and across the site, are defined in the glossary. If a finding in your report does not match what you see, tell us. A real person reads it.