How DeployRadar decides what an AI crawler can read
Every finding in a report is the end of the same chain: what your robots.txt says, what our requests received, how the two compare, how sure we can be, and what a block on that crawler changes. This page walks through each link, and through what the chain cannot establish from outside your site.
First: was the scan itself reliable?
Before any crawler verdict, a scan records its own health: whether robots.txt answered, whether an ordinary browser could load the page, and how many crawler requests completed. From those it sets a confidence of high, medium or low.
A robots.txt that answered 503 is the case this exists for. Crawlers treat that as a temporary block on everything, so every policy below it describes the file and not what crawlers are applying right now. The report shows that as a banner above the results, and the API puts health first for the same reason: a verdict under low confidence is provisional, whatever colour it is.
1. Policy: what your robots.txt declares
We read robots.txt the way RFC 9309 says a crawler must, for each of 26 crawler tokens across 13 operators. For each token the result is allow, disallow, or no rule, together with the rule that decided it and whether that rule named the crawler or was a blanket User-agent: *. A file that allows your home page but disallows /blog/ is reported as the partial block it is, not as clear.
Policy is a statement of intent. It says nothing about whether a request is actually let through, which is why it is only the first link.
2. Request: what our requests received
For 5 crawlers, GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and meta-externalagent, we request your home page under that crawler’s user agent. In the same scan we request it as an ordinary browser, and under our own bot name as a control. A refusal is asked once more, so a one-off error is not reported as a block. The scanner page lists every request a scan makes and how to block it.
The other tokens are not requested, for one of three reasons. Some are robots.txt tokens that no crawler sends, such as Google-Extended. Googlebot and Bingbot are never impersonated, because a request claiming to be them trips the fake-bot rules that protect sites from exactly that, and the answer would describe the rule rather than the crawler. The rest sit outside the small set we request, since every request costs the site something. For all of them, the verdict rests on robots.txt alone and says so.
3. Enforcement: how the answers compare
The crawler’s answer is read against the browser’s. Each requested crawler ends in one of these outcomes:
| No user-agent block detected | The crawler’s request reached the page. Never reported as “allowed”: our request comes from our address, not the operator’s published ranges, and an edge that checks addresses may treat the real crawler differently. |
|---|---|
| Blocked by user agent | The browser got the page and the crawler was refused, and asking once more did not get through either. A rule somewhere matches that user agent. |
| Possibly blocked at the edge | The crawler was refused behind a bot-management CDN while robots.txt allows it. Our request is itself an unverified bot, so this may be the CDN turning away impostors rather than the real crawler. The control request narrows it down where it can. |
| Pay-per-crawl | The site answered 402 with a crawl price, which is a definite signature of a paid crawling arrangement. |
| Inconclusive | A server error, a timeout, a refusal that did not repeat, or a browser that could not load the page either. Neither a block nor access, and reported as neither. |
| Not tested | No request was sent as this crawler, for one of the reasons above. |
4. Evidence: how each finding is established
Every finding carries one of four evidence levels, and the report prints it beside the verdict:
| Verified | We sent the request and saw the answer ourselves. |
|---|---|
| Declared | It rests on what robots.txt or the operator’s documentation says. |
| Inferred | A response pointed at a cause we could not attribute for certain. |
| Unknown | Nothing that can be checked from outside the site decides it. |
5. Impact: what a block on that crawler changes
A block matters according to what the crawler is for, and that comes from the operator’s own description of it. For each crawler we state whether a block affects AI search results, citations in AI answers, model training, and live fetches for a person asking about your page. A block on GPTBot changes training and nothing else. A block on OAI-SearchBot changes ChatGPT search and citations. Where an operator runs one crawler for several purposes without saying which, the impact is reported as unknown rather than guessed.
We show the consequence. Which crawlers to allow is your decision, and the report does not score it.
JavaScript: what a non-rendering crawler receives
Most AI crawlers do not run JavaScript. We load the page in a real browser and compare its text with the text in the initial HTML response, and report the share that is present without JavaScript. A page whose content only appears after scripts run is, to those crawlers, mostly empty.
What a scan cannot know
- How the real crawler is treated when it arrives from its operator’s own addresses. We can detect a rule keyed on the user agent. A rule keyed on the address, we cannot.
- Whether a crawler has visited, or what it did with the page. Your server logs know; we do not see them.
- Pages other than the home page, beyond what robots.txt says about their paths. We read your sitemap to measure how much a rule excludes, and never request the URLs in it.
- Whether an operator follows its own documentation. The registry records what each one says about obeying robots.txt, and marks the ones known not to.
How the crawler registry is kept
Every crawler fact above comes from one registry. Each entry is sourced from the operator’s documentation where one exists and carries the date it was last checked against that source. Entries are re-checked every 91 days, and a scheduled job opens an issue when one is due. The registry was last reviewed 19 September 2026. Read it as a page at the crawler reference, or as data at /crawlers.json.
The terms used on this page, and across the site, are defined in the glossary. If a finding in your report does not match what you see, tell us. A real person reads it.