AI crawler access API

The same scan, returned as JSON: robots.txt policy and CDN enforcement for each AI crawler, Google indexing basics, and how much of the page survives without JavaScript. No key, sign-up or account. Use it in scripts, CI or a scheduled check when a report page is not enough.

One endpoint

GET /api/v1/scan?url=example.com

url is the only parameter. Pass a hostname or full URL. We scan the origin, so the scheme, path and query are ignored.

curl 'https://deployradar.dev/api/v1/scan?url=example.com'

CORS is open, so it can be called from a browser. There is no authentication and no cookie-based session.

To import it into Postman, Insomnia or a client generator, use the OpenAPI 3.1 document at /openapi.json. It points at the same JSON Schema this page describes below.

What comes back

{
  "apiVersion": 1,
  "target": "https://example.com",
  "scannedAt": "2026-09-15T09:12:44.102Z",
  "cached": false,
  "registryVersion": "2026-09-15",
  "health": {
    "robots": "ok",
    "baseline": "ok",
    "probes": "complete",
    "confidence": "high",
    "notes": []
  },
  "groups": [
    {
      "groupId": "chatgpt_search",
      "label": "ChatGPT Search",
      "status": "blocked",
      "summary": "Your site will not appear in or be cited by ChatGPT search answers.",
      "findings": [
        {
          "token": "OAI-SearchBot",
          "operator": "OpenAI",
          "policy": "deny",
          "policySource": "token",
          "matchedRule": {
            "type": "disallow",
            "path": "/",
            "agent": "oai-searchbot"
          },
          "enforcement": "ua_block_confirmed",
          "confidence": "high",
          "consequence": "Search/citation crawler. Blocking this while allowing GPTBot is almost always a mistake.",
          "notes": [],
          "evidence": {
            "verdict": "blocked",
            "level": "verified",
            "robots": { "read": true, "allowed": false },
            "request": { "tested": true, "status": 403, "baselineStatus": 200 }
          }
        }
      ]
    }
  ],
  "findings": [],
  "site": [],
  "robotsStatus": 200,
  "google": [],
  "parity": {
    "status": "ok",
    "textParity": 0.98,
    "summary": "98% of the rendered text is present in the initial HTML."
  },
  "parityReason": null,
  "docs": "https://deployradar.dev/api"
}

Trimmed above: findings, site and google are arrays with real entries in a live response. The fields that matter:

healthHow much of the scan actually ran, and how far to trust the rest. Read this first: a robots.txt that answered 503 makes every verdict below provisional, and a caller that skips it will publish a green light for a site nobody could reach.
groupsOne entry per AI system rather than per token, with a status of ok, info, warning or blocked and a sentence saying what it means. This is the layer to build an alert on.
findingsThe per-token detail behind those groups: the policy, which rule decided it, whether the rule named the crawler or was a blanket rule, and what a block costs. Each finding also carries an evidence object: a plain verdict (accessible, restricted, blocked, possibly_blocked or unknown), how it was established (verified by a request we saw, declared by robots.txt, inferred, or unknown), and the HTTP status from the crawler's request next to the status from an ordinary browser. When no request was sent, untestedReason explains why. evidence.reason gives the verdict as a code to switch on (primary, such as ROBOTS_DISALLOW, UA_EDGE_BLOCK or PROBE_TIMEOUT) with the observations behind it (supporting, such as CRAWLER_HTTP_403 or CDN_CLOUDFLARE). impact says what a block on that crawler changes: search, citations, training and agents, each true, false or unknown.
evidenceThe raw observations: every request the scan sent, the user agent it carried, its status, any redirects, and a short list of diagnostic headers. No page content. For callers that want to draw their own conclusions.
planWhat the scan set out to request once robots.txt was read, and maxFetches, the most it was allowed to send. The limit is enforced, so this is a ceiling on what the scan cost the site.
provenanceThe scanner, classifier and registry versions the result was made under. If a verdict changes while classifierVersion also changed, the change may be in how we read the site rather than in the site.
siteFacts about the robots.txt itself: whether a Content-Signal or a Content-Usage (IETF AIPREF) preference is declared, whether the two contradict each other, and whether the file was written by Cloudflare rather than by the site's owner.
googleWhether Google can index the page at all: Googlebot rules, noindex directives, title, canonical. Judged against the rendered HTML, because Googlebot runs JavaScript.
parityHow much of the rendered text survives in the initial HTML response. null when the render could not run, with parityReason saying why. render says what the browser spent (requests, bytes, time) and whether it hit its cap.

registryVersion tells you which crawler registry the result used. It is included because crawler policies change. The current registry is 2026-09-19, covering 26 tokens across 13 operators.

Caching, and why you want it

Results are cached for 24 hours per domain. Cached responses are immediate, do not consume your rate limit, and do not send another request to the site being scanned. This endpoint cannot force a fresh scan.

Every response carries cached and scannedAt, so you can see how old an answer is rather than guessing. Responses are also sent with Cache-Control: public, max-age=600.

Limits

New scans5 per hour, per address. Cache hits are free and uncounted, so polling one domain costs nothing after the first call.
Shared ceilingA process-wide hourly cap on brand-new scans exists as well. Hitting it returns at_capacity, which is about us rather than about you.
BudgetSeparate from the one the website uses, so calling the API never leaves you rate limited on deployradar.dev, or the other way round.

Errors

Every failure is { "error": { "code", "message" } }. Match on code and show message. The sentences are written for people and will be reworded, so anything matching on prose is something we would break by improving an error.

CodeHTTPMeans
missing_url400No ?url= was given.
invalid_target400The URL could not be resolved, or resolves to a private or loopback address. The message says which.
rate_limited429You have used this hour's scans. Cached results still answer immediately and do not count.
at_capacity503Our own hourly ceiling on new scans, not yours. Retry-After says when to come back.
scan_failed502The scan started and did not finish, usually a timeout on the way to the target.

Versioning

The version is in the path because the response shape is now somebody else’s dependency. Fields will be added without a version bump; nothing will be removed or change meaning under /v1. apiVersion is on every response so a client can assert what it is reading.

The response shape is published as a JSON Schema at /schemas/scan-v1.json and linked from every response as schema. Allow unknown properties when you validate, since new fields arrive without a version bump.

What it will not do

It will not scan a private or loopback address. Every target is resolved and checked before a request goes anywhere, so this cannot be used to map an internal network. It will not force a fresh scan past the cache. It will not fetch more than the handful of pages /bot lists, and it honours a robots.txt that excludes our scanner, on the target’s side as well as ours.

And it reports no score out of 100. Per-crawler access, per-crawler enforcement and a parity percentage are measurements; a single number over them is a thing that moves 20% between runs and means nothing.

If you would rather not write code

The checker answers the robots.txt half with a paste, the generator writes a file, and the scanner does the whole thing for one domain in about ten seconds.

Watching rather than checking

Polling and diffing these responses yourself has edge cases. A temporary 503 can look like every policy changed at once. The CLI handles that comparison, distinguishing regressions, improvements and unreliable scans.

← DeployRadar