Glossary

The words used in our reports and guides, explained in plain English. If a finding mentions something new to you, you will probably find it here.

Crawler
A program that requests pages from websites on someone's behalf. AI companies run several, with different jobs: some collect text for training, some fetch pages to cite in answers, and some open a page only when a person asks about it. Every AI crawler, and what blocking it costs →
User agent
The name a crawler or browser sends with every request, such as GPTBot or an ordinary Chrome string. Anyone can send any user agent, so the name is a claim about who is asking, not proof.

Why it matters: A request carrying a crawler's name may not come from that crawler, which is why published IP ranges matter.

robots.txt token
The name a crawler looks for in your robots.txt, written after User-agent:. Usually it matches the crawler's user agent. A few, like Google-Extended, are tokens only: nothing ever visits under that name.
robots.txt
A text file at the root of your site that tells crawlers, by name, which paths they may fetch. It states what you want. Well-behaved crawlers follow it, but nothing forces the others to. Check what your robots.txt says →
Search and citation crawler
A crawler that fetches pages so an assistant can quote and link them in answers, such as OAI-SearchBot for ChatGPT search. Blocking one keeps your pages out of those answers.
Training crawler
A crawler that collects text to train future models, such as GPTBot or CCBot. Blocking one is a policy choice about your content. By itself, it does not affect whether AI search can cite you.
User-triggered fetcher
A fetcher that opens one page because a person asked an assistant about it, such as ChatGPT-User or Claude-User. Some operators document that these may not follow robots.txt, because a person made the request.
Grounding
Feeding an AI model fresh content when it answers, so it can use current pages rather than relying only on what it was trained on. For Gemini, Google-Extended controls it.
noindex
A directive in a meta tag or an X-Robots-Tag header that tells search engines not to list a page. It removes the page from Google Search, including AI Overviews. A leftover one from a staging site is a common accident.

Why it matters: One leftover tag removes a page from Google and AI Overviews, while the page looks fine to every visitor.

Snippet controls
nosnippet, data-nosnippet and max-snippet limit how much of a page Google may quote. They are what decide what AI Overviews can show, and they shape ordinary search snippets at the same time. What decides AI Overviews →

Why it matters: They are the only way to limit what AI Overviews quotes, and they limit ordinary snippets at the same time.

Canonical
A link in a page's head naming the URL that should represent it in search. If it points to the wrong page, or to another site, it can hand that page's place to the URL it names.

Why it matters: A wrong one quietly gives your page's place in search to a different URL.

CDN
A content delivery network, such as Cloudflare, that sits in front of your server and answers requests at the edge. It can turn a crawler away before the request reaches your site, whatever your robots.txt says. Is Cloudflare blocking AI crawlers? →

Why it matters: It can block an AI crawler even when robots.txt allows it, and nothing on your site shows it.

WAF and bot management
Firewall and bot rules, often run by a CDN, that challenge or block traffic that looks automated. They are a common reason a crawler is refused while robots.txt says it is welcome.

Why it matters: A setting switched on to stop scrapers can also turn away the crawlers that put you in AI answers.

Verified IP ranges
The address lists some operators publish so a site can confirm a request really came from them. A request with a crawler's name from anywhere else is unverified, which is why our own test requests can be treated differently from the real crawler.
Server rendering
Sending a page's content in the HTML the server returns, so it is there before any JavaScript runs. Crawlers that do not run JavaScript can read it.
Client-side rendering
Building a page's content in the browser with JavaScript. Googlebot runs JavaScript and usually copes, but most AI crawlers do not, so they may receive an empty shell.

Why it matters: Google may see the whole page while ChatGPT and Claude see an empty one.

Content-Signal and Content-Usage
Lines in robots.txt that state how you would like your content used, for search, training or as AI input. Content-Signal is Cloudflare's; Content-Usage is an IETF draft. Both are preferences, not controls.
Pay-Per-Crawl
A Cloudflare feature that answers a crawler with HTTP 402 and a price instead of the page. The crawler is asked to pay rather than turned away.
llms.txt
A proposed Markdown file describing a site for language models. It is a nice-to-have: it does not control access, and Google has said it does not use it. What llms.txt does and doesn't do →

← Guides