llms.txt: what it does and doesn't do
Short answer: it’s a nice-to-have, not a lever. It won’t get you cited, it won’t block anyone, and it won’t fix a site AI can’t read. It’s still worth ten minutes, for a smaller reason than people claim. Registry verified 19 September 2026.
What it is
A Markdown file at /llms.txt that describes your site for a language model: a title, a one-line summary, then lists of your important pages with a sentence on each. Think of it as a table of contents written for a reader that would otherwise have to guess the shape of your site from whichever page it landed on.
Jeremy Howard proposed it in September 2024. It is still a proposal, not a standard, and nobody is obliged to read it.
What it doesn’t do
- It doesn’t control access. That’s robots.txt’s job, and your CDN’s. An llms.txt that says “please read everything” does nothing for a crawler your robots.txt or Cloudflare turns away.
- Google doesn’t use it. Google’s Gary Illyes said in July 2025 that Google doesn’t support it and isn’t planning to. It has no effect on Search or AI Overviews.
- No AI crawler documents using it. Every operator in our registry points to robots.txt when it explains how to control its crawler. Some server logs show the file being fetched, but fetching a file is not the same as using it, and nobody has shown it changes who gets cited.
What it’s actually good for
The moment someone hands your URL to an assistant and asks about you. Then a clean summary and a list of the pages that matter beat the assistant piecing you together from your home page’s nav. That’s the job of the user-triggered fetchers, and it’s also why docs sites love the file: coding assistants read it.
We ship one, here, for exactly that reason. It’s generated from the same data as the site, so it can’t go stale. A hand-written one that describes pages you’ve since deleted is worse than none.
Do these first
If AI can’t see you, llms.txt isn’t why. In order:
- Your robots.txt lets the search and citation crawlers in. Check it.
- Your CDN isn’t turning them away. Here’s how to check.
- Your content is in the HTML, not only after JavaScript runs.
- Then, if you like, add an llms.txt. It’s the cherry, not the cake.
The scan covers the first three in about ten seconds.
Where to go next
- AI SEO setup: the five things worth checking
Five layers you can check from outside, each capable of hiding a site. We go through them in the order they tend to break. - Which AI crawlers should you allow?
The tokens have four jobs. Which ones you want depends on what you are trying to achieve, and the costly mistake is blocking the citation crawler when you meant to block the training one. - Does Google-Extended block AI Overviews?
A common robots.txt mix-up, and the controls that actually decide what AI Overviews can quote.