meta-webindexer: what it does, and what blocking it means
Meta's search crawler, the token that decides whether Meta AI can cite you. It is not meta-externalagent, and blocking the two has opposite consequences.
| Operator | Meta |
|---|---|
| What it does | Search & citations |
| Obeys robots.txt | Yes, per its operator's docs |
| Governs | Meta AI |
| Has a crawler | Yes |
| Last verified | 15 September 2026 |
| Next review | 15 December 2026. We re-check every entry against the operator’s documentation each quarter. |
What happens if you block meta-webindexer?
Blocking it affects
- Meta AI answers and citations
It does not affect
- Meta AI training
Is meta-webindexer the same as meta-externalagent?
Meta documents five crawlers under five separate robots.txt tokens, and only one of them is about training. meta-externalagent collects training data. meta-externalfetcher fetches a single link when a user asks for it. meta-externalads reads the landing pages behind your own ads. facebookexternalhit builds link previews when someone shares you. meta-webindexer builds the index Meta AI answers from, so this is the token that decides whether your pages can be surfaced and linked in those answers.
The split is the same shape as OpenAI's GPTBot and OAI-SearchBot, and it goes wrong the same way: a "block Meta AI" snippet lists every meta- token its author could find, and the site leaves Meta AI's index along with its training set. Meta's own documentation says plainly that allowing this token is what helps Meta AI cite and link your content.
Meta publishes no address ranges for any of its crawlers, so a request carrying this user agent cannot be confirmed from outside. We do not send it at your site either. With nothing to verify against, a 403 could not be told apart from ordinary bot management, and we would be reporting our own block as yours.
How do you spot meta-webindexer in your logs?
To find meta-webindexer in your logs, match the User-Agent shown on this page. The robots.txt token is what the crawler reads, not what appears in the request.
meta-webindexer/1.1 (+https://developers.facebook.com/documentation/sharing/webmasters/web-crawlers)Anyone can send that string, and Meta publishes no address ranges for meta-webindexer, so there is no way to confirm from outside that a request carrying it is really theirs.
How do you block or allow meta-webindexer in robots.txt?
To block meta-webindexer:
User-agent: meta-webindexer
Disallow: /To allow it explicitly, use an Allow: / rule if your file has a blanket User-agent: * disallow:
User-agent: meta-webindexer
Allow: /With neither rule, meta-webindexer is allowed by default.
Not sure how meta-webindexer fits with the others? Build the whole file, or see which AI crawlers to allow.
Is your site blocking meta-webindexer?
A CDN or bot-management rule can still block meta-webindexer even when robots.txt says it is welcome. You cannot see that mismatch from the front of the site.
We request your robots.txt and homepage. We read your sitemap only when robots.txt hides pages from a crawler and we need to measure how much is affected. No other page.
Official sources
How is meta-webindexer different from Meta’s other crawlers?
| meta-webindexer | meta-externalagent | meta-externalfetcher | meta-externalads | facebookexternalhit | |
|---|---|---|---|---|---|
| Its job | Search & citations | Model training | Fetches when a user asks | Reference | Reference |
| Blocking it affects | Meta AI answers and citations | Meta AI training and direct indexing | Meta AI opening a link someone asks about (it may ignore the rule) | Meta reading the landing pages behind your own ads | Link previews on Facebook, Instagram, WhatsApp and Messenger |
| Blocking it leaves alone | Meta AI training | Link previews when people share you; The landing pages behind your Meta ads | Meta AI training | Meta AI answers and citations; Meta AI training | Meta AI answers and citations; Meta AI training |
| Obeys robots.txt | Yes, per its operator's docs | Yes, per its operator's docs | May bypass, per its operator's docs | Yes, per its operator's docs | Generally, with caveats |
| Verifiable by IP | No, no published ranges | No, no published ranges | No, no published ranges | No, no published ranges | No, no published ranges |
| Last verified | 15 September 2026 | 19 September 2026 | 19 September 2026 | 19 September 2026 | 19 September 2026 |