MistralAI-Index: what it does, and what blocking it means
Mistral's search crawler. Blocking it takes you out of Le Chat's answers and citations, which is almost never what a training opt-out was meant to do.
| Operator | Mistral |
|---|---|
| What it does | Search & citations |
| Obeys robots.txt | Yes, per its operator's docs |
| Governs | Other AI assistants |
| Has a crawler | Yes |
| Last verified | 15 September 2026 |
| Next review | 15 December 2026. We re-check every entry against the operator’s documentation each quarter. |
What happens if you block MistralAI-Index?
Blocking it affects
- Le Chat's search results and citations
It does not affect
- Mistral model training
Does blocking MistralAI-Index stop Mistral training?
Mistral is one of the few operators that splits its crawlers cleanly and documents all three: MistralAI-Training for model training, MistralAI-Index for search, MistralAI-User for pages Le Chat opens because somebody asked. One token, one job, no overlap to reason about.
That makes this an unusually easy decision. If you want out of training, block MistralAI-Training on its own and this one keeps working. Blocking this one instead is the accidental version: Le Chat stops being able to find and cite you, and nothing about your training exposure changes.
Mistral publishes address ranges for this crawler, so a request claiming to be it can actually be checked rather than taken on trust.
How do you spot MistralAI-Index in your logs?
To find MistralAI-Index in your logs, match the User-Agent shown on this page. The robots.txt token is what the crawler reads, not what appears in the request.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Index/1.0; +https://docs.mistral.ai/robots)Anyone can send that User-Agent, so treat it as a claim. Mistral's published IP ranges are the stronger way to verify the request.
How do you block or allow MistralAI-Index in robots.txt?
To block MistralAI-Index:
User-agent: MistralAI-Index
Disallow: /To allow it explicitly, use an Allow: / rule if your file has a blanket User-agent: * disallow:
User-agent: MistralAI-Index
Allow: /With neither rule, MistralAI-Index is allowed by default.
Not sure how MistralAI-Index fits with the others? Build the whole file, or see which AI crawlers to allow.
Is your site blocking MistralAI-Index?
A CDN or bot-management rule can still block MistralAI-Index even when robots.txt says it is welcome. You cannot see that mismatch from the front of the site.
We request your robots.txt and homepage. We read your sitemap only when robots.txt hides pages from a crawler and we need to measure how much is affected. No other page.
Official sources
- Mistral’s crawler documentation
- Published IP ranges is the only way to verify a request claiming this user agent really came from Mistral
How is MistralAI-Index different from Mistral’s other crawlers?
| MistralAI-Index | MistralAI-Training | MistralAI-User | |
|---|---|---|---|
| Its job | Search & citations | Model training | Fetches when a user asks |
| Blocking it affects | Le Chat's search results and citations | Mistral model training | Le Chat opening your page when someone asks about it |
| Blocking it leaves alone | Mistral model training | Le Chat's search and the pages it opens for users | Le Chat's search results; Mistral model training |
| Obeys robots.txt | Yes, per its operator's docs | Yes, per its operator's docs | Yes, per its operator's docs |
| Verifiable by IP | Yes, ranges are published | No, no published ranges | Yes, ranges are published |
| Last verified | 15 September 2026 | 15 September 2026 | 15 September 2026 |