MistralAI-Training
Mistral's training opt-out. Blocking it keeps your content out of Mistral model training.
| Operator | Mistral |
|---|---|
| What it does | Model training |
| Obeys robots.txt | Yes, per its operator's docs |
| Governs | AI model training |
| Has a crawler | Yes |
| Last verified | 13 August 2026 |
What blocking it costs you
What to know
Mistral splits its crawlers cleanly, which makes this an easy decision: MistralAI-Index handles Mistral's search, and MistralAI-User fetches pages for Le Chat. Blocking the training token leaves both of those working.
The robots.txt rules
To block MistralAI-Training:
User-agent: MistralAI-Training
Disallow: /To allow it explicitly — worth doing when your file also contains a blanket User-agent: * disallow, since the most specific matching group wins and a token with no group of its own falls back to the wildcard:
User-agent: MistralAI-Training
Allow: /If neither rule is present, MistralAI-Training is allowed. That is the default, and it is why an accidental block is nearly always something that was added rather than something that was forgotten — usually a snippet copied from a blog post that listed more tokens than the author intended to block.
Is it blocked on your site?
Your robots.txt is only the half you can read. A CDN or bot-management rule can turn MistralAI-Training away while your file still says it is welcome, and that mismatch is invisible from the front of the site. We check both layers separately and tell you which one is doing what.
We request two things: your robots.txt and your homepage. No other pages, ever.