AI crawler · Mistral AI
MistralAI-Training: what it is and how to block it
MistralAI-Training crawls web content to build datasets for training Mistral's AI models. Mistral says it isn't used for search or to answer live questions.
BrandVector Editorial · Last verified Oct 8, 2026
- Operator
- Mistral AI
- Type
- AI training
- robots.txt token
MistralAI-Training- Follows robots.txt
- Yes, per Mistral AI
- IP ranges
- Not published
- Blocked by name
- 0.5% of top sites
What is MistralAI-Training?
MistralAI-Training is Mistral AI's AI training crawler: it collects content that may be used to train AI models. In Mistral AI's words:
“MistralAI-Training crawls web content to help build datasets for training Mistral generative AI models.”
“This crawler is not used for search indexing or to answer live user queries in Vibe.”
MistralAI-Training user agent string
Mistral AI documents this user agent string for MistralAI-Training:
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Training/1.0; +https://docs.mistral.ai/robots)
In robots.txt, use the token MistralAI-Training, not the full string. Crawlers match on the token, and version numbers in the full string can change.
How many websites block MistralAI-Training?
In our Oct 8, 2026 check of 817 of the most-visited websites, 0.5% block MistralAI-Training by name for their whole site (4 sites).
| robots.txt rule for MistralAI-Training | Sites | Share |
|---|---|---|
| Blocks the whole site by name | 4 | 0.5% |
| Blocks part of the site by name | 0 | 0.0% |
| Names it but doesn't block it | 0 | 0.0% |
| Blocked only by a catch-all (*) rule | 37 | 4.5% |
For comparison, 0.1% of the same sites block Googlebot by name. The sites are the top 1,000 origins in the Chrome UX Report for August 2026; see how we measured.
How to block MistralAI-Training in robots.txt
Add this group to the robots.txt file at the root of your domain to block MistralAI-Training from your whole site:
User-agent: MistralAI-Training Disallow: /
To block only part of your site, list those paths instead:
User-agent: MistralAI-Training Disallow: /private/
A group that names MistralAI-Training replaces your User-agent: * rules for it entirely (RFC 9309). If your catch-all group blocks paths that MistralAI-Training should also stay out of, repeat them in its group. To let MistralAI-Training in while a catch-all rule blocks others:
User-agent: * Disallow: / User-agent: MistralAI-Training Allow: /
“Webmasters can disallow this user agent in their robots.txt file.”
What happens if you block MistralAI-Training?
Mistral AI doesn't document what blocking MistralAI-Training changes.
How to verify requests from MistralAI-Training
Mistral AI doesn't publish IP ranges or a verification method for MistralAI-Training. Any client can send its user agent string, so treat requests that claim to be MistralAI-Training as unverified.
Other Mistral AI crawlers
- MistralAI-Index · Search
- MistralAI-User · User-triggered
See every AI crawler and how often top sites block it
Sources
Rather have an expert do it?
Tell us about your site and goals, and we'll match you with an AI visibility (GEO) specialist. Free, no obligation.