AI crawler · Mistral AI

MistralAI-Training: what it is and how to block it

MistralAI-Training crawls web content to build datasets for training Mistral's AI models. Mistral says it isn't used for search or to answer live questions.

BrandVector Editorial · Last verified Oct 8, 2026

Operator
Mistral AI
Type
AI training
robots.txt token
MistralAI-Training
Follows robots.txt
Yes, per Mistral AI
IP ranges
Not published
Blocked by name
0.5% of top sites

What is MistralAI-Training?

MistralAI-Training is Mistral AI's AI training crawler: it collects content that may be used to train AI models. In Mistral AI's words:

“MistralAI-Training crawls web content to help build datasets for training Mistral generative AI models.”

“This crawler is not used for search indexing or to answer live user queries in Vibe.”

MistralAI-Training user agent string

Mistral AI documents this user agent string for MistralAI-Training:

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Training/1.0; +https://docs.mistral.ai/robots)

In robots.txt, use the token MistralAI-Training, not the full string. Crawlers match on the token, and version numbers in the full string can change.

How many websites block MistralAI-Training?

In our Oct 8, 2026 check of 817 of the most-visited websites, 0.5% block MistralAI-Training by name for their whole site (4 sites).

robots.txt rule for MistralAI-TrainingSitesShare
Blocks the whole site by name40.5%
Blocks part of the site by name00.0%
Names it but doesn't block it00.0%
Blocked only by a catch-all (*) rule374.5%

For comparison, 0.1% of the same sites block Googlebot by name. The sites are the top 1,000 origins in the Chrome UX Report for August 2026; see how we measured.

How to block MistralAI-Training in robots.txt

Add this group to the robots.txt file at the root of your domain to block MistralAI-Training from your whole site:

User-agent: MistralAI-Training
Disallow: /

To block only part of your site, list those paths instead:

User-agent: MistralAI-Training
Disallow: /private/

A group that names MistralAI-Training replaces your User-agent: * rules for it entirely (RFC 9309). If your catch-all group blocks paths that MistralAI-Training should also stay out of, repeat them in its group. To let MistralAI-Training in while a catch-all rule blocks others:

User-agent: *
Disallow: /

User-agent: MistralAI-Training
Allow: /

“Webmasters can disallow this user agent in their robots.txt file.”

What happens if you block MistralAI-Training?

Mistral AI doesn't document what blocking MistralAI-Training changes.

How to verify requests from MistralAI-Training

Mistral AI doesn't publish IP ranges or a verification method for MistralAI-Training. Any client can send its user agent string, so treat requests that claim to be MistralAI-Training as unverified.

See every AI crawler and how often top sites block it

Sources

  1. Mistral AI: Robots
  2. RFC 9309: Robots Exclusion Protocol
  3. Chrome UX Report top 1,000 origins (global), August 2026

Rather have an expert do it?

Tell us about your site and goals, and we'll match you with an AI visibility (GEO) specialist. Free, no obligation.