AI crawler · Meta

Meta-ExternalAgent: what it is and how to block it

Meta-ExternalAgent crawls the web for uses that Meta says include training its foundation AI models. One token covers both training and indexing content for Meta's products.

BrandVector Editorial · Last verified Oct 8, 2026

Operator
Meta
Type
AI training
robots.txt token
Meta-ExternalAgent
Follows robots.txt
Yes, per Meta
IP ranges
Not published
Blocked by name
9.2% of top sites

What is Meta-ExternalAgent?

Meta-ExternalAgent is Meta's AI training crawler: it collects content that may be used to train AI models. In Meta's words:

“The Meta-ExternalAgent crawler crawls the web for use cases such as training foundation AI models or improving products by indexing content directly.”

“Please allow up to 24 hours for changes to robots.txt to take effect because crawlers may cache the contents of robots.txt for up to 24 hours.”

Meta-ExternalAgent user agent string

Meta documents these user agent strings for Meta-ExternalAgent:

meta-externalagent/1.1 (+/documentation/sharing/webmasters/web-crawlers)
meta-externalagent/1.1

In robots.txt, use the token Meta-ExternalAgent, not the full string. Crawlers match on the token, and version numbers in the full string can change.

How many websites block Meta-ExternalAgent?

In our Oct 8, 2026 check of 817 of the most-visited websites, 9.2% block Meta-ExternalAgent by name for their whole site (75 sites).

robots.txt rule for Meta-ExternalAgentSitesShare
Blocks the whole site by name759.2%
Blocks part of the site by name182.2%
Names it but doesn't block it30.4%
Blocked only by a catch-all (*) rule293.5%

For comparison, 0.1% of the same sites block Googlebot by name. The sites are the top 1,000 origins in the Chrome UX Report for August 2026; see how we measured.

How to block Meta-ExternalAgent in robots.txt

Add this group to the robots.txt file at the root of your domain to block Meta-ExternalAgent from your whole site:

User-agent: Meta-ExternalAgent
Disallow: /

To block only part of your site, list those paths instead:

User-agent: Meta-ExternalAgent
Disallow: /private/

A group that names Meta-ExternalAgent replaces your User-agent: * rules for it entirely (RFC 9309). If your catch-all group blocks paths that Meta-ExternalAgent should also stay out of, repeat them in its group. To let Meta-ExternalAgent in while a catch-all rule blocks others:

User-agent: *
Disallow: /

User-agent: Meta-ExternalAgent
Allow: /

“In order to block these crawlers, add a disallow for the relevant crawler to robots.txt.”

What happens if you block Meta-ExternalAgent?

Meta doesn't document what blocking Meta-ExternalAgent changes.

How to verify requests from Meta-ExternalAgent

Meta doesn't publish IP ranges or a verification method for Meta-ExternalAgent. Any client can send its user agent string, so treat requests that claim to be Meta-ExternalAgent as unverified.

See every AI crawler and how often top sites block it

Sources

  1. Meta for Developers: Meta web crawlers
  2. RFC 9309: Robots Exclusion Protocol
  3. Chrome UX Report top 1,000 origins (global), August 2026

Rather have an expert do it?

Tell us about your site and goals, and we'll match you with an AI visibility (GEO) specialist. Free, no obligation.