Technical reference

AI crawler list: every AI user agent and who blocks it

31 AI crawlers and robots.txt tokens from 15 companies: what each one does, how to block it, and how often the most-visited websites block it. Every fact comes from the operator's own documentation.

BrandVector Editorial · Last verified Oct 8, 2026

Key findings · Oct 8, 2026

  • 29% of the 817 top websites we checked name at least one AI crawler in robots.txt.
  • The most-blocked is CCBot (14% block it by name for their whole site), followed by GPTBot (14%) and Bytespider (13%).
  • Sites block AI companies' training crawlers far more often than their search crawlers: GPTBot 14% vs OAI-SearchBot 5.1%; ClaudeBot 12% vs Claude-SearchBot 5.1%.

All AI crawlers

CrawlerOperatorTypeBlocked by name
CCBotCommon CrawlOpen dataset14%
GPTBotOpenAIAI training14%
BytespiderByteDanceSearch13%
ClaudeBotAnthropicAI training12%
Google-ExtendedGoogleControl token11%
DiffbotDiffbotSearch9.9%
Meta-ExternalAgentMetaAI training9.2%
PerplexityBotPerplexitySearch8.8%
YouBotYou.comSearch8.0%
AmazonbotAmazonAI training7.6%
Applebot-ExtendedAppleControl token7.3%
ChatGPT-UserOpenAIUser-triggered7.0%
Perplexity-UserPerplexityUser-triggered5.4%
AI2BotAi2AI training5.4%
DuckAssistBotDuckDuckGoSearch5.3%
OAI-SearchBotOpenAISearch5.1%
Claude-SearchBotAnthropicSearch5.1%
Meta-ExternalFetcherMetaUser-triggered5.1%
Claude-UserAnthropicUser-triggered4.9%
Google-CloudVertexBotGoogleOwner-requested4.5%
MistralAI-UserMistral AIUser-triggered4.0%
Meta-WebIndexerMetaSearch2.8%
Amzn-SearchBotAmazonSearch1.5%
Amzn-UserAmazonUser-triggered1.3%
Diffbot-UserDiffbotUser-triggered0.7%
ApplebotAppleSearch0.6%
MistralAI-TrainingMistral AIAI training0.5%
ExaSearchBotExaSearch0.4%
MistralAI-IndexMistral AISearch0.2%
OAI-AdsBotOpenAIAd review0.0%
Google-AgentGoogleUser-triggered–

“Blocked by name” is the share of top websites whose robots.txt names the crawler and disallows their whole site. See the methodology.

Types of AI crawlers

AI training

These crawl the web to collect content that may be used to train AI models. Blocking them doesn't affect whether you appear in AI search answers, per the operators that run separate search crawlers.

GPTBot (OpenAI), ClaudeBot (Anthropic), MistralAI-Training (Mistral AI), Meta-ExternalAgent (Meta), Amazonbot (Amazon), AI2Bot (Ai2)

Search

These index pages for search results, including the search that AI assistants use to find and cite sources. Blocking one can keep your pages out of that service's results.

OAI-SearchBot (OpenAI), Claude-SearchBot (Anthropic), Applebot (Apple), MistralAI-Index (Mistral AI), PerplexityBot (Perplexity), Meta-WebIndexer (Meta), Amzn-SearchBot (Amazon), Bytespider (ByteDance), DuckAssistBot (DuckDuckGo), YouBot (You.com), ExaSearchBot (Exa), Diffbot (Diffbot)

User-triggered

These fetch a page only when a person asks an AI assistant to, for example by pasting a link. Some operators say robots.txt may not apply to them, because a user started the request.

ChatGPT-User (OpenAI), Claude-User (Anthropic), Google-Agent (Google), MistralAI-User (Mistral AI), Perplexity-User (Perplexity), Meta-ExternalFetcher (Meta), Amzn-User (Amazon), Diffbot-User (Diffbot)

Control token

These aren't crawlers. They're tokens you name in robots.txt, and the operator's existing crawlers read them to decide how your content may be used.

Google-Extended (Google), Applebot-Extended (Apple)

Open dataset

These build open archives of the web that anyone can download, including AI developers training models.

CCBot (Common Crawl)

Owner-requested

These crawl a site only when its owner sets up an AI product, such as a custom search agent, that uses it.

Google-CloudVertexBot (Google)

Ad review

These visit the landing pages of ads to check them against the ad platform's policies.

OAI-AdsBot (OpenAI)

AI Overviews and Copilot don't have their own crawlers

Some AI search features use their company's main search crawler, so there's no AI-specific token to block. Blocking Googlebot or Bingbot would remove you from regular search too, so both companies point to page-level controls instead.

Google AI Overviews and AI Mode

“AI is built into Search and integral to how Search functions, which is why robots.txt directives for Googlebot is the control for site owners to manage access to how their sites are crawled for Search.”

“To limit the information shown from your pages in Search, use nosnippet, data-nosnippet, max-snippet, or noindex controls.”

Microsoft Copilot

“Bing and Copilot search experiences rely on the same core crawling, indexing, and ranking foundation as traditional search.”

“NOARCHIVE prevents content from being used in Copilot responses and grounding results.”

How to block AI training but stay in AI search

This robots.txt group blocks the crawlers and tokens whose operators say they're used for AI training, and leaves search crawlers such as OAI-SearchBot, Claude-SearchBot, and PerplexityBot untouched:

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: MistralAI-Training
User-agent: Meta-ExternalAgent
User-agent: Amazonbot
User-agent: CCBot
User-agent: AI2Bot
Disallow: /

Consecutive User-agent lines share the rules that follow them. Each operator decides what its token controls, so check the crawler's page before relying on a block, and remember that robots.txt is a request: it doesn't stop a crawler that ignores it.

Methodology

Websites. The 1,000 most-visited origins worldwide in the Chrome UX Report for August 2026, via the CrUX top lists archive. Origins are counted separately, so country versions of one site (such as Google's) each count.

Collection. On Oct 8, 2026, we requested /robots.txt once from each origin, following redirects, with a user agent that identifies BrandVector. 759 sites returned a robots.txt file and 58 had none, which under RFC 9309 means every crawler is allowed. We excluded 183 sites we couldn't read: 103 refused the request (usually bot protection), 38 were unreachable, 1 returned server errors, and 41 returned a web page instead of a robots.txt file. Percentages are of the 817 sites we could check.

Rules. We applied RFC 9309: all groups that name a crawler apply to it (matched case-insensitively on its token), a User-agent: * group applies only when none do, and the longest matching rule wins. “Blocks the whole site” means the root is disallowed with no allowed exceptions.

Limitations. robots.txt states a site's preference. It doesn't show firewall or CDN blocks, or whether crawlers comply, and sites that refused our request may block more crawlers than the ones we could read. We repeat the study monthly.

Sources

  1. OpenAI: Overview of OpenAI crawlers
  2. Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?
  3. Google: Google's common crawlers
  4. Google Search Central: AI features and your website
  5. Google: Google's user-triggered fetchers
  6. Google: Web Bot Auth
  7. Google Cloud: Prepare data for ingesting
  8. Apple: About Applebot
  9. Mistral AI: Robots
  10. Perplexity: Perplexity crawlers
  11. Perplexity Help Center: How does Perplexity follow robots.txt?
  12. Meta for Developers: Meta web crawlers
  13. Amazon: Amazonbot and Amazon's other crawlers
  14. Toutiao Search Webmaster Platform: About Bytespider (in Chinese)
  15. Common Crawl: CCBot
  16. Common Crawl: FAQ
  17. Common Crawl: About
  18. DuckDuckGo: DuckAssistBot
  19. You.com: YouBot
  20. Exa: ExaSearchBot
  21. Diffbot: Does Diffbot respect robots.txt?
  22. Ai2: AI2 Bot
  23. RFC 9309: Robots Exclusion Protocol
  24. Chrome UX Report top 1,000 origins (global), August 2026

Rather have an expert do it?

Tell us about your site and goals, and we'll match you with an AI visibility (GEO) specialist. Free, no obligation.