AI crawler · Ai2

AI2Bot: what it is and how to block it

AI2Bot is the crawler of Ai2 (the Allen Institute for AI), which collects web content to train open language models. Ai2 doesn't say whether it follows robots.txt.

BrandVector Editorial · Last verified Oct 8, 2026

Operator
Ai2
Type
AI training
robots.txt token
AI2Bot
Follows robots.txt
Ai2 doesn't say
IP ranges
Not published
Blocked by name
5.4% of top sites

What is AI2Bot?

AI2Bot is Ai2's AI training crawler: it collects content that may be used to train AI models. In Ai2's words:

“The AI2 Bot explores certain domains to find web content. This web content is used to train open language models.”

“This user agent string can be used to filter or reject traffic from our crawler if desired.”

AI2Bot user agent string

Ai2 documents this user agent string for AI2Bot:

Mozilla/5.0 (compatible) AI2Bot (+https://www.allenai.org/crawler)

In robots.txt, use the token AI2Bot, not the full string. Crawlers match on the token, and version numbers in the full string can change.

How many websites block AI2Bot?

In our Oct 8, 2026 check of 817 of the most-visited websites, 5.4% block AI2Bot by name for their whole site (44 sites).

robots.txt rule for AI2BotSitesShare
Blocks the whole site by name445.4%
Blocks part of the site by name70.9%
Names it but doesn't block it00.0%
Blocked only by a catch-all (*) rule374.5%

For comparison, 0.1% of the same sites block Googlebot by name. The sites are the top 1,000 origins in the Chrome UX Report for August 2026; see how we measured.

How to block AI2Bot in robots.txt

Add this group to the robots.txt file at the root of your domain to block AI2Bot from your whole site:

User-agent: AI2Bot
Disallow: /

To block only part of your site, list those paths instead:

User-agent: AI2Bot
Disallow: /private/

A group that names AI2Bot replaces your User-agent: * rules for it entirely (RFC 9309). If your catch-all group blocks paths that AI2Bot should also stay out of, repeat them in its group. To let AI2Bot in while a catch-all rule blocks others:

User-agent: *
Disallow: /

User-agent: AI2Bot
Allow: /

Ai2 doesn't say whether AI2Bot follows robots.txt, so a robots.txt block is a request it may not honor.

What happens if you block AI2Bot?

Ai2 doesn't document what blocking AI2Bot changes.

How to verify requests from AI2Bot

Ai2 doesn't publish IP ranges or a verification method for AI2Bot. Any client can send its user agent string, so treat requests that claim to be AI2Bot as unverified.

See every AI crawler and how often top sites block it

Sources

  1. Ai2: AI2 Bot
  2. RFC 9309: Robots Exclusion Protocol
  3. Chrome UX Report top 1,000 origins (global), August 2026

Rather have an expert do it?

Tell us about your site and goals, and we'll match you with an AI visibility (GEO) specialist. Free, no obligation.