AI crawler · Ai2
AI2Bot: what it is and how to block it
AI2Bot is the crawler of Ai2 (the Allen Institute for AI), which collects web content to train open language models. Ai2 doesn't say whether it follows robots.txt.
BrandVector Editorial · Last verified Oct 8, 2026
- Operator
- Ai2
- Type
- AI training
- robots.txt token
AI2Bot- Follows robots.txt
- Ai2 doesn't say
- IP ranges
- Not published
- Blocked by name
- 5.4% of top sites
What is AI2Bot?
AI2Bot is Ai2's AI training crawler: it collects content that may be used to train AI models. In Ai2's words:
“The AI2 Bot explores certain domains to find web content. This web content is used to train open language models.”
“This user agent string can be used to filter or reject traffic from our crawler if desired.”
AI2Bot user agent string
Ai2 documents this user agent string for AI2Bot:
Mozilla/5.0 (compatible) AI2Bot (+https://www.allenai.org/crawler)
In robots.txt, use the token AI2Bot, not the full string. Crawlers match on the token, and version numbers in the full string can change.
How many websites block AI2Bot?
In our Oct 8, 2026 check of 817 of the most-visited websites, 5.4% block AI2Bot by name for their whole site (44 sites).
| robots.txt rule for AI2Bot | Sites | Share |
|---|---|---|
| Blocks the whole site by name | 44 | 5.4% |
| Blocks part of the site by name | 7 | 0.9% |
| Names it but doesn't block it | 0 | 0.0% |
| Blocked only by a catch-all (*) rule | 37 | 4.5% |
For comparison, 0.1% of the same sites block Googlebot by name. The sites are the top 1,000 origins in the Chrome UX Report for August 2026; see how we measured.
How to block AI2Bot in robots.txt
Add this group to the robots.txt file at the root of your domain to block AI2Bot from your whole site:
User-agent: AI2Bot Disallow: /
To block only part of your site, list those paths instead:
User-agent: AI2Bot Disallow: /private/
A group that names AI2Bot replaces your User-agent: * rules for it entirely (RFC 9309). If your catch-all group blocks paths that AI2Bot should also stay out of, repeat them in its group. To let AI2Bot in while a catch-all rule blocks others:
User-agent: * Disallow: / User-agent: AI2Bot Allow: /
Ai2 doesn't say whether AI2Bot follows robots.txt, so a robots.txt block is a request it may not honor.
What happens if you block AI2Bot?
Ai2 doesn't document what blocking AI2Bot changes.
How to verify requests from AI2Bot
Ai2 doesn't publish IP ranges or a verification method for AI2Bot. Any client can send its user agent string, so treat requests that claim to be AI2Bot as unverified.
See every AI crawler and how often top sites block it
Sources
Rather have an expert do it?
Tell us about your site and goals, and we'll match you with an AI visibility (GEO) specialist. Free, no obligation.