AI crawler · Common Crawl
CCBot: what it is and how to block it
CCBot builds Common Crawl's free, open archive of the web. Common Crawl says the archive has become one of the most widely used sources of training data for large language models.
BrandVector Editorial · Last verified Oct 8, 2026
- Operator
- Common Crawl
- Type
- Open dataset
- robots.txt token
CCBot- Follows robots.txt
- Yes, per Common Crawl
- IP ranges
- Published
- Blocked by name
- 14% of top sites
What is CCBot?
CCBot is Common Crawl's open dataset crawler: it builds an open archive of the web that others, including AI developers, use. In Common Crawl's words:
“CCBot is an automated crawler, checking first the robots.txt, and if crawling a page is allowed, fetches pages using HTTP GET requests.”
“has become one of the most widely used sources of training data for large language models”
“We obey the Crawl-delay parameter for robots.txt.”
“Please note that we are aware of crawlers falsely identifying themselves as CCBot.”
CCBot user agent string
Common Crawl documents this user agent string for CCBot:
CCBot/2.0 (https://commoncrawl.org/faq/)
In robots.txt, use the token CCBot, not the full string. Crawlers match on the token, and version numbers in the full string can change.
How many websites block CCBot?
In our Oct 8, 2026 check of 817 of the most-visited websites, 14% block CCBot by name for their whole site (113 sites).
| robots.txt rule for CCBot | Sites | Share |
|---|---|---|
| Blocks the whole site by name | 113 | 14% |
| Blocks part of the site by name | 35 | 4.3% |
| Names it but doesn't block it | 6 | 0.7% |
| Blocked only by a catch-all (*) rule | 30 | 3.7% |
For comparison, 0.1% of the same sites block Googlebot by name. The sites are the top 1,000 origins in the Chrome UX Report for August 2026; see how we measured.
How to block CCBot in robots.txt
Add this group to the robots.txt file at the root of your domain to block CCBot from your whole site:
User-agent: CCBot Disallow: /
To block only part of your site, list those paths instead:
User-agent: CCBot Disallow: /private/
A group that names CCBot replaces your User-agent: * rules for it entirely (RFC 9309). If your catch-all group blocks paths that CCBot should also stay out of, repeat them in its group. To let CCBot in while a catch-all rule blocks others:
User-agent: * Disallow: / User-agent: CCBot Allow: /
What happens if you block CCBot?
“Add these lines to your robots.txt file and our crawler will stop crawling your website”
How to verify requests from CCBot
Common Crawl publishes the IP addresses CCBot uses. Check a request's IP address against that list before trusting its user agent, which any client can fake.
“CCBot is now run on dedicated IP address ranges with reverse DNS (except over IPv6 where reverse DNS is not yet supported.) This allows webmasters to verify whether a logged request stems from the real CCBot”
See every AI crawler and how often top sites block it
Sources
Rather have an expert do it?
Tell us about your site and goals, and we'll match you with an AI visibility (GEO) specialist. Free, no obligation.