AI crawler · Common Crawl

CCBot: what it is and how to block it

CCBot builds Common Crawl's free, open archive of the web. Common Crawl says the archive has become one of the most widely used sources of training data for large language models.

BrandVector Editorial · Last verified Oct 8, 2026

Operator
Common Crawl
Type
Open dataset
robots.txt token
CCBot
Follows robots.txt
Yes, per Common Crawl
IP ranges
Published
Blocked by name
14% of top sites

What is CCBot?

CCBot is Common Crawl's open dataset crawler: it builds an open archive of the web that others, including AI developers, use. In Common Crawl's words:

“CCBot is an automated crawler, checking first the robots.txt, and if crawling a page is allowed, fetches pages using HTTP GET requests.”

“has become one of the most widely used sources of training data for large language models”

“We obey the Crawl-delay parameter for robots.txt.”

“Please note that we are aware of crawlers falsely identifying themselves as CCBot.”

CCBot user agent string

Common Crawl documents this user agent string for CCBot:

CCBot/2.0 (https://commoncrawl.org/faq/)

In robots.txt, use the token CCBot, not the full string. Crawlers match on the token, and version numbers in the full string can change.

How many websites block CCBot?

In our Oct 8, 2026 check of 817 of the most-visited websites, 14% block CCBot by name for their whole site (113 sites).

robots.txt rule for CCBotSitesShare
Blocks the whole site by name11314%
Blocks part of the site by name354.3%
Names it but doesn't block it60.7%
Blocked only by a catch-all (*) rule303.7%

For comparison, 0.1% of the same sites block Googlebot by name. The sites are the top 1,000 origins in the Chrome UX Report for August 2026; see how we measured.

How to block CCBot in robots.txt

Add this group to the robots.txt file at the root of your domain to block CCBot from your whole site:

User-agent: CCBot
Disallow: /

To block only part of your site, list those paths instead:

User-agent: CCBot
Disallow: /private/

A group that names CCBot replaces your User-agent: * rules for it entirely (RFC 9309). If your catch-all group blocks paths that CCBot should also stay out of, repeat them in its group. To let CCBot in while a catch-all rule blocks others:

User-agent: *
Disallow: /

User-agent: CCBot
Allow: /

What happens if you block CCBot?

“Add these lines to your robots.txt file and our crawler will stop crawling your website”

How to verify requests from CCBot

Common Crawl publishes the IP addresses CCBot uses. Check a request's IP address against that list before trusting its user agent, which any client can fake.

“CCBot is now run on dedicated IP address ranges with reverse DNS (except over IPv6 where reverse DNS is not yet supported.) This allows webmasters to verify whether a logged request stems from the real CCBot”

See every AI crawler and how often top sites block it

Sources

  1. Common Crawl: CCBot
  2. Common Crawl: FAQ
  3. Common Crawl: About
  4. RFC 9309: Robots Exclusion Protocol
  5. Chrome UX Report top 1,000 origins (global), August 2026

Rather have an expert do it?

Tell us about your site and goals, and we'll match you with an AI visibility (GEO) specialist. Free, no obligation.