AI crawler · Diffbot
Diffbot: what it is and how to block it
Diffbot crawls the web to build a general search engine and its Knowledge Graph. Diffbot says it isn't used for AI training, and that robots.txt can be overridden in specific cases.
BrandVector Editorial · Last verified Oct 8, 2026
- Operator
- Diffbot
- Type
- Search
- robots.txt token
Diffbot- Follows robots.txt
- Partly, per Diffbot
- IP ranges
- Not published
- Blocked by name
- 9.9% of top sites
What is Diffbot?
Diffbot is Diffbot's search crawler: it indexes pages for search results. In Diffbot's words:
“General, proactive web crawling for building a general search engine. This allows websites to be discovered and cited from the Diffbot Knowledge Graph and web search services in response to keyword queries. It is not used for AI training.”
Diffbot user agent string
Diffbot doesn't publish a full user agent string for Diffbot. In robots.txt, use the token Diffbot.
How many websites block Diffbot?
In our Oct 8, 2026 check of 817 of the most-visited websites, 9.9% block Diffbot by name for their whole site (81 sites).
| robots.txt rule for Diffbot | Sites | Share |
|---|---|---|
| Blocks the whole site by name | 81 | 9.9% |
| Blocks part of the site by name | 5 | 0.6% |
| Names it but doesn't block it | 1 | 0.1% |
| Blocked only by a catch-all (*) rule | 35 | 4.3% |
For comparison, 0.1% of the same sites block Googlebot by name. The sites are the top 1,000 origins in the Chrome UX Report for August 2026; see how we measured.
How to block Diffbot in robots.txt
Add this group to the robots.txt file at the root of your domain to block Diffbot from your whole site:
User-agent: Diffbot Disallow: /
To block only part of your site, list those paths instead:
User-agent: Diffbot Disallow: /private/
A group that names Diffbot replaces your User-agent: * rules for it entirely (RFC 9309). If your catch-all group blocks paths that Diffbot should also stay out of, repeat them in its group. To let Diffbot in while a catch-all rule blocks others:
User-agent: * Disallow: / User-agent: Diffbot Allow: /
Diffbot says robots.txt doesn't fully control Diffbot. To stop it completely, block it at your firewall or CDN using the verification details below.
“In specific cases — typically because of a partnership or agreement you have with the site to be crawled — the robots.txt instruction can be ignored/overridden.”
What happens if you block Diffbot?
Diffbot doesn't document what blocking Diffbot changes.
How to verify requests from Diffbot
Diffbot doesn't publish IP ranges or a verification method for Diffbot. Any client can send its user agent string, so treat requests that claim to be Diffbot as unverified.
Other Diffbot crawlers
- Diffbot-User · User-triggered
See every AI crawler and how often top sites block it
Sources
Rather have an expert do it?
Tell us about your site and goals, and we'll match you with an AI visibility (GEO) specialist. Free, no obligation.