AI crawler · OpenAI

GPTBot: what it is and how to block it

GPTBot is OpenAI's crawler for collecting content that may be used to train its generative AI models. It's separate from OAI-SearchBot, so blocking it doesn't keep you out of ChatGPT search.

BrandVector Editorial · Last verified Oct 8, 2026

Operator
OpenAI
Type
AI training
robots.txt token
GPTBot
Follows robots.txt
Yes, per OpenAI
IP ranges
Published
Blocked by name
14% of top sites

What is GPTBot?

GPTBot is OpenAI's AI training crawler: it collects content that may be used to train AI models. In OpenAI's words:

“GPTBot is used to make our generative AI foundation models more useful and safe. It is used to crawl content that may be used in training our generative AI foundation models.”

“Each setting is independent of the others”

“If your site has allowed both bots, we may use the results from just one crawl for both use cases to avoid duplicative crawling.”

GPTBot user agent string

OpenAI documents this user agent string for GPTBot:

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot

In robots.txt, use the token GPTBot, not the full string. Crawlers match on the token, and version numbers in the full string can change.

How many websites block GPTBot?

In our Oct 8, 2026 check of 817 of the most-visited websites, 14% block GPTBot by name for their whole site (111 sites).

robots.txt rule for GPTBotSitesShare
Blocks the whole site by name11114%
Blocks part of the site by name415.0%
Names it but doesn't block it111.3%
Blocked only by a catch-all (*) rule212.6%

For comparison, 0.1% of the same sites block Googlebot by name. The sites are the top 1,000 origins in the Chrome UX Report for August 2026; see how we measured.

How to block GPTBot in robots.txt

Add this group to the robots.txt file at the root of your domain to block GPTBot from your whole site:

User-agent: GPTBot
Disallow: /

To block only part of your site, list those paths instead:

User-agent: GPTBot
Disallow: /private/

A group that names GPTBot replaces your User-agent: * rules for it entirely (RFC 9309). If your catch-all group blocks paths that GPTBot should also stay out of, repeat them in its group. To let GPTBot in while a catch-all rule blocks others:

User-agent: *
Disallow: /

User-agent: GPTBot
Allow: /

“OpenAI uses OAI-SearchBot and GPTBot robots.txt tags to enable webmasters to manage how their sites and content work with AI.”

What happens if you block GPTBot?

“Disallowing GPTBot indicates a site's content should not be used in training generative AI foundation models.”

How to verify requests from GPTBot

OpenAI publishes the IP addresses GPTBot uses. Check a request's IP address against that list before trusting its user agent, which any client can fake.

See every AI crawler and how often top sites block it

Sources

  1. OpenAI: Overview of OpenAI crawlers
  2. RFC 9309: Robots Exclusion Protocol
  3. Chrome UX Report top 1,000 origins (global), August 2026

Rather have an expert do it?

Tell us about your site and goals, and we'll match you with an AI visibility (GEO) specialist. Free, no obligation.