AI crawler · Google

Google-Extended: what it is and how to block it

Google-Extended is a robots.txt token, not a crawler. It controls whether content Google crawls may be used to train Gemini models and to ground Gemini's answers. Google says it doesn't affect Google Search, which includes AI Overviews.

BrandVector Editorial · Last verified Oct 8, 2026

Operator
Google
Type
Control token
robots.txt token
Google-Extended
Follows robots.txt
Yes, per Google
IP ranges
Not published
Blocked by name
11% of top sites

What is Google-Extended?

Google-Extended is Google's control token robots.txt token: it is a robots.txt token, not a crawler: it controls how already-crawled content is used. In Google's words:

“Google-Extended is a standalone product token that web publishers can use to manage whether content Google crawls from their sites may be used for training future generations of Gemini models that power Gemini Apps and Vertex AI API for Gemini and for grounding (providing content from the Google Search index to the model at prompt time to improve factuality and relevancy) in Gemini Apps and Grounding with Google Search on Vertex AI.”

“AI is built into Search and integral to how Search functions, which is why robots.txt directives for Googlebot is the control for site owners to manage access to how their sites are crawled for Search.”

Google-Extended user agent string

Google-Extended has no user agent string of its own. It's a token you name in robots.txt, and Google's existing crawlers read it when deciding how your content may be used.

How many websites block Google-Extended?

In our Oct 8, 2026 check of 817 of the most-visited websites, 11% block Google-Extended by name for their whole site (88 sites).

robots.txt rule for Google-ExtendedSitesShare
Blocks the whole site by name8811%
Blocks part of the site by name384.7%
Names it but doesn't block it131.6%
Blocked only by a catch-all (*) rule202.4%

For comparison, 0.1% of the same sites block Googlebot by name. The sites are the top 1,000 origins in the Chrome UX Report for August 2026; see how we measured.

How to block Google-Extended in robots.txt

Add this group to the robots.txt file at the root of your domain to block Google-Extended from your whole site:

User-agent: Google-Extended
Disallow: /

To block only part of your site, list those paths instead:

User-agent: Google-Extended
Disallow: /private/

A group that names Google-Extended replaces your User-agent: * rules for it entirely (RFC 9309). If your catch-all group blocks paths that Google-Extended should also stay out of, repeat them in its group. To let Google-Extended in while a catch-all rule blocks others:

User-agent: *
Disallow: /

User-agent: Google-Extended
Allow: /

“Crawling is done with existing Google user agent strings; the robots.txt user-agent token is used in a control capacity.”

What happens if you block Google-Extended?

“Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search.”

How to verify requests from Google-Extended

Google doesn't publish IP ranges or a verification method for Google-Extended. Any client can send its user agent string, so treat requests that claim to be Google-Extended as unverified.

See every AI crawler and how often top sites block it

Sources

  1. Google: Google's common crawlers
  2. Google Search Central: AI features and your website
  3. RFC 9309: Robots Exclusion Protocol
  4. Chrome UX Report top 1,000 origins (global), August 2026

Rather have an expert do it?

Tell us about your site and goals, and we'll match you with an AI visibility (GEO) specialist. Free, no obligation.