Google

google.com

Training crawler

Google-Extended

Google-Extended is a robots.txt product control token from Google, not a fetching crawler user-agent. Robots.txt product token that controls eligible use of Google-crawled content for Gemini training and grounding.

Source checked 18 Aug 2026Officially documentedActive

Purpose

Training crawler

Evidence

Officially documented

Status

Active

robots.txt

Honors robots.txt
Last source check: 18 Aug 2026 · First verified in this record: 18 Aug 2026
[01]

What is Google-Extended?

Google-Extended is a robots.txt product control token from Google, not a fetching crawler user-agent. Robots.txt product token that controls eligible use of Google-crawled content for Gemini training and grounding.

[02]

Control token

No official HTTP user-agent is published for this record. Google-Extended is a robots.txt control token.

[03]

Allow or block in robots.txt

Use the narrow token below only after deciding whether this crawler's purpose fits your policy. An allow rule makes that choice explicit. A block rule asks compliant automated crawlers not to fetch matching paths.

Allow Google-Extended
User-agent: Google-Extended
Allow: /
Block Google-Extended
User-agent: Google-Extended
Disallow: /

Robots.txt is a request policy. It does not authenticate the sender, remove content already collected, or guarantee that a model will use or not use content. User-triggered fetchers can have different behavior from automated crawlers.

[04]

Verify a request before attributing it

01

Match the full user-agent

Use the published format where one exists. Treat the match as a clue, not proof.

02

Look for operator verification

No operator-published IP, reverse-DNS, or signature check is attached to this record.

03

Keep attribution separate from policy

Log what you observed, what the operator documents, and what you inferred as three separate fields.

[05]

JavaScript behavior

Rendering is not documented

The cited operator material does not verify JavaScript rendering. Serve useful HTML before client-side JavaScript where possible.

[06]

Related crawlers and distinctions

Google-Extended controls AI training use, not whether Google Search can index you. If the practical question is whether Google’s assistant recommends your brand, monitor brand mentions in Gemini separately.

To see why Google-Extended access is worth checking rather than assuming, read one site’s before and after, from blocked crawlers to first-place citations.

[08]

Frequently asked questions

Google-Extended is a robots.txt product control token from Google, not a fetching crawler user-agent. Robots.txt product token that controls eligible use of Google-crawled content for Gemini training and grounding.

Spoofing warning: any client can copy a user-agent string. Use operator-published IP ranges, reverse DNS, or request signatures where available. A name match alone does not verify ownership.

See which crawler signatures reach your site

Paste or upload a server log to find known AI crawlers, status codes, and requested pages. The analysis stays in your browser.