◀ All articles

AI crawlers

Google-Extended: what blocking it does and does not do

September 2, 2026 · 5 min read

Google-Extended is about Gemini training, not Search crawling. Blocking it is a real choice — just know which door you closed.


Your legal team wants Google-Extended blocked. Your SEO lead swears that will tank Search. Both are reacting to the same robots.txt line — and usually to different doors.

Google-Extended is about Gemini-related training and grounding use, not about Googlebot indexing your pages for classic Search. Blocking it is a real product choice. Just know which door you closed.

What Google-Extended actually is

Google publishes Google-Extended as a product token you can target in robots.txt. Teams use it to opt out of certain Google generative-AI uses of their content (training and related Gemini-facing uses), while leaving Googlebot free to crawl for Search.

That split matters. For years, "block Google" meant "block Googlebot" — which also meant disappearing from Search. Google-Extended exists so you can say no to one product surface without saying no to the other.

If you only remember one sentence: blocking Google-Extended is not the same as blocking Google Search crawling.

For a wider map of which agents do what, see AI crawlers explained. Extended is a control plane for a Google product family — not a synonym for every Googlebot hit in your logs.

What blocking it does

When you disallow Google-Extended:

EffectHappens?Notes
Stops classic Googlebot from indexing for SearchNo (by design)Keep Googlebot allowed unless you intend to leave Search
Opts out of Google-Extended-governed generative usesYesTraining / Gemini-facing uses covered by that token
Guarantees you never appear in Gemini answersNoLive retrieval, Search grounding, and other paths can still surface you
Fixes a low BrandKnown readiness score by itselfNoReadiness is website checks, not citation likelihood
Replaces a missing About page or messy brand namingNoConsent ≠ entity clarity

You closed a specific consent door. You did not uninstall Google from the internet.

What blocking it does not do

It does not:

  • Remove you from Google Search results (if Googlebot remains allowed)
  • Stop users from pasting your URL into Gemini
  • Guarantee competitors who allow Google-Extended will "win" every AI answer
  • Replace a clear Organization schema, a readable About page, or consistent brand naming

It also does not fix the quieter failure mode we see more often: a WAF or Bot Fight rule that already 403s AI agents while robots.txt looks welcoming. Those are different layers. See when your WAF blocks the bots if fetch access is the real problem.

Pew Research's March 2025 Google browsing study found AI Overviews on roughly 18% of searches, with traditional-result click rates lower when a summary was present (about 8% with a summary vs 15% without). That is a Search-results-page story. Extended is a robots token story. Do not mash them into one panic.

How to decide (a short decision table)

Your priorityGoogle-ExtendedGooglebot
Maximize Search + okay with Gemini training usesAllowAllow
Stay in Search, opt out of Extended usesDisallowAllow
Leave Search entirely (rare, deliberate)Disallow or irrelevantDisallow
Unsure, B2B SaaS with public docsUsually Allow, revisit quarterlyAllow

Northstar Analytics (our fictional B2B SaaS) chose Allow for both while their docs were the primary acquisition channel. Six months later, after a content-licensing review, they flipped Extended to Disallow and left Googlebot alone. Search traffic did not cliff. Their readiness score barely moved — because the score was never a Gemini forecast.

If you are still weighing a broader "should we block AI?" policy, read should you block AI crawlers before you copy a disallow-everything gist from a forum.

How to set the robots rule

  1. Open the robots.txt at your site root (https://example.com/robots.txt).
  2. Add a clear group for Google-Extended — do not bury it inside a catch-all User-agent: * block if you want different rules for Googlebot.
  3. Keep Googlebot's group separate and intentional.
  4. Deploy, then fetch the live file with curl (not only the CMS preview).
  5. Re-check any CDN or edge config that might serve a different robots file by host.
  6. Write the decision down: date, owner, legal/SEO sign-off.

Example pattern (illustrative — verify against Google's current docs before you ship):

User-agent: Google-Extended
Disallow: /

User-agent: Googlebot
Allow: /

If you want the opposite policy, swap the Disallow/Allow lines for Extended only. Do not copy a blog template that also Disallows Googlebot "for good measure."

Common mistakes

Mistake 1: Blocking Googlebot when you meant Extended.
One wrong user-agent string and you opted out of Search. Proofread the token.

Mistake 2: Assuming Disallow means "invisible to all Google AI."
Retrieval and product surfaces change. Treat Extended as the control Google documents — not as a universal invisibility cloak.

Mistake 3: Fighting robots while your firewall already blocks the fetch.
A polite Disallow is different from a hard 403. Diagnose which layer is answering.

Mistake 4: Setting it once and never reviewing.
Policy, product names, and your content mix change. Put a calendar reminder on the robots file the same way you review privacy policy vendors.

Mistake 5: Using Extended as a substitute for quality controls.
If the worry is inaccurate model output about your plans, fix the public pricing page and About page facts first. Robots cannot patch vague HTML.

How to verify

  • Fetch /robots.txt from production and confirm the Google-Extended group matches intent.
  • Confirm Googlebot is still allowed if Search matters.
  • Spot-check that important HTML still returns 200 to normal crawlers (robots does not override a WAF deny).
  • Record the decision in your runbook: date, owner, why.

A BrandKnown scan (~60 seconds) will flag crawler-access and identity issues on the site itself. It will not tell you whether Gemini will cite you next week — and neither will a robots line. Use Extended for consent. Use schema, clear naming, and fetch access for readiness.

Blocking Google-Extended is legitimate. Just close the door you meant to close.

See how your own site scores

One scan checks your homepage, robots.txt, llms.txt, About page and JSON-LD, then hands you the copy-paste fixes. Free, no account needed for the first run.

Keep reading