◀ All articles

AI crawlers

A robots.txt template for AI search in 2026

September 5, 2026 · 4 min read

A commented template that separates training crawlers from answer crawlers, with a decision table for each bot.


In 2026, a single Disallow: / aimed at "the AI bots" is how teams accidentally block the crawler that fetches live answers while thinking they only opted out of training. Your robots.txt needs to separate training preferences from answer-time fetch preferences — and still tell the truth about what your WAF will allow.

This is a commented template plus a decision table. It is not legal advice, and product user-agents change; verify strings against vendor documentation before you paste into production.

What robots.txt can and cannot do

robots.txt is a voluntary convention. Well-behaved crawlers honor it. It does not:

  • Stop a determined scraper
  • Override a WAF challenge (you can "allow" in robots and still serve 403)
  • Guarantee citation if you allow a bot
  • Replace authentication on private docs

It does give clear signals to GPTBot, ChatGPT-User, PerplexityBot, Google-Extended, and others that publish robots guidance. For broader context see AI crawlers explained and should you block AI crawlers.

Decision table by bot class

User-agent (verify current)Typical roleCommon choice for public marketing sitesNotes
GPTBotOpenAI training / model crawlingAllow or deny based on training policyNot the same as live ChatGPT browsing
ChatGPT-UserUser-triggered fetch for answersOften allow if you want to be citable in browsing contextsConfusing with GPTBot is a frequent mistake — see ChatGPT-User vs GPTBot
PerplexityBotPerplexity crawling / citation pipelineAllow if you want Perplexity visibilityPair with WAF allowlist; PerplexityBot guide
Google-ExtendedGemini / Google AI training extensionSeparate from Googlebot SearchBlocking Extended ≠ blocking Search
GooglebotClassic SearchAllow (unless you have a rare reason)Do not casually Disallow
CCBotCommon CrawlPolicy callTraining corpora; not an "answer button"
*Everyone elseKeep existing SEO rulesAvoid new blanket Disallows you do not understand

Pew Research (March 2025) found AI Overviews on roughly 18% of searches in their study, with 88% of those summaries citing three or more sources. Allowing fetch access does not put you in that set — but blocking fetch access is an effective way to stay out.

Commented template

Adjust paths to your site. Comments starting with # are for humans.

# BrandKnown-style example robots.txt for AI search readiness (2026)
# Verify user-agent tokens with each vendor before production use.

User-agent: *
Disallow: /admin/
Disallow: /api/private/
Allow: /

# --- OpenAI ---
# GPTBot: training-oriented crawl. Deny if you opt out of that use.
User-agent: GPTBot
Disallow: /

# ChatGPT-User: fetches pages when users ask ChatGPT to browse.
# Allow public content if you want live-answer eligibility.
User-agent: ChatGPT-User
Allow: /

# --- Perplexity ---
User-agent: PerplexityBot
Allow: /

# --- Google AI training extension (not Googlebot Search) ---
# Uncomment one policy. Blocking Extended does not remove you from Search.
# User-agent: Google-Extended
# Disallow: /

# User-agent: Google-Extended
# Allow: /

# --- Optional: Common Crawl ---
# User-agent: CCBot
# Disallow: /

Sitemap: https://www.example.com/sitemap.xml

Northstar Analytics might flip GPTBot to Allow if training inclusion is acceptable. The point of the template is explicit rows, not a moral stance we pretend is universal.

How to roll out safely

  1. Inventory current robots.txt and any CDN overrides.
  2. List AI agents you care about; look up today's official tokens.
  3. Decide per row: training vs answer fetch (do not merge them mentally).
  4. Deploy to staging; curl robots.txt.
  5. Fetch a key URL with the relevant user-agent; confirm 200 and body text.
  6. Check WAF/Bot Fight rules — robots allow + WAF deny still equals deny (when your WAF blocks the bots).
  7. Document the policy in your runbook so a future "block all AI" ticket does not wipe ChatGPT-User.

Pair with llms.txt only if you maintain it

An llms.txt file can hint at preferred pages for LLM consumption (llms.txt what it is). It does not replace robots rules. If you cannot keep llms.txt current, skip it rather than shipping a stale map.

Mistakes that cause outages of visibility

  • Disallow: / under User-agent: * while debugging, then forgetting
  • Blocking ChatGPT-User because a blog said "block GPTBot"
  • Allowing bots in robots while serving cookie walls with zero HTML content
  • Different robots.txt on www vs apex
  • Assuming Similarweb / TechCrunch (June 2025) AI referral growth (~1.13B referrals to top sites, ChatGPT >80% of AI referrals) will accrue to you while your robots file denies the fetchers

How to verify

CheckPass looks like
robots reachable200 at /robots.txt
SyntaxNo accidental global Disallow
Fetch testAnswer bots get 200 + HTML with brand name
WAFNo challenge page for allowlisted agents
Sitemap lineURL resolves

Re-scan after changes. A BrandKnown-style readiness check will catch "robots deny" findings quickly (~60-second scan); it still will not cite you into answers by itself.

Honest ceiling

A perfect robots template cannot compensate for empty JavaScript shells, inconsistent brand names, or a WAF that ignores your intentions. It also cannot force Perplexity or OpenAI to trust you. It removes self-inflicted blocks — the floor, not the ceiling.

Start with explicit allow/deny rows for the two OpenAI agents and PerplexityBot. Add Google-Extended as a conscious training decision. Leave Googlebot alone unless you have a specific Search reason. Then go fix the HTML those bots will actually read.

See how your own site scores

One scan checks your homepage, robots.txt, llms.txt, About page and JSON-LD, then hands you the copy-paste fixes. Free, no account needed for the first run.

Keep reading