A commented template that separates training crawlers from answer crawlers, with a decision table for each bot.
In 2026, a single Disallow: / aimed at "the AI bots" is how teams accidentally block the crawler that fetches live answers while thinking they only opted out of training. Your robots.txt needs to separate training preferences from answer-time fetch preferences — and still tell the truth about what your WAF will allow.
This is a commented template plus a decision table. It is not legal advice, and product user-agents change; verify strings against vendor documentation before you paste into production.
What robots.txt can and cannot do
robots.txt is a voluntary convention. Well-behaved crawlers honor it. It does not:
- Stop a determined scraper
- Override a WAF challenge (you can "allow" in robots and still serve 403)
- Guarantee citation if you allow a bot
- Replace authentication on private docs
It does give clear signals to GPTBot, ChatGPT-User, PerplexityBot, Google-Extended, and others that publish robots guidance. For broader context see AI crawlers explained and should you block AI crawlers.
Decision table by bot class
| User-agent (verify current) | Typical role | Common choice for public marketing sites | Notes |
|---|---|---|---|
GPTBot | OpenAI training / model crawling | Allow or deny based on training policy | Not the same as live ChatGPT browsing |
ChatGPT-User | User-triggered fetch for answers | Often allow if you want to be citable in browsing contexts | Confusing with GPTBot is a frequent mistake — see ChatGPT-User vs GPTBot |
PerplexityBot | Perplexity crawling / citation pipeline | Allow if you want Perplexity visibility | Pair with WAF allowlist; PerplexityBot guide |
Google-Extended | Gemini / Google AI training extension | Separate from Googlebot Search | Blocking Extended ≠ blocking Search |
Googlebot | Classic Search | Allow (unless you have a rare reason) | Do not casually Disallow |
CCBot | Common Crawl | Policy call | Training corpora; not an "answer button" |
* | Everyone else | Keep existing SEO rules | Avoid new blanket Disallows you do not understand |
Pew Research (March 2025) found AI Overviews on roughly 18% of searches in their study, with 88% of those summaries citing three or more sources. Allowing fetch access does not put you in that set — but blocking fetch access is an effective way to stay out.
Commented template
Adjust paths to your site. Comments starting with # are for humans.
# BrandKnown-style example robots.txt for AI search readiness (2026)
# Verify user-agent tokens with each vendor before production use.
User-agent: *
Disallow: /admin/
Disallow: /api/private/
Allow: /
# --- OpenAI ---
# GPTBot: training-oriented crawl. Deny if you opt out of that use.
User-agent: GPTBot
Disallow: /
# ChatGPT-User: fetches pages when users ask ChatGPT to browse.
# Allow public content if you want live-answer eligibility.
User-agent: ChatGPT-User
Allow: /
# --- Perplexity ---
User-agent: PerplexityBot
Allow: /
# --- Google AI training extension (not Googlebot Search) ---
# Uncomment one policy. Blocking Extended does not remove you from Search.
# User-agent: Google-Extended
# Disallow: /
# User-agent: Google-Extended
# Allow: /
# --- Optional: Common Crawl ---
# User-agent: CCBot
# Disallow: /
Sitemap: https://www.example.com/sitemap.xml
Northstar Analytics might flip GPTBot to Allow if training inclusion is acceptable. The point of the template is explicit rows, not a moral stance we pretend is universal.
How to roll out safely
- Inventory current
robots.txtand any CDN overrides. - List AI agents you care about; look up today's official tokens.
- Decide per row: training vs answer fetch (do not merge them mentally).
- Deploy to staging;
curlrobots.txt. - Fetch a key URL with the relevant user-agent; confirm 200 and body text.
- Check WAF/Bot Fight rules — robots allow + WAF deny still equals deny (when your WAF blocks the bots).
- Document the policy in your runbook so a future "block all AI" ticket does not wipe ChatGPT-User.
Pair with llms.txt only if you maintain it
An llms.txt file can hint at preferred pages for LLM consumption (llms.txt what it is). It does not replace robots rules. If you cannot keep llms.txt current, skip it rather than shipping a stale map.
Mistakes that cause outages of visibility
Disallow: /underUser-agent: *while debugging, then forgetting- Blocking
ChatGPT-Userbecause a blog said "block GPTBot" - Allowing bots in robots while serving cookie walls with zero HTML content
- Different robots.txt on www vs apex
- Assuming Similarweb / TechCrunch (June 2025) AI referral growth (~1.13B referrals to top sites, ChatGPT >80% of AI referrals) will accrue to you while your robots file denies the fetchers
How to verify
| Check | Pass looks like |
|---|---|
| robots reachable | 200 at /robots.txt |
| Syntax | No accidental global Disallow |
| Fetch test | Answer bots get 200 + HTML with brand name |
| WAF | No challenge page for allowlisted agents |
| Sitemap line | URL resolves |
Re-scan after changes. A BrandKnown-style readiness check will catch "robots deny" findings quickly (~60-second scan); it still will not cite you into answers by itself.
Honest ceiling
A perfect robots template cannot compensate for empty JavaScript shells, inconsistent brand names, or a WAF that ignores your intentions. It also cannot force Perplexity or OpenAI to trust you. It removes self-inflicted blocks — the floor, not the ceiling.
Start with explicit allow/deny rows for the two OpenAI agents and PerplexityBot. Add Google-Extended as a conscious training decision. Leave Googlebot alone unless you have a specific Search reason. Then go fix the HTML those bots will actually read.
