One fetches for live answers; one trains models. Confusing them is how teams block the bot they meant to welcome.
Two OpenAI user-agents show up in logs. Someone pastes a "block GPTBot" snippet from a 2023 thread. A week later the team asks why ChatGPT browsing never cites the docs.
Because GPTBot and ChatGPT-User are not the same door. One is oriented around model crawling/training pipelines. The other fetches pages when a user (or agent) needs live content for an answer. Confusing them is how you block the bot you meant to welcome.
The short answer
| Agent | What it is for (plain terms) | If you want live ChatGPT citations/browsing | If you want to opt out of OpenAI training crawl |
|---|---|---|---|
GPTBot | OpenAI's crawler associated with training / model data collection policies | Not the primary lever | Often the agent people Disallow |
ChatGPT-User | Fetches pages in user-initiated browsing / retrieval contexts | Usually Allow on public content | Separate decision — do not assume denying GPTBot denies this |
Always confirm current tokens and meanings in OpenAI's published documentation before shipping rules. Names and policies evolve; your robots.txt comments should include a "verified on DATE" line.
Why the confusion happens
Security blogs, privacy guides, and Slack screenshots flatten "the OpenAI bot" into one villain. Meanwhile product behavior split:
- Training-time collection preferences → GPTBot-style rules
- Answer-time fetch when someone asks ChatGPT to look at the web → ChatGPT-User
If your policy is "we are fine being quoted in answers, but we do not want to donate the blog to training," you may Allow ChatGPT-User and Disallow GPTBot. If your policy is "no OpenAI systems may fetch us," you deny both — and you should say that explicitly so marketing stops promising ChatGPT visibility.
See also should you block AI crawlers for the broader tradeoff, and robots.txt template for AI 2026 for a full file shape.
What allowing ChatGPT-User does not guarantee
Allowing the agent means well-behaved fetches are permitted. It does not mean:
- ChatGPT will name Northstar Analytics for category queries
- You will win against review sites and Wikipedia in retrieval
- Your JavaScript-only homepage will suddenly be readable
- Your WAF will cooperate
Similarweb / TechCrunch (June 2025) reported AI platforms sending about 1.13B referrals to the top 1,000 sites (up 357% YoY), with ChatGPT representing more than 80% of those AI referrals — while Google Search still sent about 191B in the same month. ChatGPT matters in the AI-referral mix; it is still not a substitute for Search. Access is the ticket to the room, not the trophy.
Similarweb Gen AI Landscape 2025 (US desktop, Sep 2025) also reported stronger average engagement on ChatGPT referrals (~15 minutes on site, ~12 pages/session, ~7% conversion on transactional sites) versus Google (~8 minutes, ~9 pages, ~5%). Treat those as landscape averages, not your forecast.
robots.txt patterns
Want live fetch, opt out of GPTBot training crawl (example):
User-agent: GPTBot
Disallow: /
User-agent: ChatGPT-User
Allow: /
Want neither:
User-agent: GPTBot
Disallow: /
User-agent: ChatGPT-User
Disallow: /
Want both allowed:
User-agent: GPTBot
Allow: /
User-agent: ChatGPT-User
Allow: /
Do not rely on a single User-agent: * Allow to express a nuanced OpenAI policy — be explicit per agent.
WAF and middleware gotchas
robots.txt can allow ChatGPT-User while Cloudflare Bot Fight Mode, AWS WAF bot scores, or custom middleware return 403. Symptoms:
- Browser fetch: 200
curlwith ChatGPT-User: 403 or challenge HTML- Marketing: "but robots says Allow"
Fix the allowlist for the user-agent (and IP ranges if the vendor publishes them) without disabling all bot protection. Details overlap with when your WAF blocks the bots.
Also watch for edge middleware that branches on "bot" heuristics and serves empty shells. If ChatGPT-User receives <div id="root"></div> and no brand text, you allowed a fetch of nothing (JavaScript rendering and AI crawlers).
How to test in fifteen minutes
- Read production
/robots.txt. Note GPTBot and ChatGPT-User rules separately. - Fetch homepage and one docs URL:
curl -sI -A "ChatGPT-User" https://www.example.com/docs curl -sL -A "ChatGPT-User" https://www.example.com/docs | head -n 50 - Repeat with
GPTBotif you care about that policy path. - Confirm the brand name and the answer content appear in the body.
- Log status codes in your CDN for both agents for 24 hours after changes.
Decision guide for stakeholders
| Business preference | GPTBot | ChatGPT-User |
|---|---|---|
| Fine with training + want answers | Allow | Allow |
| No training crawl; still want answer fetch | Disallow | Allow |
| Maximize privacy from OpenAI fetches | Disallow | Disallow |
| Unsure | Pause; do not ship a blunt Disallow: / under * | Read policies; decide in writing |
Put the decision in the runbook. The worst outcome is a rotating cast of contractors applying whichever Hacker News snippet is trending.
Honest ceilings
You cannot control OpenAI product UX, ranking inside ChatGPT, or whether a user's question triggers browsing at all. Pew Research (March 2025) showed how often AI Overviews appear and how rarely users click citations in that Google study (~1% of visits clicked a citation in the summary) — different product, same lesson: citation and click economics are harsh. Your job on-site is narrower: do not refuse the fetch you claim to want, and make the HTML worth fetching.
Gartner (Feb 2024 forecast) suggested traditional search volume may drop 25% by 2026 due to AI agents — forecast, not a measurement of your traffic. Use it for planning scenarios, not as proof that blocking ChatGPT-User is "safe because Search is dying."
How BrandKnown thinks about this
A readiness score that marks AI crawler access should distinguish answer-oriented agents from training-oriented ones when the product can. BrandKnown frames the score as website checks — including crawler access — not citation likelihood. Agency plans ($49/mo) are for teams that re-scan after robots/WAF changes; the free first scan is enough to see if you accidentally denied the wrong OpenAI agent.
Allow the bot that serves your policy. Document which one that is. Then stop treating "GPT" as a single switch.
