◀ All articles

AI crawlers

Every AI crawler in your robots.txt, explained

August 7, 2026 · 4 min read

GPTBot, ChatGPT-User, OAI-SearchBot, PerplexityBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot — what each one is for and what blocking it costs.


Your robots.txt probably contains rules for user-agents nobody on your team chose. They arrive in CMS defaults, in "block AI scrapers" plugins, and in copy-pasted snippets from posts written during a news cycle. The result is that a lot of sites are blocking the crawler that would have cited them while cheerfully allowing the one that trains on them.

So: what each agent is, who runs it, and what blocking it actually costs.

The agents

User-agentOperatorWhat it doesBlocking it means
GPTBotOpenAIBulk crawl, historically for training and model groundingYour content is excluded from that corpus. No effect on live ChatGPT browsing.
ChatGPT-UserOpenAIFetches a page because a user asked, in-sessionChatGPT cannot open your page when someone asks it to. This is a user request, not a crawl.
OAI-SearchBotOpenAIBuilds the index behind ChatGPT searchYou are not in the index that ChatGPT search results come from. This is the expensive one.
PerplexityBotPerplexityIndexes for Perplexity's answer engineYou will not be cited in Perplexity answers.
ClaudeBotAnthropicCrawls for ClaudeExcluded from Claude's crawl.
Google-ExtendedGoogleControl token, not a crawler. Governs Gemini and AI grounding useYour pages stay in Google Search and are excluded from Gemini training/grounding. Search ranking is unaffected.
Applebot-ExtendedAppleSame idea, for Apple IntelligenceExcluded from Apple's generative use. Applebot for Siri and Spotlight is separate.
CCBotCommon CrawlOpen web archive many datasets are built fromExcluded from a corpus that feeds a great many downstream models and research projects.

Two of these are worth reading twice.

Google-Extended is not a crawler. There is no bot with that name. It is a token you put in robots.txt to opt out of generative use while staying in search. Blocking it costs you nothing in rankings, which makes it the safest "no" on the list.

ChatGPT-User is a person. When someone pastes your URL into ChatGPT and asks what it says, this is the agent that fetches it. Blocking it does not stop training — it stops a prospective customer from reading your page. It is close to the worst block on the list and it is very common, because a lot of blanket rules catch it.

The distinction that matters

Split them into two groups:

  • Training crawlers — GPTBot, ClaudeBot, CCBot, and the -Extended tokens. They take your content into a corpus. The benefit to you is indirect and slow, arriving whenever a model is next trained.
  • Answer and citation crawlers — OAI-SearchBot, PerplexityBot, ChatGPT-User. They fetch your page in order to answer a question now, and they name sources when they do.

Blocking the first group is a defensible business decision — publishers make it every day, for good reasons. Blocking the second group is almost always an accident, and it is the accident that costs you visibility. We work through the decision here.

Writing the rules

robots.txt matches user-agents by longest match and does not merge groups. That means a specific group for GPTBot replaces the * group entirely for that agent — it does not add to it. Sites get caught by this constantly.

An additive allow, one group per agent:

User-agent: OAI-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Google-Extended
Disallow: /

User-agent: CCBot
Disallow: /

That is a coherent position: let the answer engines read and cite you, keep your content out of the bulk training corpora. Change the halves to suit your own view; the shape is the point.

Three ways to get it wrong:

  1. Disallow: with nothing after it means allow everything. Disallow: / blocks everything. One character apart, opposite meanings.
  2. Assuming your * rules still apply to an agent that has its own group. They do not.
  3. Putting the file anywhere but the root. It must be at /robots.txt, served as plain text, with a 200.

The bigger caveat

robots.txt is a request, not a control. Well-behaved crawlers honour it; the ones you are most worried about often do not. And it works the other way too — an Allow rule does not help if your CDN is returning 403 to the agent before the file is ever consulted. That mismatch is common enough to deserve its own post.

If you want a hard block, it belongs at the WAF. If you want an honest signal of intent, it belongs here.

See how your own site scores

One scan checks your homepage, robots.txt, llms.txt, About page and JSON-LD, then hands you the copy-paste fixes. Free, no account needed for the first run.

Keep reading