There is a real answer, and it is not the same for every bot. Split them into training crawlers and answer crawlers, then decide once for each group.
"Block all AI bots" and "allow everything" are both defaults dressed up as decisions. The useful version of this question is asked once per group, and it has different answers depending on what your business is.
Split the bots first
Nothing else in this post makes sense until you have separated two things that get lumped together.
Training crawlers take your content into a corpus that a future model learns from. GPTBot, ClaudeBot, CCBot, and the Google-Extended and Applebot-Extended tokens. You get no attribution, no link, and no traffic. What you get, eventually, is a model that knows your company exists.
Answer crawlers fetch your page to answer a question being asked right now, and they cite what they used. OAI-SearchBot, PerplexityBot, ChatGPT-User. This is the modern equivalent of being in the index. Full list and what each does.
The honest summary: blocking training crawlers is a real trade-off with arguments on both sides. Blocking answer crawlers is almost always self-harm.
The case for allowing training crawlers
You are not in an assistant's answer if the assistant has never heard of you. For a company that is not a household name, being in the training corpus is how a model comes to know your category placement at all. Live retrieval helps you at answer time, but a model that has absorbed your existence is a model that thinks to look for you.
If your content is marketing — your homepage, product pages, docs, an educational blog like this one — the content is for being read and repeated. Blocking a crawler from your marketing pages is like refusing to hand out business cards.
The case for blocking them
If your content is the product, the calculus flips entirely. A publisher, a research firm, a course business, a documentation-as-a-moat company: a model that has absorbed your archive can answer the questions people used to pay you to answer. There is no attribution and no compensation, and "exposure" does not pay for a newsroom.
The leverage argument is real too. Several large publishers blocked first and negotiated licensing deals afterwards. That option only exists if you have content worth licensing and are prepared to hold the line.
And there are cases where it is not even a business decision: content you licensed from someone else and have no right to sublicense, user-generated content where the community expects it not to be scraped, or anything under a regulatory obligation.
The decision, by business type
| You are | Training crawlers | Answer crawlers |
|---|---|---|
| B2B SaaS, agency, local service business | Allow | Allow |
| E-commerce | Allow | Allow — emphatically, product data is what gets cited |
| Publisher, media, research | Block or license | Allow — you still want the citation and the click |
| Course or paid-content business | Block the paid content, allow the marketing pages | Allow the marketing pages |
| Community platform with UGC | Depends on what you promised your users | Allow |
| Anyone with a compliance obligation | Ask your lawyer, not a blog | Ask your lawyer |
The pattern in that table: the answer column says "allow" almost everywhere, and the disagreement is entirely in the training column.
The half-measure worth knowing
You do not have to answer at the whole-site level. robots.txt takes paths:
User-agent: GPTBot
Allow: /
Disallow: /research/
Disallow: /members/
Marketing pages open, archive closed. For most publishers this is a better position than a blanket block, because it keeps the model aware that you exist and that you are the authority in your field, without handing over the thing you sell.
What blocking does not do
Be clear about the limits, because people over-estimate this file badly.
- It does not remove you from a model already trained. If your content is in GPT-4-era training data, blocking today changes nothing about that. It affects future crawls only.
- It does not stop bad actors.
robots.txtis a request. Scrapers that ignore it were always going to ignore it. If you need enforcement, that is a WAF rule, and it is a different tool. - It does not affect Google Search ranking — as long as you block
Google-Extendedand notGooglebot. Confusing the two is a genuinely expensive mistake. - It does not stop a human copy-pasting your page into a chat window. Nothing does.
If you decide to block
Do it deliberately, write down why, and put a comment in the file so the next person does not undo it or double it:
# Training corpora: excluded by decision of 2026-03, see /internal/ai-policy
User-agent: GPTBot
Disallow: /
# Answer engines: allowed, we want the citation
User-agent: OAI-SearchBot
Allow: /
The thing to avoid is the accidental version: a plugin that blocked eleven agents you never evaluated, including the two that would have named you in an answer this morning.
