◀ All articles

AI crawlers

Should you block AI crawlers?

August 4, 2026 · 4 min read

There is a real answer, and it is not the same for every bot. Split them into training crawlers and answer crawlers, then decide once for each group.


"Block all AI bots" and "allow everything" are both defaults dressed up as decisions. The useful version of this question is asked once per group, and it has different answers depending on what your business is.

Split the bots first

Nothing else in this post makes sense until you have separated two things that get lumped together.

Training crawlers take your content into a corpus that a future model learns from. GPTBot, ClaudeBot, CCBot, and the Google-Extended and Applebot-Extended tokens. You get no attribution, no link, and no traffic. What you get, eventually, is a model that knows your company exists.

Answer crawlers fetch your page to answer a question being asked right now, and they cite what they used. OAI-SearchBot, PerplexityBot, ChatGPT-User. This is the modern equivalent of being in the index. Full list and what each does.

The honest summary: blocking training crawlers is a real trade-off with arguments on both sides. Blocking answer crawlers is almost always self-harm.

The case for allowing training crawlers

You are not in an assistant's answer if the assistant has never heard of you. For a company that is not a household name, being in the training corpus is how a model comes to know your category placement at all. Live retrieval helps you at answer time, but a model that has absorbed your existence is a model that thinks to look for you.

If your content is marketing — your homepage, product pages, docs, an educational blog like this one — the content is for being read and repeated. Blocking a crawler from your marketing pages is like refusing to hand out business cards.

The case for blocking them

If your content is the product, the calculus flips entirely. A publisher, a research firm, a course business, a documentation-as-a-moat company: a model that has absorbed your archive can answer the questions people used to pay you to answer. There is no attribution and no compensation, and "exposure" does not pay for a newsroom.

The leverage argument is real too. Several large publishers blocked first and negotiated licensing deals afterwards. That option only exists if you have content worth licensing and are prepared to hold the line.

And there are cases where it is not even a business decision: content you licensed from someone else and have no right to sublicense, user-generated content where the community expects it not to be scraped, or anything under a regulatory obligation.

The decision, by business type

You areTraining crawlersAnswer crawlers
B2B SaaS, agency, local service businessAllowAllow
E-commerceAllowAllow — emphatically, product data is what gets cited
Publisher, media, researchBlock or licenseAllow — you still want the citation and the click
Course or paid-content businessBlock the paid content, allow the marketing pagesAllow the marketing pages
Community platform with UGCDepends on what you promised your usersAllow
Anyone with a compliance obligationAsk your lawyer, not a blogAsk your lawyer

The pattern in that table: the answer column says "allow" almost everywhere, and the disagreement is entirely in the training column.

The half-measure worth knowing

You do not have to answer at the whole-site level. robots.txt takes paths:

User-agent: GPTBot
Allow: /
Disallow: /research/
Disallow: /members/

Marketing pages open, archive closed. For most publishers this is a better position than a blanket block, because it keeps the model aware that you exist and that you are the authority in your field, without handing over the thing you sell.

What blocking does not do

Be clear about the limits, because people over-estimate this file badly.

  • It does not remove you from a model already trained. If your content is in GPT-4-era training data, blocking today changes nothing about that. It affects future crawls only.
  • It does not stop bad actors. robots.txt is a request. Scrapers that ignore it were always going to ignore it. If you need enforcement, that is a WAF rule, and it is a different tool.
  • It does not affect Google Search ranking — as long as you block Google-Extended and not Googlebot. Confusing the two is a genuinely expensive mistake.
  • It does not stop a human copy-pasting your page into a chat window. Nothing does.

If you decide to block

Do it deliberately, write down why, and put a comment in the file so the next person does not undo it or double it:

# Training corpora: excluded by decision of 2026-03, see /internal/ai-policy
User-agent: GPTBot
Disallow: /

# Answer engines: allowed, we want the citation
User-agent: OAI-SearchBot
Allow: /

The thing to avoid is the accidental version: a plugin that blocked eleven agents you never evaluated, including the two that would have named you in an answer this morning.

See how your own site scores

One scan checks your homepage, robots.txt, llms.txt, About page and JSON-LD, then hands you the copy-paste fixes. Free, no account needed for the first run.

Keep reading