◀ All articles

AI crawlers

Cloudflare Bot Fight Mode vs the AI crawlers you want

September 1, 2026 · 4 min read

Bot Fight Mode will 403 the agents your robots.txt just allowed. How to tell, and how to allowlist without disabling protection.


You updated robots.txt to allow ChatGPT-User and PerplexityBot. The CMS preview looked fine. Then an assistant fetched your pricing page and got a 403 — because Bot Fight Mode never read your robots file.

Cloudflare's bot mitigations sit in front of origin. They can challenge or block automated clients your robots.txt just welcomed. That is useful against scrapers. It is a silent foot-gun when you want answer engines to retrieve your pages.

What Bot Fight Mode is doing

Bot Fight Mode (and related Cloudflare bot products) score and challenge traffic that looks automated. AI answer crawlers are automated. Many will fail a JS challenge or land in a "definitely bot" bucket even when you intended to allow them.

robots.txt is advisory for well-behaved agents. A WAF 403 is not advisory. The agent never sees your carefully worded Organization schema if the edge returns an error page.

LayerWhat it controlsOverrides robots?
robots.txtPolite crawl permissionNo — it is not a firewall
Bot Fight / bot scoreChallenges, blocks, JS checksYes — can 403 allowed UAs
Custom WAF rulesExplicit allow/denyYes
Origin app authLogged-in HTMLYes — empty shell for bots
CDN cache of challenge pagesStale interstitialsYes — can linger after you "fix"

If you have been through WAF blocks for bots before, this is the same class of bug with a Cloudflare-shaped UI.

How to tell Bot Fight is the problem

  1. From a clean network, curl -A "ChatGPT-User" -I https://yoursite.com/important-page (use the real user-agent string you care about).
  2. Note status: 403, 429, or a challenge interstitial HTML are red flags; 200 with real body is the goal.
  3. Repeat with a normal browser UA. If browsers pass and AI UAs fail, you are in bot-mitigation territory — not a content bug.
  4. In Cloudflare: Security → Events (or Bot analytics). Filter for the path and look for Bot Fight / bot score actions.
  5. Compare to /robots.txt. If robots Allows the agent and Events shows a block, the edge is lying relative to your published policy.
  6. Try a second path (/about, /pricing). One allowed marketing URL does not prove the site is open.

Hypothetical: say you allow three answer crawlers in robots and Bot Fight still challenges two of them. Your "AI-ready" checklist is fiction until those Events rows go green.

Allowlist without turning protection off

You rarely need to disable Bot Fight Mode entirely. Prefer narrow exceptions.

ApproachWhen to useRisk
Skip/allow specific verified bots Cloudflare already knowsFirst stop if the agent appears in Cloudflare's bot listLow if scoped
WAF custom rule: allow by exact User-Agent (and IP ranges if published)Agents not in the default listUA spoofing — keep rules tight
Lower sensitivity on marketing paths only (/blog, /docs, /pricing)You must keep Fight Mode strict on /app and /apiMis-scoped paths
Turn Bot Fight off globallyLast resort, short window while debuggingHigh — scrapers return

Practical sequence:

  1. Inventory which AI user-agents you actually want (live answer fetchers vs training crawlers). See ChatGPT-User vs GPTBot and PerplexityBot allowlisting.
  2. Confirm robots.txt Allows those agents on the URLs that matter — a 2026 robots template helps keep groups readable.
  3. In Cloudflare, add allow exceptions for those agents (product UI names change — look for skip Bot Fight / allow verified bots / custom WAF allow).
  4. Re-test with curl using the exact UA.
  5. Leave login, checkout, and admin routes protected. Answer engines need public HTML, not your app shell.
  6. Document the rule IDs so the next security pass does not "clean up" your allowlist.

Next.js / Vercel note

If the site is on Vercel behind Cloudflare, you have two edges. A Cloudflare allowlist that still hits a Vercel firewall or middleware bot check will look "fixed" in one dashboard and broken in the other. Test the public hostname end-to-end. Our Next.js and Vercel crawler guide covers the app-layer half.

Mistakes that keep the 403 alive

Allowing in robots, blocking at the edge.
Robots is not a WAF rule.

Allowlisting only GPTBot when you meant ChatGPT-User.
Training vs live fetch are different agents. Confusing them is how teams "allow OpenAI" and still starve answers.

Challenges on HTML you need cited.
JS challenges that return empty or interstitial HTML fail non-browser crawlers even when the eventual destination is public.

Fixing production once, forgetting staging DNS.
Preview hosts often have stricter bot rules. Fine — just do not debug robots against staging and declare prod healthy.

Allowlisting by ASN once and never reviewing UA changes.
Vendors rotate strings. Recheck quarterly.

How to verify after the change

  • curl with each allowed UA against homepage, About, pricing, and one article.
  • Confirm 200 and that the brand name appears in the raw HTML (not only after hydration).
  • Watch Cloudflare Events for 24 hours for residual blocks on those paths.
  • Re-run a readiness-style crawl checklist. BrandKnown's ~60-second scan is built for website access and identity checks — not for promising citations.

Bot Fight Mode is not the enemy. Unexamined Bot Fight Mode on pages you invited answer crawlers to read is. Align the edge with the robots file, keep apps locked down, and re-test with the user-agent you claim to welcome.

See how your own site scores

One scan checks your homepage, robots.txt, llms.txt, About page and JSON-LD, then hands you the copy-paste fixes. Free, no account needed for the first run.

Keep reading