◀ All articles

AI crawlers

How to allowlist AI user agents in your WAF

August 5, 2026 · 4 min read

Cloudflare, AWS WAF, and generic bot scores — patterns that let answer crawlers through without opening every scanner on earth.


Your robots.txt welcomes answer crawlers. Your WAF still grades them as "automated" and drops them. That mismatch is one of the most common reasons AI never sees a brand that "did everything right" on paper.

Allowlisting is not "turn off security." It is telling the firewall which agents you already decided to let read public pages.

Why WAFs fight the bots you invited

Most WAFs score traffic with signals: user-agent, ASN, TLS fingerprint, request rate, JS challenge success. Official AI crawlers often look automated because they are automated. Bot Fight Mode, managed bot rules, and aggressive "block likely bots" policies will 403 or challenge them even when robots allows them.

LayerWhat it thinksWhat you want for public marketing pages
robots.txtPolicy for polite crawlersAllow answer agents you care about
WAF / CDNThreat scoreAllow those same agents without disabling all bot defense
Origin appAuth / middlewareDo not require login for public HTML

If you are still diagnosing a hard 403, walk how to fix 403 errors for AI bots first, then come back here for the allowlist patterns. Background on the agents themselves is in AI crawlers explained.

Decision table: what to allowlist

Agent (examples)Typical roleSensible default for public sites
GPTBotOpenAI crawling / training-related fetchAllow if you want OpenAI systems to read public pages; block if you refuse that use
ChatGPT-UserUser-initiated browsing fetchesAllow if you want live answers to reach your pages
PerplexityBotPerplexity fetch for answersAllow if you want Perplexity citations
Google-ExtendedGemini training (not classic Googlebot)Separate choice from Search; see product docs before blocking
Generic scrapers / unknown UAsUnknownKeep blocked or challenged

Do not confuse Googlebot (Search) with Google-Extended (training-related). Blocking one does not automatically configure the other. For the Google-Extended decision specifically, read Google-Extended: what blocking it does and does not do. Separating ChatGPT-User vs GPTBot prevents allowlisting the training agent while still starving live answers — or the reverse.

How to allowlist (practical patterns)

Cloudflare

  1. Confirm Bot Fight Mode or Super Bot Fight Mode is on — that is often the smoker. See also Cloudflare Bot Fight Mode vs the AI crawlers you want.
  2. Create a WAF custom rule (or exception) that skips bot fight / managed challenges when the user-agent matches the agents you allow and the path is public content.
  3. Prefer http.user_agent contains "..." patterns that tolerate version suffixes over brittle exact matches.
  4. Put allow/skip rules in an order that actually runs before the block.
  5. Optionally restrict by published bot IP ranges when the vendor provides them — UA-only rules are easier to spoof, but IP+UA is tighter when available.
  6. Purge cache, then verify with curl -I -A "...".

Example logic in plain language: if user-agent contains GPTBot or PerplexityBot or ChatGPT-User, skip Bot Fight / allow. Keep the rest of your bot rules for everyone else.

AWS WAF

  1. Identify which rule group returns 403 (AWS managed bot control, rate-based, or custom).
  2. Add a rule with higher priority that allows matching user-agents (or labels them so later rules skip).
  3. Scope the allow to hostnames and paths that are meant to be public — not /admin, not APIs that mutate data.
  4. Deploy to staging first if you have it; then production.
  5. Check sampled requests in AWS WAF logs for the bot UAs after deploy.

Generic / host panels (cPanel, Sucuri, Wordfence-style)

  1. Find "bot protection," "firewall," or "block empty user-agents / bad bots."
  2. Add exceptions for the exact strings your vendors document.
  3. Disable only the rule that blocks legitimate crawlers — not the entire firewall.
  4. Re-test from an external network, not only from your office IP.

Step checklist before you call it done

  1. robots.txt allows the agent on the URLs you care about.
  2. WAF allow/skip rule exists and is ordered correctly.
  3. No origin middleware challenges the same UA.
  4. curl with bot UA returns 200 and real HTML (not a challenge page).
  5. Logs show 2xx for that UA on homepage + one deep URL.
  6. You did not create a global "allow all bots" rule.

Mistakes that reopen the floodgates

  • Allowlisting bot or crawl as a substring — you just invited junk.
  • Disabling Bot Fight Mode entirely because one agent failed.
  • Copy-pasting allowlists from old blog posts with dead user-agent names.
  • Allowlisting on the CDN but forgetting a second WAF at the origin.
  • Testing only with browser DevTools, which never sends GPTBot.
  • Allowlisting staging but not production (or only the marketing subdomain while docs live elsewhere).

What allowlisting will not do

Allowlisting does not improve your readiness score's identity or evidence sections by itself. It only removes a false "blocked" access result so the rest of the checklist can run. Similarweb / TechCrunch (June 2025) reported AI platforms sending roughly 1.13B referrals to the top 1,000 sites that month (up 357% YoY), with ChatGPT accounting for more than 80% of those AI referrals — while Google Search still sent on the order of 191B. Access matters; it is not the whole game.

A BrandKnown scan takes about a minute and checks website readiness, not citation likelihood. After you allowlist, re-scan to confirm access findings cleared. Keep your other bot defenses on for the traffic you never invited.

See how your own site scores

One scan checks your homepage, robots.txt, llms.txt, About page and JSON-LD, then hands you the copy-paste fixes. Free, no account needed for the first run.

Keep reading