Cloudflare, AWS WAF, and generic bot scores — patterns that let answer crawlers through without opening every scanner on earth.
Your robots.txt welcomes answer crawlers. Your WAF still grades them as "automated" and drops them. That mismatch is one of the most common reasons AI never sees a brand that "did everything right" on paper.
Allowlisting is not "turn off security." It is telling the firewall which agents you already decided to let read public pages.
Why WAFs fight the bots you invited
Most WAFs score traffic with signals: user-agent, ASN, TLS fingerprint, request rate, JS challenge success. Official AI crawlers often look automated because they are automated. Bot Fight Mode, managed bot rules, and aggressive "block likely bots" policies will 403 or challenge them even when robots allows them.
| Layer | What it thinks | What you want for public marketing pages |
|---|---|---|
| robots.txt | Policy for polite crawlers | Allow answer agents you care about |
| WAF / CDN | Threat score | Allow those same agents without disabling all bot defense |
| Origin app | Auth / middleware | Do not require login for public HTML |
If you are still diagnosing a hard 403, walk how to fix 403 errors for AI bots first, then come back here for the allowlist patterns. Background on the agents themselves is in AI crawlers explained.
Decision table: what to allowlist
| Agent (examples) | Typical role | Sensible default for public sites |
|---|---|---|
| GPTBot | OpenAI crawling / training-related fetch | Allow if you want OpenAI systems to read public pages; block if you refuse that use |
| ChatGPT-User | User-initiated browsing fetches | Allow if you want live answers to reach your pages |
| PerplexityBot | Perplexity fetch for answers | Allow if you want Perplexity citations |
| Google-Extended | Gemini training (not classic Googlebot) | Separate choice from Search; see product docs before blocking |
| Generic scrapers / unknown UAs | Unknown | Keep blocked or challenged |
Do not confuse Googlebot (Search) with Google-Extended (training-related). Blocking one does not automatically configure the other. For the Google-Extended decision specifically, read Google-Extended: what blocking it does and does not do. Separating ChatGPT-User vs GPTBot prevents allowlisting the training agent while still starving live answers — or the reverse.
How to allowlist (practical patterns)
Cloudflare
- Confirm Bot Fight Mode or Super Bot Fight Mode is on — that is often the smoker. See also Cloudflare Bot Fight Mode vs the AI crawlers you want.
- Create a WAF custom rule (or exception) that skips bot fight / managed challenges when the user-agent matches the agents you allow and the path is public content.
- Prefer
http.user_agent contains "..."patterns that tolerate version suffixes over brittle exact matches. - Put allow/skip rules in an order that actually runs before the block.
- Optionally restrict by published bot IP ranges when the vendor provides them — UA-only rules are easier to spoof, but IP+UA is tighter when available.
- Purge cache, then verify with
curl -I -A "...".
Example logic in plain language: if user-agent contains GPTBot or PerplexityBot or ChatGPT-User, skip Bot Fight / allow. Keep the rest of your bot rules for everyone else.
AWS WAF
- Identify which rule group returns 403 (AWS managed bot control, rate-based, or custom).
- Add a rule with higher priority that allows matching user-agents (or labels them so later rules skip).
- Scope the allow to hostnames and paths that are meant to be public — not
/admin, not APIs that mutate data. - Deploy to staging first if you have it; then production.
- Check sampled requests in AWS WAF logs for the bot UAs after deploy.
Generic / host panels (cPanel, Sucuri, Wordfence-style)
- Find "bot protection," "firewall," or "block empty user-agents / bad bots."
- Add exceptions for the exact strings your vendors document.
- Disable only the rule that blocks legitimate crawlers — not the entire firewall.
- Re-test from an external network, not only from your office IP.
Step checklist before you call it done
- robots.txt allows the agent on the URLs you care about.
- WAF allow/skip rule exists and is ordered correctly.
- No origin middleware challenges the same UA.
curlwith bot UA returns 200 and real HTML (not a challenge page).- Logs show 2xx for that UA on homepage + one deep URL.
- You did not create a global "allow all bots" rule.
Mistakes that reopen the floodgates
- Allowlisting
botorcrawlas a substring — you just invited junk. - Disabling Bot Fight Mode entirely because one agent failed.
- Copy-pasting allowlists from old blog posts with dead user-agent names.
- Allowlisting on the CDN but forgetting a second WAF at the origin.
- Testing only with browser DevTools, which never sends
GPTBot. - Allowlisting staging but not production (or only the marketing subdomain while docs live elsewhere).
What allowlisting will not do
Allowlisting does not improve your readiness score's identity or evidence sections by itself. It only removes a false "blocked" access result so the rest of the checklist can run. Similarweb / TechCrunch (June 2025) reported AI platforms sending roughly 1.13B referrals to the top 1,000 sites that month (up 357% YoY), with ChatGPT accounting for more than 80% of those AI referrals — while Google Search still sent on the order of 191B. Access matters; it is not the whole game.
A BrandKnown scan takes about a minute and checks website readiness, not citation likelihood. After you allowlist, re-scan to confirm access findings cleared. Keep your other bot defenses on for the traffic you never invited.
