A 403 from Cloudflare beats a robots.txt rule every time. How to tell a bot-management block from a crawler problem, and how to allow the agents you chose.
You added the allow rules. You checked robots.txt renders. Weeks later nothing has changed, and the reason is that your robots.txt was never the thing making the decision.
A WAF or bot-management layer sits in front of your origin and answers requests before your application sees them. If it decides a request is a bot, it returns a 403, a challenge page, or a JavaScript interstitial — and it does that regardless of what your robots.txt politely says. The crawler never gets far enough to read the file.
Recognising it
The symptom is a mismatch between what you see in a browser and what a crawler reports. Reproduce it with a user-agent:
# What a normal browser gets
curl -sI https://yoursite.example
# What GPTBot gets
curl -sI -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.2; +https://openai.com/gptbot" https://yoursite.example
Compare the status lines. What each result means:
| Response | Diagnosis |
|---|---|
200 both times | Not a blocking problem. Look at rendering or robots.txt next. |
403 for the bot only | Bot management. This post. |
503 with a cf-mitigated header | Cloudflare challenge — same class of problem. |
200 but a tiny body with Just a moment... | A JS interstitial. Fatal for a crawler that does not execute JavaScript. |
429 | Rate limiting, not blocking. Different fix — usually raising a threshold. |
| Times out for the bot only | Often a firewall dropping rather than rejecting. |
A 403 that only appears for one user-agent is close to conclusive.
Where the rule usually lives
Cloudflare. Two places. Security → Bots has "AI Scrapers and Crawlers" as a one-click block, which is on by default for some plans — this is the most common cause we see, and many site owners do not know it was enabled. Separately, Security → WAF → Custom rules may have a hand-written user-agent rule from years ago. Check both. Cloudflare also publishes a verified-bots list, and allowing a category is safer than allowing a user-agent string.
AWS WAF. The AWSManagedRulesBotControlRuleSet managed rule group. Its CategoryAI and generic bot labels catch these agents. You add an allow rule scoped by label rather than disabling the group.
Akamai, Fastly, Imperva. All have equivalent bot-management products with a category for AI crawlers, all with the same shape of fix.
Your own server. Do not overlook a .htaccess or nginx if ($http_user_agent ~* ...) block added by a previous developer during a scraping incident. These outlive the incident and everyone who remembers it.
Your hosting platform. Some managed hosts apply bot rules you do not administer. If you cannot find the rule, ask support whether one exists above you.
Deciding what to allow
Do not simply turn bot protection off. Allow the specific agents you decided you wanted, and leave the rest of the protection in place.
A reasonable position for most businesses: allow OAI-SearchBot, ChatGPT-User and PerplexityBot, since those fetch in order to answer and cite. Decide separately about GPTBot, ClaudeBot and CCBot, which take content for training. And note that a WAF block and a robots.txt block are different instruments — use robots.txt to state intent to well-behaved crawlers, and the WAF only where you need actual enforcement.
Verify against user-agent, not IP alone
User-agent strings are trivially forged. Anyone can send a request claiming to be GPTBot, and if your allow rule is a naive string match you have just built a bypass for your own bot protection.
The operators publish IP ranges for exactly this reason — OpenAI, Anthropic and Perplexity each list theirs, and Cloudflare's verified-bot categories do the verification for you. Match on user-agent and source range, or use the platform's verified list. A string-only allow rule is worse than no rule.
After you change it
Re-run the curl check with each user-agent you allowed. Expect 200 and a real body, not a challenge page.
Then wait. Crawlers back off from hosts that returned 403 repeatedly, and re-crawl on their own schedule — days to weeks, not minutes. There is no submit-for-reconsideration button. This is one more reason to fix access before spending anything on content: the feedback loop is slow, and everything downstream is blocked behind it.
On the top plan we fetch your homepage with the GPTBot user-agent during a scan and report exactly what came back — status, body length, whether it was an empty shell, and a specific hint when the response looks like a WAF 403. It is the same check as the curl above; the value is that it runs alongside everything else and tells you which of your findings are downstream of a block.
