Confirm the 403, find whether robots or WAF is lying, allowlist the agent, re-test with curl. A numbered path.
You allow GPTBot and PerplexityBot in robots.txt. A week later your logs still show 403s. The crawler never saw your about page, your schema, or your carefully worded product description. From the assistant's side, you do not exist.
A 403 for an AI bot is usually not a content problem. It is an access problem: robots said yes, the edge said no. Fix the edge, then prove it with a real request.
What a 403 means (and what it does not)
A 403 means the server understood the request and refused it. It is not a soft miss. It is not "come back later." For answer crawlers that fetch live pages, a 403 is a closed door.
| Status | What the bot learns | Typical cause |
|---|---|---|
| 200 with HTML | Page is usable | Healthy path |
| 301/302 to HTTPS or apex | Follow and retry | Normal; watch loops |
| 403 | Explicitly blocked | WAF, bot score, IP reputation, auth |
| 404 | URL does not exist | Bad sitemap or dead link |
| 429 / 503 | Temporary throttle | Rate limits; often recoverable |
A BrandKnown readiness score is a website checklist, not a citation forecast. If access fails, identity and evidence checks never get a fair look.
Step-by-step: fix 403s for AI bots
1. Confirm the 403 yourself
Do not trust a dashboard screenshot alone. Reproduce the request with the bot's user-agent.
curl -I -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.0; +https://openai.com/gptbot" \
https://www.example.com/
Repeat for other agents you care about (for example PerplexityBot, ChatGPT-User). Save status, server, cf-ray or similar edge headers, and any www-authenticate clues.
If curl with a normal browser UA returns 200 and the bot UA returns 403, you have a bot-specific block — not a site outage.
2. Check robots.txt before you touch the WAF
Open https://www.example.com/robots.txt. Confirm the agent is not Disallow: / for the path you tested. Teams often fix the firewall while robots.txt still blocks training or browse agents. Confusing ChatGPT-User vs GPTBot is a common reason the "wrong" bot stays blocked after a "fix."
If robots already allows the agent, move on. Do not keep editing robots when the 403 is coming from Cloudflare, AWS WAF, or a host bot filter.
3. Decide: robots lying, or WAF lying?
| Signal | Likely culprit | Next move |
|---|---|---|
| Bot UA 403, browser UA 200, robots allows | WAF / bot fight / IP rules | Allowlist or lower score threshold for that agent |
| Bot UA 403 and robots Disallow | robots.txt | Allow the agent for public content |
| Everyone gets 403 | Auth, IP allowlist, geo block | Fix site-wide access first |
| Intermittent 403 | Challenge pages, rate limits | Check challenge logs and rate rules |
If you already know Cloudflare Bot Fight Mode is in play, see Cloudflare Bot Fight Mode vs the AI crawlers you want and when your WAF blocks the bots.
4. Allowlist the agent in the WAF
Prefer user-agent and known IP ranges from the vendor when they publish them. Prefer narrow allow rules over turning off all bot protection. A practical pattern for Cloudflare, AWS WAF, and generic scores is in how to allowlist AI user agents in your WAF.
Do not open every scanner on earth. Allow the answer crawlers you intend to welcome; keep training-only and abuse traffic under separate rules if that matches your policy.
5. Clear caches and re-test with curl
After the rule change, purge the edge cache for the homepage and one deep content URL. Re-run the same curl -I -A "..." commands. You want:
- HTTP 200 (or a clean redirect to a 200)
- HTML body that contains your brand name without waiting for JavaScript (see JavaScript rendering and AI crawlers)
If you still get a challenge interstitial (CAPTCHA HTML), the allowlist did not stick. Check rule order: a later "block high bot score" rule can override an earlier allow.
6. Spot-check logs for the next 24–48 hours
Look for the bot user-agents hitting 200 on the URLs that matter: homepage, about, product or service pages, and key articles. One successful homepage fetch is not enough if your sitemap points them at soft 404s.
Common mistakes
- Fixing only production while staging still blocks, then testing on staging.
- Allowlisting GPTBot but forgetting PerplexityBot (or the reverse).
- Matching user-agent with a too-strict exact string that breaks on version bumps.
- Assuming a 403 in Search Console is an AI-bot problem — Googlebot and answer bots are different pipelines.
- Declaring victory because robots.txt is green while the WAF still challenges.
How to verify without guessing
curlwith each target UA returns 200 on homepage and one article.- Response body includes your canonical brand string in the first HTML payload.
- Access logs show those UAs with 2xx, not 403, over a quiet day.
- Re-run a site scan after fixes. A BrandKnown scan (~60 seconds) is a website checklist — it can confirm crawler access checks improved; it will not promise more citations.
Honest ceiling
You can remove the 403. You cannot force ChatGPT, Perplexity, or AI Overviews to cite you afterward. Access is table stakes. Identity clarity, citable third-party pages, and useful on-page answers still have to do their jobs — and zero-click search remains common (SparkToro / Similarweb, Jan–Apr 2026, put Google zero-click near ~68%). Fix the door first. Then work the rooms behind it.
