◀ All articles

AI crawlers

How to fix 403 errors for AI bots, step by step

August 6, 2026 · 4 min read

Confirm the 403, find whether robots or WAF is lying, allowlist the agent, re-test with curl. A numbered path.


You allow GPTBot and PerplexityBot in robots.txt. A week later your logs still show 403s. The crawler never saw your about page, your schema, or your carefully worded product description. From the assistant's side, you do not exist.

A 403 for an AI bot is usually not a content problem. It is an access problem: robots said yes, the edge said no. Fix the edge, then prove it with a real request.

What a 403 means (and what it does not)

A 403 means the server understood the request and refused it. It is not a soft miss. It is not "come back later." For answer crawlers that fetch live pages, a 403 is a closed door.

StatusWhat the bot learnsTypical cause
200 with HTMLPage is usableHealthy path
301/302 to HTTPS or apexFollow and retryNormal; watch loops
403Explicitly blockedWAF, bot score, IP reputation, auth
404URL does not existBad sitemap or dead link
429 / 503Temporary throttleRate limits; often recoverable

A BrandKnown readiness score is a website checklist, not a citation forecast. If access fails, identity and evidence checks never get a fair look.

Step-by-step: fix 403s for AI bots

1. Confirm the 403 yourself

Do not trust a dashboard screenshot alone. Reproduce the request with the bot's user-agent.

curl -I -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.0; +https://openai.com/gptbot" \
  https://www.example.com/

Repeat for other agents you care about (for example PerplexityBot, ChatGPT-User). Save status, server, cf-ray or similar edge headers, and any www-authenticate clues.

If curl with a normal browser UA returns 200 and the bot UA returns 403, you have a bot-specific block — not a site outage.

2. Check robots.txt before you touch the WAF

Open https://www.example.com/robots.txt. Confirm the agent is not Disallow: / for the path you tested. Teams often fix the firewall while robots.txt still blocks training or browse agents. Confusing ChatGPT-User vs GPTBot is a common reason the "wrong" bot stays blocked after a "fix."

If robots already allows the agent, move on. Do not keep editing robots when the 403 is coming from Cloudflare, AWS WAF, or a host bot filter.

3. Decide: robots lying, or WAF lying?

SignalLikely culpritNext move
Bot UA 403, browser UA 200, robots allowsWAF / bot fight / IP rulesAllowlist or lower score threshold for that agent
Bot UA 403 and robots Disallowrobots.txtAllow the agent for public content
Everyone gets 403Auth, IP allowlist, geo blockFix site-wide access first
Intermittent 403Challenge pages, rate limitsCheck challenge logs and rate rules

If you already know Cloudflare Bot Fight Mode is in play, see Cloudflare Bot Fight Mode vs the AI crawlers you want and when your WAF blocks the bots.

4. Allowlist the agent in the WAF

Prefer user-agent and known IP ranges from the vendor when they publish them. Prefer narrow allow rules over turning off all bot protection. A practical pattern for Cloudflare, AWS WAF, and generic scores is in how to allowlist AI user agents in your WAF.

Do not open every scanner on earth. Allow the answer crawlers you intend to welcome; keep training-only and abuse traffic under separate rules if that matches your policy.

5. Clear caches and re-test with curl

After the rule change, purge the edge cache for the homepage and one deep content URL. Re-run the same curl -I -A "..." commands. You want:

If you still get a challenge interstitial (CAPTCHA HTML), the allowlist did not stick. Check rule order: a later "block high bot score" rule can override an earlier allow.

6. Spot-check logs for the next 24–48 hours

Look for the bot user-agents hitting 200 on the URLs that matter: homepage, about, product or service pages, and key articles. One successful homepage fetch is not enough if your sitemap points them at soft 404s.

Common mistakes

  • Fixing only production while staging still blocks, then testing on staging.
  • Allowlisting GPTBot but forgetting PerplexityBot (or the reverse).
  • Matching user-agent with a too-strict exact string that breaks on version bumps.
  • Assuming a 403 in Search Console is an AI-bot problem — Googlebot and answer bots are different pipelines.
  • Declaring victory because robots.txt is green while the WAF still challenges.

How to verify without guessing

  1. curl with each target UA returns 200 on homepage and one article.
  2. Response body includes your canonical brand string in the first HTML payload.
  3. Access logs show those UAs with 2xx, not 403, over a quiet day.
  4. Re-run a site scan after fixes. A BrandKnown scan (~60 seconds) is a website checklist — it can confirm crawler access checks improved; it will not promise more citations.

Honest ceiling

You can remove the 403. You cannot force ChatGPT, Perplexity, or AI Overviews to cite you afterward. Access is table stakes. Identity clarity, citable third-party pages, and useful on-page answers still have to do their jobs — and zero-click search remains common (SparkToro / Similarweb, Jan–Apr 2026, put Google zero-click near ~68%). Fix the door first. Then work the rooms behind it.

See how your own site scores

One scan checks your homepage, robots.txt, llms.txt, About page and JSON-LD, then hands you the copy-paste fixes. Free, no account needed for the first run.

Keep reading