Perplexity citations need fetch access. Here is the user-agent, the robots rule, and the WAF checks that still block you.
Perplexity shows citations. That is the whole product vibe. If PerplexityBot cannot fetch your page, you are arguing about citation strategy for a URL the system never successfully retrieved.
Allowing PerplexityBot is straightforward in robots.txt and easy to undo with a WAF. This guide covers the user-agent, the robots rule, and the checks that still block you after you "allowlisted" everything.
What you are allowing
PerplexityBot is Perplexity's published crawler identity for retrieving web content used in their answer experience. Exact token spelling and any companion agents can change — confirm against Perplexity's current docs before production.
Allowing it means: public pages may be fetched by that agent when Perplexity's systems request them. It does not mean every answer in your category will cite you. It does not mean your competitors stop getting cited.
For how brands show up in assistants generally, see how assistants choose brands. This post is the access plumbing.
robots.txt rule
Minimal allow:
User-agent: PerplexityBot
Allow: /
If you use broad Disallows under User-agent: *, remember that more specific user-agent blocks are matched separately — but path Disallows under * can still apply depending on how you structure the file. Prefer an explicit PerplexityBot group. Keep private paths disallowed consistently:
User-agent: PerplexityBot
Disallow: /admin/
Disallow: /app/
Allow: /
Do not "open the floodgates" by disabling all bot protection globally. You are allowlisting one agent (and related infrastructure), not inviting every scanner on earth.
Pair this with the fuller robots.txt template for AI 2026.
Decision table: allow vs tighten
| Situation | Recommendation | Why |
|---|---|---|
| Public marketing site, want citations | Allow PerplexityBot on public content | Citations need fetch access |
| Docs partly public | Allow public docs; Disallow auth-only paths | Avoid serving login walls as "content" |
| Highly sensitive unpublished research | Disallow those paths (or keep them authed) | robots is not auth |
| Already drowning in scrape traffic | Allowlist agent + rate limits; do not confuse with blanket allow | Precision over panic |
| Legal hold / temporary takedown | Disallow + remove content; do not rely on robots alone | robots is voluntary |
WAF checks that still block you
After robots says Allow, verify the real response:
- Bot Fight / managed bot scores — challenge or 403 for datacenter UAs.
- Geo or ASN blocks — overly broad rules.
- Rate limiting — shared crawler IPs trip thresholds.
- JS challenge interstitial — body is a puzzle page, not your article.
- Wrong host — apex allowlisted, www not (or vice versa).
Test:
curl -sI -A "PerplexityBot" https://www.example.com/pricing
curl -sL -A "PerplexityBot" https://www.example.com/pricing | head -n 40
You want status 200 and HTML that includes your brand name, plan facts, and main answer content. If you see a CAPTCHA title or empty root div, fix infrastructure before debating schema.
More on this failure class: when your WAF blocks the bots.
Content worth fetching once access works
Perplexity-style citation favors pages that look like evidence:
- Clear definitions and comparison tables
- Pricing that states numbers plainly (Product schema for SaaS can reinforce offers)
- Docs with steps (HowTo schema when truly procedural)
- Third-party profiles that match your entity (sameAs links)
SparkToro / Similarweb (Jan–Apr 2026) estimated ~68% of Google searches were zero-click. Answer engines that show citations are part of how brands still earn attention when classic blue links get fewer clicks. Access plus citable pages is the boring sequence that works.
How to allowlist without lowering the drawbridge
- Publish explicit
PerplexityBotAllow for public paths in robots.txt. - In Cloudflare (or equivalent), add a WAF exception for that user-agent or vendor IP ranges if published — prefer vendor guidance over random forum IP lists.
- Keep
/admin, app dashboards, and cart endpoints protected. - Monitor 403 rates for PerplexityBot in CDN logs for a week.
- Re-test after every Bot Fight or WAF rule change (these regress silently).
- Re-scan site readiness so "crawler access" findings clear.
BrandKnown's checks treat crawler access as part of the readiness score (website checks, not a prediction that Perplexity will cite you). A ~60-second scan is useful after WAF edits; Agency ($49/mo) fits agencies repeating this across clients.
Measuring without fooling yourself
- Watch referrals from Perplexity hostnames in analytics when volume exists.
- Do not declare victory from one anecdotal citation screenshot.
- Similarweb / TechCrunch (June 2025) AI referral totals (~1.13B to top 1,000 sites; ChatGPT >80% of AI referrals) show ChatGPT still dominated AI referrals in that snapshot — Perplexity may be smaller in your logs even when citations appear. Size expectations accordingly.
- Pew Research (March 2025) click behavior around AI Overviews (8% vs 15% traditional CTR with/without summaries; ~1% citation clicks) is Google-specific, but it is a reminder that being cited ≠ being clicked. Track both if you care about traffic.
Common mistakes
- Allowing PerplexityBot while blocking all AI via a wildcard rule elsewhere
- Serving different content to "bots" that strips the citable paragraph
- Assuming FAQ schema alone gets you cited (FAQ schema)
- Updating robots on staging only
- Forgetting soft 404s — a 200 "not found" page can be cited as nonsense
Honest ceiling
You cannot force Perplexity to prefer your domain over a strong aggregator. You can stop failing the fetch pretest. Do that first, then improve entity clarity and evidence pages, then measure referrals with a calm sample size.
Allowlist the agent, confirm 200s with real HTML, keep private surfaces closed, and treat every WAF change as a reason to re-test. That is the whole floodgate strategy: one controlled gate, not an open harbor.
