◀ All articles

SEO

Soft 404s that confuse AI crawlers

August 1, 2026 · 4 min read

A 200 with "page not found" text is worse than a real 404. How to find them and what status to return instead.


A real 404 is honest: this URL does not exist. A soft 404 is a 200 OK that serves "Page not found," an empty shell, or a homepage clone. Humans shrug and hit Back. Crawlers index confusion. Answer engines may treat the empty page as a valid source candidate — or waste a fetch on nothing useful.

If your CMS loves soft 404s, AI readiness work on those URLs is theater.

Why soft 404s are worse than hard 404s for AEO

ResponseMachine-readable meaningTypical AEO impact
404 / 410GoneClear; remove from sitemap
301 to a real pageMovedOK if target is correct
200 with not-found copy"Success" with failure contentConfuses discovery and citations
200 with near-empty app shell"Success" without substanceEspecially bad for non-JS crawlers

JavaScript-heavy apps often return 200 with a spinner root and no brand copy in the first HTML. That is a cousin of the soft 404 problem — see JavaScript rendering and AI crawlers and SSR vs CSR for AEO.

How soft 404s show up in the wild

  • Deleted blog posts that render the theme's 404 component without changing status
  • Unknown paths routed to the SPA homepage
  • "No results" search and filter pages that look like thin content with 200
  • Staging error templates accidentally mapped to production unknown routes
  • Localized paths that fail translation lookup but still 200

Northstar Analytics, our fictional B2B example, once shipped a docs migration where old /guides/foo URLs rendered a React "Not found" view with status 200. Perplexity-style fetchers would have received a polite shrug packaged as success.

How to find soft 404s

  1. Sample the sitemap. Pull 50 random URLs from your XML sitemap. curl -I for status, then curl for body snippets.
  2. Hit known-bad paths. Request /this-should-not-exist-brandknown-test-404 and note status + body.
  3. Search Console / Bing. Coverage reports often label soft 404 suspects — treat them as leads, not gospel.
  4. Log scrapes. Find URLs with 200 responses whose HTML contains phrases like "not found," "page doesn’t exist," or "keep browsing."
  5. Bot UA check. Repeat a few fetches with GPTBot or PerplexityBot user-agents in case your edge serves different bodies to bots (fix 403s if they never arrive).
curl -s -o /tmp/body.html -w "%{http_code}" \
  -A "Mozilla/5.0 (compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)" \
  https://www.example.com/old-deleted-post
# Inspect /tmp/body.html for not-found copy when code is 200

What status to return instead

CaseReturnAlso do
Truly gone, no replacement404 or 410Remove from sitemap; update internal links
Moved permanently301Canonical and sitemap point at target
Temporarily missing503 with Retry-After if appropriateDo not soft-200
Unauthorized401/403Not a not-found problem
Gone but you want a helpful HTML bodyStill 404; custom 404 template is fineHelpful body ≠ 200

Custom 404 pages are good UX. Keep them as 404. The status code is the contract.

Step-by-step fix in a CMS or SPA

  1. Map unknown routes to a true 404 status at the server or edge — not only a client-side route.
  2. For SPAs on Vercel/Netlify/nginx, configure fallback so missing paths do not rewrite to index.html with 200 unless you intentionally handle status in a server function. Next.js teams should also skim serving AI crawlers on Next.js and Vercel.
  3. Purge CDN cache for fixed URLs (cached soft 404s linger).
  4. Delete soft-404 URLs from sitemaps and resubmit.
  5. Add a monitor: weekly request to a random invented path must not return 200 with a full site chrome and "not found" only in a React root.
  6. Re-check internal links and ads that still point at deleted paths — send them to successors with 301s when you have a true replacement.

Mistakes

  • "We have a nice 404 page" while curl -I shows 200
  • Redirecting every unknown path to the homepage (soft 404's aggressive cousin)
  • Leaving soft 404s in the sitemap for months after a migration
  • Fixing HTML copy but not status codes
  • Assuming Google's soft 404 detection means every other bot agrees — still fix the status
  • Treating a soft 404 as an SEO "keep the link juice" hack by forcing 200

How to verify

TestExpected
Invented path404/410
Deleted post URL404/410 or 301 to successor
Live article200 with real content in raw HTML
Sitemap sampleNo not-found bodies
Bot UA on deleted URLSame honest status, not a special soft 200

Honest ceiling

Cleaning soft 404s will not guarantee citations. Gartner (Feb 2024 forecast) suggested traditional search volume may drop 25% by 2026 due to AI agents — a forecast, not a measured fact — which is one reason teams care about answer surfaces at all. Soft 404 cleanup is still table stakes: you cannot be the third source in an AI Overview (Pew Research, March 2025: 88% of summaries cited 3+ sources) if the URL you hoped to cite is an empty 200.

BrandKnown's score is a website checklist. Soft 404s undermine access and evidence quality. Fix status codes, re-scan, and keep the sitemap honest.

See how your own site scores

One scan checks your homepage, robots.txt, llms.txt, About page and JSON-LD, then hands you the copy-paste fixes. Free, no account needed for the first run.

Keep reading