◀ All articles

SEO

Serving AI crawlers on Next.js and Vercel

August 31, 2026 · 4 min read

SSR, static HTML, and middleware gotchas that make a Next site look empty to non-JS crawlers.


A Next.js site can look perfect in Chrome and empty to a crawler that does not run your client bundle. Assistants that fetch HTML once — no hydration, no click — will quote whatever arrived in the first response. If that response is a spinner shell, your brand never entered the room.

This is the SSR/CSR problem with a Vercel-shaped deployment checklist.

Why Next apps surprise teams

Next.js supports static pages, server-rendered pages, and client-heavy islands. Vercel hosts them well. None of that helps if the route that holds your positioning, pricing, or About copy is client-only.

AI answer crawlers vary, but many behave closer to "fetch HTML" than "full Chrome." If the brand string, plan names, or definitional sentence only appear after useEffect, you are invisible for that fetch. Related background: JavaScript rendering and AI crawlers and the broader SSR vs CSR decision table.

Route patternWhat the first HTML often containsAEO risk
Static / SSGFull contentLow
SSR (server components / getServerSideProps-style)Full contentLow
CSR page ('use client' fetching CMS in effect)Empty layout + loading UIHigh
Middleware-gated "bot challenge"Interstitial or 403High
Preview-only data on prod hostnameWrong or emptyMedium
Edge personalization that strips defaultsGeneric or blank heroMedium–high

Decision table: how to serve the pages that matter

Page typePreferAcceptable fallback
Homepage, About, PricingSSR or staticPrerender / ISR
Docs and blog postsStatic or SSRISR with real lastmod
App dashboard behind authCSR fineDo not expect citations
Marketing landing with personalizationStatic default + enhanceNever personalize away the brand name

If you only fix three routes, fix / , /about , and /pricing (or your equivalents). Those are the pages assistants reuse when someone asks what you are and what you cost.

Step-by-step: make a Next/Vercel site fetchable

  1. Pick the canonical host. Apex vs www should 301 one way. Split hosts split entities.
  2. Confirm important routes are not client-only. In the App Router, prefer Server Components for marketing content. In the Pages Router, avoid shipping critical copy only via client fetch.
  3. View page source (not DevTools Elements). Search for your brand name and one unique sentence. If it is missing from source, crawlers that skip JS miss it too.
  4. Check middleware.ts. Bot blocking, geo walls, or "soft auth" that returns empty HTML will hit AI user-agents. Allow public marketing paths.
  5. Check Vercel Deployment Protection / password walls on production. Preview protection is fine; prod should be publicly readable where you want citations.
  6. Align Cloudflare or other CDN in front of Vercel — Bot Fight can 403 agents you allowed in robots. See Cloudflare Bot Fight vs AI crawlers.
  7. Ship Organization (and Article) JSON-LD in the server HTML, not injected only after mount.
  8. Re-test with curl using a plain UA and an AI UA you care about.
  9. Watch logs for 403/429 on those UAs for a day after deploy.
curl -sL https://www.example.com/about | head -n 80
curl -sL -A "ChatGPT-User" -o /tmp/about.html -w "%{http_code}" https://www.example.com/about

You want 200 and readable text in /tmp/about.html.

Vercel-specific gotchas

Deployment Protection on production.
A login wall means assistants never see the page. Keep protection on previews.

Edge Config / middleware experiments.
Feature flags that serve an empty shell to "unknown" clients will treat crawlers as unknown.

ISR that never rebuilt.
Stale is better than empty — but a forever-stale pricing page still misleads. When plans change, invalidate.

Image-only hero text.
If the H1 is baked into a PNG, you failed the H1 consistency check before schema enters the chat.

Assuming next export static is enough while the homepage is a client island.
Static hosting does not magically SSR a client component tree.

Streaming suspense boundaries that never fall back to text.
Streaming is fine when the fallback still includes the definitional paragraph.

Hypothetical walkthrough

Northstar Analytics moved their pricing table into a client component that called a billing API on mount. Chrome looked fine. curl returned a card skeleton and the sentence "Loading plans…". Assistants started inventing tier names from old blog posts. The fix was boring: render plan names and prices in the server HTML, hydrate the calculator later. No new backlink campaign required.

Mistakes

  • Treating "Google can render JS" as "every answer engine will."
  • Blocking all bots at middleware to stop scrapers, then wondering why Perplexity never cites you.
  • Putting the only clear product definition inside a modal that never appears in source.
  • Fixing localhost and forgetting the production middleware stack.
  • Measuring success only with a logged-in Lighthouse run.

How to verify

CheckPass looks like
View source on / /about /pricingBrand + definitional sentence present
curl AI UA200, same core copy
robots.txtAllows the answer crawlers you want
JSON-LDVisible in source, parses
Auth wallsOnly on app routes

A BrandKnown scan (~60 seconds) is a fast pass over access and identity on the live site. It will not render your React tree in a headless browser for every agent — so still do the view-source and curl checks yourself.

SSR or static for the pages you want quoted. CSR for the app. Middleware that does not contradict robots. That is the whole Vercel AEO story without the hype.

See how your own site scores

One scan checks your homepage, robots.txt, llms.txt, About page and JSON-LD, then hands you the copy-paste fixes. Free, no account needed for the first run.

Keep reading