◀ All articles

SEO

Your JavaScript site may be invisible to AI crawlers

July 21, 2026 · 3 min read

AI crawlers generally do not execute JavaScript. If your name, description and schema only appear after hydration, they were never there at all.


There is a specific failure that produces a perfect-looking site and a completely empty scan, and it catches good engineering teams more often than bad ones.

Your site is a single-page app. Open it in a browser and everything is there — the headline, the description, the schema block, the footer. Open it with curl and you get this:

<!doctype html>
<html><head><title>Northstar</title></head>
<body><div id="root"></div><script src="/bundle.js"></script></body></html>

Googlebot renders JavaScript and will eventually see the real page. AI crawlers generally do not. They fetch HTML, parse it, and move on. To them, that empty <div> is your entire website.

Why they do not render

Rendering is expensive — a headless browser per page instead of an HTTP request, seconds instead of milliseconds, orders of magnitude more compute across a crawl. Google built that infrastructure over a decade because search is their business. A crawler collecting text for a corpus, or fetching three pages to answer one question, has no reason to pay that cost.

Assume no execution. If you are wrong, you lost nothing; if you assume the opposite and are wrong, you are invisible.

Checking in thirty seconds

The critical thing: do not use DevTools' Elements panel. It shows the live DOM after JavaScript has run, which is exactly the thing you are trying to look past. It will tell you everything is fine.

Use one of these instead:

curl -sL https://yoursite.example | head -100
# Is the schema actually in the served HTML?
curl -sL https://yoursite.example | grep -c 'application/ld+json'

Or view-source:https://yoursite.example in the address bar, which shows the served document rather than the DOM. Search it for your company name, your description, and ld+json. Whatever is not in there does not exist as far as an AI crawler is concerned.

What has to be in the HTML

Not everything needs to be server-rendered. This list does:

  • <title> and the meta description
  • the <h1> and the main body copy of the page
  • the Organization JSON-LD block
  • your navigation links, as real <a href> elements
  • your About page content
  • product names, prices and descriptions on commerce pages

Dashboards, charts, interactive configurators and anything behind a login can stay client-side. Nobody is crawling your app shell.

Fixing it

Next.js. App Router components are server components by default, which means they render to HTML unless you opt out. The usual culprit is a "use client" boundary drawn too high, or content fetched in a useEffect. Move data fetching into the server component and keep "use client" for the leaf that actually needs interactivity. JSON-LD goes in the server component, in a <script type="application/ld+json"> tag.

Nuxt, SvelteKit, Astro, Remix. All server-render by default. Check you have not disabled SSR for the route, and that your schema is not being injected by a client-side plugin.

Create React App, Vite SPA, plain Vue. No server rendering at all. Options, in order of effort: prerender the marketing routes at build time (vite-plugin-ssr, react-snap, or just a static export of the handful of pages that matter), migrate the marketing site to a framework that renders, or — the pragmatic one — split the marketing site off from the app entirely. Your app can be a SPA. Your homepage should not be.

WordPress and other server-rendered CMSs. You are usually fine. The exception is a headless setup with a JS front end, which puts you back in the SPA case.

The minimum viable fix

If a rendering migration is not happening this quarter, you can still fix the entity layer today. Server-render four things:

  1. the Organization JSON-LD block
  2. <title> and the meta description
  3. the <h1>
  4. a real, static About page — hand-written HTML if necessary

That is enough for a crawler to identify you correctly even if the rest of the site is a shell. It is not the whole job, but it takes the difference between "unknown entity" and "known company with a thin site," and that gap is the one that matters.

The scan behaviour, for what it is worth

When we fetch a homepage and find almost no text with a script bundle in it, we flag it as a JS shell rather than reporting an empty result. It is a distinct finding because the fix is different: you do not have a schema problem or a content problem, you have a delivery problem, and adding more markup to a page that is never parsed does nothing.

See how your own site scores

One scan checks your homepage, robots.txt, llms.txt, About page and JSON-LD, then hands you the copy-paste fixes. Free, no account needed for the first run.

Keep reading