AI crawlers generally do not execute JavaScript. If your name, description and schema only appear after hydration, they were never there at all.
There is a specific failure that produces a perfect-looking site and a completely empty scan, and it catches good engineering teams more often than bad ones.
Your site is a single-page app. Open it in a browser and everything is there — the headline, the description, the schema block, the footer. Open it with curl and you get this:
<!doctype html>
<html><head><title>Northstar</title></head>
<body><div id="root"></div><script src="/bundle.js"></script></body></html>
Googlebot renders JavaScript and will eventually see the real page. AI crawlers generally do not. They fetch HTML, parse it, and move on. To them, that empty <div> is your entire website.
Why they do not render
Rendering is expensive — a headless browser per page instead of an HTTP request, seconds instead of milliseconds, orders of magnitude more compute across a crawl. Google built that infrastructure over a decade because search is their business. A crawler collecting text for a corpus, or fetching three pages to answer one question, has no reason to pay that cost.
Assume no execution. If you are wrong, you lost nothing; if you assume the opposite and are wrong, you are invisible.
Checking in thirty seconds
The critical thing: do not use DevTools' Elements panel. It shows the live DOM after JavaScript has run, which is exactly the thing you are trying to look past. It will tell you everything is fine.
Use one of these instead:
curl -sL https://yoursite.example | head -100
# Is the schema actually in the served HTML?
curl -sL https://yoursite.example | grep -c 'application/ld+json'
Or view-source:https://yoursite.example in the address bar, which shows the served document rather than the DOM. Search it for your company name, your description, and ld+json. Whatever is not in there does not exist as far as an AI crawler is concerned.
What has to be in the HTML
Not everything needs to be server-rendered. This list does:
<title>and the meta description- the
<h1>and the main body copy of the page - the
OrganizationJSON-LD block - your navigation links, as real
<a href>elements - your About page content
- product names, prices and descriptions on commerce pages
Dashboards, charts, interactive configurators and anything behind a login can stay client-side. Nobody is crawling your app shell.
Fixing it
Next.js. App Router components are server components by default, which means they render to HTML unless you opt out. The usual culprit is a "use client" boundary drawn too high, or content fetched in a useEffect. Move data fetching into the server component and keep "use client" for the leaf that actually needs interactivity. JSON-LD goes in the server component, in a <script type="application/ld+json"> tag.
Nuxt, SvelteKit, Astro, Remix. All server-render by default. Check you have not disabled SSR for the route, and that your schema is not being injected by a client-side plugin.
Create React App, Vite SPA, plain Vue. No server rendering at all. Options, in order of effort: prerender the marketing routes at build time (vite-plugin-ssr, react-snap, or just a static export of the handful of pages that matter), migrate the marketing site to a framework that renders, or — the pragmatic one — split the marketing site off from the app entirely. Your app can be a SPA. Your homepage should not be.
WordPress and other server-rendered CMSs. You are usually fine. The exception is a headless setup with a JS front end, which puts you back in the SPA case.
The minimum viable fix
If a rendering migration is not happening this quarter, you can still fix the entity layer today. Server-render four things:
- the
OrganizationJSON-LD block <title>and the meta description- the
<h1> - a real, static About page — hand-written HTML if necessary
That is enough for a crawler to identify you correctly even if the rest of the site is a shell. It is not the whole job, but it takes the difference between "unknown entity" and "known company with a thin site," and that gap is the one that matters.
The scan behaviour, for what it is worth
When we fetch a homepage and find almost no text with a script bundle in it, we flag it as a JS shell rather than reporting an empty result. It is a distinct finding because the fix is different: you do not have a schema problem or a content problem, you have a delivery problem, and adding more markup to a page that is never parsed does nothing.
