Fresh lastmod, real URLs, no soft 404s — a sitemap checklist that helps humans and machines find the pages worth citing.
Sitemaps feel like 2012 SEO homework. They still matter in 2026 because answer engines and classic crawlers share a boring need: a clean list of URLs that deserve a fetch.
If your sitemap is stale, full of soft 404s, or missing the pages that define your brand, you are asking machines to guess. Guessing is how the wrong URL becomes "the" Northstar Analytics page in a citation.
What a sitemap is for (AEO edition)
An XML sitemap does not rank you. It announces candidates. For AI discovery, the useful jobs are:
- Surface canonical product, about, docs, and comparison URLs
- Advertise freshness via honest
lastmodvalues - Keep dead and duplicate URLs out of the candidate set
| Sitemap quality | Effect on discovery | Effect on trust |
|---|---|---|
| Fresh, canonical, 200 OK URLs | Easier refetch of the right pages | Neutral to positive |
Stale lastmod everywhere | Recrawl priority becomes noise | Looks neglected |
| Soft 404s and parameter junk | Wastes crawl; confuses entity | Harmful |
| Missing about / pricing / docs | Brand definition pages stay obscure | Harmful for AEO |
Pair sitemap hygiene with canonical tags and the one-URL problem so you do not advertise four addresses for one entity. Access still comes first: a perfect sitemap cannot help if AI crawlers get 403s.
Sitemap checklist
Use this as a working checklist, not a vibes pass.
| Check | Pass looks like | Fail looks like |
|---|---|---|
| Location | Linked from robots.txt (Sitemap: https://…/sitemap.xml) | Only in Search Console, never declared |
| URLs | Absolute HTTPS canonicals | HTTP, www/apex mix, session IDs |
| Status | Spot-check returns 200 | Soft 404 HTML with 200, or hard 404 |
| lastmod | Changes when content changes | Same timestamp on every URL for years |
| Size | Index + child sitemaps within common limits | One giant broken file |
| Scope | Indexable marketing and docs URLs | Cart, account, filtered faceted URLs |
| Images/video | Optional; only if real assets | Empty noise entries |
Which URLs earn a place
Think in entity terms, not "every CMS node."
- Homepage and about (definition of the brand)
- Core product or service URLs
- Pricing or plans if public
- Docs or guides that answer real questions
- Comparison or integration pages you maintain honestly
- Location or contact pages when local entity matters
Deprioritize infinite filters, thank-you pages, and tag archives that only exist for internal browsing. Say you run 20 product marketing URLs and 2,000 tag combinations — the sitemap should champion the 20.
How to improve your sitemap in one sitting
- Open robots.txt and confirm a
Sitemap:line points at the live index. - Download the index and one child sitemap. Count URLs. Spot-check ten random ones with
curl -I. - Kill soft 404s. If a URL returns 200 with "not found" copy, fix status or remove it from the sitemap. See soft 404s that confuse AI crawlers.
- Normalize hosts. Pick apex or
www, pick trailing-slash policy, enforce with redirects and canonicals — then list only the winner. - Prioritize entity pages. Make sure those pages also match the story on your about page.
- Set honest
lastmod. If your CMS cannot do real dates, omit fake precision rather than stamp today's date on unchanged pages. - Resubmit in Google Search Console (and Bing if you use it). Re-fetch a sample as an AI bot UA after access rules are correct.
Example shape (keep it boring)
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://www.northstaranalytics.example/about/</loc>
<lastmod>2026-07-12</lastmod>
</url>
<url>
<loc>https://www.northstaranalytics.example/product/dashboards/</loc>
<lastmod>2026-08-01</lastmod>
</url>
</urlset>
No keyword stuffing in paths. No twenty near-duplicate "best analytics tool 2024/2025/2026" clones unless those pages are truly distinct and maintained.
After a migration
Migrations create the worst sitemap debt: old locs that 404, new locs missing, lastmod frozen at cutover day. Budget an hour post-launch to regenerate from the canonical URL set, compare against a crawl, and delete anything that is not 200. If the edge still challenges AI agents, fix WAF allowlisting before you obsess over lastmod.
Mistakes that waste the hour
- Including URLs you
noindex - Listing both
/pageand/page/as separate entries while they 200 without consolidating - Auto-generating sitemaps from every CMS revision URL
- Celebrating "submitted" in Search Console while
curlstill gets 403 from AI agents — discovery cannot help if your WAF blocks the bots - Treating sitemap priority tags as a ranking lever (most systems ignore or barely use them)
- Shipping a sitemap index that points at 404 child files after a CMS migration
How this ties to AI answers
Pew Research (March 2025) found AI Overviews on roughly 18% of Google searches in their study, with traditional-result clicks lower when a summary appeared (8% with vs 15% without). When summaries do cite, 88% cited three or more sources. Being in that citation set starts with being fetchable and findable — sitemap included — then being worth citing.
You still cannot control whether Perplexity or ChatGPT picks your URL tomorrow. You can control whether the URL they would pick is in a clean sitemap, returns 200, and matches your canonical entity address.
A quick BrandKnown scan will not replace Search Console coverage reports, but it will flag access and identity issues that make sitemap work pointless. Fix discovery and access together.
