◀ All articles

Playbooks

What to steal from a competitor's entity setup

June 26, 2026 · 3 min read

A rival taking your shortlist slot usually has three or four specific things you do not. Here is how to find them by reading their public markup.


A visibility check that names three competitors is more useful than one that names you, because those three are a worked example. Whatever they are doing has convinced a model to shortlist them, and almost all of it is published on their own website in plain view.

This is how to read it. Everything here uses public markup, and none of it involves anything you would be uncomfortable explaining.

The ten-minute pass

Open the competitor's homepage and run these six checks in order.

1. The served HTML.

curl -sL https://competitor.example > comp.html
wc -c comp.html

A few hundred bytes means a JS shell and they have the same problem you might. Fifty kilobytes of real markup means they render server-side, and that alone may be the difference.

2. Their schema.

grep -o 'application/ld+json' comp.html | wc -l

Then pull the block out and read it. What @type did they choose? Do they have sameAs, and what is in it? Is there a description, and is it a definition or a slogan? Do they use @id to link nodes together? This is the single highest-information artifact on their site.

3. Their robots.txt. competitor.example/robots.txt. Which AI agents are allowed, which are blocked. If they allow OAI-SearchBot and you block it, you have found your gap and it is one line long.

4. Their llms.txt. competitor.example/llms.txt. Most sites 404 here. If theirs exists, read the summary line — it is their own best one-sentence description of themselves, and it tells you how they want to be categorised.

5. Their About page. Read the first two hundred words as a machine would. Can you extract category, audience, size, location, founding year? Compare against yours honestly.

6. Their title tag and H1. Does the brand name appear identically in both, and in the schema name? Consistency is not glamorous and it is frequently the whole difference.

What you are looking for

After three competitors you will have a table like this, and the pattern will be obvious:

YouComp AComp BComp C
Server-rendered HTML✗✓✓✓
Organization schema✓✓✓✓
sameAs profiles0647
Answer crawlers allowed✗✓✓✓
Wikidata item✗✓✗✓
One-sentence descriptionvaguecrispcrispcrisp

The useful gaps are the columns where all three of them agree and you differ. Those are not preferences, they are the working configuration for your category.

The parts you cannot copy

Be clear-eyed about this, because it is where the ten-minute pass stops helping.

Their third-party presence. The round-ups they appear in, the review profiles with two hundred reviews, the Reddit threads recommending them. This is usually the majority of why they are being named, and it is months of work, not an afternoon.

Their age. A company that has existed for eight years is in more training data than one that launched last year. Nothing fixes that except time.

Their actual product. If they are named because they are genuinely the better fit for the question asked, no markup changes that, and it should not.

The honest split we see: entity setup explains a meaningful share of the gap for companies that have done none of it, and almost none of the gap for companies that have. If you have already done the access and identity work and they are still taking every slot, the answer is off-site and it is not technical.

Doing it without being weird about it

Everything above is a public file served to anyone who asks. Reading a competitor's robots.txt is not espionage.

What would be a problem: hammering their site, scraping at volume, or anything behind a login. One fetch of six public URLs is a person looking at a website. A thousand is a load test, and it is both rude and, under most terms of service, prohibited.

Our competitor diff runs the comparison automatically on the Continue and Boss Mode plans — same public resources, one fetch each, and it ranks the gaps by which are worth closing first. The manual version above costs you ten minutes and tells you the same thing; the tool mainly saves you from doing it three times.

See how your own site scores

One scan checks your homepage, robots.txt, llms.txt, About page and JSON-LD, then hands you the copy-paste fixes. Free, no account needed for the first run.

Keep reading