Entity optimization for AI search: the audit, not the markup

The short answer
Entity optimization is the work of making a model confident about what you are, so it names you without hedging. Almost everyone sells it as a schema exercise, and that is the part Google has explicitly told you to stop over-doing.
Key takeaways
- Google’s May 2026 guidance names structured-data over-optimisation as unnecessary for its AI features, alongside llms.txt, content chunking and AI-specific rewrites.
- Entity clarity is mostly built off your own domain. 84 percent of AI citations come from third-party sources, so the sentence other people use to describe you matters more than the markup on your page.
- Markup still matters, at the floor. Get Organization schema right once, then stop. There is no bonus for a fifth nested property.
- The useful diagnostic is the gap between how often engines name you and how often they cite a page you own. Those are two different failures with two different fixes.
- Run the audit before buying anything. Thirty prompts, three engines, one afternoon, and you will know which of the three problems you have.
Where this comes from. We run organic for more than 250 enterprises at Pepper and track over 10 million prompts across every major engine. The audit below is the one we run in week one of an engagement, and it is written out in full because a client who understands their own gap is a better client. Where we think the standard advice is wrong, we say so.
What is entity optimization?
An entity is a thing a machine can identify and hold facts about: a company, a product, a person, a place. Entity optimization is the work of making your entity unambiguous, consistently described, and connected to other entities the model already trusts.
Three things have to be true before a model will name you confidently.
Identity. The model can tell which thing you are, and does not confuse you with a similarly named company.
Attributes. It knows what you do, who you serve and what category you sit in, in the words buyers use rather than the words your positioning deck uses.
Corroboration. Credible third parties describe you the same way. This is the one nobody budgets for, and it is the one that carries the most weight.
The related terminology is untangled in AEO vs GEO vs AIO vs LLMO, and the underlying model in what GEO is.
See where you show up. Pepper’s GEO platform tracks Brand Visibility, Domain Prompt Presence and Share of Voice across ChatGPT, Perplexity, Gemini, Claude and Google AI Overviews. Your team can log in, connect Search Console and GA4, manage the prompt set and run your own agents in the Agent Atlas. A growth team works the same account alongside you. Book a growth audit or see where you show up.
The thing the guides get wrong
Read the current top results for this keyword and you will find the same recommendation in eight variations: add more structured data. Nest your Organization schema. Add sameAs links. Mark up everything.
On 15 May 2026 Google published its first official guidance on optimising for its AI features and filed it under SEO fundamentals. It named several widely sold tactics as unnecessary, and structured-data over-optimisation was one of them, alongside llms.txt files, content chunking, AI-specific rewrites and inauthentic mentions.
Read that carefully, because it is easy to over-correct in the other direction. Google did not say schema is worthless. Schema is how you state facts unambiguously, and stating them once, correctly, is worth doing. What Google said is that piling on more of it does not buy additional favour in its AI features.

So if markup is a floor rather than a lever, where does the real gain sit?
Muck Rack analysed more than 25 million cited links across ChatGPT, Claude and Gemini in May 2026. Earned media accounted for 84 percent of AI citations. Journalism alone made up 27 percent. Paid and advertorial content accounted for 0.3 percent.
That is the answer. Your entity is defined mostly by other people’s sentences about you, not by your own markup. Which is uncomfortable, because markup is a ticket you can close in a sprint and third-party description is not.

The entity signal stack
Work from the bottom. Each layer is wasted effort if the one beneath it is broken.

Most programmes invert this. They start at identity, because markup is tractable and sits in a sprint, then never reach corroboration, because it does not. The stack is worth reading as a budget allocation rather than a checklist.
The audit: five steps, one afternoon
This is the diagnostic we run in week one. It costs nothing and it decides where the budget goes.

1. Write 30 prompts in your buyers’ words
Not keywords. The questions a buyer would actually type. A mix of category questions, comparison questions and questions naming you directly, so you can separate the three failure modes.
2. Run them across three engines
Google, ChatGPT and Perplexity at minimum. Log two things separately for every prompt: is your brand named, and is a page on your domain cited. Two columns, not one.
3. Read the gap between those columns
This is the whole point of the exercise.
Named and cited. The entity is working. Move on to share of voice against the category leader.
Named but never cited. The model knows you exist and does not consider your pages worth quoting. That is a citability problem: the content is not extractable, or the authority is not there.
Neither named nor cited. The model does not have a confident entity for you. That is where identity and corroboration work goes.
Named wrongly. The worst outcome and the easiest to miss if you only count mentions. If the model describes you in the wrong category, no amount of visibility work helps until the description is corrected.
4. Check what the model thinks you are
Ask three engines directly: “what does [your company] do”, “who are [your company]’s main competitors”, “what category is [your company] in”. Compare the answers against how you would describe yourself.
We have found companies described as the wrong category, competing against companies they have never met, using product names they retired two years ago. One of those was a currently published article on our own site describing a competitor as a workflow builder when it was an AI search platform, so this is not a problem other people have.
5. Trace the sources
For every answer where a competitor is cited and you are not, open the sources. You are looking for which third parties the model trusts in your category. That list is your earned-media target list, and it is more useful than any keyword report.
The mechanics of working those sources are in LLM seeding, and the structural side in how to structure content for AI citation.
What the fixes actually are
For identity problems. Get Organization schema right once: legal name, URL, logo, sameAs pointing at profiles you control and that actually exist. Make sure your own site says what you are in the first sentence of the page a model is most likely to retrieve. Then stop. Our breakdown of the markup itself is in schema markup and Organization schema.
For attribute problems. Describe yourself in buyer language, in one consistent sentence, everywhere. The mismatch we see most often is a positioning line that means something internally and nothing to a model.
For corroboration problems. This is the 84 percent. Get credible third parties in your category to describe you accurately. Directories, analyst notes, journalism, comparison pages you do not own. Slow, unglamorous, and the highest-return work available.
For citability problems. Being named is not being cited. Answer-led structure, extractable passages, real tables rather than images. See how to get cited in LLMs and the LLM trust hierarchy.
What this looked like on a real account
Acceldata is the clearest example we have of corroboration and depth beating volume.
The work was not markup. It was building content deep enough to be worth citing, on a defined set of topics, so the entity became unambiguous in a technical category. Organic traffic grew 6X. Top-three keyword rankings went from 85 to more than 300. The site added over 100,000 new organic users at a 47 percent engagement rate. A single hero guide produced more than 260,000 impressions on its own.
That last number is the entity argument in miniature. One page that genuinely owned its topic outperformed a quarter of scheduled posts, because depth is what makes a model confident enough to quote you.
SalesHood shows the same pattern on the visibility side: AI Overview visibility went from 14 keywords to 97, alongside technical fixes and a content engine rather than a schema project.
Our methodology: how we weighted the entity signals
| Signal | Weight | Why it carries this much |
|---|---|---|
| Third-party corroboration | 40% | 84% of AI citations come from sources you do not own, per a 25M-link corpus study |
| Consistent self-description | 25% | The cheapest fix available, and the one most often inconsistent across a site |
| Retrievability and structure | 20% | Being named is not being cited, and this is what closes that specific gap |
| Structured data, done once | 15% | A floor, not a lever. Google names over-optimisation as unnecessary |
Entity work at a glance: what to do and what it costs
| Job | What it involves | Typical 2026 cost | Where it falls short |
|---|---|---|---|
| Organization schema, done once | Legal name, URL, logo, real sameAs links | A few hours of developer time | It is a floor. Doing more of it buys nothing further |
| Consistent self-description | One buyer-language sentence, applied everywhere | Internal, a day of editing | Does nothing on its own if nobody credible repeats it |
| The 30-prompt audit | Named versus cited, three engines, logged | Free, about two hours a month | Tells you the problem, not the fix |
| Third-party corroboration | Directories, analysts, journalism, comparison pages | The largest line item, and slow | Cannot be bought outright. Paid placements are 0.3% of citations |
| Depth on fewer topics | Pages that genuinely own a subject | Real content investment | Slower than publishing volume, and harder to staff |
| Multi-engine tracking platform | Automated per-engine measurement | Roughly $99 to $400 a month | Diagnoses, does not fix |
—
How to choose where to spend
Almost nobody has budget for all six rows above. The audit tells you which one to start with, and the answer is rarely markup.
The scorecard
Score your entity out of 100. The lowest area is your starting point.
| Area | Weight | How to score it honestly |
|---|---|---|
| Third-party corroboration | 35% | Credible sources in your category describe you accurately and in the right category, without you having paid for the placement |
| Description consistency | 25% | Your homepage, your about page and your LinkedIn say the same thing about what you are, in buyer language rather than positioning language |
| Retrievability | 20% | Headings read as real questions, the first two sentences answer them, and tables are text rather than images |
| Structured data floor | 10% | Organization schema is present and correct once. Full marks for correct, no extra marks for elaborate |
| Measurement | 10% | You track named and cited as separate numbers, per engine, on a fixed prompt set |
The live test, over 90 days
Fix 30 prompts in your buyers’ words and hold them steady. Real questions, such as “what does Acceldata do”, “who are the main data observability vendors”, “which sales enablement platform is best for mid-market” and “is there a free way to check AI search visibility”.
Run all 30 across Google, ChatGPT and Perplexity. Log named and cited separately, plus the sources cited instead of you. Then rerun the identical set monthly for 90 days.
Three readings is the shortest honest window. One is a snapshot, two could be noise, three shows direction. Change the prompt set and you have reset the comparison, so resist the urge to tidy it.
Weak approach versus strong approach
The weaker approach: commission a schema audit, implement every recommendation, nest the markup, add sameAs links to profiles nobody maintains, then wonder why nothing moved. Google has now explicitly named this as unnecessary for its AI features.
The stronger approach: fix the markup once in an afternoon, then spend the rest of the quarter making sure the sources a model already trusts describe you correctly.
The difference is where the effort goes. Markup is inside your control, which is exactly why it is over-sold. Corroboration is outside it, which is why it is under-sold and why it works.
Red flags
- “We will build your knowledge graph.” Nobody sells you a position in Google’s knowledge graph. Entities are earned through corroboration.
- A schema audit as the whole engagement. Google named structured-data over-optimisation as unnecessary in May 2026.
- llms.txt as a headline recommendation. Named in the same guidance. Our own llms.txt explainer reaches the same conclusion.
- A single AI visibility score. It hides the named-versus-cited gap, which is the diagnostic.
- Mentions counted without sentiment or category. Being named in the wrong category is worse than not being named.
- Guaranteed knowledge panel or citation. Nobody controls generated output, and Google warns against providers guaranteeing rankings.
Five questions worth asking
- “What in Google’s May 2026 guidance changes your recommendations?” Anyone unfamiliar with it is not current.
- “Show me a query where a competitor is cited and explain why.” Separates retrieval understanding from markup reselling.
- “How much of your scope is off our domain?” If none, it addresses the minority of the citation signal.
- “How will you tell a naming problem from a citation problem?” The answer should be two separate numbers.
- “What would you tell us not to bother with?” A good partner cuts something.
Reduced to one principle: fix the markup once, then spend the quarter on what other people say about you. Nearly every other decision follows from that.
One closing note that costs us something. Corroboration work is slow, it does not photograph well in a monthly report, and it is genuinely harder to sell than a schema audit, ours included. Any partner who only proposes the fast visible work is proposing the wrong work.
What nobody should promise you
A knowledge panel, a knowledge graph entry, or a guaranteed citation. Nobody controls generated output.
That more schema produces more AI visibility. Google’s own 2026 guidance names structured-data over-optimisation as unnecessary for its AI features.
A specific percentage improvement in entity recognition. The figures circulating on this topic do not trace to a primary study with a stated method.
A single visibility score that means something on its own. There is no universal good number. Trend on a fixed prompt set and distance from the category leader carry the information.
When you should skip this entirely
The answer that costs us the sale. If your important pages are not indexed, or your site does not clearly state what you sell on the page a model would retrieve first, you do not need an entity programme yet and we would tell you not to buy one.
Fix indexing and the plain description of what you do. Both are free. Entity work layered on a site engines cannot read or parse is paying to amplify a problem you already have.
Frequently asked questions
What is entity optimization?
It is the work of making your company unambiguous to a model: clear identity, accurate attributes in buyer language, and consistent description by credible third parties. The aim is that an engine names you confidently and in the right category.
Is entity optimization just schema markup?
No, and this is the most common mistake. Schema states facts unambiguously and is worth doing once, correctly. Google’s May 2026 guidance names structured-data over-optimisation as unnecessary for its AI features, and 84 percent of AI citations come from sources you do not own.
How do I run an AI search visibility audit?
Write 30 buyer-phrased prompts, run them across Google, ChatGPT and Perplexity, and log two things separately for each: whether your brand is named, and whether a page on your domain is cited. The gap between those two columns tells you which problem you have.
What is the difference between being named and being cited?
Being named means the engine mentions your brand in its answer. Being cited means it links a page on your domain as a source. Named but not cited is an authority or extractability problem, and no amount of entity markup fixes it.
Does structured data still matter for AI search?
Yes, as a floor. Getting Organization schema correct once helps a model state facts about you without guessing. Adding more of it beyond that point is what Google has named as unnecessary.
How long does entity optimization take to work?
Plan on three monthly readings of a fixed prompt set before judging anything, so about 90 days. Corroboration work is the slowest part because it depends on third parties publishing, which you do not control.
Can I pay for entity mentions?
You can, and the evidence says it barely helps. Paid and advertorial content accounted for just 0.3 percent of AI citations in a study of more than 25 million cited links, against 84 percent for earned media. Budget for being described accurately by credible third parties instead of for placement.
What is the cheapest way to start?
The 30-prompt audit. It costs an afternoon and nothing else, and it tells you whether you have an identity problem, a citability problem or a corroboration problem before you spend anything.
Where to go next
Run the five-step audit. It takes an afternoon and it will tell you which of the three problems you have, which is more than a schema report will.
Most teams discover the answer is corroboration, which is the slowest thing to fix and the reason to start this quarter rather than next.
Book a growth audit · Read the case studies · Explore Pepper’s platform
Sources and further reading
- Google Search Central. First official AI search optimisation guidance, published 15 May 2026, filed under SEO fundamentals. Names llms.txt, content chunking, AI-specific rewrites, inauthentic mentions and structured-data over-optimisation as unnecessary for its AI features, and warns against guaranteed-ranking claims.
- Muck Rack. “What Is AI Reading?” May 2026. More than 25 million cited links across ChatGPT, Claude and Gemini, 17 industries. Source of the 84 percent earned media, 27 percent journalism and 0.3 percent paid figures. Link
- Pepper. Acceldata case study and SalesHood case study. Metrics as published on those pages.
- Pepper. Entity optimization and LLM brand recognition, the companion piece on how models form brand associations.
- Pepper. Visibility, Citability and Retrievability, the framework the audit above maps onto.
- Deliberately excluded: the “40 percent more accurate entity recognition” figure, which we could not trace to a primary study with a stated sample and method. Also the 2024 academic GEO tactics research, which predates 2026 and tested a simulated engine.
Latest Blogs
The short answer Structure your pages so an answer engine can find and lift a single passage. That means one front-loaded H1, question-shaped H2s that stand alone, and a direct 40 to 60 word answer under each one. That work buys retrievability. It does not buy citation. Confusing the two is the most expensive mistake […]
Every guide on AI search monitoring tools assumes the answer is yes and jumps to which one. This one gives you the arithmetic instead: what these platforms genuinely catch that Search Console cannot, what their entry tiers actually meter, why engine counts are the least comparable number on the page, and the honest test for when you should not buy one yet.
Entity optimization is sold as a schema exercise, and that is the part Google explicitly told you to stop over-doing in May 2026. Here is what actually makes a model confident about what you are, plus the five-step audit we run in week one of every engagement.