GEO / AI Search

Entity optimization for AI search: the audit, not the markup

Team Pepper
Posted on 7/09/2612 min read
Entity optimization for AI search: the audit, not the markup

The short answer

Entity optimization is the work of making a model confident about what you are, so it names you without hedging. Almost everyone sells it as a schema exercise, and that is the part Google has explicitly told you to stop over-doing.

Key takeaways

  • Google’s May 2026 guidance names structured-data over-optimisation as unnecessary for its AI features, alongside llms.txt, content chunking and AI-specific rewrites.
  • Entity clarity is mostly built off your own domain. 84 percent of AI citations come from third-party sources, so the sentence other people use to describe you matters more than the markup on your page.
  • Markup still matters, at the floor. Get Organization schema right once, then stop. There is no bonus for a fifth nested property.
  • The useful diagnostic is the gap between how often engines name you and how often they cite a page you own. Those are two different failures with two different fixes.
  • Run the audit before buying anything. Thirty prompts, three engines, one afternoon, and you will know which of the three problems you have.

Where this comes from. We run organic for more than 250 enterprises at Pepper and track over 10 million prompts across every major engine. The audit below is the one we run in week one of an engagement, and it is written out in full because a client who understands their own gap is a better client. Where we think the standard advice is wrong, we say so.


What is entity optimization?

An entity is a thing a machine can identify and hold facts about: a company, a product, a person, a place. Entity optimization is the work of making your entity unambiguous, consistently described, and connected to other entities the model already trusts.

Three things have to be true before a model will name you confidently.

Identity. The model can tell which thing you are, and does not confuse you with a similarly named company.

Attributes. It knows what you do, who you serve and what category you sit in, in the words buyers use rather than the words your positioning deck uses.

Corroboration. Credible third parties describe you the same way. This is the one nobody budgets for, and it is the one that carries the most weight.

The related terminology is untangled in AEO vs GEO vs AIO vs LLMO, and the underlying model in what GEO is.

See where you show up. Pepper’s GEO platform tracks Brand Visibility, Domain Prompt Presence and Share of Voice across ChatGPT, Perplexity, Gemini, Claude and Google AI Overviews. Your team can log in, connect Search Console and GA4, manage the prompt set and run your own agents in the Agent Atlas. A growth team works the same account alongside you. Book a growth audit or see where you show up.


The thing the guides get wrong

Read the current top results for this keyword and you will find the same recommendation in eight variations: add more structured data. Nest your Organization schema. Add sameAs links. Mark up everything.

On 15 May 2026 Google published its first official guidance on optimising for its AI features and filed it under SEO fundamentals. It named several widely sold tactics as unnecessary, and structured-data over-optimisation was one of them, alongside llms.txt files, content chunking, AI-specific rewrites and inauthentic mentions.

Read that carefully, because it is easy to over-correct in the other direction. Google did not say schema is worthless. Schema is how you state facts unambiguously, and stating them once, correctly, is worth doing. What Google said is that piling on more of it does not buy additional favour in its AI features.

Three cards summarising Google's May 2026 guidance: what to stop over-doing including structured-data over-optimisation, what to still do once, and where the work has moved
Figure 1: Google’s May 2026 guidance separates the floor from the lever. Source: Google Search Central, 15 May 2026.

So if markup is a floor rather than a lever, where does the real gain sit?

Muck Rack analysed more than 25 million cited links across ChatGPT, Claude and Gemini in May 2026. Earned media accounted for 84 percent of AI citations. Journalism alone made up 27 percent. Paid and advertorial content accounted for 0.3 percent.

That is the answer. Your entity is defined mostly by other people’s sentences about you, not by your own markup. Which is uncomfortable, because markup is a ticket you can close in a sprint and third-party description is not.

Three statistics on the source of AI citations: 84 percent from earned media, 27 percent from journalism alone, 0.3 percent from paid and advertorial content
Figure 2: You cannot buy your way into an answer. Source: Muck Rack, May 2026, 25M+ cited links.

The entity signal stack

Work from the bottom. Each layer is wasted effort if the one beneath it is broken.

Five stacked layers of entity signals: indexability at the base, then identity, attributes, retrievability, and corroboration at the top
Figure 3: The stack, bottom-up. Only the top layer carries most of the citation weight, and it is the one furthest outside your control. Sources: Google Search Central, 15 May 2026; Muck Rack, May 2026.

Most programmes invert this. They start at identity, because markup is tractable and sits in a sprint, then never reach corroboration, because it does not. The stack is worth reading as a budget allocation rather than a checklist.


The audit: five steps, one afternoon

This is the diagnostic we run in week one. It costs nothing and it decides where the budget goes.

Five-stage chevron of the entity audit: write 30 buyer-phrased prompts, run three engines, read the named-versus-cited gap, check what the model thinks you are, trace which sources win instead
Figure 4: The week-one diagnostic. Source: Pepper, run on every new engagement.

1. Write 30 prompts in your buyers’ words

Not keywords. The questions a buyer would actually type. A mix of category questions, comparison questions and questions naming you directly, so you can separate the three failure modes.

2. Run them across three engines

Google, ChatGPT and Perplexity at minimum. Log two things separately for every prompt: is your brand named, and is a page on your domain cited. Two columns, not one.

3. Read the gap between those columns

This is the whole point of the exercise.

Named and cited. The entity is working. Move on to share of voice against the category leader.

Named but never cited. The model knows you exist and does not consider your pages worth quoting. That is a citability problem: the content is not extractable, or the authority is not there.

Neither named nor cited. The model does not have a confident entity for you. That is where identity and corroboration work goes.

Named wrongly. The worst outcome and the easiest to miss if you only count mentions. If the model describes you in the wrong category, no amount of visibility work helps until the description is corrected.

4. Check what the model thinks you are

Ask three engines directly: “what does [your company] do”, “who are [your company]’s main competitors”, “what category is [your company] in”. Compare the answers against how you would describe yourself.

We have found companies described as the wrong category, competing against companies they have never met, using product names they retired two years ago. One of those was a currently published article on our own site describing a competitor as a workflow builder when it was an AI search platform, so this is not a problem other people have.

5. Trace the sources

For every answer where a competitor is cited and you are not, open the sources. You are looking for which third parties the model trusts in your category. That list is your earned-media target list, and it is more useful than any keyword report.

The mechanics of working those sources are in LLM seeding, and the structural side in how to structure content for AI citation.


What the fixes actually are

For identity problems. Get Organization schema right once: legal name, URL, logo, sameAs pointing at profiles you control and that actually exist. Make sure your own site says what you are in the first sentence of the page a model is most likely to retrieve. Then stop. Our breakdown of the markup itself is in schema markup and Organization schema.

For attribute problems. Describe yourself in buyer language, in one consistent sentence, everywhere. The mismatch we see most often is a positioning line that means something internally and nothing to a model.

For corroboration problems. This is the 84 percent. Get credible third parties in your category to describe you accurately. Directories, analyst notes, journalism, comparison pages you do not own. Slow, unglamorous, and the highest-return work available.

For citability problems. Being named is not being cited. Answer-led structure, extractable passages, real tables rather than images. See how to get cited in LLMs and the LLM trust hierarchy.


What this looked like on a real account

Acceldata is the clearest example we have of corroboration and depth beating volume.

The work was not markup. It was building content deep enough to be worth citing, on a defined set of topics, so the entity became unambiguous in a technical category. Organic traffic grew 6X. Top-three keyword rankings went from 85 to more than 300. The site added over 100,000 new organic users at a 47 percent engagement rate. A single hero guide produced more than 260,000 impressions on its own.

That last number is the entity argument in miniature. One page that genuinely owned its topic outperformed a quarter of scheduled posts, because depth is what makes a model confident enough to quote you.

SalesHood shows the same pattern on the visibility side: AI Overview visibility went from 14 keywords to 97, alongside technical fixes and a content engine rather than a schema project.


Our methodology: how we weighted the entity signals

SignalWeightWhy it carries this much
Third-party corroboration40%84% of AI citations come from sources you do not own, per a 25M-link corpus study
Consistent self-description25%The cheapest fix available, and the one most often inconsistent across a site
Retrievability and structure20%Being named is not being cited, and this is what closes that specific gap
Structured data, done once15%A floor, not a lever. Google names over-optimisation as unnecessary

Entity work at a glance: what to do and what it costs

JobWhat it involvesTypical 2026 costWhere it falls short
Organization schema, done onceLegal name, URL, logo, real sameAs linksA few hours of developer timeIt is a floor. Doing more of it buys nothing further
Consistent self-descriptionOne buyer-language sentence, applied everywhereInternal, a day of editingDoes nothing on its own if nobody credible repeats it
The 30-prompt auditNamed versus cited, three engines, loggedFree, about two hours a monthTells you the problem, not the fix
Third-party corroborationDirectories, analysts, journalism, comparison pagesThe largest line item, and slowCannot be bought outright. Paid placements are 0.3% of citations
Depth on fewer topicsPages that genuinely own a subjectReal content investmentSlower than publishing volume, and harder to staff
Multi-engine tracking platformAutomated per-engine measurementRoughly $99 to $400 a monthDiagnoses, does not fix

How to choose where to spend

Almost nobody has budget for all six rows above. The audit tells you which one to start with, and the answer is rarely markup.

The scorecard

Score your entity out of 100. The lowest area is your starting point.

AreaWeightHow to score it honestly
Third-party corroboration35%Credible sources in your category describe you accurately and in the right category, without you having paid for the placement
Description consistency25%Your homepage, your about page and your LinkedIn say the same thing about what you are, in buyer language rather than positioning language
Retrievability20%Headings read as real questions, the first two sentences answer them, and tables are text rather than images
Structured data floor10%Organization schema is present and correct once. Full marks for correct, no extra marks for elaborate
Measurement10%You track named and cited as separate numbers, per engine, on a fixed prompt set

The live test, over 90 days

Fix 30 prompts in your buyers’ words and hold them steady. Real questions, such as “what does Acceldata do”, “who are the main data observability vendors”, “which sales enablement platform is best for mid-market” and “is there a free way to check AI search visibility”.

Run all 30 across Google, ChatGPT and Perplexity. Log named and cited separately, plus the sources cited instead of you. Then rerun the identical set monthly for 90 days.

Three readings is the shortest honest window. One is a snapshot, two could be noise, three shows direction. Change the prompt set and you have reset the comparison, so resist the urge to tidy it.

Weak approach versus strong approach

The weaker approach: commission a schema audit, implement every recommendation, nest the markup, add sameAs links to profiles nobody maintains, then wonder why nothing moved. Google has now explicitly named this as unnecessary for its AI features.

The stronger approach: fix the markup once in an afternoon, then spend the rest of the quarter making sure the sources a model already trusts describe you correctly.

The difference is where the effort goes. Markup is inside your control, which is exactly why it is over-sold. Corroboration is outside it, which is why it is under-sold and why it works.

Red flags

  • “We will build your knowledge graph.” Nobody sells you a position in Google’s knowledge graph. Entities are earned through corroboration.
  • A schema audit as the whole engagement. Google named structured-data over-optimisation as unnecessary in May 2026.
  • llms.txt as a headline recommendation. Named in the same guidance. Our own llms.txt explainer reaches the same conclusion.
  • A single AI visibility score. It hides the named-versus-cited gap, which is the diagnostic.
  • Mentions counted without sentiment or category. Being named in the wrong category is worse than not being named.
  • Guaranteed knowledge panel or citation. Nobody controls generated output, and Google warns against providers guaranteeing rankings.

Five questions worth asking

  1. “What in Google’s May 2026 guidance changes your recommendations?” Anyone unfamiliar with it is not current.
  2. “Show me a query where a competitor is cited and explain why.” Separates retrieval understanding from markup reselling.
  3. “How much of your scope is off our domain?” If none, it addresses the minority of the citation signal.
  4. “How will you tell a naming problem from a citation problem?” The answer should be two separate numbers.
  5. “What would you tell us not to bother with?” A good partner cuts something.

Reduced to one principle: fix the markup once, then spend the quarter on what other people say about you. Nearly every other decision follows from that.

One closing note that costs us something. Corroboration work is slow, it does not photograph well in a monthly report, and it is genuinely harder to sell than a schema audit, ours included. Any partner who only proposes the fast visible work is proposing the wrong work.


What nobody should promise you

A knowledge panel, a knowledge graph entry, or a guaranteed citation. Nobody controls generated output.

That more schema produces more AI visibility. Google’s own 2026 guidance names structured-data over-optimisation as unnecessary for its AI features.

A specific percentage improvement in entity recognition. The figures circulating on this topic do not trace to a primary study with a stated method.

A single visibility score that means something on its own. There is no universal good number. Trend on a fixed prompt set and distance from the category leader carry the information.


When you should skip this entirely

The answer that costs us the sale. If your important pages are not indexed, or your site does not clearly state what you sell on the page a model would retrieve first, you do not need an entity programme yet and we would tell you not to buy one.

Fix indexing and the plain description of what you do. Both are free. Entity work layered on a site engines cannot read or parse is paying to amplify a problem you already have.


Frequently asked questions

What is entity optimization?
It is the work of making your company unambiguous to a model: clear identity, accurate attributes in buyer language, and consistent description by credible third parties. The aim is that an engine names you confidently and in the right category.

Is entity optimization just schema markup?
No, and this is the most common mistake. Schema states facts unambiguously and is worth doing once, correctly. Google’s May 2026 guidance names structured-data over-optimisation as unnecessary for its AI features, and 84 percent of AI citations come from sources you do not own.

How do I run an AI search visibility audit?
Write 30 buyer-phrased prompts, run them across Google, ChatGPT and Perplexity, and log two things separately for each: whether your brand is named, and whether a page on your domain is cited. The gap between those two columns tells you which problem you have.

What is the difference between being named and being cited?
Being named means the engine mentions your brand in its answer. Being cited means it links a page on your domain as a source. Named but not cited is an authority or extractability problem, and no amount of entity markup fixes it.

Does structured data still matter for AI search?
Yes, as a floor. Getting Organization schema correct once helps a model state facts about you without guessing. Adding more of it beyond that point is what Google has named as unnecessary.

How long does entity optimization take to work?
Plan on three monthly readings of a fixed prompt set before judging anything, so about 90 days. Corroboration work is the slowest part because it depends on third parties publishing, which you do not control.

Can I pay for entity mentions?
You can, and the evidence says it barely helps. Paid and advertorial content accounted for just 0.3 percent of AI citations in a study of more than 25 million cited links, against 84 percent for earned media. Budget for being described accurately by credible third parties instead of for placement.

What is the cheapest way to start?
The 30-prompt audit. It costs an afternoon and nothing else, and it tells you whether you have an identity problem, a citability problem or a corroboration problem before you spend anything.


Where to go next

Run the five-step audit. It takes an afternoon and it will tell you which of the three problems you have, which is more than a schema report will.

Most teams discover the answer is corroboration, which is the slowest thing to fix and the reason to start this quarter rather than next.

Book a growth audit · Read the case studies · Explore Pepper’s platform


Sources and further reading

  • Google Search Central. First official AI search optimisation guidance, published 15 May 2026, filed under SEO fundamentals. Names llms.txt, content chunking, AI-specific rewrites, inauthentic mentions and structured-data over-optimisation as unnecessary for its AI features, and warns against guaranteed-ranking claims.
  • Muck Rack. “What Is AI Reading?” May 2026. More than 25 million cited links across ChatGPT, Claude and Gemini, 17 industries. Source of the 84 percent earned media, 27 percent journalism and 0.3 percent paid figures. Link
  • Pepper. Acceldata case study and SalesHood case study. Metrics as published on those pages.
  • Pepper. Entity optimization and LLM brand recognition, the companion piece on how models form brand associations.
  • Pepper. Visibility, Citability and Retrievability, the framework the audit above maps onto.
  • Deliberately excluded: the “40 percent more accurate entity recognition” figure, which we could not trace to a primary study with a stated sample and method. Also the 2024 academic GEO tactics research, which predates 2026 and tested a simulated engine.