How to benchmark AI visibility against competitors, honestly

The short answer
You cannot see a competitor’s analytics, so every AI visibility benchmark is built the same way: run a set of prompts, count who gets named, and compare.
That makes the prompt set the whole instrument. Change it and the ranking changes, which has three consequences worth knowing before you build one.
- A benchmark measures your prompt set, not the market. Pick prompts where you already do well and you will produce a flattering chart that predicts nothing.
- Share of voice is not a standard definition. Vendors count mentions, first mentions, or weighted position differently. Two tools will disagree, and both can be right.
- Sampling depth decides whether a gap is real. The same prompt returns different answers between runs, so a five-point difference on three observations is noise.
The fix is not a better tool. It is building the prompt set honestly, fixing it in writing before you look at the results, and reporting the comparison alongside its own limits. That is what this page walks through.
We run this work daily. Pepper is an agentic organic growth engine, not an SEO agency. This page comes out of what we run: organic for more than 250 enterprises across eight years, more than 10 million tracked prompts, and the questions buyers put to us in client reviews. Customers log in and run the platform themselves, with a Pepper growth team attached. Book a growth audit and we will run a competitor benchmark on a prompt set you approve first.
Key takeaways
Key takeaways from the benchmarking questions buyers bring us most often.
- Write the prompt set before you look at any result. Choosing prompts after seeing the data is how benchmarks become marketing.
- Include prompts you expect to lose. A set with no losses in it is not a benchmark, it is a highlight reel.
- Fix the definition of share of voice in writing, because vendors do not agree on one.
- Count three things separately: who is mentioned, who is cited, and who appears first. They rank differently.
- Sample deep enough to see past variance. Fifty prompts run daily produce roughly 1,500 observations a month against 50 on a monthly cadence.
- Benchmark quarterly, not weekly. Weekly competitor charts mostly measure answer-to-answer noise.
What is an AI visibility benchmark, and what is it actually measuring?
An AI visibility benchmark is a comparison of how often you and named competitors appear in AI answers across a defined set of prompts.
The definition matters because of that last clause. In traditional SEO you benchmark against a keyword universe that exists independently of you: search volume is a fact about the world. In AI search there is no equivalent. You assemble the prompt set, so the instrument is partly your own construction.
That is not a reason to skip benchmarking. It is a reason to treat the prompt set as the most important artefact in the exercise, and to document it the way you would document a methodology.
Where it falls short. Even a well-built benchmark tells you about relative presence in answers, not about pipeline. A competitor can be named more often than you and sell less. Treat the benchmark as a diagnostic of visibility, not as a proxy for commercial performance, and pair it with the GEO ROI model if you need the money question answered.

Step 1: Build the prompt set before you look at anything
This is the step that decides whether the rest is worth doing. Write the list, get it agreed, and date it.
- Start from buyer language, not brand language. Category-shaped questions, phrased the way someone who has never heard of you would ask them.
- Cover the funnel deliberately. A handful of definitional prompts, a larger block of comparison and selection prompts, and a few that name competitors directly.
- Include prompts you expect to lose. If every prompt is one you already win, the benchmark cannot detect a problem, which is the only thing it exists to do.
- Aim for forty to eighty. Enough to cover the category, few enough to sample deeply within a sensible budget.
- Freeze it and date it. Adding prompts mid-quarter breaks comparability, and removing a losing prompt quietly is the single most common way these exercises go wrong.
Step 2: Choose competitors you actually lose to
Three to five, not ten. The list should be uncomfortable.
- Pick who buyers name, not who you consider a peer, and if you are weighing whether to run this in house at all, should you buy a GEO platform or hire a team covers that split. Ask your sales team who comes up on calls, because that list is usually different from the marketing one.
- Include one incumbent and one insurgent. The incumbent tells you what authority looks like in your category. The insurgent tells you what is changing.
- Include the aggregators. In many categories the thing beating you in AI answers is not a competitor at all, it is a review site or a listicle. Counting those separately is more useful than ignoring them.
- Do not include everyone. Every competitor you add dilutes the chart and adds sampling cost without adding insight.
Step 3: Count three things, separately
A single share-of-voice number hides the diagnosis. Count these three and you can tell what to do.

- Mention rate. How often each brand is named at all. This is presence in the category conversation.
- Citation rate. How often an engine links a page from that brand’s own domain. This is whether the brand is the source or merely the subject.
- First position. How often each brand appears first in the answer rather than in a trailing list.
Why they separate. A competitor can be mentioned more often than you while you are cited more often, which means they have more awareness and you have better pages. Those call for opposite responses. The gap between mention and citation is the diagnosis, and our framework for acting on it is Visibility, Citability and Retrievability.
Step 4: Fix the definition of share of voice in writing
Vendors do not agree on what this means, so decide yours and write it down.
- Mentions over total mentions. Simplest and most common. Your named appearances divided by all brand appearances across the set.
- Prompts where you appear, over prompts run. A coverage measure rather than a volume one. Less sensitive to an answer that lists fifteen brands.
- Weighted by position. First mention counts more than fifth. Closer to commercial reality and harder to compute consistently.
Pick one, state it in the report, and do not change it mid-year. A share of voice figure without its definition attached is not comparable to anything, including its own value last quarter. We work through the wider measurement set in AI search visibility metrics and KPIs.
Step 5: Sample deeply enough to see past the noise
The same prompt does not return the same answer every time. That variance is the floor your comparison sits on.

- On a monthly cadence, a quarter gives you three observations per prompt. A five-point gap between you and a competitor is indistinguishable from noise.
- On a daily cadence, the same quarter gives you about ninety, and small differences become readable.
- Prompt count is not sample size. Fifty prompts sampled daily produce far more evidence than 350 sampled monthly, at a lower price.
- Ask the vendor for the figure. Many do not publish it, and at least one withdrew a disclosure it used to make. The platforms that do publish it are compared in top AI visibility platforms.
Step 6: Report it with its own limits attached
A benchmark that hides its construction gets discounted the moment someone asks how it was built. Put the construction in the report.
- The prompt set size, and the date it was frozen.
- Your sampling cadence, and the resulting observation count per prompt.
- Which share of voice definition you used, in one line.
- Name the prompts you lose, not just the aggregate. The losing prompts are the actionable part.
- Say what it cannot tell you: nothing about pipeline, nothing about why, and nothing about prompts outside the set.


| Criterion | Weight | What earns full marks |
|---|---|---|
| Prompt set frozen before results seen | 30 | Written, dated and agreed in advance |
| Sampling depth stated | 25 | Observations per prompt, not just prompt count |
| Three counts kept separate | 20 | Mention, citation and first position reported apart |
| Share of voice definition stated | 15 | One definition, unchanged across periods |
| Losing prompts named | 10 | The specific prompts, not just the total |
Total 100. The first two carry more than half because they are what separate a benchmark from a chart, and both are things you control rather than things you buy.
What stops a benchmark of AI visibility against competitors being useful
These are the ways competitor benchmarks go wrong in practice, and all four are avoidable.
- Prompt shopping. Choosing the prompt set after seeing which ones you win. The tell is a set with no losses in it.
- Comparing across tools. Two vendors counting differently will disagree, and neither is wrong. Pick one instrument and stay on it.
- Weekly reporting. Answer-to-answer variance dominates at that frequency, so you end up explaining noise to your executives every Monday.
- Treating aggregators as competitors. When a review site outranks everyone, the answer is usually to get listed well on it, not to try to outrank it.
When not to benchmark at all
And here is the answer that loses us the sale. If you have fewer than about twenty prompts worth tracking, or nobody who will act on a losing prompt when they see one, you do not need a competitor benchmark and you do not need us to run one. Run your ten most important prompts by hand once a quarter, write down who got named, and spend the budget on the pages and coverage those answers are already reaching for.
A benchmark tells you that a competitor is winning. It does not tell you why, and it never does the work.
What nobody should promise you
- An objective market-wide benchmark. Every benchmark is built on a chosen prompt set, so none of them is the market. Anyone claiming otherwise is not describing how this works.
- A share of voice figure comparable across vendors. Definitions differ and most are unpublished.
- A guaranteed improvement in your ranking against a named competitor. Nobody controls what a model outputs, and we will not promise it either.
- A benchmark that explains why. It tells you the gap exists. Finding the cause is a separate piece of work.
- That a weekly chart means anything. At most sampling cadences, weekly movement is variance.
Frequently asked questions
How do I benchmark my AI visibility against competitors?
Freeze a prompt set of forty to eighty category questions before you look at any result, pick three to five competitors buyers actually name, then count mention rate, citation rate and first position separately across enough samples to see past variance.
Why do two tools give me different competitor rankings?
Because they run different prompt sets, sample at different depths and define share of voice differently. Two disagreeing benchmarks is the expected outcome, not evidence that either is broken.
Building the benchmark
How many prompts should a competitor benchmark use?
Forty to eighty for most categories. Enough to cover the buying questions, few enough to sample each one deeply within a sensible budget. Sampling depth matters more than prompt count.
Which competitors should I include?
Three to five that buyers actually name on sales calls, including one incumbent and one insurgent. Count review sites and aggregators separately, because in many categories they are what is really beating you.
Should I include prompts I expect to lose?
Yes, and a set without them is not a benchmark. The losing prompts are the only actionable part of the exercise, and excluding them produces a chart that cannot detect the problem it exists to find.
Reading the results
What is a good share of voice in AI search?
There is no cross-industry answer, because the figure depends entirely on your prompt set and the definition used. It is meaningful against your own history with both held constant, and against nothing else.
How often should I re-run the benchmark?
Quarterly for most teams. Weekly reporting mostly surfaces answer-to-answer variance rather than real movement, and monthly is only readable if your sampling cadence is daily. Re-run sooner only if you have changed something substantial and want to check it landed.
What if a review site beats every brand in my category?
That is common, and it changes the task. Getting listed and reviewed well on that source is usually a faster path than trying to outrank it with your own pages.
Sources
This page is a method, not a study. Where it rests on our judgement or our own product, we say so.
Research
- Huang, Goyal, Saha and Chandrasekharan, Answer Bubbles: Information Exposure in AI-Mediated Search, arXiv, 17 March 2026. 11,000 real search queries across five systems. Finds Wikipedia and longer third-party sources cited far more often than their share of the web would suggest. Academic and independent. Used here for the point that aggregators and reference sources often outrank brands in AI answers.
Google documentation
- Google Search Central, AI features and your website, 15 May 2026. Places AEO and GEO inside SEO, and advises against providers guaranteeing rankings.
Our own published work
- Pepper, Top AI visibility platforms. Records which platforms publish sampling depth and which do not, including one vendor that withdrew its published response volumes during 2026.
- Pepper, our fourteen AI search measurements. The fourteen measurements this page draws its three counts from.
- Pepper, How to track brand mentions in AI search. The five measurement methods, of which this page develops one in depth.
On what this page is
We have no survey data on how organisations benchmark AI visibility, and we are not aware of one. The prompt set sizes, competitor counts and cadence recommendations are our judgement from running this work, and the weighting table is ours rather than an industry standard. Use it, change it, or build your own, but write down whichever version you use before you look at the results.
Latest Blogs
Google now reports AI search two ways, and neither one joins to the other. Search Console shows impressions with no clicks. GA4 shows arrivals from a source list Google does not publish. Here is how to read both
You do not need a study to find out which sources AI cites in your category. You need forty prompts and an afternoon. Here is the procedure, and why the published figures are worth less than your own list.
A competitor benchmark in AI search measures your prompt set, not the market. Change the prompts and the ranking changes. Here is how to build one that survives that fact.