Artificial Intelligence

How to measure AI search visibility without expensive tools

Dhriti
Posted on 17/08/2613 min read
How to measure AI search visibility without expensive tools

The short answer

You can measure AI search visibility properly with a spreadsheet, a fixed set of buyer questions, and about two hours a week. We run a version of this before we run anything else at Pepper, on every account, including the ones that later move onto our platform.

Key takeaways

  • Google now reports generative AI performance for free. In June 2026 it added Search Generative AI performance reports to Search Console, covering AI Overviews, AI Mode and generative AI in Discover.
  • Manual prompt tracking is not a poor substitute. A disciplined prompt set run on a schedule is the same underlying method a paid platform automates.
  • Measure two different things. Whether engines name you, and whether they cite a page from your domain. The gap between them tells you what to fix.
  • Run it twice. Engines are probabilistic, so a single run on a single engine tells you almost nothing.
  • Free measurement stops working at a specific, identifiable point, and we say exactly where below rather than hurrying you past it.

A note on where this comes from. We run organic for more than 250 enterprises at Pepper and track over 10 million prompts across every major engine, so this method is the manual version of what we automate. We are describing it in full because a team that measures well makes a better client, and because the alternative, buying a platform before you know which question you are asking, wastes money we would rather you spent on the work.


What is AI search visibility?

AI search visibility is how often generative engines mention your brand and cite your pages when people ask questions in your category. It splits into two measurements that behave differently and need tracking separately.

Brand Visibility is whether an engine names you in the answer at all. It is the awareness number.

Domain Prompt Presence is whether the engine cites a page from your domain as a source. It is the capture number.

Those are the names Pepper’s platform uses, and they matter here because the gap between them is the most useful diagnostic you can get for free. A brand named constantly but cited rarely has an authority problem, not a content problem, and no amount of publishing will fix it. Our Visibility, Citability and Retrievability framework sets out the three levers in full.


What free actually gets you in 2026

Two things changed this year, and both are worth taking before you spend anything.

Google Search Console now reports generative AI performance. On 3 June 2026, Google introduced Search Generative AI performance reports in Search Console, giving dedicated views of impressions within generative AI features on Search, including AI Overviews and AI Mode, as well as generative AI features in Discover (Google Search Central, June 2026).

That is genuinely significant. Before it, AI Overview traffic sat inside the standard Performance report mixed with everything else, so isolating it was guesswork.

Two caveats worth holding. The reporting is impression-oriented rather than a full click and query breakdown, so it tells you where you appeared more clearly than what happened next. And rollout has been staged rather than universal, so check what is actually live in your own property rather than assuming. If the report is not there yet, the manual method below covers you in the meantime.

Bing Webmaster Tools has an AI Performance report, in public preview since February 2026, showing cited URLs and the queries that triggered them. Worth connecting even if Bing is not a priority engine, because first-party citation data is rare.

Between them, that is two free first-party sources, and both are worth connecting before you look at paid monitoring tools. Neither covers ChatGPT, Perplexity or Claude, which is where the manual method earns its place.

See where you show up. Pepper’s GEO platform tracks Brand Visibility, Domain Prompt Presence and Share of Voice across ChatGPT, Perplexity, Gemini, Claude and Google AI Overviews. Your team can log in, connect Search Console and GA4, manages the prompt set and runs your own agents in the Agent Atlas, with a growth team working the same account alongside you. Book a growth audit or see where you show up.


Where the engines differ, and why one number misleads

Before you build the tracker, know what you are sampling.

Muck Rack’s May 2026 study of more than 25 million cited links across ChatGPT, Claude and Gemini found ChatGPT carries a citation in 96 percent of responses, Gemini in 82 percent, and Claude in just 55 percent (Muck Rack, May 2026). ChatGPT averages around five citations per answer, Gemini around eight.

So the same page can look strong on one engine and absent on another, and a blended average across engines describes none of them. Track each engine in its own column. This is the single most common mistake we see in homemade trackers, and it is free to avoid.

Bar chart of citation rates by engine: ChatGPT carries a citation in 96 percent of responses, Gemini in 82 percent and Claude in 55 percent.
Figure 1: A 41-point spread is why one blended number describes none of them. Source: Muck Rack, May 2026.

The same study found earned media accounts for 84 percent of AI citations, with paid and advertorial at 0.3 percent. Which tells you where to look when your numbers are bad: usually not at your own site, and usually at the sources that get cited instead.


Free tools at a glance

SourceWhat it coversWhat it costsWhere it stops
Google Search ConsoleAI Overviews, AI Mode, generative AI in DiscoverFreeImpression-oriented, staged rollout, Google surfaces only
Bing Webmaster ToolsCited URLs and triggering queriesFreeMicrosoft surfaces, not ChatGPT or Perplexity
Manual prompt trackingChatGPT, Perplexity, Gemini, ClaudeFree, about two hours a monthNo history you did not record, no attribution
Free visibility checkersA quick read on a small prompt setFree tierPrompt set you did not choose, shallow analysis
Tracking platformAll of the above, automated, with historyFrom roughly $99 a monthCosts money, and needs someone to act on it

The method: five steps, about two hours a week

Step 1: Write the prompt set

Between 30 and 50 questions your buyers genuinely ask. Not your brand name, which surfaces your brand and proves nothing.

Mix three types. Category questions like “best platforms for this use case”. Comparison questions like “alternatives to our closest competitor”. Problem questions phrased the way a buyer would type them, with constraints attached: “how do I handle this specific problem with a small team”.

Freeze the list. The whole method depends on asking the same questions over time, and a set you keep editing produces movement you cannot interpret.

Five-stage chevron showing the method: write 30 to 50 buyer prompts, build the sheet, run it twice on different days, read the gap between mentions and citations, repeat monthly.
Figure 2: The whole method fits in about two hours a month.

Step 2: Build the sheet

One tab, one row per prompt per engine per run. Nine columns is enough:

date, engine, prompt, brand mentioned yes or no, our domain cited yes or no, which of our URLs, competitors mentioned, sources cited, and a notes field for anything odd.

That last column matters more than it looks. Wrong claims about your product, outdated pricing, a competitor being described as the category leader: these are the findings that change what you do, and they never show up in an aggregate number.

Step 3: Run it, twice

Work through the prompt set on each engine you care about. ChatGPT, Perplexity, Gemini and Claude cover most B2B buying, and Google AI Overviews is covered by Search Console.

Then run it again on a different day. Engines are probabilistic and answers move between runs. A single pass produces a number you will over-interpret; two passes tell you what is stable and what is noise. If the two runs disagree wildly on a prompt, that prompt is contested rather than lost, which is a different problem.

Step 4: Read the gap

Two pivot tables. Brand mention rate by engine, and domain citation rate by engine.

If both are low, you have a visibility problem and more of your own strong content genuinely helps. If mentions are healthy but citations are not, the engines know you and are citing somebody else, which is an authority problem living off your domain. That distinction determines everything you do next, and it costs nothing to establish.

Then look at the sources column. The domains that keep appearing across your category are your citation core, usually ten to twenty of them. Our note on citation analysis covers how to read them. Most brands have never written that list down, and it is the most valuable artefact this exercise produces.

Step 5: Repeat monthly, not weekly

Monthly is the right cadence for a manual set. Weekly produces noise you will mistake for movement, and it burns the discipline you need to keep this going past month three.

Add prompts sparingly, three to five at a time, and give them two months before judging them. That is the same guidance we give inside our own platform, because the statistical problem is identical whether you are counting by hand or not.


Where free measurement stops

Being straight about this is more useful than selling past it.

Free works well for direction: are we present, roughly where, who is winning, what sources decide our category. For a team starting out, that is most of the value, and we would rather you did this properly than bought something you cannot yet interpret.

It stops working at four specific points.

Scale. Fifty prompts across four engines twice a month is 400 manual checks. At 200 prompts it stops being feasible, and the discipline breaks before the method does.

History. A spreadsheet gives you what you recorded. It cannot tell you what an engine said in March if you were not looking in March, and competitive movement is only visible in hindsight.

Attribution. Manual tracking shows visibility. Connecting that to sessions, then pipeline, needs your analytics joined to the citation data, which is where a platform earns its cost.

Consistency. Two people scoring “was the brand mentioned” will disagree more than you expect, especially on partial or indirect mentions.

The honest read: do this manually until one of those four becomes your binding constraint. When it does, you will know exactly what you need from a platform, which makes you a much harder buyer to oversell. That is a better position than the alternative, and it is why we tell clients to start here.

Grouped bar chart comparing a manual method against a tracking platform across direction and diagnosis, engine breadth, historical trend and attribution to pipeline.
Figure 3: Manual matches paid on diagnosis. It loses on history and attribution.

See where you show up. When manual tracking hits its ceiling, Pepper’s GEO platform tracks Brand Visibility, Domain Prompt Presence and Share of Voice across ChatGPT, Perplexity, Gemini, Claude and Google AI Overviews, with a growth team working the account alongside your own. The full case study library shows how that plays out. Book a growth audit or see where you show up.


What this looks like when it works

SalesHood is the clearest example we can show of measurement turning into movement.

They were losing organic traffic and disappearing from AI answers while competing against better-funded rivals. The work started with exactly the diagnostic above: a strategic audit mapping competitor performance and keyword gaps, then GEO and AEO tracking across ChatGPT, Gemini and Claude, alongside technical SEO that resolved 10 issues including sitemap and schema work, and a content engine producing 20 new pieces and refreshing 8.

AI Overview visibility went from 14 keywords to 97. Blog impressions went from 200,000 to 417,000. Rich results grew 291 percent in five months, and clicks rose 20 percent against an industry-wide decline.

The number that matters most is the first one, because it is the one this article’s method would have surfaced. Their CMO, Elay Cohen, described the outcome in business terms rather than metrics: he could see a direct connection between the investment and lead flow from LLMs that closed.

Worth being precise: that engagement had a platform and a team behind it, not a spreadsheet. The point is that the diagnostic is the same, and the spreadsheet version tells you whether the problem is worth solving before you commit to solving it.

Three results from the SalesHood case study: AI Overview visibility from 14 to 97 keywords, blog impressions from 200,000 to 417,000, and rich results up 291 percent with clicks up 20 percent.
Figure 4: SalesHood, as published on the case study page.

Our methodology: how we tested this

We did not invent the manual method for this article. It is the audit we run at the start of engagements, simplified so it works without our platform.

CriterionWeightWhat that means here
Works with tools you already have35%Search Console, a browser and a spreadsheet, no trial signups
Produces a decision, not a number30%The output has to tell you what to fix next
Survives contact with a real week20%Two hours monthly is sustainable, two hours weekly is not
Honest about its own limits15%Stated where it breaks rather than implied

What this costs, honestly

ApproachCostWhat you get
Fully manualFree, plus about two hours a month of someone’s timeDirection, diagnosis and your citation core
Manual plus free first-party reportsFreeThe above, plus Google and Bing impression and citation data
Entry tracking platformFrom roughly $99 a monthAutomation and history, usually one engine at that tier
Mid-tier platformRoughly $300 to $1,500 a monthMulti-engine coverage and competitor share of voice
Platform with a teamCustom, usually annualMeasurement plus the execution capacity to act on it

The two hours a month is the real cost of the free option, and it is the one that actually fails. Not because the method is hard, but because it is nobody’s job.

How to choose when to upgrade

The decision is not free versus paid. It is whether you have hit one of the four ceilings above.

The scorecard I would use

Score your situation out of 100. Above 60 and a platform will pay for itself. Below 40 and you are buying convenience you cannot yet use.

AreaWeightPoints toward buying when
Prompt set size25%Above roughly 100 prompts, where manual runs stop completing
Engines you need20%More than two matter commercially, since effort multiplies per engine
Need for history20%You need to show change over quarters, not describe this month
Attribution pressure20%Someone is asking what visibility did to pipeline, and a spreadsheet cannot answer
Who runs it15%Nobody reliably has the two hours, which is the most common real reason

Run this test before you buy anything

Do one full manual cycle first. Thirty prompts, two engines, two runs, one afternoon.

Then take the result to any platform you are evaluating and ask them five questions: where do we appear across these prompts, which competitors appear instead, which sources are influencing those answers, why are those sources winning, and what would you change in the first 90 days.

You now have your own answer to compare theirs against. Almost nobody does this, and it changes the conversation completely.

The weak playbook against the strong one

The weaker version: buy a platform, watch a dashboard, wait for a number to move, cancel in month four when nobody can say what changed.

The stronger version: establish a baseline manually, identify whether the gap is visibility or citability, fix the specific thing that gap points at, and only then automate the measurement so you can see it compound.

The difference matters because a dashboard tells you the score. It does not tell you which of the two problems you have, and buying before you know that is how teams end up with excellent reporting on work nobody prioritised.

Red flags when you do start evaluating

  • “We guarantee citations.” Nobody controls a generative engine’s output, and Google gives the same warning about guaranteed rankings.
  • A single composite visibility score. It hides the mention versus citation gap, which is the useful part.
  • Branded prompts in the default set. Your own name surfaces your own brand.
  • One engine. Citation rates differ by 41 points between ChatGPT and Claude.
  • No stated method for prompt selection. Then the reporting is arbitrary.
  • No competitor share of voice. Your number rising while a competitor rises faster is a loss reported as a win.

The five questions I would ask any vendor

  1. “Show me the raw prompt runs, not the summary.” You should be able to read what the engine actually said.
  2. “How do you handle variance between runs?” Listen for a method, not reassurance.
  3. “Can I export everything?” Your prompt set and history should be yours.
  4. “How does this connect to analytics?” Attribution is the main thing you are buying over a spreadsheet.
  5. “What does this not tell me?” Anyone who says nothing is either inexperienced or selling.

One more thing worth checking, because almost nobody does. Ask a vendor to run your prompt set twice, on different days, and show you both results side by side. Any platform sampling properly will show variance, and a vendor comfortable showing you that variance is telling you something good about their methodology. A vendor whose two runs are suspiciously identical is either caching results or showing you a summary rather than the underlying data.

Reduced to one principle: measure by hand until measuring by hand is the bottleneck. You will buy better, and you will use what you buy.

One honest closing note. Very few platforms are equally strong across measurement breadth, execution and attribution, ours included, and the good ones will tell you which is their weakest without being asked.


What nobody should promise you

A guaranteed citation or position in an AI answer. Nobody controls the output.

A single composite AI visibility score sold as a ranking. Ask a model the same question repeatedly and the answers move.

A universal benchmark. There is no good number that holds across categories: 30 percent brand visibility might be excellent in a fragmented market and poor in a concentrated one. Track your own trend and the category leader.

And clean attribution from citation to closed revenue. Every link in the chain is observable, but it is not a straight line.


You may not need to measure this yet

The answer that costs us a sale. If your site is not indexed properly, your product pages are inaccurate, or nobody has written down who your buyer is, you do not need to measure AI visibility yet. You will produce a baseline that says you are absent, and the reason will be something you already knew. Fix indexing, fix the facts on your pages, write the buyer questions down, and then measure. That order costs nothing and saves a quarter.


Frequently asked questions

Can you measure AI search visibility for free?
Yes. Google Search Console now reports generative AI performance, Bing Webmaster Tools shows AI citations, and a fixed prompt set run manually across engines covers ChatGPT, Perplexity and Claude. A spreadsheet handles the rest.

What does Google Search Console show about AI search?
Since June 2026 it includes Search Generative AI performance reports covering AI Overviews, AI Mode and generative AI in Discover. Rollout has been staged, so check your own property rather than assuming availability.

How many prompts should I track?
Thirty to fifty is enough to be directional and few enough to sustain manually. Add three to five at a time and give new prompts two months before judging them, or you cannot attribute movement.

How often should I run the check?
Monthly for a manual set. Weekly produces variance you will mistake for trend, and it exhausts the discipline the method depends on. Always run each cycle at least twice on different days.

Why do the engines give different answers?
Because they cite at different rates and draw on different sources. Muck Rack found ChatGPT carries a citation in 96 percent of responses, Gemini 82 percent and Claude 55 percent, so blending them into one number hides the real picture.

What is the difference between being mentioned and being cited?
Being mentioned means the engine names your brand. Being cited means it links a page from your domain as a source. Healthy mentions with weak citations signal an authority problem rather than a content problem.

When is a paid tool actually worth it?
When one of four things binds: prompt volume beyond manual capacity, a need for historical data you did not capture, pressure to attribute visibility to pipeline, or nobody reliably having the time to run it.

Do free AI visibility checkers work?
They are useful for a quick read and less useful as a system, since most sample a small prompt set you did not choose. A prompt set you wrote yourself will always tell you more about your own category.


Where to go next

Do one cycle this week. Thirty prompts, two engines, two runs, one spreadsheet. You will finish the afternoon knowing whether your problem is that engines do not know you or that they know you and cite someone else.

Those are different problems with different fixes, and almost nobody can answer that question about their own brand today.

Book a growth audit · See where you show up


Sources and further reading

  • Google Search Central. “Introducing Search Generative AI performance reports in Search Console.” June 2026. Link
  • Muck Rack. “What Is AI Reading?” May 2026 edition. More than 25 million cited links across ChatGPT, Claude and Gemini, 17 industries. Link
  • Bing Webmaster Tools AI Performance report, public preview from February 2026.
  • Pepper. SalesHood case study. Metrics as published on the page.

Similar Posts