GEO / AI Search

How to run a GEO and AI visibility audit

Rishabh Shekhar
Posted on 10/09/2614 min read
How to run a GEO and AI visibility audit

Most audits sold as GEO audits are technical SEO audits with a section about schema bolted on. They check whether Google can crawl you, which matters, and then stop at your domain boundary, which is where the interesting part starts.

A GEO audit answers a different question. Not “can search engines reach my pages” but “when someone asks a question in my category, what does the engine say, whose pages does it cite, and why theirs rather than mine.” Most of that evidence sits on sites you do not own.

Pepper runs this on every new account, and the framework below is what our growth teams actually work through. Five stages, sequenced so the cheap checks that invalidate everything downstream come first.


The short answer

Run it in five stages, in order:

  1. Access. Can each engine’s crawler fetch and render your pages at all
  2. Baseline. What do the engines actually say about your category today
  3. Sources. Which domains are feeding those answers
  4. Gap. Why those sources rather than yours
  5. Plan. What changes in the next 90 days, and who does it

Stage 1 is free and takes ten minutes. If it fails, stages 2 to 5 are wasted, which is why it goes first.

Key takeaways

  • Four of the five stages look outside your own site. That is the main difference from a technical SEO audit.
  • Start with crawler access, always. Blocking the wrong bot is the most common fault we find, and it silently zeroes everything else.
  • Never average across engines. Independent research found identical queries produce structurally different answers per system, so an average destroys the finding.
  • The output is a source list, not a score. A ranked list of domains feeding your category is the only genuinely actionable artefact an audit produces.
  • One run proves nothing. Engines are probabilistic. Run each prompt several times, on more than one engine.

A note on where this comes from. We run organic for more than 250 enterprises and track over 10 million prompts across every major engine. What follows is shaped by that, by the client reviews we sit in weekly, and by the audits we inherit that were never implemented.


What is a GEO audit, and how is it different?

A GEO audit is a diagnostic of whether AI engines can reach your content, whether they cite it, and which sources they use instead. GEO means generative engine optimization; AEO, answer engine optimization, describes the same practice.

The distinction from a technical SEO audit is scope, not thoroughness.

A technical SEO audit examines your site: crawling, rendering, indexation, architecture, structured data. It is essential, and our enterprise SEO audit framework sets out the 47 checks we run at scale. Everything in it happens inside your domain.

A GEO audit starts with a subset of that, because a page an engine cannot fetch cannot be cited, and then spends most of its time on questions your own site cannot answer. What do engines say. Who do they cite. Why.

Google’s May 2026 guidance is worth holding alongside this. Google states that AEO and GEO are part of SEO, and that AI Overviews and AI Mode run on core Search ranking systems with no separate AI index. That is true for Google, and Google is one engine. ChatGPT and Perplexity retrieve differently, which is why the audit covers all of them separately.

If you want a lighter version to run yourself first, our seven-point AI search audit template is the quick pass. This page is the full framework.


The five stages at a glance

StageWhat it answersTimeCost
1. AccessCan each engine fetch and render our pages10 minutesFree
2. BaselineWhat do engines say about our category today1 dayFree on a free tier
3. SourcesWhich domains feed those answersHalf a dayFree, from stage 2 output
4. GapWhy those sources rather than ours1 dayFree
5. PlanWhat changes in 90 days, and who does itHalf a dayFree to produce, expensive to execute

The whole audit costs nothing but time. What it produces is a work plan whose execution is the expensive part, which is worth knowing before you commission one from anybody.

Five ordered layers showing the GEO audit sequence from crawler access through to the 90 day plan
Figure 1: Run in this order. Each stage assumes the one above it passed.

If you want this run against your category rather than done from memory, book a growth audit and we will work through the five stages with your team.


How we weight the five stages

Horizontal bar chart of the five weighted criteria for evaluating a GEO audit proposal, summing to 100
Figure 4: How we weight an audit proposal. Weights match the decision-section table exactly.
AreaWeightWhat decides the score
Access30Whether every engine’s search crawler can fetch and render your key templates. Binary, and it gates everything below
Source analysis30Whether the audit produces a frequency-ranked list of the domains actually feeding answers in your category
Baseline quality20Whether the prompt set is real commercial questions, run enough times to separate signal from variance
Gap diagnosis15Whether it explains why those sources win, rather than only listing them
Plan and ownership5Whether the output names actions, owners and hours rather than recommendations

Access and source analysis carry equal top weight for opposite reasons. Access is cheap and invalidates everything if it fails. Source analysis is the only part that produces something you can act on. Our GEO agency ranking methodology explains how we build weightings like this.


Stage 1 · Access

Ten minutes, free, and it decides whether the rest of the audit means anything.

Every major AI company runs two crawlers: one that controls whether you appear in answers, and one that only collects training data. Blocking the training crawler costs nothing. Blocking the search crawler removes you entirely, and the tokens look similar enough that this goes wrong constantly.

The checklist:

  • OAI-SearchBot allowed. This is what surfaces you in ChatGPT search. GPTBot is training only and safe to block
  • PerplexityBot allowed. Used for results, not for training. Perplexity publishes IP ranges as JSON if you want to verify by address rather than user agent
  • Googlebot allowed. AI Overviews and AI Mode run on core Search ranking systems, so ordinary access is the whole requirement. Google-Extended controls Gemini training only, and Google states it “does not impact a site’s inclusion in Google Search”
  • Rendered against raw HTML compared on every key template. Content that exists only after JavaScript runs may not be seen
  • Server response health checked. Latency and 5xx errors affect how much gets crawled
  • Key pages return 200 to a crawler, not a redirect chain or a soft 404

Where this stage falls short: it tells you access is possible, not that anything is good. A perfectly crawlable thin page is still thin.

Our engine-by-engine guide covers the crawler tokens in full, with each one checked at the vendor’s own documentation.


Stage 2 · Baseline

One day, and free if you use one of the tools with a genuine free tier.

Build the prompt set. Thirty real commercial questions, written the way a buyer types them. Not “what is generative engine optimization” but “which agency should we hire for AI search” and “best tool for tracking brand mentions in ChatGPT”. Branded questions do not count, because those answers were already yours.

Run them properly:

  • Across every engine your buyers use, at minimum ChatGPT, Perplexity and Google AI Mode
  • More than once each. Engines are probabilistic and one run establishes nothing
  • On separate days, so a single bad day does not read as a trend
  • Recorded verbatim. Save the answer text, not a summary

Record four things per prompt, per engine: whether you appear, which competitors appear instead, which specific URLs are cited, and which domains those URLs sit on.

Where this stage falls short: thirty prompts is enough to establish a pattern and not enough to run a category. It is a diagnostic sample, not a monitoring programme.

Comparison matrix showing the four things to record per prompt and the mistake each one prevents
Figure 2: What to record, and the specific error each field prevents.

Stage 3 · Sources

Half a day, and this is where the audit becomes useful.

Sort every cited domain by frequency. Across all prompts and all engines, count how often each domain appears. That list is the single most valuable artefact the audit produces, and almost nothing else you do will change your visibility as much as acting on it.

  • Rank domains by citation frequency, not by whether you have heard of them
  • Split by engine. Independent research found identical queries produce structurally different results per system, so a merged list hides real differences
  • Separate your own domain from everything else, and calculate what share of citations you hold
  • Flag community and reference sources. Research across 11,000 queries found Wikipedia and longer sources disproportionately overrepresented in AI answers
  • Mark which domains you could plausibly reach this quarter. Count honestly. If that number is near zero, you have found your real constraint

Where this stage falls short: it tells you which sources win, not how to get into them. That is a separate capability and usually the hardest one to build.


Stage 4 · Gap

One day. Why those sources and not yours.

Two independent studies give you the criteria to check against, and both are academic rather than vendor research.

Zhang, He and Yao measured 21,143 citations across ChatGPT, Google AI Overview and Perplexity in April 2026 and found the pages engines actually absorb are longer, more structured, semantically aligned, and richer in extractable evidence: definitions, numerical facts, comparisons and procedural steps. They also concluded that citation counts alone inadequately measure effectiveness.

Compare your page against the ones being cited:

  • Is there a standalone definition an engine could lift without rewriting
  • Are there specific numbers, sourced and dated, rather than general claims
  • Is there a genuine comparison, structured as a table rather than prose
  • Are there procedural steps where the question implies a process
  • Is the key claim liftable as a single self-contained passage
  • Is the page longer and more structured than the ones outranking it, or thinner
  • Are dates visible and accurate on anything time-sensitive

Then check the two Pepper-native numbers. Brand Visibility is how often engines mention you by name. Domain Prompt Presence is how often they cite a page from your domain. A wide gap between them is the diagnosis: engines know who you are and trust somebody else to describe you. That is a citability problem, and it is solved off your own site, not on it.

Where this stage falls short: it explains the gap, it does not close it. Closing it usually needs earned media, which is the most expensive capability in this discipline.


Stage 5 · Plan

Half a day, and the reason most audits fail.

An audit that ends in findings is a document. An audit that ends in a sequenced plan with named owners is a piece of work. The difference is entirely in this stage.

  • Sequence by dependency, not by severity. Access, then rendering, then extractability, then off-site
  • Name an owner per item, a person rather than a team
  • Estimate hours per item, so the plan can be resourced honestly
  • State what comes off that person’s plate to create the time
  • Set the 90-day checkpoint and what would count as success
  • Identify the single binding constraint. Usually it is earned media, occasionally engineering time, rarely content

Where this stage falls short, including for us: we can produce the plan. Whether anyone executes it depends on capacity we do not control, and that is the most common reason a good audit changes nothing.


What a GEO audit costs

Stacked bar chart showing the audit itself is a small fraction of total cost compared with executing the findings
Figure 3: The audit is the cheap part. Illustrative shape, assumptions on the chart.

The audit itself is close to free if you run it yourself. Measurement tooling now starts at zero, since several platforms have genuine free tiers, covered in the best answer engine optimization tools.

What costs money is stage 5 being executed. Rendering fixes need engineering time. Restructuring pages needs editorial time. Getting onto the sources from stage 3 needs earned media, which is the largest line and cannot be done from your own site.

Budget the execution before you commission the diagnosis. Our cost breakdown of platforms against hiring models the tradeoff, and SEO audit agencies covers what to pay when buying one in.


What nobody should promise you

A single AI visibility score. Composite numbers hide the gap between Brand Visibility and Domain Prompt Presence, which is the finding you could act on.

Guaranteed citations after remediation. Google advises against providers guaranteeing rankings, because third parties cannot access internal ranking systems. The same holds for AI answers.

A meaningful reading from one run. Engines are probabilistic. A single run on a single engine is a sample of one presented as a fact.

An audit as the deliverable. It is a diagnosis. If nobody has hours to act on it, you have bought a document.


How to evaluate a GEO audit before you commission one

I would judge the proposal on where it spends its time, because that single question separates a real diagnostic from a technical audit with an AI section attached.

Google’s May 2026 guidance anchors the scope. AEO and GEO are part of SEO, and AI Overviews and AI Mode run on core Search ranking systems with no separate AI index. So the technical half is work a competent SEO team already does, and any proposal treating it as an entirely new discipline is overstating its novelty. What is genuinely new is the off-site half: which sources feed the answers, and why.

Score the proposal before you sign.

AreaWeightWhat a strong proposal demonstrates
Time spent off-site30Most of the engagement analysing sources feeding your category, not auditing your own site again
Named prompt methodology25How the prompt set is built, how many runs per prompt, and across which engines, stated before starting
Per-engine reporting20Results kept separate by engine rather than averaged, with the reason understood
Raw evidence handed over15You receive the verbatim answers and cited URLs, not only a score and a summary
Plan with owners and hours10Findings arrive sequenced by dependency with named owners, not sorted by severity

Those weights sum to 100. Score your own readiness too, because an audit landing on a team with no capacity produces a backlog rather than a result.

The live test, before you commission anything. Give every shortlisted provider the same 30 real commercial prompts from your category. Four worked examples: “which GEO agency should we hire”, “how do we get cited by ChatGPT in our category”, “best AI visibility platform for a mid-market company”, “how do we run a GEO audit ourselves”.

Ask for five things back. Where do we appear and where do we not. Which competitors appear instead, consistently. Which sources influence those answers, sorted by frequency. Why are those sources winning. What would you change in the next 90 days, and who does it.

A provider who returns a frequency-ranked domain list is running a real audit. A provider who returns a composite score and a list of on-page fixes is running a technical SEO audit with a new label.

The weak audit against the strong one. The weaker version crawls your site, checks schema, recommends FAQ blocks, and reports an AI visibility score. Every finding is inside your domain, which is the smaller half of the problem, and the score cannot be audited because the prompt set is undisclosed.

The stronger version spends stage 1 on access, then most of its time outside your site: what engines say, who they cite, why those sources win. It ends in a ranked list of domains and a sequenced plan. It usually produces fewer findings and more change.

Red flags, each one something you will genuinely hear. “Here is your AI visibility score”, offered as a composite with no prompt set disclosed. “We guarantee citations.” “We tracked your branded prompts”, which always looks healthy because those answers were already yours. “We ran it on ChatGPT”, where one engine stands in for the category. “Add FAQ schema and you will get cited.” “The audit is free if you sign the retainer”, which tells you what it is designed to conclude. And the quiet one: a scope that never mentions robots.txt or crawler access at all.

Five questions for the first call.

  1. How many prompts, how many runs each, and on which engines? A good answer is three numbers. A weak one is that they run a comprehensive analysis.
  2. Will we get the verbatim answers and cited URLs? A good answer is yes with a format. A weak one explains why the score is more useful.
  3. What proportion of the engagement is off-site analysis? A good answer is most of it. A weak one describes a site crawl.
  4. Will results be reported per engine or averaged? A good answer knows why averaging is wrong.
  5. Does the output include owners and hours? A good answer is developer-ready. A weak one is a slide deck.

It all comes down to one principle: a GEO audit that never leaves your own website has audited the smaller half of the problem. Check access first, then go and look at what the engines are actually citing.

One closing note that costs us something. Very few providers are equally strong across measurement, execution, earned authority and attribution, ours included, and the good ones will tell you which of the four is their weakest. If stage 1 turns out to be your whole problem, fix the robots.txt and commission nothing.


Where Pepper fits

Pepper is an agentic organic growth engine, which matters here because stages 1 to 4 are diagnosis and stage 5 is work.

  • What you log into: workspace setup, brand profile, competitors and personas. GA4 and Search Console connected. Themes and prompts managed, GEO analytics read directly, and your own agents built and run in the Agent Atlas
  • Engines tracked: 6, including ChatGPT, Perplexity, Gemini and Google AI Overviews, reported separately rather than averaged
  • And a growth team is attached to the account, doing the stage 5 work alongside your people: access and rendering fixes, restructuring, and the earned media that closes the source gap
  • Metrics reported: Brand Visibility, Domain Prompt Presence and Share of Voice, with the gap between the first two as the citability diagnostic
  • Best for: teams who want the audit acted on rather than delivered
  • Where it falls short: we do not publish pricing, so you cannot size it before a call. A standalone audit document is not what we sell, so if that is genuinely all you need, run the five stages yourself using this page

Frequently asked questions

What is a GEO audit?
A diagnostic of whether AI engines can reach your content, whether they cite it, and which sources they use instead. It differs from a technical SEO audit because four of its five stages examine sources outside your own website.

How is a GEO audit different from an SEO audit?
An SEO audit examines your site: crawling, rendering, indexation, structure. A GEO audit starts with a subset of that, then spends most of its time on what engines say about your category and which domains they cite.

How long does a GEO audit take?
Around three days of work if run properly: ten minutes on access, a day on the baseline, half a day on sources, a day on the gap analysis, and half a day producing the plan. Executing the plan takes considerably longer.

What does a GEO audit cost?
The audit itself is close to free, since measurement tools now have genuine free tiers. The cost sits in executing the findings, particularly earned media, which cannot be done from your own site and is usually the largest line.

How many prompts should a GEO audit use?
Around 30 real commercial questions, run several times each across every engine your buyers use. Fewer than that is a sample too small to separate a pattern from ordinary run-to-run variance.

Can I run a GEO audit myself?
Yes, and you should run stage 1 today regardless. Checking your robots.txt for six crawler tokens takes ten minutes and is the highest-return check in this discipline.

What is the most common problem a GEO audit finds?
Blocked crawler access, usually because a team blocked all bots with AI in the name without realising the search crawler and the training crawler are separate tokens.

What should the audit produce?
A frequency-ranked list of the domains feeding answers in your category, kept separate per engine, plus a sequenced plan with named owners. A composite score is not an output you can act on.


Where to go next

Run stage 1 now. Open your robots.txt and check six tokens: OAI-SearchBot, GPTBot, PerplexityBot, Perplexity-User, Googlebot, Google-Extended. If a search crawler is blocked, nothing else on this page matters until that changes.

Then build the thirty-prompt set and run stage 2 on a free tier. Most teams discover their real constraint within a day.

Our engine-by-engine optimisation guide covers the access checks in full, the seven-point AI search audit template is the lighter version to run yourself, and the Visibility, Citability and Retrievability framework explains which lever is stuck.

For the technical half at scale, our enterprise SEO audit framework sets out 47 checks. For tooling, the best answer engine optimization tools covers which have free tiers. Our Acceldata case study shows the work on a technical B2B account.

The honest exit. If your crawler access is clean and you have fewer than about 30 meaningful commercial prompts in your category, you do not need a formal audit. Run the five stages informally once a quarter and spend the budget on earned media instead.


Sources

Crawler tokens were checked at each company’s own documentation on 10 September 2026. Every study cited was published in 2026 and checked at the original source.

  • OpenAI, Bots and crawlers documentation, checked 10 September 2026. OAI-SearchBot surfaces sites in ChatGPT search; GPTBot is for training.
  • Perplexity, Crawler documentation, checked 10 September 2026. PerplexityBot surfaces results and is not used for training; Perplexity-User generally ignores robots.txt.
  • Google Search Central, Google common crawlers, checked 10 September 2026. Google-Extended “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.”
  • Google Search Central, AI features and your website, 15 May 2026. States AEO and GEO are part of SEO, and advises against providers guaranteeing rankings.
  • Zhang Kai, He Xinyue and Yao Jingang, From Citation Selection to Citation Absorption, arXiv, 28 April 2026. 602 prompts, 21,143 citations, 18,151 fetched pages. Academic and independent.
  • Huang, Goyal, Saha and Chandrasekharan, Answer Bubbles: Information Exposure in AI-Mediated Search, arXiv, 17 March 2026. 11,000 real search queries across five systems. Academic and independent.
  • Muck Rack, Earned media still drives 84% of AI citations, 7 May 2026. More than 25 million cited links, 17 industries. Muck Rack sells PR software, so read it as an interested party with a large dataset.