How to run a GEO and AI visibility audit

Most audits sold as GEO audits are technical SEO audits with a section about schema bolted on. They check whether Google can crawl you, which matters, and then stop at your domain boundary, which is where the interesting part starts.
A GEO audit answers a different question. Not “can search engines reach my pages” but “when someone asks a question in my category, what does the engine say, whose pages does it cite, and why theirs rather than mine.” Most of that evidence sits on sites you do not own.
Pepper runs this on every new account, and the framework below is what our growth teams actually work through. Five stages, sequenced so the cheap checks that invalidate everything downstream come first.
The short answer
Run it in five stages, in order:
- Access. Can each engine’s crawler fetch and render your pages at all
- Baseline. What do the engines actually say about your category today
- Sources. Which domains are feeding those answers
- Gap. Why those sources rather than yours
- Plan. What changes in the next 90 days, and who does it
Stage 1 is free and takes ten minutes. If it fails, stages 2 to 5 are wasted, which is why it goes first.
Key takeaways
- Four of the five stages look outside your own site. That is the main difference from a technical SEO audit.
- Start with crawler access, always. Blocking the wrong bot is the most common fault we find, and it silently zeroes everything else.
- Never average across engines. Independent research found identical queries produce structurally different answers per system, so an average destroys the finding.
- The output is a source list, not a score. A ranked list of domains feeding your category is the only genuinely actionable artefact an audit produces.
- One run proves nothing. Engines are probabilistic. Run each prompt several times, on more than one engine.
A note on where this comes from. We run organic for more than 250 enterprises and track over 10 million prompts across every major engine. What follows is shaped by that, by the client reviews we sit in weekly, and by the audits we inherit that were never implemented.
What is a GEO audit, and how is it different?
A GEO audit is a diagnostic of whether AI engines can reach your content, whether they cite it, and which sources they use instead. GEO means generative engine optimization; AEO, answer engine optimization, describes the same practice.
The distinction from a technical SEO audit is scope, not thoroughness.
A technical SEO audit examines your site: crawling, rendering, indexation, architecture, structured data. It is essential, and our enterprise SEO audit framework sets out the 47 checks we run at scale. Everything in it happens inside your domain.
A GEO audit starts with a subset of that, because a page an engine cannot fetch cannot be cited, and then spends most of its time on questions your own site cannot answer. What do engines say. Who do they cite. Why.
Google’s May 2026 guidance is worth holding alongside this. Google states that AEO and GEO are part of SEO, and that AI Overviews and AI Mode run on core Search ranking systems with no separate AI index. That is true for Google, and Google is one engine. ChatGPT and Perplexity retrieve differently, which is why the audit covers all of them separately.
If you want a lighter version to run yourself first, our seven-point AI search audit template is the quick pass. This page is the full framework.
The five stages at a glance
| Stage | What it answers | Time | Cost |
|---|---|---|---|
| 1. Access | Can each engine fetch and render our pages | 10 minutes | Free |
| 2. Baseline | What do engines say about our category today | 1 day | Free on a free tier |
| 3. Sources | Which domains feed those answers | Half a day | Free, from stage 2 output |
| 4. Gap | Why those sources rather than ours | 1 day | Free |
| 5. Plan | What changes in 90 days, and who does it | Half a day | Free to produce, expensive to execute |
The whole audit costs nothing but time. What it produces is a work plan whose execution is the expensive part, which is worth knowing before you commission one from anybody.

If you want this run against your category rather than done from memory, book a growth audit and we will work through the five stages with your team.
How we weight the five stages

| Area | Weight | What decides the score |
|---|---|---|
| Access | 30 | Whether every engine’s search crawler can fetch and render your key templates. Binary, and it gates everything below |
| Source analysis | 30 | Whether the audit produces a frequency-ranked list of the domains actually feeding answers in your category |
| Baseline quality | 20 | Whether the prompt set is real commercial questions, run enough times to separate signal from variance |
| Gap diagnosis | 15 | Whether it explains why those sources win, rather than only listing them |
| Plan and ownership | 5 | Whether the output names actions, owners and hours rather than recommendations |
Access and source analysis carry equal top weight for opposite reasons. Access is cheap and invalidates everything if it fails. Source analysis is the only part that produces something you can act on. Our GEO agency ranking methodology explains how we build weightings like this.
Stage 1 · Access
Ten minutes, free, and it decides whether the rest of the audit means anything.
Every major AI company runs two crawlers: one that controls whether you appear in answers, and one that only collects training data. Blocking the training crawler costs nothing. Blocking the search crawler removes you entirely, and the tokens look similar enough that this goes wrong constantly.
The checklist:
OAI-SearchBotallowed. This is what surfaces you in ChatGPT search.GPTBotis training only and safe to blockPerplexityBotallowed. Used for results, not for training. Perplexity publishes IP ranges as JSON if you want to verify by address rather than user agentGooglebotallowed. AI Overviews and AI Mode run on core Search ranking systems, so ordinary access is the whole requirement.Google-Extendedcontrols Gemini training only, and Google states it “does not impact a site’s inclusion in Google Search”- Rendered against raw HTML compared on every key template. Content that exists only after JavaScript runs may not be seen
- Server response health checked. Latency and 5xx errors affect how much gets crawled
- Key pages return 200 to a crawler, not a redirect chain or a soft 404
Where this stage falls short: it tells you access is possible, not that anything is good. A perfectly crawlable thin page is still thin.
Our engine-by-engine guide covers the crawler tokens in full, with each one checked at the vendor’s own documentation.
Stage 2 · Baseline
One day, and free if you use one of the tools with a genuine free tier.
Build the prompt set. Thirty real commercial questions, written the way a buyer types them. Not “what is generative engine optimization” but “which agency should we hire for AI search” and “best tool for tracking brand mentions in ChatGPT”. Branded questions do not count, because those answers were already yours.
Run them properly:
- Across every engine your buyers use, at minimum ChatGPT, Perplexity and Google AI Mode
- More than once each. Engines are probabilistic and one run establishes nothing
- On separate days, so a single bad day does not read as a trend
- Recorded verbatim. Save the answer text, not a summary
Record four things per prompt, per engine: whether you appear, which competitors appear instead, which specific URLs are cited, and which domains those URLs sit on.
Where this stage falls short: thirty prompts is enough to establish a pattern and not enough to run a category. It is a diagnostic sample, not a monitoring programme.

Stage 3 · Sources
Half a day, and this is where the audit becomes useful.
Sort every cited domain by frequency. Across all prompts and all engines, count how often each domain appears. That list is the single most valuable artefact the audit produces, and almost nothing else you do will change your visibility as much as acting on it.
- Rank domains by citation frequency, not by whether you have heard of them
- Split by engine. Independent research found identical queries produce structurally different results per system, so a merged list hides real differences
- Separate your own domain from everything else, and calculate what share of citations you hold
- Flag community and reference sources. Research across 11,000 queries found Wikipedia and longer sources disproportionately overrepresented in AI answers
- Mark which domains you could plausibly reach this quarter. Count honestly. If that number is near zero, you have found your real constraint
Where this stage falls short: it tells you which sources win, not how to get into them. That is a separate capability and usually the hardest one to build.
Stage 4 · Gap
One day. Why those sources and not yours.
Two independent studies give you the criteria to check against, and both are academic rather than vendor research.
Zhang, He and Yao measured 21,143 citations across ChatGPT, Google AI Overview and Perplexity in April 2026 and found the pages engines actually absorb are longer, more structured, semantically aligned, and richer in extractable evidence: definitions, numerical facts, comparisons and procedural steps. They also concluded that citation counts alone inadequately measure effectiveness.
Compare your page against the ones being cited:
- Is there a standalone definition an engine could lift without rewriting
- Are there specific numbers, sourced and dated, rather than general claims
- Is there a genuine comparison, structured as a table rather than prose
- Are there procedural steps where the question implies a process
- Is the key claim liftable as a single self-contained passage
- Is the page longer and more structured than the ones outranking it, or thinner
- Are dates visible and accurate on anything time-sensitive
Then check the two Pepper-native numbers. Brand Visibility is how often engines mention you by name. Domain Prompt Presence is how often they cite a page from your domain. A wide gap between them is the diagnosis: engines know who you are and trust somebody else to describe you. That is a citability problem, and it is solved off your own site, not on it.
Where this stage falls short: it explains the gap, it does not close it. Closing it usually needs earned media, which is the most expensive capability in this discipline.
Stage 5 · Plan
Half a day, and the reason most audits fail.
An audit that ends in findings is a document. An audit that ends in a sequenced plan with named owners is a piece of work. The difference is entirely in this stage.
- Sequence by dependency, not by severity. Access, then rendering, then extractability, then off-site
- Name an owner per item, a person rather than a team
- Estimate hours per item, so the plan can be resourced honestly
- State what comes off that person’s plate to create the time
- Set the 90-day checkpoint and what would count as success
- Identify the single binding constraint. Usually it is earned media, occasionally engineering time, rarely content
Where this stage falls short, including for us: we can produce the plan. Whether anyone executes it depends on capacity we do not control, and that is the most common reason a good audit changes nothing.
What a GEO audit costs

The audit itself is close to free if you run it yourself. Measurement tooling now starts at zero, since several platforms have genuine free tiers, covered in the best answer engine optimization tools.
What costs money is stage 5 being executed. Rendering fixes need engineering time. Restructuring pages needs editorial time. Getting onto the sources from stage 3 needs earned media, which is the largest line and cannot be done from your own site.
Budget the execution before you commission the diagnosis. Our cost breakdown of platforms against hiring models the tradeoff, and SEO audit agencies covers what to pay when buying one in.
What nobody should promise you
A single AI visibility score. Composite numbers hide the gap between Brand Visibility and Domain Prompt Presence, which is the finding you could act on.
Guaranteed citations after remediation. Google advises against providers guaranteeing rankings, because third parties cannot access internal ranking systems. The same holds for AI answers.
A meaningful reading from one run. Engines are probabilistic. A single run on a single engine is a sample of one presented as a fact.
An audit as the deliverable. It is a diagnosis. If nobody has hours to act on it, you have bought a document.
How to evaluate a GEO audit before you commission one
I would judge the proposal on where it spends its time, because that single question separates a real diagnostic from a technical audit with an AI section attached.
Google’s May 2026 guidance anchors the scope. AEO and GEO are part of SEO, and AI Overviews and AI Mode run on core Search ranking systems with no separate AI index. So the technical half is work a competent SEO team already does, and any proposal treating it as an entirely new discipline is overstating its novelty. What is genuinely new is the off-site half: which sources feed the answers, and why.
Score the proposal before you sign.
| Area | Weight | What a strong proposal demonstrates |
|---|---|---|
| Time spent off-site | 30 | Most of the engagement analysing sources feeding your category, not auditing your own site again |
| Named prompt methodology | 25 | How the prompt set is built, how many runs per prompt, and across which engines, stated before starting |
| Per-engine reporting | 20 | Results kept separate by engine rather than averaged, with the reason understood |
| Raw evidence handed over | 15 | You receive the verbatim answers and cited URLs, not only a score and a summary |
| Plan with owners and hours | 10 | Findings arrive sequenced by dependency with named owners, not sorted by severity |
Those weights sum to 100. Score your own readiness too, because an audit landing on a team with no capacity produces a backlog rather than a result.
The live test, before you commission anything. Give every shortlisted provider the same 30 real commercial prompts from your category. Four worked examples: “which GEO agency should we hire”, “how do we get cited by ChatGPT in our category”, “best AI visibility platform for a mid-market company”, “how do we run a GEO audit ourselves”.
Ask for five things back. Where do we appear and where do we not. Which competitors appear instead, consistently. Which sources influence those answers, sorted by frequency. Why are those sources winning. What would you change in the next 90 days, and who does it.
A provider who returns a frequency-ranked domain list is running a real audit. A provider who returns a composite score and a list of on-page fixes is running a technical SEO audit with a new label.
The weak audit against the strong one. The weaker version crawls your site, checks schema, recommends FAQ blocks, and reports an AI visibility score. Every finding is inside your domain, which is the smaller half of the problem, and the score cannot be audited because the prompt set is undisclosed.
The stronger version spends stage 1 on access, then most of its time outside your site: what engines say, who they cite, why those sources win. It ends in a ranked list of domains and a sequenced plan. It usually produces fewer findings and more change.
Red flags, each one something you will genuinely hear. “Here is your AI visibility score”, offered as a composite with no prompt set disclosed. “We guarantee citations.” “We tracked your branded prompts”, which always looks healthy because those answers were already yours. “We ran it on ChatGPT”, where one engine stands in for the category. “Add FAQ schema and you will get cited.” “The audit is free if you sign the retainer”, which tells you what it is designed to conclude. And the quiet one: a scope that never mentions robots.txt or crawler access at all.
Five questions for the first call.
- How many prompts, how many runs each, and on which engines? A good answer is three numbers. A weak one is that they run a comprehensive analysis.
- Will we get the verbatim answers and cited URLs? A good answer is yes with a format. A weak one explains why the score is more useful.
- What proportion of the engagement is off-site analysis? A good answer is most of it. A weak one describes a site crawl.
- Will results be reported per engine or averaged? A good answer knows why averaging is wrong.
- Does the output include owners and hours? A good answer is developer-ready. A weak one is a slide deck.
It all comes down to one principle: a GEO audit that never leaves your own website has audited the smaller half of the problem. Check access first, then go and look at what the engines are actually citing.
One closing note that costs us something. Very few providers are equally strong across measurement, execution, earned authority and attribution, ours included, and the good ones will tell you which of the four is their weakest. If stage 1 turns out to be your whole problem, fix the robots.txt and commission nothing.
Where Pepper fits
Pepper is an agentic organic growth engine, which matters here because stages 1 to 4 are diagnosis and stage 5 is work.
- What you log into: workspace setup, brand profile, competitors and personas. GA4 and Search Console connected. Themes and prompts managed, GEO analytics read directly, and your own agents built and run in the Agent Atlas
- Engines tracked: 6, including ChatGPT, Perplexity, Gemini and Google AI Overviews, reported separately rather than averaged
- And a growth team is attached to the account, doing the stage 5 work alongside your people: access and rendering fixes, restructuring, and the earned media that closes the source gap
- Metrics reported: Brand Visibility, Domain Prompt Presence and Share of Voice, with the gap between the first two as the citability diagnostic
- Best for: teams who want the audit acted on rather than delivered
- Where it falls short: we do not publish pricing, so you cannot size it before a call. A standalone audit document is not what we sell, so if that is genuinely all you need, run the five stages yourself using this page
Frequently asked questions
What is a GEO audit?
A diagnostic of whether AI engines can reach your content, whether they cite it, and which sources they use instead. It differs from a technical SEO audit because four of its five stages examine sources outside your own website.
How is a GEO audit different from an SEO audit?
An SEO audit examines your site: crawling, rendering, indexation, structure. A GEO audit starts with a subset of that, then spends most of its time on what engines say about your category and which domains they cite.
How long does a GEO audit take?
Around three days of work if run properly: ten minutes on access, a day on the baseline, half a day on sources, a day on the gap analysis, and half a day producing the plan. Executing the plan takes considerably longer.
What does a GEO audit cost?
The audit itself is close to free, since measurement tools now have genuine free tiers. The cost sits in executing the findings, particularly earned media, which cannot be done from your own site and is usually the largest line.
How many prompts should a GEO audit use?
Around 30 real commercial questions, run several times each across every engine your buyers use. Fewer than that is a sample too small to separate a pattern from ordinary run-to-run variance.
Can I run a GEO audit myself?
Yes, and you should run stage 1 today regardless. Checking your robots.txt for six crawler tokens takes ten minutes and is the highest-return check in this discipline.
What is the most common problem a GEO audit finds?
Blocked crawler access, usually because a team blocked all bots with AI in the name without realising the search crawler and the training crawler are separate tokens.
What should the audit produce?
A frequency-ranked list of the domains feeding answers in your category, kept separate per engine, plus a sequenced plan with named owners. A composite score is not an output you can act on.
Where to go next
Run stage 1 now. Open your robots.txt and check six tokens: OAI-SearchBot, GPTBot, PerplexityBot, Perplexity-User, Googlebot, Google-Extended. If a search crawler is blocked, nothing else on this page matters until that changes.
Then build the thirty-prompt set and run stage 2 on a free tier. Most teams discover their real constraint within a day.
Our engine-by-engine optimisation guide covers the access checks in full, the seven-point AI search audit template is the lighter version to run yourself, and the Visibility, Citability and Retrievability framework explains which lever is stuck.
For the technical half at scale, our enterprise SEO audit framework sets out 47 checks. For tooling, the best answer engine optimization tools covers which have free tiers. Our Acceldata case study shows the work on a technical B2B account.
The honest exit. If your crawler access is clean and you have fewer than about 30 meaningful commercial prompts in your category, you do not need a formal audit. Run the five stages informally once a quarter and spend the budget on earned media instead.
Sources
Crawler tokens were checked at each company’s own documentation on 10 September 2026. Every study cited was published in 2026 and checked at the original source.
- OpenAI, Bots and crawlers documentation, checked 10 September 2026.
OAI-SearchBotsurfaces sites in ChatGPT search;GPTBotis for training. - Perplexity, Crawler documentation, checked 10 September 2026.
PerplexityBotsurfaces results and is not used for training;Perplexity-Usergenerally ignores robots.txt. - Google Search Central, Google common crawlers, checked 10 September 2026.
Google-Extended“does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.” - Google Search Central, AI features and your website, 15 May 2026. States AEO and GEO are part of SEO, and advises against providers guaranteeing rankings.
- Zhang Kai, He Xinyue and Yao Jingang, From Citation Selection to Citation Absorption, arXiv, 28 April 2026. 602 prompts, 21,143 citations, 18,151 fetched pages. Academic and independent.
- Huang, Goyal, Saha and Chandrasekharan, Answer Bubbles: Information Exposure in AI-Mediated Search, arXiv, 17 March 2026. 11,000 real search queries across five systems. Academic and independent.
- Muck Rack, Earned media still drives 84% of AI citations, 7 May 2026. More than 25 million cited links, 17 industries. Muck Rack sells PR software, so read it as an interested party with a large dataset.
Latest Blogs
Every list of AEO agencies for B2B tells you who is good. None of them tells you who will survive your procurement process. So we audited what ten agencies publish about themselves against the five things an enterprise buyer has to produce internally before signing. All ten name enterprise clients. Two publish a price. None publishes team size, contract terms, or anything a security review would accept.
Almost no agency publishes what a standalone SEO audit costs. Two do, and they differ by 3.4 times on price and nine times on turnaround while covering a similar number of pages. So we computed the metric nobody publishes, cost per page audited, and set out what a purchased audit has to contain before it is worth buying at any price.
Nobody outside a platform can measure its data accuracy without running a controlled test, and nobody in this category publishes one. So we ranked 14 platforms on the thing that actually decides whether their numbers can be accurate: how many times each prompt is sampled, computed from each vendor’s own published allowances. Only five publish enough to work it out, and the spread between them is thirty-fold.