GEO / AI Search

How AI search actually works, from query to cited answer

janvi
•
Posted on 5/10/26•16 min read
How AI search actually works, from query to cited answer

The short answer

How AI search actually works is less tidy than the usual five steps: you ask a question, the engine searches, it reads what it finds, it writes an answer, and it cites its sources. Every one of those steps behaves differently in practice. The engine rewrites your question into several of its own, it frequently skips searching altogether, retrieval rather than reasoning causes most of its mistakes, and it decides whether to link you and whether to name you as two separate things. So the pipeline is real, but almost nothing in it works the way the diagram suggests.

Key takeaways

  • Stage one often does not happen. In a July 2026 pilot of 48 buying prompts, the engine searched the web on only 42% of them. On the rest it answered from what the model already held, so nothing you published could have changed the answer.
  • Your question is not the question it searches. The same pilot captured 6.3 sub-queries per searching prompt, and our own tracking puts the range at eight to twelve. You wrote none of them.
  • Retrieval is where it breaks, not reasoning. Across 2,100 questions and six chatbots, more than 70% of all errors came from retrieval. The authors put it plainly: when a model finds the right source, it usually extracts the right answer.
  • There is no separate AI index. Google states that its AI features run on core ranking systems, so the crawl and index you already have is the one being used.
  • Selection favours earned media heavily. Of more than 25 million cited links, 84% came from earned sources and 0.3% from paid.
  • Writing the answer introduces its own errors. The same six chatbots scored above 90% on multiple choice and lost 11 to 13 points under free response, and a false premise in the question dropped accuracy as low as 19%.
  • Attribution is two decisions, not one. Only 13.2% of brand appearances produced both a citation and a mention, and engines split in opposite directions on which they give you.
  • Pepper is an agentic organic growth engine and an organic growth partner. Agent Atlas puts the agents in your team’s hands. Pepper’s GEO platform reports Brand Visibility, Domain Prompt Presence and Share of Voice across six engines, including ChatGPT, Perplexity, Gemini and Google AI Overviews. A growth team works alongside yours. Eight years, more than 250 enterprises, more than 10 million tracked prompts.

A note on where this comes from. I spend my time on how these systems read, rank and retrieve, and most of the arguments I have about AI search come down to people picturing a different machine. Pepper runs organic growth for more than 250 enterprises over eight years and tracks more than 10 million prompts across every major engine. In practice, once a team sees the fan-out and the no-search rate, the strategy argument usually resolves itself.

Disclosure: Pepper sells services that help brands appear in AI answers, so a complicated pipeline with many levers suits us commercially. This article says that one of the five stages frequently does not run, that a second is mostly your existing crawl and index, and that the most useful work is unglamorous. Every figure traces to a named study with its method. We name competitors but never link them.

What is the AI search pipeline?

The AI search pipeline is the sequence an engine runs between receiving a question and showing an answer with sources. Five stages describe it well enough to be useful.

  • Stage 1, query interpretation. The engine decides what you are really asking, and whether it needs to look anything up.
  • Stage 2, retrieval. It fetches candidate documents, from a live index, a cached crawl, or both.
  • Stage 3, selection. It narrows those candidates to the handful it will actually use.
  • Stage 4, synthesis. It writes the answer from the selected material plus whatever the model already holds.
  • Stage 5, attribution. It decides which sources to link, and separately whether to name any brand in the text.
Figure 1: the five stages, with the failure rate or skip rate measured at each one.

Two things make this different from classic search. First, stages 3 and 4 collapse what used to be ten blue links into one answer. The competition is therefore for inclusion rather than position. Second, stage 1 is not guaranteed to run at all, which has no equivalent in a search engine that always returns results.

Stage 1: the engine rewrites your question, and may not search

This stage does the most damage to people’s mental model, so it is worth being precise.

First, the engine decides whether to search. An EMGI pilot published July 2026 ran 48 SaaS buying prompts through GPT-5.2, with live web search enabled. The model searched the web on only 42% of them. On the other 58% it answered from parametric memory. No retrieval event happened at all. Nothing you published last quarter could have influenced those answers, and no monitoring tool can observe them, ours included.

Second, when it does search, it does not search for what you asked. The same pilot captured 127 fan-out queries, averaging 6.3 per searching prompt. Our own tracking of engine sub-queries puts the typical range at eight to twelve. The engine decomposes one question into several narrower ones, runs them, and assembles from the results.

Figure 2: the share of prompts where no retrieval event happened at all.

So the practical consequence is uncomfortable. If you track 25 prompts, roughly ten trigger a search, each fanning out to six or more, which is about sixty actual queries, none of which you wrote. Prompt coverage therefore reports a floor rather than a measurement, and anyone presenting it as a percentage of reality is overstating what it can show.

If you want help working out which of your queries even trigger a search, book a growth audit and bring twenty questions your buyers actually ask.

Stage 2: retrieval, where most of the failures live

Retrieval sounds like plumbing and it is the single most consequential stage.

A study published 21 May 2026 evaluated six chatbots over fourteen days, among them Gemini 3 Pro, Grok 4, Claude 4.5 Sonnet and GPT-5. It used 2,100 factual questions drawn from same-day BBC reporting in six languages. Its central finding is the one worth memorising:

Retrieval, not reasoning, drives more than 70% of all errors.

The authors spell out why that matters. When a model lands on the correct source, it usually extracts the correct answer. The hard part is landing on the right source in the first place, which is a retrieval problem and therefore partly your problem.

There is no separate AI index to get into. Google’s own guidance, last updated 10 July 2026, states that its AI features run on core ranking systems. Consequently the crawl and index you already have is the one being used, and there is no second door.

The crawlers are not one crawler. OpenAI runs OAI-SearchBot, ChatGPT-User and GPTBot, each with a different job. Perplexity runs PerplexityBot for scheduled indexing and Perplexity-User for on-demand fetches. Blocking one fails differently from blocking the other. Our technical guide to AI crawlers and our note on PerplexityBot work through the robots.txt consequences.

Retrieval also carries a language bias worth knowing about. In the same evaluation, every model scored lowest on Hindi, at 79% against 89 to 91% elsewhere. Models answering Hindi queries also cited English Wikipedia more often than any Hindi outlet. If your market is not Anglophone, retrieval is working against you before anything else happens.

Stage 3: selection, and what gets picked

Retrieval produces candidates. Selection decides which few become the answer, and the pattern here is stable enough to plan against.

Muck Rack’s May 2026 analysis of more than 25 million cited links across ChatGPT, Claude and Gemini found 84% of citations came from earned media and 0.3% from paid or advertorial. Earned media here covers journalism, academic and government sources, encyclopedic sites and third-party corporate content. That share has held between 82% and 89% across three editions since mid-2025.

So two conclusions follow, and the second is the useful one.

  • Your own pages are a minority of the citation surface, which is uncomfortable but not fixable by publishing more of them.
  • Any budget line spent on placement is buying three tenths of one percent of that surface. Use this to end any proposal that sells sponsored coverage as a citation strategy.

Stage 4: synthesis, where the answer gets written

The model now writes, using the selected sources plus whatever it already holds. Two findings from the 2,100-question evaluation describe this stage well.

Format changes accuracy. The same systems scored above 90% on multiple choice and lost 11 to 13 points under free response, with a 16 to 17 point drop across the whole cohort. Recognising the right answer is easier than producing it, which is exactly what you would expect and rarely what gets reported.

A false premise is catastrophic. Models scoring 88 to 96% on well-formed questions fell to between 19% and 70% when the question carried a subtle false premise. The most vulnerable model accepted fabricated facts 64% of the time. Therefore if a competitor’s page asserts something untrue about your category, the engine may simply carry it forward.

Figure 3: the same six chatbots, on well-formed questions and on questions carrying a false premise.

This is also where the no-search path rejoins the story. On the 58% of prompts with no retrieval, stages 2 and 3 never ran, and the answer came entirely from training. Those answers are shaped by what the web said about you months ago, which is slow to change and impossible to monitor directly.

Stage 5: attribution, which is two decisions

Most people treat this as one step. It is two, and they do not move together.

A study published 9 June 2026 logged 3,981 domain appearances across 115 prompts and 14 countries. It found:

  • 61.7% were citations with no brand name in the answer, a pattern now called a ghost citation.
  • 25.1% were mentions with no citation.
  • Only 13.2% were both.

The split also inverts by engine. Gemini named brands in 83.7% of appearances and cited sources in 21.4%. ChatGPT did close to the reverse, naming brands 20.7% of the time and citing sources 87% of the time. So the final stage produces a different output depending on which engine ran it. A blended visibility score hides that completely. Our page on zero-click search and AI citation covers what each outcome is worth.

How we weighted the five stages

If you can only influence some of this, here is how we rank the stages by how much your work moves them.

Figure 4: Pepper’s weighting, which favours the stages you can actually change.
CriterionWeightWhy it carries that weight
How much your own work moves it35%Retrieval responds to crawlability and content. The decision to search at all does not respond to you.
How often the stage runs25%A stage that skips on 58% of prompts is worth less attention than one that always runs.
How much of the failure it causes25%Retrieval drives more than 70% of errors, which makes it the highest-yield place to work.
How measurable it is15%You can observe citations. You cannot observe a no-search answer, at any price.

## How AI search actually works, at a glance

StageWhat actually happensMeasured failure or skip rateCan you influence itCost to work on
1. Query interpretationRewrites into 6 to 12 sub-queries, or skips search entirelyNo search on 58% of buying promptsBarelyNot purchasable
2. RetrievalFetches candidates via core index and crawlersOver 70% of all errorsYes, most of allFree to moderate, technical work
3. SelectionNarrows to a handful, heavily favouring earned media84% earned, 0.3% paidYes, slowlyHigh, earned media takes quarters
4. SynthesisWrites the answer, from sources plus memoryLoses 11 to 13 points on free responseIndirectly, via source qualityFree, correct your facts at source
5. AttributionDecides citation and mention separatelyOnly 13.2% produce bothPartlyLow, framing and naming in content

## What this means for the work

Five stages, and the practical advice collapses into a short list.

  1. Fix retrieval first. It causes most of the errors and it responds to you. Check indexation by template, crawler access in robots.txt, and whether your pages render without JavaScript.
  2. Stop optimising for a separate AI index. There is not one. Google says its AI features run on core ranking systems.
  3. Treat prompt coverage as a floor. With fan-out and the no-search rate, your tracked prompts are a sample of a much larger set you did not write.
  4. Shift budget towards earned media if citations are the goal, because that is 84% of the surface and paid is 0.3%.
  5. Correct false claims at the source. A false premise can cut accuracy to 19%, and the fix is usually a third-party page rather than your own.
  6. Report citations and mentions separately, per engine, since the two invert.

What this costs

Diagnosing where you sit in this pipeline costs nothing but time. Indexation and crawler checks are an afternoon. Reading which sources an engine cites on twenty prompts is about an hour.

The expensive stage is selection, because earned media is slow and you do not control it. The cheap stages are retrieval and attribution, which are technical and editorial work you can start this week. Published entry pricing for monitoring tools starts around $29 a month for very limited prompt counts, though note that no tool can observe the no-search path, so some of this pipeline is unmeasurable at any price.

How Pepper fits

Pepper is an agentic organic growth engine and an organic growth partner, which means three things working together rather than one product.

Pepper’s GEO platform is the self-serve workspace. You set up a brand profile, add competitors and personas, connect GA4 and Search Console, and define the themes and prompts you care about. It reports Brand Visibility, Domain Prompt Presence and Share of Voice across six engines, including ChatGPT, Perplexity, Gemini and Google AI Overviews, with Platform Breakdown, New Mentions and Citation Analysis underneath. That maps onto stage 5, and partly onto stage 3.

Agent Atlas is where your team builds, versions and runs its own agents, with quick runs for one input and sheet runs for bulk. This is the practical way to pull the cited sources out of a few hundred answers and find which third-party pages keep deciding your category.

The growth team is attached to the account and works alongside yours, which matters most at stage 3, where the work is earning coverage rather than publishing more pages.

Where it falls short: we cannot see stage 1 at all. When the model answers from memory, no retrieval event exists, so no platform can observe it and we will not pretend otherwise. We also report presence rather than persuasion, and we do not publish run-to-run variance. By the measurement standard published in August 2026, that makes our numbers directional rather than decision-grade, along with everybody else’s.

Eight years, more than 250 enterprises, more than 10 million tracked prompts. You can see the shape of the work in the Acceldata case study, in how we run it for B2B SaaS brands, and across the case study library.

How to choose where to spend your effort

The decision is which stage to work on, and the honest answer is that most teams pick the wrong one because it is the most visible. So here are the criteria, weighted.

The weighted scorecard

Score each row from 1 to 5, multiply by the weight, and total out of 100.

CriterionWeightScore 1 meansScore 5 means
Technical retrieval health35Clean index, crawlers allowed, server-renderedTemplates unindexed, crawlers blocked, JS-dependent
Third-party coverage of your category25Well covered by journalism and reference sitesNobody independent writes about you
Accuracy of what the web says about you20Current everywhereStale prices and retired features on third-party pages
Clarity of brand naming inside your content20Every quotable sentence names youGeneric prose anyone could have written

Under 40, you are in good shape and the remaining gains are slow. From 40 to 70, fix retrieval and naming first, because both are cheap. Above 70, you have a structural problem and no monitoring tool will help until it is fixed.

The weaker playbook against the stronger one

The weaker approach is to publish more pages and watch a visibility score. That feels productive because output is visible. It also addresses stage 3 only, which is the slowest stage and the one least responsive to volume. The stronger approach is to spend the first month entirely on stage 2. That means indexation by template, crawler access, render checks, and the handful of third-party pages the engines keep citing. One adds to the pile. The other removes the reason you were not retrieved.

Run a live test before you commit budget

Take 20 prompts that reflect how your buyers actually search, written in their words rather than yours, such as “best data observability tool for a mid-size team”, “how do I stop duplicate records in my CRM”, or “alternatives to the incumbent for a regulated business”. For each one, log three things. Did the engine appear to search, which sources did it cite, and did it name you? Repeat at 30 days and again at 90 days. If the cited sources barely change while your own publishing continues, your constraint is stage 3 and not stage 2.

Red flags

  • Anyone offering to get you into “the AI index”, which does not exist
  • A proposal built on sponsored placement, against 0.3% of the citation surface
  • Prompt coverage presented as a percentage of reality rather than a floor
  • Any claim to measure answers where the model never searched
  • Advice to restructure content for AI specifically, which Google’s own guidance lists as unnecessary
  • A single blended visibility score with no citation and mention split
  • Certainty about ranking factors in systems nobody outside the vendor can inspect

Five questions worth asking any vendor

  1. How do you handle prompts where the model did not search, and are they counted in my score?
  2. Which crawlers do you check access for, and do you distinguish scheduled indexing from on-demand fetching?
  3. Do you report citations and mentions separately, per engine?
  4. How many runs per prompt sit behind a reported change, and what is your run-to-run variance?
  5. What would you fix first on my site, and can you point at the evidence for it?

The reducing principle. It comes down to one question: can the engine fetch, parse and trust your page? If it cannot, nothing downstream matters, because stages 3, 4 and 5 never get the chance to include you. If it can, your remaining problem is that other people’s pages are more convincing than yours. That is slower and more expensive to fix. Everything in this article is a refinement of that split.

The honest closing note. If your templates are indexed, your crawlers are allowed and your pages render, you do not need us for stage 2, and that is the stage with the highest yield. Do it yourself, then come back when the constraint is coverage, because that is the part that takes a team and a year rather than an afternoon.

What nobody should promise you

  • Entry into a separate AI index. Google says its AI features use core ranking systems, so there is nothing separate to enter.
  • Visibility on prompts where the model never searched. No tool can see those answers, ours included.
  • A complete picture from prompt tracking. Fan-out means most of the queries run were never in your set.
  • A fast fix at stage 3. Earned media moves over quarters, because it depends on pages you do not own.
  • Known ranking factors. These systems are not inspectable from outside, and anyone claiming otherwise is guessing confidently.

Where this stops working, including for us

The evidence here is good and narrow, so treat the stages as better established than the exact numbers.

The 42% search rate comes from a pilot of 48 prompts whose authors describe it as a pilot rather than a population study, and note that API behaviour may differ from the consumer app. The 2,100-question evaluation is strong, but it covers news questions over fourteen days. News is unusually time-sensitive, so retrieval may matter more there than in a stable commercial category. The attribution study publishes no collection date range and no limitations section, and a vendor produced it using its own tool.

Engine behaviour also changes without notice. Every figure here is a snapshot from 2026, and a model update could move any of them.

And our own position needs the same scepticism. A five-stage pipeline with many levers is a good thing for us to sell, and this article argues that one stage cannot be influenced, another is mostly your existing technical hygiene, and the highest-yield work is something you can do without us. Check the sources yourself. We have linked them all.

Where to go next

Frequently asked questions

How does AI search actually work?
An engine interprets your question and decides whether to search. It then retrieves candidate pages, selects a handful, and writes an answer from them plus what it already knows. Finally it decides which sources to link, and separately whether to name any brand. Five stages, each with its own failure mode.

Does the engine search the web every time?
No, and this surprises most people. In a July 2026 pilot of 48 buying prompts, the engine searched on only 42% of them. On the rest it answered from training data, so no retrieval event existed for any tool to observe.

What is query fan-out?
Fan-out is the engine decomposing one question into several narrower searches of its own. A 2026 pilot measured 6.3 sub-queries per searching prompt, and our own tracking puts the typical range at eight to twelve. You wrote none of them.

Where do AI search answers go wrong most often?
Retrieval. Across 2,100 questions and six chatbots, more than 70% of errors came from failing to land on the right source rather than from faulty reasoning. When models retrieve correctly, they usually extract the correct answer.

Is there a separate AI index I need to get into?
No. Google’s guidance states that its AI features run on core ranking systems, so your existing crawl and index are what matter. Advice to build for a separate AI index is selling you something that does not exist.

Which sources do AI engines prefer?
Earned media, overwhelmingly. Across more than 25 million cited links, 84% came from earned sources including journalism, academic and reference sites, while paid and advertorial accounted for 0.3%. That split has been stable for over a year.

Why does an engine cite my page but never name my brand?
Because citation and mention are separate decisions. In one 2026 study, 61.7% of appearances were citations with no brand name in the answer. Engines also differ sharply: Gemini names far more often than it cites, and ChatGPT does the reverse.

What should I fix first?
Retrieval, because it causes most of the failures and responds to your work. Check indexation by template, crawler access in robots.txt, and whether your revenue pages render without JavaScript. That is an afternoon and it costs nothing.

Sources and further reading

  • Suzgun, Shen, Bianchi, Spangher, Icard, Ho, Jurafsky and Zou, Evaluating Commercial AI Chatbots as News Intermediaries, arXiv:2605.22785, submitted 21 May 2026. A fourteen-day evaluation (9 to 22 February 2026) of six chatbots on 2,100 factual questions from same-day BBC reporting across six regional services. Source of the 70% retrieval-failure finding, the multiple-choice versus free-response gap, the Hindi accuracy gap and the false-premise results. Limitation: news questions are unusually time-sensitive.
  • EMGI AI search retrieval pilot, published July 2026. 48 SaaS buying prompts through GPT-5.2 with live web search via the DataForSEO LLM API. Source of the 42% search rate and the 6.3 fan-out average. Its authors describe it as a pilot rather than a population study.
  • Semrush with Kevin Indig and Growth Memo, Why 62% of AI citations don’t lead to brand mentions, published 9 June 2026. 3,981 domain appearances across 115 prompts and 14 countries. Source of the citation and mention split and the per-engine rates. The publication states no collection date range and no limitations, and a vendor produced it using its own tool.
  • Muck Rack, “What Is AI Reading?”, third edition, May 2026. More than 25 million links from ChatGPT, Claude and Gemini across 17 industries. Source of the 84% earned and 0.3% paid shares.
  • Google Search Central, guide to optimising for generative AI features, page last updated 10 July 2026. Source of the core-ranking-systems statement and the list of unnecessary AI-specific tactics. Google describes Google Search only.
  • Pepper, our own fan-out range, 24 June 2026.
  • Pepper, the crawler and robots.txt guide, for stage 2 in practice.
  • Pepper, PerplexityBot and on-demand fetching, for why blocking one crawler fails differently.

A note on sources. Only studies published in 2026 are cited here. Where a figure comes from a vendor’s own research, we have said so in the line that uses it.