GEO / AI Search

How to find which sources AI cites in your category

Pranay Batta
Posted on 23/09/268 min read
How to find which sources AI cites in your category

The short answer

Do not look for the study. Build the list yourself, because it takes an afternoon and the result is specific to you.

The procedure is simple enough to describe in one paragraph. Take forty prompts a buyer in your category would actually type. Run them across the engines you care about. Record every cited URL. Count domains. The domains that recur are your citation core, and that list is the authority target for everything else you do.

Three things make this better than any published average.

  • Categories differ enormously. The sources an engine reaches for when asked about compliance software are not the ones it reaches for on running shoes. An industry-wide figure describes neither.
  • You can check your own work. Somebody else’s methodology is usually a paragraph. Yours is a spreadsheet you built.
  • The widely quoted numbers are mostly vendor research. We went looking for the studies behind the citation-share figures circulating in B2B, and the two we found were published by agencies selling services in this category, and both returned errors when we tried to read them.

What follows is the procedure, how to score the list once you have it, and what to do with the result.

We run this work daily. Pepper is an agentic organic growth engine, not an SEO agency. This page comes out of what we run: organic for more than 250 enterprises across eight years, more than 10 million tracked prompts, and the questions buyers put to us in client reviews. Customers log in and run the platform themselves, with a Pepper growth team attached. Book a growth audit and we will build the citation core for your category with you.

Key takeaways

Key takeaways from doing this exercise across client categories.

  • Forty prompts and an afternoon produces a better list than any published average, because it is yours.
  • Count domains, not pages. The unit that matters is which site keeps getting reached for.
  • Expect third-party sources to dominate. Independent academic work across 11,000 queries found Wikipedia and longer third-party sources cited far more often than their share of the web would suggest.
  • Separate the sources you can influence from the ones you cannot. A review site you can get listed on is a different task from an encyclopaedia entry.
  • Run it per engine. Engines disagree about sources, and an averaged list hides the one you are losing.
  • Re-run quarterly. The set moves, though more slowly than prompt-level results do.

What is a citation core, and why does it matter more than a ranking?

A citation core is the small set of domains an AI engine reaches for repeatedly when answering questions in your category.

It matters because of where citations actually come from. In most categories the engine is not quoting brands. It is quoting the places that write about brands: review sites, documentation, community threads, trade publications and reference works. Your own domain is competing for a minority of the available citations, which is why a strategy confined to your own site addresses the smaller half of the problem.

Where it falls short. A citation core tells you which sources have authority in your category right now. It does not tell you how to get onto them, what that costs, or whether the effort is worth it against your other options. Those are separate judgements and this page does not make them for you.

Panel showing the three kinds of source in a citation core and what you can do about each
Figure 1: Three kinds of source, three different levels of control.

Step 1: Write forty prompts a buyer would actually type

Same discipline as any prompt set. Category language, not brand language.

  • Lead with selection questions. “Best X for Y”, “cheapest alternative to Z”, “X vs Y for small teams”. These produce the richest citation sets because the engine has to compare.
  • Add problem questions. “How do I fix X”, “why does Y keep happening”. These surface documentation and community sources rather than review sites.
  • Add a few definitional ones. “What is X”. These surface reference works, which behave differently again.
  • Avoid your own brand name. A prompt containing your brand will cite your site, which tells you nothing about the category.
  • Freeze and date the list. You will want to re-run it unchanged next quarter.

Step 2: Run them and harvest every cited URL

This is the manual part, and it is genuinely an afternoon for forty prompts across two engines.

  • Pick your engines deliberately. Start with the two that matter most to your buyers rather than trying to cover everything. The engine-by-engine differences are in how to optimize for ChatGPT and Perplexity.
  • Record the full URL, the domain, and the prompt it came from. Three columns is enough.
  • Run each prompt more than once. Answers vary between runs, so a single pass will miss sources. Three passes per prompt is a reasonable floor by hand.
  • Note the position. A source cited first carries more weight than the fifth link in a list.
  • Keep the raw answers. You will want to re-read them when a number looks strange.

If forty prompts across three passes and two engines sounds like a lot of clicking, it is: that is 240 runs. This is the point at which most teams move to a platform, and the options are compared in top AI visibility platforms. The manual version is still worth doing once, because it teaches you what the data looks like before you trust a dashboard to summarise it.

Step 3: Count domains and sort into three buckets

Now the analysis, which is simpler than it sounds.

Bar chart showing the typical shape of a citation core with a small number of dominant domains
Figure 2: The shape this exercise usually produces. A short head and a very long tail.
  • Count how many distinct prompts each domain was cited for, not how many total citations it received. One page cited twenty times for one prompt is less interesting than a domain cited once each across twenty prompts.
  • Sort descending. In most categories a handful of domains account for a large share of the citations, and then it flattens fast.
  • Bucket them. Sources you own, sources you can influence, and sources you cannot.
  • Note your own position. How far down the list does your domain appear, and for which prompts?

The three buckets, and why they matter. Owned sources are your own site, where the fix is structural and fast. Influenceable sources are review sites, directories, trade press and communities, where the fix is relationship and effort. Uninfluenceable sources are reference works and editorially independent publications, where the honest answer is usually that you cannot do much directly.

Step 4: Score the list so you know where to spend

A ranked list is not yet a plan. Score each influenceable source on four things.

CriterionWeightWhat earns full marks
Prompt coverage35Cited across many of your prompts, not just one
Realistic access30You can get listed, reviewed or published there
Position when cited20Tends to appear first in the answer, not in the tail
Cross-engine presence15Cited by more than one engine, not just one

Total 100. Prompt coverage and access carry nearly two thirds between them, because a source that appears everywhere but is closed to you is a fact rather than an opportunity, and a source you can easily reach that appears once is not worth the effort.

Bar chart of the four weighted criteria for scoring sources in a citation core
Figure 3: How we would score an influenceable source. Coverage and access carry two thirds.

The output is a short ranked list of places worth investing in, with a reason attached to each. That is what a citation core is for, and it is more actionable than any visibility score. Our wider framework for turning this into work is Visibility, Citability and Retrievability.

Step 5: Check the same list per engine

An averaged list hides the engine you are losing, and engines genuinely disagree about sources.

  • Build one column per engine. The same forty prompts, counted separately.
  • Look for sources that appear in one and not others. That asymmetry is usually the most actionable finding in the whole exercise.
  • Check where your own domain sits in each. Being cited by one engine and not another is a different problem from being cited by none.
  • Do not average them into a single ranking. The point of running per engine is to see the disagreement.
Matrix showing how the same prompt set produces different source rankings across engines
Figure 4: Same prompts, different engines, different sources. Averaging hides this.

What to do once you know which sources AI cites

Four moves, in rough order of speed.

  • Fix your own pages first, because it is fastest and you control it. If your domain appears in the list but low, the pages exist and are not being reached for. The practices are in our GEO best practices playbook.
  • Get listed properly on the top influenceable sources. Complete profiles, current information, real reviews. Unglamorous and frequently the highest-return work on the list.
  • Pitch the trade publications that recur. You now have evidence that engines reach for them in your category, which is a better pitch than a generic one. Deciding who owns that outreach is covered in the AI search team ownership model.
  • Leave the uninfluenceable ones alone. Trying to get into a reference work through the front door usually fails and occasionally backfires.

What nobody should promise you

  • An industry-wide citation source list that applies to you. Categories differ, and a general figure describes nobody’s category in particular.
  • A guaranteed citation from any named source. Nobody controls what a model reaches for, and we will not promise it either.
  • That the list is stable. It moves. Re-run it quarterly rather than treating one snapshot as permanent.
  • A shortcut into reference works. These are editorially governed, and attempts to game them tend to be visible and counterproductive.
  • That your own site can carry this alone. Independent research keeps finding third-party sources cited far more often than their share of the web would suggest.

Frequently asked questions

How do I find which sources AI cites in my category?
Run forty category prompts across your two most important engines, three passes each, and record every cited URL with its domain and prompt. Count how many distinct prompts each domain was cited for, then sort descending.

Why not just use a published study?
Because the sources an engine reaches for vary enormously by category, so an industry-wide figure describes nobody in particular. Several widely quoted figures are also vendor research, and two we tried to verify could not be read at all.

Running the exercise

How many prompts do I need?
Around forty for most categories, run three times each. Fewer than twenty will miss sources because answers vary between runs, and more than eighty gets expensive by hand without adding much to the head of the list.

Should I count pages or domains?
Domains, and specifically how many distinct prompts each domain was cited for. One page cited repeatedly for a single prompt matters less than a domain that keeps appearing across different questions.

How often should I re-run it?
Quarterly. The citation core moves more slowly than prompt-level results do, so monthly re-runs mostly show noise while annual ones miss real shifts in which sources engines favour.

Acting on the results

What if a review site dominates my category?
That is common and it changes the task. Getting listed and reviewed well on that source is usually faster and cheaper than trying to outrank it with your own pages.

What if my own domain does not appear at all?
Then you have a coverage problem rather than a citability one. Your pages are not being reached for in this category, and third-party presence is likely to move the needle faster than rewriting your site.

Can a platform do this for me?
Yes, and at forty prompts across several engines with repeat passes it quickly becomes the sensible option. Do it manually once first, so you know what the underlying data looks like before trusting a summary of it.

Sources

This page is a method. Where it rests on published research we cite it, and where it rests on our judgement we say so.

Research

  • Huang, Goyal, Saha and Chandrasekharan, Answer Bubbles: Information Exposure in AI-Mediated Search, arXiv, 17 March 2026. 11,000 real search queries across five systems. Finds Wikipedia and longer third-party sources cited far more often than their share of the web would suggest. Academic and independent.
  • Zhang Kai, He Xinyue and Yao Jingang, From Citation Selection to Citation Absorption, arXiv, 28 April 2026. 602 set prompts, 21,143 search-layer citations, 18,151 fetched pages across ChatGPT, Google AI Overview and Perplexity. Finds absorbed pages are longer, more structured and richer in extractable evidence. Academic and independent.

Our own published work