GEO / AI Search

Where ChatGPT and Perplexity get their information about companies

janvi
•
Posted on 1/10/26•14 min read
Where ChatGPT and Perplexity get their information about companies

The short answer

Where ChatGPT and Perplexity get their information about companies differs more than most teams assume, and it is measurable rather than a matter of opinion. Across 27 million citations analysed in January 2026, Perplexity drew 19.4% of its citations from social sources against ChatGPT’s 5%. ChatGPT took 15% from institutional sources against Perplexity’s 5%. They also search differently before they cite anything. 91% of ChatGPT’s retrieval queries were unique, against 14% of Perplexity’s. One rewrites your question, the other searches it almost verbatim. Consequently the same content can perform very differently on each.

Key takeaways

  • The source mix is close to opposite at the edges. Perplexity takes 19.4% of citations from social sources, nearly four times ChatGPT’s 5%. ChatGPT takes 15% from institutional sources, three times Perplexity’s 5%.
  • They do not search the same way either. 91% of ChatGPT’s queries were unique, meaning it almost never searches the same way twice. Perplexity’s figure was 14%.
  • Perplexity searches your words. ChatGPT does not. Word overlap between the user’s prompt and the queries actually run was 88% for Perplexity and 13% for ChatGPT.
  • Perplexity cites more sources per answer, about 7.3 domains against ChatGPT’s 5.0.
  • The most quoted statistic on this subject is out of date. The widely repeated finding that only 11% of cited domains overlap between the two engines comes from a study published 1 July 2025, not 2026. It is routinely re-dated. We have left it out.
  • The two big datasets disagree about brand-owned content, and the disagreement is probably about bucketing rather than fact. Say which definition you are using before quoting either.
  • Pepper is an agentic organic growth engine and an organic growth partner. Agent Atlas puts the agents in your team’s hands. Pepper’s GEO platform reports Brand Visibility, Domain Prompt Presence and Share of Voice across six engines, including ChatGPT, Perplexity, Gemini and Google AI Overviews. A growth team works alongside yours. Eight years, more than 250 enterprises, more than 10 million tracked prompts.

A note on where this comes from. I keep being asked which engine to optimise for, and the question is usually premature, because most teams have not looked at what the two actually read. Pepper runs organic growth for more than 250 enterprises over eight years and tracks more than 10 million prompts across every major engine. In my experience the useful shift happens when a team stops asking which engine matters and starts asking which of their sources each engine can see.

Disclosure: Pepper sells services that help brands appear in AI answers, so a story about two engines needing two strategies suits us commercially. This article says the opposite where the evidence points that way, and it excludes the single most quoted statistic in this subject because the primary source is from 2025. Two of the datasets below were published by a direct competitor, which we flag in the lines that use them. We name competitors but never link them.

What is a citation source, and why does the engine matter?

A citation source is the page an engine lists as the origin of something it said. Group those pages by what kind of thing they are, and you get a source mix, which is the useful unit here.

The grouping used in the largest 2026 dataset has four buckets, and the definitions matter more than the numbers.

  • Brand. Company-owned sites, plus competitor sites and other corporate pages.
  • Media. Earned coverage and press wire material.
  • Institution. Government, academic and standards bodies.
  • Social. Community platforms, forums and user-generated content.

So why does the engine matter? Because each one weights those buckets differently, and because each one looks for different things before it retrieves anything at all. Therefore a page that reads well to one engine may never be fetched by the other. Our framework for visibility, citability and retrievability sets out the three levers underneath this.

Where ChatGPT and Perplexity get their information, measured

The largest published breakdown comes from an analysis of 27 million citations across seven engines, published 8 January 2026 by Profound. It is a vendor’s own research, which we weigh accordingly, and it is the only per-engine category split at this scale we could find.

Figure 1: the same four buckets, and two engines that diverge most where it matters least to most marketers.
Source categoryChatGPTPerplexity
Brand54%57%
Media25%19%
Institution15%5%
Social5%19.4%

Two things stand out, and neither is the brand row.

Perplexity is a community engine. At 19.4% social against ChatGPT’s 5%, nearly a fifth of what Perplexity says about your category comes from forums and user posts. Is your category discussed on Reddit? If you are absent from that discussion, Perplexity builds its picture of you without you.

ChatGPT is an institutional engine. At 15% institutional against Perplexity’s 5%, it leans three times harder on government, academic and standards sources alone. For regulated categories that is an advantage you can earn. For consumer categories it mostly means ChatGPT reaches for a reference site before a forum.

The brand row is the one to be careful with. At 54% and 57% it looks as though company-owned pages dominate. However, that bucket includes competitor sites and other corporate pages, not just your own. So read it as “corporate web” rather than “your website”.

If you want this run on your own category rather than an aggregate, book a growth audit and bring twenty questions your buyers ask.

Two large datasets disagree, and the reason is definitional

This is worth slowing down on. Both figures circulate, and they cannot both be read the way people read them.

  • The 27-million-citation analysis puts the brand bucket at 54% to 73% across engines, which reads as “most citations are corporate pages”.
  • Muck Rack’s May 2026 analysis of more than 25 million links across ChatGPT, Claude and Gemini puts earned media at 84% and paid at 0.3%, which reads as “most citations are not yours”.

Both are 2026, both are large, and they appear to contradict. However, the likely explanation is bucketing rather than error. Muck Rack’s earned category explicitly includes third-party corporate content, which the four-bucket grouping files under Brand. So a supplier’s page describing your product lands in “earned” in one dataset and “brand” in the other.

The practical rule: never quote either share without saying whose definition you are using. The one number both datasets support is that your own website is a minority of the citation surface, which is the finding that should change a content plan.

How each engine searches before it cites

Source mix tells you what gets cited. It does not tell you how the engine got there, and the retrieval behaviour is where the two differ most sharply.

A study published 30 April 2026 tracked 10,000 randomly sampled prompts over a fourteen-day window in late March and April 2026. It covered ChatGPT, Perplexity and Copilot, capturing the complete set of queries each engine sent to its retrieval layer.

Figure 2: the share of retrieval queries that were unique, by engine.
  • ChatGPT almost never searches the same way twice. 91% of its queries were unique. Copilot came in at 47%, and Perplexity at 14%.
  • Perplexity searches close to your literal question. Word overlap between the original prompt and the queries actually run was 88% for Perplexity, against 50% for Copilot and 13% for ChatGPT.
  • Perplexity also preserves intent better on commercial queries, holding the original intent on 74% of product and shopping prompts.
  • Perplexity cites more per answer, roughly 7.3 domains against ChatGPT’s 5.0 and Copilot’s 2.5.

So the two engines are doing different jobs. Perplexity behaves like a fast, literal research assistant. It takes your phrasing, searches it, and shows its working across more sources. ChatGPT behaves like an interpreter. It decides what you meant, writes its own queries, and rarely repeats them.

What follows for your content. Matching a buyer’s literal phrasing helps on Perplexity. That phrasing is close to what actually gets searched. On ChatGPT it helps far less. The query it runs shares only 13% of its words with the question. There, breadth of coverage across a topic matters more than any single page’s wording.

The statistic everyone quotes is from 2025

The single most repeated claim in this subject is that only 11% of cited domains overlap between ChatGPT and Perplexity. The figure usually comes with 37.4% exclusive to ChatGPT and 51.6% to Perplexity.

We traced it to its primary source and found it published on 1 July 2025. Several 2026 write-ups present it as current. At least one attributes it to a March 2026 analysis.

We have excluded it, under the same rule we apply to everything else: 2026 sources only, read at the original. The reason is not pedantry. July 2025 predates the current model generation on both engines, and this is the fastest-moving measurement in the category. If the figure has been reproduced on 2026 data, we could not find it.

The pattern here is familiar. The larger and rounder the number, the more likely it has been passed along without anyone opening the source.

Where to go next

Our deeper pages on each side of this:

How we weighted the source types

Not every source type is worth the same effort, and the weights below are ours to argue about.

Figure 4: Pepper’s weighting, which favours sources you can realistically influence this year.
CriterionWeightWhy it carries that weight
Share of citations it carries30%A bucket worth 5% of the surface cannot justify a programme, whatever else is true of it.
Whether you can influence it at all30%Institutional sources are nearly impossible to earn quickly. Community presence is not.
How differently the engines treat it25%Social swings from 5% to 19.4% between engines, so it is a per-engine decision rather than a general one.
Time to see any movement15%Community and media move in quarters. Institutional moves in years, if ever.

The two engines at a glance

Figure 3: the two engines side by side, on the dimensions that change what you should do.
ChatGPTPerplexity
Social share of citations5%19.4%
Institutional share15%5%
Media share25%19%
Brand bucket share54%57%
Domains cited per answerabout 5.0about 7.3
Unique retrieval queries91%14%
Word overlap with your prompt13%88%
Behaves likeAn interpreter that rewrites the questionA literal researcher that searches your words
What helps mostBreadth of coverage across a topicMatching your buyer’s exact phrasing
Cost to influenceSlow, earned and institutional workFaster, community presence and clear phrasing

What this means for the work

  1. Decide per engine, not in general. A single content plan optimised for “AI search” is optimising for an average of two engines that diverge fourfold on one bucket.
  2. If your buyers use Perplexity, your category’s community presence is a measurable gap. Nearly a fifth of what it cites is social.
  3. If your buyers use ChatGPT, breadth beats phrasing. The query it runs shares 13% of its words with the question, so chasing exact phrases is low yield.
  4. Do not plan around your own website carrying the answer. Both datasets agree it is a minority of the surface, whatever the bucket labels say.
  5. Check which engine your buyers actually use before any of this. Nothing above matters if you are optimising for an engine your market does not open.

What this costs

Finding out where you currently sit costs nothing but time. Run twenty buyer questions on both engines. Log every cited domain, then group them into the four buckets. That is about two hours and it tells you your own mix rather than the industry’s.

Changing the mix costs differently by bucket. Community presence is the cheapest and fastest, and it is the one that moves Perplexity most. Media takes quarters. Institutional sources take years and are not a realistic target for most brands. Monitoring tools start around $29 a month at entry tiers with very limited prompt counts, though you can do the two-hour version first and should.

How Pepper fits

Pepper is an agentic organic growth engine and an organic growth partner, which means three things working together rather than one product.

Pepper’s GEO platform is the self-serve workspace: brand profile, competitors, personas, GA4 and Search Console connected, themes and prompts defined. It reports Brand Visibility, Domain Prompt Presence and Share of Voice across six engines, including ChatGPT, Perplexity, Gemini and Google AI Overviews. Citation Analysis is the part that matters here, because it shows which domains an engine actually used rather than whether you appeared.

Agent Atlas is where your team builds, versions and runs its own agents, with quick runs for one input and sheet runs for bulk. This is how you do the two-hour exercise above across several hundred prompts instead of twenty, and keep it current.

The growth team is attached to the account and works alongside yours. That matters most on the media and community buckets, where the work is earning presence on pages you do not own.

Where it falls short: we report which domains were cited, not which source category they belong to. The bucketing in this article came from a competitor’s published research, not from our product, and we are not going to pretend otherwise. We also do not publish run-to-run variance. By the measurement standard released in August 2026 that makes our numbers directional rather than decision-grade, along with everyone else’s.

Eight years, more than 250 enterprises, more than 10 million tracked prompts. You can see the shape of the work in the Acceldata case study, in how we run it for B2B SaaS brands, and across the case study library.

How to choose which engine to build for

The decision is which engine’s source mix to design against first, and for most teams it is one rather than both. So here are the criteria, weighted.

The weighted scorecard

Score each row from 1 to 5, multiply by the weight, and total out of 100.

CriterionWeightScore 1 meansScore 5 means
Share of your buyers on one engine30Spread evenly across fourAlmost all on one
How much your category is discussed in communities25Nobody posts about itActive subreddits and forums
Whether institutional sources exist for your category20None, consumer productRegulated, with standards bodies
How distinctive your buyers’ phrasing is15Generic category languageSpecific, repeatable phrasing
Existing earned media footprint10NoneRegular independent coverage

Under 40, do the generic work: fix retrieval, publish clearly, and revisit in two quarters. From 40 to 70, pick one engine and design for its mix. Above 70, you have a clear target and should build the whole programme around it.

The weaker playbook against the stronger one

The weaker approach is to write “AI search optimisation” into a plan and produce more pages. That treats two engines with a fourfold divergence on one bucket as a single destination. It also defaults to the only bucket you control, which both datasets say is a minority of the surface. The stronger approach is to find out which engine your buyers use, measure your own four-bucket mix on twenty prompts, and spend against the gap. One produces output. The other changes what the engine reads.

Run a live test before you commit

Take 20 prompts written in your buyers’ language rather than yours. Good examples look like “best observability platform for a regulated bank”, “how do I cut onboarding time for new engineers”, or “alternatives to the market leader for a small team”. Run each on both engines, log every cited domain, and bucket them. Repeat at 30 days and again at 90 days. If your mix on one engine barely moves while the other shifts, you have found which engine your work is actually reaching.

Red flags

  • The 11% domain overlap figure quoted as current, when its primary source is July 2025
  • A source-mix percentage quoted without saying whose bucket definitions it uses
  • Any “AI search” plan that does not name an engine
  • Advice to match exact buyer phrasing for ChatGPT, where word overlap is 13%
  • Vendor research presented without noting that the vendor sells the remedy
  • A single blended visibility score across engines that behave this differently
  • Anyone promising institutional citations on a quarterly timeline

Five questions worth asking any vendor

  1. Do you report cited domains per engine, and can I export them?
  2. Whose source-category definitions do you use, and where are they published?
  3. How do you handle prompts where the engine returned no citations at all?
  4. How many runs per prompt sit behind a reported change, and what is your variance?
  5. What is my current four-bucket mix, and how does it differ between ChatGPT and Perplexity?

The reducing principle. It comes down to one question: which engine do your buyers actually open? Everything else in this article is a refinement of that, because the two read different internets and designing for the average produces a plan that fits neither. If you genuinely do not know, that is the measurement to buy first, and it is cheaper than any content programme.

The honest closing note. For a lot of brands the right answer is that your category has no meaningful community presence and no institutional sources. That leaves media and your own pages, which is ordinary content and PR work. You do not need us to tell you that. We would rather say so than sell a two-engine strategy to a brand that needs one good set of pages.

What nobody should promise you

  • A current figure for domain overlap between the engines. The quoted one is from 2025 and we could not find a 2026 replication.
  • One strategy that works equally on both. They diverge fourfold on social and threefold on institutional.
  • Institutional citations on a short timeline. Those sources take years to earn, if they are available at all.
  • That your own website will carry the answer. Both large datasets put owned content in the minority.
  • Stable numbers. Every figure here is a 2026 snapshot and engine behaviour moves without notice.

Where this stops working, including for us

Two of the three datasets here come from a direct competitor, and we have said so each time we used them. They are the only per-engine splits at this scale we could find. Neither publishes a collection date range or a limitations section. That is the same gap we would criticise in anyone else, so discount accordingly and treat the direction as firmer than the decimals.

The four-bucket grouping is also a choice rather than a fact. A different grouping produces different headlines from the same citations. That is precisely what the disagreement above demonstrates.

And the retrieval-behaviour study covers three engines over fourteen days, which is a short window for systems that change weekly.

Our own limitation is specific. We report cited domains, not source categories, so the central framing of this article is not something our product outputs today. We also do not publish run-to-run variance. Check the sources yourself where we have linked them, and ask us for the raw domains where we have not.

Frequently asked questions

Where does ChatGPT get its information about companies?
Across 27 million citations analysed in January 2026, ChatGPT took 54% from the brand bucket, 25% from media, 15% from institutional sources and 5% from social. Its institutional share is three times Perplexity’s, so reference and standards sources matter more there.

Where does Perplexity get its information?
From a noticeably more social mix. In the same analysis Perplexity drew 19.4% of citations from community and user-generated sources, nearly four times ChatGPT’s 5%, alongside 57% brand, 19% media and 5% institutional.

Do the two engines cite the same sources?
Much less than you would expect, though the exact figure is unreliable. The widely quoted 11% domain overlap comes from a study published in July 2025, and we could not find a 2026 replication. Treat the direction as established and the number as stale.

Why does the same content perform differently on each?
Because they search differently before they cite. Word overlap between the user’s prompt and the queries actually run was 88% for Perplexity and 13% for ChatGPT. So matching exact phrasing helps on one engine and barely registers on the other.

Does my own website matter at all?
Yes, but less than most plans assume. The brand bucket sits above 50% on both engines, yet it includes competitor and other corporate pages, not just yours. Both large 2026 datasets agree your own site is a minority of the surface.

Which engine should I optimise for first?
Whichever your buyers actually open. If that is genuinely unknown, measuring it is cheaper than any content programme, and designing for the average of two engines this different produces a plan that suits neither.

Is Reddit really that important?
For Perplexity, community sources are close to a fifth of citations, so if your category is discussed in forums and you are absent, that shows up. For ChatGPT the same bucket is 5%, so the answer genuinely depends on the engine.

How do I check my own source mix?
Run twenty buyer questions on both engines, log every cited domain, and group them into brand, media, institution and social. Two hours, no tool required, and it gives you your own mix rather than an industry average.

Sources and further reading

  • Profound, citation category analysis, published 8 January 2026. 27 million citations across seven engines: ChatGPT, Gemini, Google AI Overviews, Google AI Mode, Perplexity, Claude and Microsoft Copilot. Source of the four-bucket source mix and all per-engine category shares. Published by a direct competitor, states no collection date range and no limitations.
  • Profound, retrieval behaviour analysis, published 30 April 2026. 10,000 randomly sampled prompts over a fourteen-day window, late March to mid April 2026, across ChatGPT, Perplexity and Copilot, capturing the complete query sets sent to retrieval. Source of the query uniqueness, word overlap, intent preservation and domains-per-answer figures. Same publisher caveat applies.
  • Muck Rack, “What Is AI Reading?”, third edition, May 2026. More than 25 million links from ChatGPT, Claude and Gemini across 17 industries. Source of the 84% earned and 0.3% paid shares, and of the bucketing disagreement discussed above.
  • Excluded on the 2026 rule: the widely quoted finding that 11% of cited domains overlap between ChatGPT and Perplexity, with 37.4% and 51.6% exclusive. Its primary source was published 1 July 2025 and is frequently re-dated to 2026 in secondary coverage.
  • Pepper, the trust pyramid of AI search, for the hierarchy behind source selection.
  • Pepper, Perplexity vs ChatGPT for brand visibility, for the commercial prioritisation question.
  • Pepper, getting cited on Perplexity, for the engine-specific tactics.
  • Pepper, the three levers behind citability, for where source selection sits.
  • Pepper, the crawler and robots.txt guide, for whether either engine can fetch you at all.

A note on sources. Only studies published in 2026 are cited. Where a figure comes from a vendor’s own research, we have said so in the line that uses it as well as here. Two of the three datasets in this article were published by a company we compete with.

Similar Posts