Third-party sources and AI citations: what off-page decides, and what it does not

The short answer
Off-page work decides whether an engine can find you, trust you and cite you. It does not decide whether the answer names you, and that is a separate event with a separate driver. Across 3,981 brand appearances studied in June 2026, 61.7% were citations with no brand name in the text and only 13.2% produced both. So an off-page programme can succeed completely and leave the reader never learning who you are. Worse, the two largest datasets on this disagree about how much of the citation surface is third-party at all, and the disagreement is about definitions rather than facts.
Key takeaways
- Off-page decides citation. It does not decide naming. 61.7% of appearances are citations with no brand name, and only 13.2% are both.
- The headline number is real and narrower than it sounds. 84% of AI citations come from earned media and 0.3% from paid, across more than 25 million links.
- A second large dataset appears to contradict that. Analysing 27 million citations, the brand bucket runs 54% to 73% depending on engine. Both are right, and they bucket differently.
- Never quote either share without naming whose definitions it uses. One counts third-party corporate content as earned; the other files it under brand.
- “Off-page” means different places on different engines. Perplexity draws 19.4% of citations from social and community sources against ChatGPT’s 5%, while ChatGPT takes 15% from institutional sources against Perplexity’s 5%.
- Paid placement buys three tenths of one percent of the surface. Use that figure to end any proposal built on sponsored coverage.
- The one claim both datasets support: your own website is a minority of what engines read about you.
- Pepper is an agentic organic growth engine and an organic growth partner. Agent Atlas puts the agents in your team’s hands. Pepper’s GEO platform reports Brand Visibility, Domain Prompt Presence and Share of Voice across six engines, including ChatGPT, Perplexity, Gemini and Google AI Overviews. A growth team works alongside yours. Eight years, more than 250 enterprises, more than 10 million tracked prompts.
A note on where this comes from. I spend my time on why some content gets picked up and most does not, and off-page is the part of that where the numbers in circulation are least compatible with each other. Pepper runs organic growth for more than 250 enterprises over eight years and tracks more than 10 million prompts across every major engine. The argument I have most often is not whether off-page matters. It is which of two incompatible percentages someone read first.
Disclosure: Pepper sells organic growth including earned-media work, so “off-page decides everything” is a conclusion we profit from. This article says off-page decides one of two outcomes and not the other, that the headline share depends on a definition nobody states, and that the cheapest fix in the whole subject is a sentence you can rewrite yourself. Every figure traces to a named study with its method, and two of them come from a direct competitor, which is said in the lines that use them.
What are third-party sources and AI citations, and what do they decide?
A third-party source is any page about you that you do not own. Journalism, analyst notes, review platforms, reference sites, forums, government and academic pages, and other companies’ corporate content.
They decide two of the three conditions for appearing in an AI answer, and not the third.

- Retrievability. Whether an engine can reach a page about you at all. Third-party pages are often better crawled and more frequently updated than your own.
- Citability. Whether the engine trusts the source enough to use and attribute it. This is where off-page does nearly all the work.
- Attributability. Whether the answer says your name. Off-page does almost nothing here, and it is a sentence-level choice on the page being quoted.
Our existing guide to earned authority covers the citability half in depth, including which sources matter for which questions. This article is about the boundary: what that work reaches and what it leaves behind.
The number everyone quotes, and what it actually covers
Muck Rack’s May 2026 analysis of more than 25 million cited links across ChatGPT, Claude and Gemini, spanning 17 industries, found:
- 84% of AI citations come from earned media. Journalism, academic and government sources, encyclopedic sites, and third-party corporate content.
- Journalism alone accounts for 27% of the total.
- Paid and advertorial account for 0.3%.
The stability is the useful part. Across three editions going back to mid-2025, the earned share has held between 82% and 89%, and journalism between 25% and 27%. Almost nothing else in this category holds within a few points across a year. That makes it the rare figure you can budget against.
The 0.3% is the half nobody quotes and the more actionable one. Any part of a budget spent on placement is buying three tenths of one percent of the citation surface. Use it to end a proposal, not to start one.
If you want this mapped against your own category rather than an aggregate, book a growth audit and bring twenty questions your buyers ask.
Two large datasets appear to contradict each other
Here is the part that causes the arguments, and resolving it changes where money goes.
A second analysis, published 8 January 2026, examined 27 million citations across seven engines and grouped them into four buckets. Its brand bucket, covering owned, competitor and other corporate pages, runs 54% to 73% depending on the engine.

Read quickly, those two findings cannot both be true. One says most citations are earned, the other says most are corporate pages.
They are both right, and the resolution is definitional. Muck Rack’s earned category explicitly includes third-party corporate content in its count. The four-bucket grouping files the same pages under Brand. So a supplier’s page describing your product lands in “earned” in one dataset and “brand” in the other. A competitor’s page mentioning you does the same.
The practical rule is worth making a house standard. Never quote a source-mix share without saying whose definitions it uses. The two numbers are not rival estimates of one quantity. They are two different quantities with similar names.
The one claim both support is the one that should drive a plan: your own website is a minority of what engines read about you. That holds under either grouping, and it is the finding that justifies off-page spend at all.
Off-page means different places on different engines
“Invest in third-party sources” is incomplete advice, because the engines read different third parties.

From the same 27-million-citation analysis:
- Perplexity draws 19.4% of citations from social and community sources. ChatGPT draws 5%.
- ChatGPT draws 15% from institutional sources, meaning government, academic and standards bodies. Perplexity draws 5%.
- Copilot leans hardest on media at 33%, against Gemini’s 18%.
So the honest version of the advice is engine-specific. If your buyers use Perplexity, forums and review platforms are close to a fifth of what it reads about your category. If they use ChatGPT, that bucket is a twentieth and standards bodies matter three times more. Our comparison of where each engine gets its information covers the prioritisation question.
Where off-page stops: the naming problem
Now the boundary, and it is the reason this article exists alongside our earned authority guide.
Across 3,981 brand appearances studied in June 2026:
- 61.7% were citations with no brand name in the answer.
- 25.1% were mentions with no citation.
- Only 13.2% were both.
So the thing off-page work buys you happens without your name attached nearly two thirds of the time. You can earn the coverage, get retrieved, get cited, and the reader still finishes the answer not knowing who you are.
The fix is not more off-page work. It is a sentence-level choice on the pages being quoted, including the third-party ones. A claim written as “our analysis found X” survives extraction with the attribution stripped. Written as “Pepper’s analysis of ten million prompts found X”, it carries the name into the answer.
Which means the practical off-page brief has an extra line in it. When you place a story, contribute a quote or supply data to a review platform, specify that the quotable sentence names the brand rather than the role. It is one line in the brief. It costs nothing and it is the difference between the 61.7% and the 13.2%.
How we weighted the off-page work
Four criteria, and they are not equal.

| Criterion | Weight | Why it carries that weight |
|---|---|---|
| Share of the citation surface it carries | 30% | Earned is 84% under one grouping and the largest non-brand bucket under the other. Paid is 0.3% under both. |
| Whether you can realistically influence it | 30% | Institutional sources are close to unreachable on a quarterly timeline. Review platforms are not. |
| Whether it carries your name into the answer | 25% | 61.7% of citations do not, and the fix is at the sentence rather than the placement. |
| How differently the engines treat it | 15% | Social swings from 5% to 19.4% between engines, so it is a per-engine decision. |
What off-page decides, at a glance
| Retrievability | Citability | Attributability | |
|---|---|---|---|
| Does off-page decide it | Partly | Almost entirely | Barely |
| Measured evidence | Third-party pages are often better crawled | 84% earned, 0.3% paid | 61.7% cited without a name |
| What moves it | Coverage existing at all | Coverage on sources engines trust | The wording of the quotable sentence |
| Time | Quarters | Quarters to years | Days |
| Cost | Moderate | High, and not fully controllable | Free |
| Who does it | PR and partnerships | PR, analyst relations, review platforms | Whoever writes the quote |
| Fails when | Nobody independent writes about you | Your category has no trusted sources | The quotable line says “we” |
What to do with this
- Decide whose definitions you are using before quoting any source-mix figure internally, and write it next to the number.
- Kill any line item built on paid placement. It is 0.3% of the surface under the better-documented grouping.
- Pick your third parties by engine, not in general. Community presence is worth nearly four times as much on one engine as another.
- Add the naming line to every off-page brief. The quotable sentence should carry the brand, not the role.
- Audit what is already citing you. Open the source lists on twenty buyer questions and see which third parties the engines already reach for. That is your list, and it is usually shorter than expected.
What this costs
The audit is free and takes about two hours. Run twenty buyer questions on two engines, log every cited domain, and group them.
The naming change is also free. It is a line in a brief and a sentence in a draft.
The expensive part is earning coverage that does not exist yet, which is slow and not fully in your control. Published GEO retainers run from $3,000 to $25,000 a month depending on scope, and much of what sits inside them is this work. Do the free two hours first, because it tells you whether you need the rest.
How Pepper fits
Pepper is an agentic organic growth engine and an organic growth partner, which means three things working together rather than one product.
Pepper’s GEO platform is the self-serve workspace. Brand profile, competitors, personas, GA4 and Search Console connected, themes and prompts defined, with Brand Visibility, Domain Prompt Presence and Share of Voice across six engines, including ChatGPT, Perplexity, Gemini and Google AI Overviews. Citation Analysis is the relevant part here, because it shows which domains an engine actually used, which is the only way to find your own third-party list rather than an industry average.
Agent Atlas is where your team builds, versions and runs its own agents, with quick runs for one input and sheet runs for bulk. Pulling cited domains out of several hundred answers and grouping them is exactly the repetitive work it exists for.
The growth team is attached to the account and works alongside yours, which matters most here because earning coverage on pages you do not own is relationship work rather than publishing work.
Where it falls short: we report which domains were cited, not which bucket they belong to. The four-bucket grouping in this article comes from a competitor’s published research, not from our product, and we are not going to pretend otherwise. We also measure presence rather than accuracy, so we will tell you a third-party page cited you and not whether what it said about your pricing was still true.
Eight years, more than 250 enterprises, more than 10 million tracked prompts. You can see the shape of the work in the Acceldata case study, in how we run it for B2B SaaS brands, and across the case study library.
How to choose where to spend off-page
The decision is which third parties, in which order, and the generic answer is usually wrong for a specific category. So here are the criteria, weighted.
The weighted scorecard
Score each row from 1 to 5, multiply by the weight, and total out of 100.
| Criterion | Weight | Score 1 means | Score 5 means |
|---|---|---|---|
| Existing third-party coverage of your category | 30 | Nobody independent writes about you | Regular journalism and analyst notes |
| Concentration of your buyers on one engine | 25 | Spread evenly across four | Almost all on one |
| Whether your category has public communities | 25 | No forums, no review platform of note | Active subreddits, busy review category |
| Whether your quotable claims carry your name | 20 | Generic prose anyone could have written | The brand sits inside the claim |
Under 40, the off-page lever is not available to you yet and the work is building something worth covering. From 40 to 70, fix the naming first because it is free, then pursue the one or two source types your engine favours. Above 70, this is a sustained programme with an owner rather than a project.
The weaker playbook against the stronger one
The weaker approach is to read the 84% figure, conclude that off-page is everything, and buy placement. That misreads the number twice. It ignores the 0.3% saying paid is not the route, and the 61.7% saying citation without naming is the usual result. The stronger approach is to audit which third parties already cite you, pick the source types your buyers’ engine actually favours, and add the naming line to every brief. One buys a category. The other buys the handful of pages that decide your category.
Run a live test before you commit a budget
Take 20 questions your buyers actually ask, written in their words, such as “best inventory platform for a mid-size retailer”, “how do I reduce false positives in fraud screening”, or “alternatives to the incumbent for a regulated business”. Run each on two engines and log every cited domain, grouped by type. Repeat at 30 days and again at 90 days. If the same six third-party domains keep appearing, that is your off-page list, and it is almost always shorter and more specific than a generic PR plan.
Red flags
- A source-mix percentage quoted without saying whose bucket definitions it uses
- Any proposal built on sponsored or advertorial placement, at 0.3% of the surface
- Treating a citation and a mention as the same outcome
- “Invest in off-page” with no engine named, when one bucket swings fourfold between them
- Promises of coverage on specific named publications
- A brief that does not specify how the brand should appear in the quotable sentence
- Vendor research presented without noting that the vendor sells the remedy
Five questions worth asking any agency
- Which third-party domains already cite us, and can you show me the list?
- Whose source-category definitions are you using for that percentage?
- How much of this plan is paid placement, given it is 0.3% of the surface?
- What share of our citations currently carry our name, and how will you move it?
- Which engine are you optimising for, and is it the one our buyers use?
The reducing principle. It comes down to one question: are you trying to be found, or to be remembered? Off-page work is how an engine finds and trusts you, and it does that job well. Being remembered is decided by whether the sentence it lifts contains your name, and that is a writing decision on a page that may not even be yours. Everything else in this article is a refinement of that split.
The honest closing note. If your category has no independent coverage and no public community, the off-page lever is not available to you yet and you do not need us for it. The work is building something worth writing about. That is a product and positioning problem rather than a marketing one. We would rather say that than sell an earned-media programme into a category with no earned media.
What nobody should promise you
- Coverage on a named publication. Nobody controls editorial, and anyone promising it is selling placement.
- A single number for how much of AI citation is third-party. It depends on a definition the two main datasets do not share.
- That citations will carry your name. Nearly two thirds do not.
- Results from paid placement. It is 0.3% of the citation surface.
- A fast version of earned coverage. It moves over quarters, and the slow part is not effort.
Where this stops working, including for us
Two of the three datasets here come from companies we compete with. Neither publishes a collection window or a limitations section. Their definitions are explicit, which is why they are usable, and we have flagged the conflict of interest in each line that uses them.
The four-bucket grouping is a choice rather than a fact. A different grouping produces a different headline from the same citations, which is exactly what the disagreement in this article demonstrates. Treat the direction as firmer than the decimals.
Muck Rack’s stability is its strongest feature and also a limitation. It measures which sources engines cite across the whole web, not which citations influence a purchase, and not your category specifically.
And our position is not neutral. We sell earned-media work, so “off-page decides everything” would be a convenient conclusion. This article argues it decides two of three conditions, that the third is free to fix without us, and that the lever is unavailable entirely to categories nobody writes about.
Where to go next
- For what actually earns a citation, read why third-party sources decide whether AI cites you.
- For the naming problem on its own, read the ghost citation problem.
- For which engine reads which sources, read Perplexity against ChatGPT.
- For the trust hierarchy behind source selection, read the four-tier trust model.
- For the three levers this sits inside, read Visibility, Citability and Retrievability.
Frequently asked questions
Do third-party sources really decide AI citations?
Largely, yes. Across more than 25 million cited links, 84% came from earned media and 0.3% from paid. But they decide whether you are cited, not whether the answer names you, and those are separate events with separate fixes.
Why do two studies give completely different numbers?
Because they bucket differently rather than disagree factually. One counts third-party corporate content as earned media; the other files the same pages under brand. Always state whose definitions a source-mix percentage uses.
Is paid placement worth anything for AI visibility?
Very little on the published evidence. Paid and advertorial accounted for 0.3% of citations across more than 25 million links. Use that figure to end a proposal built on sponsored coverage rather than to justify one.
Which third-party sources should I prioritise?
It depends on the engine your buyers use. Perplexity draws 19.4% of citations from community sources against ChatGPT’s 5%, while ChatGPT takes 15% from institutional sources against Perplexity’s 5%.
My pages get cited but my brand is never named. Why?
Because naming is decided separately from citation, and 61.7% of appearances are citations with no name. The fix is the wording of the quotable sentence, not more coverage. Put the brand inside the claim itself.
Does my own website still matter?
Yes, but less than most plans assume. Both large 2026 datasets agree your own site is a minority of what engines read about you, whichever bucket labels they use. Treat it as necessary rather than sufficient.
How do I find which third parties already cite me?
Run twenty buyer questions on two engines and log every cited domain. Two hours, no tool needed. The recurring six or so domains are your actual off-page list, and it is usually shorter than a generic PR plan.
What should go in an off-page brief that usually does not?
A line specifying how the brand should appear in the quotable sentence. A claim written as “our analysis found X” loses the attribution on extraction. Written as “[Brand]’s analysis found X”, it carries into the answer.
Sources and further reading
- Muck Rack, “What Is AI Reading?”, third edition, May 2026. More than 25 million links from ChatGPT, Claude and Gemini across 17 industries. Source of the 84% earned, 27% journalism and 0.3% paid shares, and of the 82% to 89% stability across three editions. Measures which sources engines cite, not which citations influence a purchase.
- Profound citation category analysis, published 8 January 2026. 27 million citations across seven engines. Source of the four-bucket grouping, the 54% to 73% brand range, and all per-engine social, institutional and media shares. Published by a direct competitor, with no stated collection window or limitations.
- Semrush with Kevin Indig and Growth Memo, Why 62% of AI citations don’t lead to brand mentions, published 9 June 2026. 3,981 domain appearances across 115 prompts and 14 countries. Source of the 61.7%, 25.1% and 13.2% split. Vendor research with no stated collection window or limitations.
- Pepper, the citability half in depth, which this article complements by marking where off-page work stops.
- Pepper, why citation and naming come apart, for the naming half in depth.
- Pepper, the source trust hierarchy, for how tiers of authority differ.
- Pepper, which engine to prioritise, for the commercial side of the engine question.
A note on sources. Only studies published in 2026 are cited. Two of the three were published by companies we compete with, which is stated in the lines that use them as well as here, and the definitional conflict between the first two is treated as the article’s subject rather than resolved in favour of whichever is more convenient.
Latest Blogs
Choosing a GEO agency for mid-market B2B used to be a shortlisting problem. It is now a procurement problem, because the published prices have largely gone. Of four agencies publishing a GEO-specific figure in September 2026, two had withdrawn their pricing pages by early October and one had moved domains. Exactly one still publishes a number a mid-market buyer can act on. So the useful question is no longer which agency is best in the abstract, but how to compare three quotes when only one of them arrived with a method attached.
Third-party sources and AI citations are tightly linked, and the link stops short of where most plans assume. Off-page work decides whether an engine retrieves and cites you. It does not decide whether the answer names you, and 61.7% of brand appearances are citations with no name attached. The two largest 2026 datasets also disagree about how much of the citation surface is third-party at all, one saying 84% and the other putting the brand bucket above half. That disagreement is definitional rather than factual, and it decides where a budget goes.
Why is my competitor showing up in ChatGPT and not me? Before accepting the premise, check it. Across 3,981 brand appearances studied in 2026, 61.7% were citations with no brand name in the answer and only 13.2% produced both. So there are three states that look identical from where you are sitting: genuinely absent, present as an unnamed source, and named less often than a rival. Each has a different cause and a different fix, and working on the wrong one is the most common way this gets expensive.