GEO / AI Search

How to get cited by Perplexity

janvi
•
Posted on 8/10/26•10 min read
How to get cited by Perplexity

Perplexity is the easiest big engine to earn a citation from. It is also the easiest to misread. Both come from the same fact: it cites widely, and leans on each source less.

The short answer

The engine documents two user agents, and they behave differently. Getting that right is most of the technical work. The content work follows the same logic as everywhere else, with one adjustment.

Four things, in order of how much they matter.

  • Two agents, blocked differently. PerplexityBot “is designed to surface and link websites in search results on Perplexity”. Block it and you are out of those results. The other agent, Perplexity-User, fetches a page when someone asks a question, and the docs say it “generally ignores robots.txt rules” because a person asked.
  • Breadth is the opening. Research across three engines found this one cites more sources, and leans on each less. ChatGPT cites fewer and leans harder. More slots, lower bar.
  • Breadth is also the trap. A citation here is worth less than one in ChatGPT on the same question. So a rising rate here is not proof of a wider win.
  • You can verify the crawler. The docs publish IP ranges for both agents. So you can check that what hit your server was the real thing.

The honest framing

Which signal are you after? Want an early signal that your content work is landing? This engine shows it first. Want proof you have won the category? It is the weakest of the big engines to prove that with, and if this is the only engine your numbers moved on, you do not need to tell your board anything yet.

Pepper is an agentic organic growth engine, not an SEO agency. This page comes out of what we run: organic for more than 250 enterprises across eight years, more than 10 million tracked prompts, and the questions buyers put to us in client reviews. Customers log in and run the platform themselves, with a Pepper growth team attached. Book a growth audit and we will show you your citation rate across engines rather than just this one.

Key takeaways

Key takeaways on what is documented and what is inferred.

  • PerplexityBot decides whether you can appear. Block it and you are out of the search results.
  • Perplexity-User generally ignores robots.txt. The fetch follows a user action, so robots rules are not a privacy control.
  • Both agents have published IP ranges. So you can check a visit rather than trust a header.
  • It cites broadly. More sources per answer, less weight on each, from controlled research across three engines.
  • A citation here is an early signal. It arrives sooner and proves less than one on a narrower engine.
  • No ranking guide exists. The docs cover crawlers, not retrieval. So everything else is inference.

What is a Perplexity citation?

A citation here means the engine used one of your pages as a source, and linked it. It shows as a numbered note next to the claim.

The engine shows its sources more openly than the others. That has two effects worth planning for.

  • Readers see the citations. The sources sit beside the answer, not under it. So a citation does more brand work here than on engines that bury the list.
  • There are more of them per answer. That lifts your odds of showing up, and cuts what each slot is worth.

Where it falls short. Source seven of nine counts as a citation in the data. It barely registers with a reader. So count citations, then read a sample to see where in the list you sit. We separate the two underlying outcomes in what is brand visibility in AI.

Figure 1: Two agents, two behaviours. Only one of them obeys robots.txt.

What the research says about Perplexity specifically

Past the crawler docs, nothing is published about how sources get picked or ranked. The best evidence is independent, and worth reading closely.

One controlled study covered 602 prompts, 21,143 search-layer citations and 18,151 fetched pages across three engines. Two findings matter here.

It cites more sources, with less weight on each. ChatGPT cites fewer and leans harder on them. That is the real split between the engines, and it changes what a citation means on each.

Pages with real influence share a shape. They run longer, carry more structure, sit closer to the question, and hold more usable evidence: definitions, numbers, comparisons and steps. The same study found citation counts alone do not measure performance well, because breadth and influence are not the same thing.

What we take from that. Build for the same qualities you would for any engine. Then read the numbers here with a discount for breadth. A rate that doubles here and moves nowhere on ChatGPT is a real signal, and a partial one.

How to get cited by Perplexity, in four steps

Step 1: Fix the crawler configuration

This is documented, free, and the only part of this page where the engine states the result itself.

The docs name two user agents, and they do different jobs.

  • PerplexityBot. Documented as “designed to surface and link websites in search results on Perplexity. It is not used to crawl content for AI foundation models.” Block it and your site will not show in those results.
  • Perplexity-User. Documented as supporting user actions: “When users ask Perplexity a question, it might visit a web page to help provide an accurate answer and include a link to the page in its response.” The docs say this one generally ignores robots.txt, because a user asked for it.

That second point surprises people, and it cuts both ways. A robots.txt disallow will not stop a fetch made for a user, so it is not a privacy or paywall control. And allowing the user agent while blocking the search one will not get you into search results. The two serve different surfaces.

Verify rather than assume. The docs publish IP ranges for both agents as JSON files on their own domain. Anyone can copy a user-agent string, so if your logs are driving the decision, check the address against the published list. The wider technical sweep is in how to audit your site for AI search.

Figure 2: The structural difference between the engines, and what it means for your reporting.

Step 2: Put the right things on the page

The engine-agnostic work applies here. We are not going to invent a trick nobody can verify.

  • Answer at the top of each section. Engines pull passages, not whole pages. The planning side of that sits in how to build an SEO content plan. A section that opens with context and answers in paragraph four hands the engine the wrong fifty words.
  • Put something usable in every section. A defined term, a number with its source, a comparison, a set of steps. The research is clear that pages rich in this stuff carry more weight when cited.
  • Make each section stand alone. A passage that leans on three paragraphs above it is harder to lift.
  • Write the commercial pages you have avoided. Pricing, alternatives and comparison pages are what these questions ask for. They are also the ones most brands leave to other people.

Where that falls short. None of it is in the docs. It comes from independent research on what cited pages look like, which is a link rather than a published rule. We would rather tell you the basis than present a formula. The page-level detail is in how to structure content for answer engines.

Step 3: Work on what sits off your site

These answers lean hard on sources you do not control. Which ones varies by category, more than any article can tell you.

Independent research across 11,000 real search queries and five systems found that identical queries produce “structurally different information realities across systems”, with some source types systematically over-represented and others under-represented. So find your own answer rather than take a list from anybody, including us.

The job takes an afternoon.

  • Run twenty category prompts and copy out every cited domain, not only the ones you know.
  • Sort by how often each domain shows up. The top few are the authority layer in your category, as this engine sees it.
  • Split what you can influence from what you cannot. Review sites, publications, docs sites and communities each need a different approach. Some need none.

We go through the full method in how to find which sources AI cites in your category.

Figure 3: Four steps, bottom up. Only the first has documented consequences.

Step 4: Measure your own citation rate

Nobody gives you this number. There is no webmaster reporting here at all. The first-party AI reporting that does exist belongs to Google and Microsoft, and covers their own surfaces.

So measure it directly.

  • Fix twenty prompts across awareness, comparison and purchase questions.
  • Run each in a fresh session, because history and settings change results. Note the country you ran from.
  • Record two columns. Did it name your brand? Did one of your pages show up in the sources? Keep them apart, because they move for different reasons.
  • Note where you sit in the source list, on a sample. Seventh of nine is not second of four, and the rate alone hides that.
  • Repeat monthly with the same prompts and watch the direction rather than the level.

Two readings give you a rate of change for your own category. That beats any benchmark anyone can sell you. The wider routine is in how to track brand mentions in AI search.

And compare across engines. Here is why: a single-engine view shows progress early and overstates it. Running the same set on a second engine costs an hour, and stops you reporting one engine’s gain as a category win. That discipline is in how to benchmark AI visibility against competitors.

What the evidence does not support

Three things get repeated here that we would not.

  • That the ranking method is known. The docs cover crawlers and nothing about retrieval or ranking. Every factor list out there is inferred from samples.
  • That blocking Perplexity-User protects your content. The docs state it generally ignores robots.txt, because the fetch follows a user action. Robots rules are not an access control here.
  • That a citation here equals one in ChatGPT. The research says otherwise. More sources, less weight each. One word, two outcomes.

And one we are wary of. Figures for how well AI-referred traffic converts, or what share of citations come from earned media, vary hugely between studies. Most are published by firms selling the service the number implies you need. We have twice failed to retrieve the underlying study behind numbers of this kind, so none appear on this page.

Figure 4: Three claims worth refusing, and what the documentation actually says.

What nobody should promise you

  • A guaranteed Perplexity citation. Nobody outside Perplexity controls the output, and Google’s own guidance warns against providers who guarantee rankings for the same structural reason.
  • A Perplexity ranking factor list. None is published. Inference is reasonable; presenting it as documented is not.
  • That robots.txt keeps your content out of answers. Perplexity-User generally ignores it, by Perplexity’s own statement.
  • That a rising Perplexity rate proves a category win. It is the broadest-citing major engine, so it moves first and proves least.

Frequently asked questions

How do you get cited by Perplexity?
Allow PerplexityBot, which Perplexity documents as the agent that surfaces and links websites in its search results, then publish pages that answer directly and contain extractable evidence: definitions, numbers with sources, comparisons and procedures.

What is the difference between PerplexityBot and Perplexity-User?
PerplexityBot surfaces and links sites in Perplexity search results and is not used to crawl content for foundation models. Perplexity-User visits a page when a user asks a question, and Perplexity states it generally ignores robots.txt rules.

Access and verification

Does blocking PerplexityBot remove me from Perplexity?
Yes, from its search results. The documentation states that blocking it prevents your site appearing there, which makes this the one configuration decision with a consequence the engine states itself. Everything else in this field is inference.

Can I verify that a visit really came from Perplexity?
Yes. Perplexity publishes IP ranges for both agents as JSON files on its own domain. Check the address against the published list rather than trusting the user-agent string, which anyone can copy.

Does robots.txt stop Perplexity reading my content?
Not reliably. PerplexityBot respects robots.txt, but Perplexity states Perplexity-User generally ignores it because the request follows a user action. Treat robots.txt as a crawling preference rather than an access control.

Content and measurement

Is Perplexity easier to get cited by than ChatGPT?
On the evidence, yes. A controlled study across three engines found Perplexity cites more sources with less influence per source, while ChatGPT cites fewer with more weight each. More slots available, and each one worth less.

How do I measure my Perplexity citation rate?
Run twenty fixed prompts in fresh sessions, record whether your brand was named and whether a page from your domain appeared in the sources, note your position in the source list on a sample, and repeat monthly.

Does Perplexity have anything like Search Console?
No. There is no publisher reporting from Perplexity. Google’s generative AI performance report and Microsoft’s AI Performance report cover their own surfaces only, so your own prompt set is the measurement.

Sources

Every claim traces to a primary source, read directly on the date shown.

Perplexity documentation, read 8 October 2026

  • Perplexity, PerplexityBot and Perplexity-User documentation. Documents PerplexityBot as “designed to surface and link websites in search results on Perplexity. It is not used to crawl content for AI foundation models”, with blocking preventing appearance in Perplexity search results. Documents Perplexity-User as supporting user actions: “When users ask Perplexity a question, it might visit a web page to help provide an accurate answer and include a link to the page in its response”, and states this agent generally ignores robots.txt rules because requests originate from user actions. Publishes IP ranges for both agents as JSON on its own domain.

Independent research

  • Zhang Kai, He Xinyue and Yao Jingang, From Citation Selection to Citation Absorption, arXiv, 28 April 2026. 602 controlled prompts, 21,143 search-layer citations, 18,151 fetched pages and 72 features across ChatGPT, Google AI Overviews and Perplexity. Finds Perplexity and Google cite more sources with less influence per source while ChatGPT cites fewer with higher influence each, that high-influence pages are longer, more structured and richer in extractable evidence, and that citation counts alone inadequately measure generative engine performance. Academic preprint, not peer-reviewed at the time of writing.
  • Huang, Goyal, Saha and Chandrasekharan, Answer Bubbles: Information Exposure in AI-Mediated Search, arXiv, 17 March 2026, revised 28 August 2026. 11,000 real search queries across five systems, finding identical queries produce “structurally different information realities across systems”. Also an academic preprint.

On what this page is not

This is not a ranking guide, because Perplexity publishes no retrieval or ranking documentation. What it does publish is the crawler behaviour, and we have quoted that directly. Everything about content is drawn from independent research into what cited pages look like, which is a correlation rather than a mechanism, and we have labelled it as such throughout. We have deliberately omitted conversion and citation-share percentages that circulate in this field, because we could not retrieve the studies behind them.