How to optimize your website for ChatGPT and Perplexity, engine by engine

Every major AI company runs two separate crawlers: one that decides whether you appear in its answers, and one that only collects training data. They are different tokens in robots.txt. Plenty of teams read a headline about AI scraping, blocked the training bot, felt protected, and never realised the search bot was a different line. Others blocked both and wondered why their visibility went to zero.
Pepper tracks over 10 million prompts across every major engine, and this is the single most common technical fault we find on a new account. The rest of this page is what actually differs between engines once access is fixed.
The short answer
Allow the search crawlers, block the training crawlers if you want to, and stop treating “AI” as one destination. Engines retrieve differently enough that a single strategy underperforms on all of them.
The three things that genuinely differ per engine: which crawler must be allowed, how many sources the engine cites, and which kinds of sources it favours.
Key takeaways
OAI-SearchBotcontrols ChatGPT visibility, notGPTBot. GPTBot is training only. Blocking it costs you nothing in ChatGPT search.Google-Extendedis the same trap at Google. Google states plainly it “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal.”- ChatGPT cites fewer sources than Perplexity or Google, but each one carries more influence in the answer. That changes what winning looks like.
- Perplexity-User generally ignores robots.txt, because a person asked for that page directly.
- Google is the simplest of the three. AI Overviews run on core Search ranking systems with no separate index, so ordinary Googlebot access is the whole requirement.
A note on where this comes from. We run organic for more than 250 enterprises and track over 10 million prompts across every major engine. What follows is shaped by that, by the client reviews we sit in weekly, and by the access faults we find on accounts we inherit. Every crawler token below was checked at the vendor’s own documentation on 10 September 2026.
What is AI visibility, and why does it differ by engine?
AI visibility is whether an engine mentions your brand and cites your pages when someone asks a question in your category. It splits into two numbers worth keeping separate: Brand Visibility, how often engines name you, and Domain Prompt Presence, how often they cite a page from your domain. A wide gap between them means engines know you but trust somebody else to describe you.
The reason it differs by engine is not marketing. It is measurable.
Huang, Goyal, Saha and Chandrasekharan ran 11,000 real search queries across five systems in March 2026 and found that identical queries produce “structurally different information realities across systems.” They also found Wikipedia and longer sources disproportionately overrepresented, and social media and negatively framed sources substantially underrepresented.
Zhang, He and Yao measured 21,143 citations in April 2026 across ChatGPT, Google AI Overview and Perplexity, and found citation breadth and depth diverge sharply: Perplexity and Google cite more sources, while ChatGPT cites fewer but each one carries more influence over the answer.
Both are academic and independent of any vendor in this space. Together they are the evidence that “optimise for AI” is too coarse an instruction to act on.
Crawler access at a glance, and the cost of getting it wrong
| Engine | Allow this to appear in answers | This one is training only | Cost to fix |
|---|---|---|---|
| ChatGPT | OAI-SearchBot | GPTBot | Free, one robots.txt line |
| Perplexity | PerplexityBot | None. It is not used for training | Free, one robots.txt line |
| Google AI Overviews and AI Mode | Googlebot | Google-Extended | Free, already allowed on most sites |
Checked at each company’s own documentation on 10 September 2026. Blocking any of the middle column costs nothing in visibility, which is the point most coverage of this topic misses.


Check yours before reading further. Open yourdomain.com/robots.txt and search for those six tokens. If OAI-SearchBot or PerplexityBot is disallowed, nothing else on this page matters until that changes.
If you want the check run across your whole site alongside your prompt set, book a growth audit and we will show you which engines can currently reach you.
How we weight the engine-by-engine work
| Area | Weight | What decides the score |
|---|---|---|
| Crawler access | 35 | Whether each engine’s search crawler can fetch and render your key templates at all. Nothing else counts until this passes |
| Extractability | 25 | Whether a single passage answers a question cleanly enough to be lifted without rewriting |
| Source authority off-site | 25 | Whether the sources each engine favours mention you, which differs by engine more than most teams expect |
| Freshness and dates | 10 | Visible, accurate dates, which matter more on engines that weight recency |
| Structured data | 5 | Helpful for machine readability, and consistently oversold as a ranking lever |
Access carries the most weight because it is binary. A page an engine cannot fetch scores zero on everything else. Our [GEO agency ranking methodology](https://www.pepper.inc/blog/geo-agency-ranking-methodology/) explains how we build weightings like this.
ChatGPT
What it runs. OAI-SearchBot surfaces your pages in ChatGPT’s search features. GPTBot collects training data for foundation models. ChatGPT-User handles fetches a person triggered directly. OAI-AdsBot validates ad landing pages.
- Allow:
OAI-SearchBot. Disallowing it prevents your content appearing in ChatGPT answers - Optional to block:
GPTBot. Blocking signals your content should not train foundation models, and does not affect ChatGPT search visibility - Citation behaviour: fewer sources per answer than Perplexity or Google, with each carrying more influence
- What that means in practice: being one of three cited sources on ChatGPT is worth more than being one of twelve elsewhere. Depth on a narrow set of questions beats breadth
- Where this falls short: because it cites fewer sources, ChatGPT is the hardest of the three to break into, and progress is lumpy rather than gradual
What to do. Pick the fifteen questions that matter most commercially and make one page the single best answer to each. Spreading thin across two hundred topics performs worse here than anywhere else.
Perplexity
What it runs. PerplexityBot surfaces and links websites in Perplexity results and is explicitly not used for model training. Perplexity-User fetches pages when a user asks something directly, and generally ignores robots.txt rules, because a person requested that page.
- Allow:
PerplexityBot. There is no training-only bot to weigh up here - Verification: Perplexity publishes IP ranges for both agents as JSON, so your infrastructure team can allowlist by IP as well as user agent rather than trusting a header
- Citation behaviour: cites more sources per answer than ChatGPT, with less weight on each
- What that means in practice: breadth pays. Being cited on many adjacent questions accumulates, where the same effort on ChatGPT might not surface at all
- Where this falls short: more citations per answer also means each one drives less attention, so Perplexity visibility converts to traffic more slowly than the citation count suggests
What to do. Cover the full question set around a topic rather than perfecting one page. Perplexity rewards a well-linked cluster more than a single flagship article.
Google AI Overviews and AI Mode
What it runs. Ordinary Googlebot. Google’s May 2026 guidance states that AI Overviews and AI Mode use core Search ranking systems with no separate AI index, and that AEO and GEO are part of SEO. Google-Extended is a separate token controlling whether your content trains Gemini models, and Google states it “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.”
- Allow:
Googlebot, which almost every site already does - Optional to block:
Google-Extended, at no cost to Search or AI Overviews visibility - Citation behaviour: cites more sources, like Perplexity
- What that means in practice: this is the one engine where your existing SEO work carries over almost entirely. Ranking well is most of the job
- Where this falls short: it is also the most competitive, because everyone doing ordinary SEO is already competing here without meaning to
What to do. Nothing engine-specific. Fix crawling, rendering and content depth, and Google AI features follow. If a page is not indexed it cannot appear in an AI Overview, so our enterprise SEO audit framework is the relevant checklist.

What is the same across all three?
More than the differences, and it is worth saying so rather than selling three separate programmes.

Retrievability. Every engine has to fetch and render the page. Retrieval failures caused more than 70% of chatbot errors in a Stanford-led evaluation of six commercial chatbots across 2,100 questions in May 2026. Content that only exists after JavaScript executes is the most common cause we find.
Extractable structure. Zhang and colleagues found the pages engines actually absorb are longer, more structured, semantically aligned, and richer in extractable evidence: definitions, numerical facts, comparisons and procedural steps. That finding holds across all three engines.
Sources you do not own. Both academic studies point at third-party sources carrying most of the answer layer, and Muck Rack’s analysis of more than 25 million cited links found earned media accounted for 84% of AI citations against 0.3% for paid. Muck Rack sells PR software, so read it as an interested party with a large dataset, but it points the same way as the independent work.
What this work costs

Crawler access: free. Editing robots.txt. The only cost is finding out, and it is the highest-return ten minutes available in this whole discipline.
Rendering fixes: engineering time. The expensive one, and the one most likely to sit in a backlog. Scope it before promising a timeline.
Extractable structure: editorial time. Rewriting existing pages so a passage can be lifted cleanly. Cheap per page, slow across a library.
Earned media: the largest line. It supplies most of the citation pool and cannot be done from your own site, so it needs either a PR function or an outside partner. Our cost breakdown of platforms against hiring models the tradeoff.
Measurement is now nearly free to start. Several AI visibility platforms have genuine free tiers, covered in the best answer engine optimization tools.
What nobody should promise you
Guaranteed citations on any engine. Google advises against providers guaranteeing rankings, because third parties cannot access internal ranking systems. Nobody has that access for ChatGPT or Perplexity either.
That blocking AI crawlers protects you at no cost. Blocking the training bots is free. Blocking the search bots removes you from the answers entirely, and the tokens look similar enough that this is genuinely easy to get wrong.
That structured data will get you cited. It makes meaning machine-readable. It does not make a weak page strong, and it is the most consistently oversold tactic in this category.
That one engine’s results predict another’s. The research says the opposite. Identical queries produce structurally different results across systems.
How to evaluate your current position, engine by engine
I would run this in a fixed order, because the steps are dependent and doing them out of sequence wastes the expensive ones.
Google’s May 2026 guidance is the right anchor for scope. Google states that AEO and GEO are part of SEO, and that AI Overviews and AI Mode run on core Search ranking systems with no separate AI index. That is correct for Google, and Google is one engine. ChatGPT and Perplexity retrieve and cite differently, which the academic work confirms. Hold both: the fundamentals carry over from work your team already does, and the per-engine layer is real but thinner than the category’s marketing suggests.
Score your position before changing anything.
| Area | Weight | What a strong position demonstrates |
|---|---|---|
| Crawler access per engine | 35 | OAI-SearchBot, PerplexityBot and Googlebot all allowed, verified in the live robots.txt rather than assumed from a deploy |
| Render parity | 25 | The content that matters exists in what a crawler receives, not only in what a browser assembles after JavaScript |
| Extractability | 25 | A single passage on each key page answers its question cleanly enough to lift without rewriting |
| Off-site presence | 10 | You appear on the third-party sources each engine actually favours, which differ between them |
| Freshness signals | 5 | Visible, accurate dates on anything time-sensitive |
Those weights sum to 100. Score each engine separately, because the answer will not be the same for all three.
The live test, and it costs nothing. Take 30 real commercial questions from your category, in your buyers’ words. Four worked examples: “how do we get cited by ChatGPT in our category”, “best tools for tracking brand mentions in AI answers”, “which vendor should we use for AI search visibility”, “how do we optimise a website for Perplexity”. Run every one on ChatGPT, on Perplexity and on Google AI Mode, more than once each, because engines are probabilistic and a single run establishes nothing.
Then record five things per engine, separately. Where do we appear and where do we not. Which competitors appear instead, consistently. Which sources influence those answers. Why are those sources winning. What would you change in the next 90 days, and who does it. Do not average across engines. The whole point of this page is that the answers differ, and averaging destroys the finding.
The weak playbook against the strong one. The weaker sequence is to add FAQ blocks and schema to a few pages, publish more articles, and call it AI optimisation. It is entirely on-site, entirely owned media, and aimed at the smallest share of the citation pool. It also skips the access check, which means it can be done perfectly on a site no engine can reach.
The stronger sequence runs in order: verify crawler access per engine, fix rendering so the content exists for a crawler, restructure key pages so a passage lifts cleanly, then build presence on the third-party sources each engine favours. Access first, because everything downstream is worthless without it.
Red flags, each one something you will genuinely hear. “We blocked the AI crawlers to protect our content”, usually meaning GPTBot, and usually without knowing OAI-SearchBot is separate. “We guarantee citations in ChatGPT.” “Add schema and you will get cited.” “Our AI visibility score went up”, offered with no competitor share of voice beside it. “We optimise for ChatGPT”, where one engine stands in for the category. “AI search is completely different from SEO”, contradicted by Google’s own published position. And the quiet one: a proposal that never mentions robots.txt at all.
Five questions for your own team.
- Which AI crawlers does our robots.txt currently allow? A good answer names the six tokens. A weak one is that we allow everything, unverified.
- Does our main template render for a crawler without JavaScript? A good answer is a rendered-versus-raw comparison. A weak one is that the site works fine in a browser.
- Which competitors appear where we do not, per engine? A good answer is three separate lists. A weak one is a single average.
- Which third-party sources feed answers in our category? A good answer is a frequency-sorted list of domains. A weak one is a guess.
- Who fixes the rendering issues, and is that time booked? A good answer names a person and a sprint.
It all comes down to one principle: check access first, then make one passage per page liftable, then go and get cited somewhere you do not own. Everything else in this discipline is detail.
One closing note that costs us something. Very few providers are equally strong across measurement, execution, earned authority and attribution, ours included, and the good ones will tell you which of the four is their weakest. If your robots.txt turns out to be the whole problem, fix it and spend nothing else this quarter.
Where Pepper fits
Pepper is an agentic organic growth engine, and the reason it is built as a platform plus a team is that this work splits across both.
- What you log into: workspace setup, brand profile, competitors and personas. GA4 and Search Console connected. Themes and prompts managed, GEO analytics read directly, and your own agents built and run in the Agent Atlas
- Engines tracked: 6, including ChatGPT, Perplexity, Gemini and Google AI Overviews, reported separately rather than averaged
- What the team does: a growth team is attached to the account and works alongside your people on the access and rendering fixes, the content restructuring, and the earned media that no software performs
- Metrics reported: Brand Visibility, Domain Prompt Presence and Share of Voice, with the gap between the first two as the citability diagnostic
- Best for: teams running organic as a long-term function who need the per-engine work done rather than only measured
- Where it falls short: we do not publish pricing, so you cannot size it before a call. A one-off audit is not what we are for, and our depth sits in content, authority and AI search rather than large technical migrations
Frequently asked questions
How do I optimize my website for ChatGPT?
Allow OAI-SearchBot in robots.txt first, since blocking it removes you from ChatGPT answers entirely. Then focus depth on your fifteen most commercially important questions, because ChatGPT cites fewer sources per answer than other engines.
Is GPTBot the same as OAI-SearchBot?
No, and confusing them is the most common fault we find. GPTBot collects training data for foundation models. OAI-SearchBot decides whether you appear in ChatGPT’s search answers. Blocking GPTBot costs you nothing in visibility.
How do I optimize for Perplexity?
Allow PerplexityBot, which is used for search results and not for training. Then cover the full question set around a topic rather than perfecting one page, because Perplexity cites more sources per answer with less weight on each.
Does blocking Google-Extended hurt my rankings?
No. Google states it “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.” It only controls whether your content trains Gemini models.
Do I need a separate strategy for each AI engine?
Partly. Crawler access and citation behaviour genuinely differ, so those need engine-specific attention. Retrievability, extractable structure and off-site authority are shared, and they account for most of the work.
Why do ChatGPT and Perplexity show different results for the same question?
Because they retrieve differently. Independent research across 11,000 queries found identical questions produce structurally different results between systems, with different source preferences and different numbers of citations per answer.
Does schema markup get me cited in AI answers?
It helps machines read your meaning, and it is consistently oversold. It will not make a weak page strong. Access, rendering and extractable structure matter considerably more.
Can Perplexity-User ignore my robots.txt?
Yes. Perplexity documents that Perplexity-User generally ignores robots.txt because the fetch was triggered by a person asking a question, rather than by automated crawling.
Where to go next
Open your robots.txt and search for six tokens: OAI-SearchBot, GPTBot, PerplexityBot, Perplexity-User, Googlebot, Google-Extended. It takes ten minutes and it is the highest-return check in this entire discipline.
Then run 30 commercial prompts across all three engines separately, and keep the results apart rather than averaging them.
Our guide to tracking brand mentions in AI search covers the five methods and what each misses, and the Visibility, Citability and Retrievability framework explains which lever is stuck. For the technical half, can AI search bots crawl my website goes deeper on access.
If you are choosing tooling, the best answer engine optimization tools covers which have free tiers, and which platform covers the most AI engines covers coverage. Our SalesHood case study documents AI Overview keywords moving from 14 to 97 across five months.
The honest exit. If your robots.txt already allows all three search crawlers and your pages render server-side, you may have no engine-specific work to do at all. Spend the quarter on earned media instead, which is where most of the citation pool sits.
Sources
Every crawler token was checked at that company’s own documentation on 10 September 2026. Every study cited was published in 2026 and checked at the original source.
- OpenAI, Bots and crawlers documentation, checked 10 September 2026.
OAI-SearchBotsurfaces sites in ChatGPT search features;GPTBotcrawls content that may be used for training;ChatGPT-Userhandles user-initiated actions;OAI-AdsBotvalidates ad landing pages. - Perplexity, Crawler documentation, checked 10 September 2026.
PerplexityBotsurfaces and links websites in results and is not used for training;Perplexity-Usersupports user-initiated requests and generally ignores robots.txt. IP ranges published as JSON for both. - Google Search Central, Google common crawlers, checked 10 September 2026.
Google-Extendedcontrols whether content trains Gemini models and “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.” - Google Search Central, AI features and your website, 15 May 2026. States AEO and GEO are part of SEO, that AI Overviews and AI Mode use core Search ranking systems with no separate index, and advises against providers guaranteeing rankings.
- Huang, Goyal, Saha and Chandrasekharan, Answer Bubbles: Information Exposure in AI-Mediated Search, arXiv, 17 March 2026. 11,000 real search queries across five systems. Academic and independent.
- Zhang Kai, He Xinyue and Yao Jingang, From Citation Selection to Citation Absorption, arXiv, 28 April 2026. 602 prompts, 21,143 citations, 18,151 fetched pages across ChatGPT, Google AI Overview and Perplexity. Academic and independent.
- Suzgun et al., Evaluating Commercial AI Chatbots as News Intermediaries, arXiv, 21 May 2026. 2,100 questions across six chatbots, on retrieval failure as a cause of answer errors. Tested on news rather than commercial queries.
- Muck Rack, Earned media still drives 84% of AI citations, 7 May 2026. More than 25 million cited links, 17 industries. Muck Rack sells PR software, so read it as an interested party with a large dataset.
Latest Blogs
Every list of AEO agencies for B2B tells you who is good. None of them tells you who will survive your procurement process. So we audited what ten agencies publish about themselves against the five things an enterprise buyer has to produce internally before signing. All ten name enterprise clients. Two publish a price. None publishes team size, contract terms, or anything a security review would accept.
Almost no agency publishes what a standalone SEO audit costs. Two do, and they differ by 3.4 times on price and nine times on turnaround while covering a similar number of pages. So we computed the metric nobody publishes, cost per page audited, and set out what a purchased audit has to contain before it is worth buying at any price.
Nobody outside a platform can measure its data accuracy without running a controlled test, and nobody in this category publishes one. So we ranked 14 platforms on the thing that actually decides whether their numbers can be accurate: how many times each prompt is sampled, computed from each vendor’s own published allowances. Only five publish enough to work it out, and the spread between them is thirty-fold.