Answer engine optimization services in 2026: scope, deliverables and how to buy

If you have read three AEO proposals, you have noticed they are converging on the same deliverable list. Priority query selection, a citation baseline, an entity page, schema rollout, listicle placements, directory work, community seeding, an llms.txt file, weekly tracking, a quarterly benchmark.
That convergence is worth pausing on for two reasons. The first is that a standard scope makes proposals genuinely comparable, which is good for you. The second is that at least one item on that standard list is something Google has publicly said does not work.
We have run organic for more than 250 enterprises across eight years, and we sell services in this category, so read what follows with that in mind. We have tried to write the guide we would want if we were buying rather than selling.
What AEO services actually cover
The work splits into four layers. Most proposals cover the first two well, the third badly, and the fourth barely at all.

Technical retrievability comes first because nothing works without it. This covers rendering, crawler access, site architecture and schema. It matters more than it used to: a Stanford-led evaluation of six commercial chatbots, covering 2,100 factual questions in May 2026, found that retrieval failures rather than reasoning errors caused more than 70% of all mistakes. Being fetchable is not hygiene any more, it is the main determinant of whether you appear at all.
The specifics worth asking about are crawler access in particular. OpenAI runs OAI-SearchBot, ChatGPT-User and GPTBot as three distinct agents with different jobs, and Perplexity runs PerplexityBot and Perplexity-User for scheduled indexing and live fetching. Blocking each one fails differently. A robots.txt decision made two years ago by somebody who has since left is a common and entirely fixable cause of invisibility.
Content structure is the layer everyone sells confidently. Answer-led openings, standalone headings, claims kept next to their evidence, comparison tables, FAQ blocks matched to real query phrasing. This work is real and it is also the easiest to do badly at volume, because the temptation is to reformat everything rather than to improve the pages that matter.
Earned authority is where most proposals go quiet, and where most of the outcome actually lives. Muck Rack analysed more than 25 million cited links across ChatGPT, Claude and Gemini in May 2026, spanning 17 industries. Earned media accounted for 84% of AI citations. Paid and advertorial content accounted for 0.3%.
Read that against a typical proposal. If 84% of what decides your category’s answers is published by other people, and your proposal is 80% on-site content work, the shape is wrong regardless of how good the content is.
Measurement is the layer that determines whether you can tell if any of it worked. It should cover per-engine tracking, a stated prompt selection methodology, competitor share of voice, and raw answers you can inspect rather than a score.
What it costs
Pricing in this category has become more transparent over the past year, and the published ranges cluster in a way that makes budgeting possible.

Entry-level tooling and light-touch service starts in the low hundreds monthly. Full retainers for mid-market companies typically run in the mid four figures to around ten thousand a month. Enterprise engagements start around ten thousand and go up, often well up, depending on how many markets and languages are in scope. Project-based audits and one-off strategy engagements generally land in the five figures.
Four things move the number more than anything else. Site size and existing technical debt come first, because a large site with a decade of accumulated redirects is a different job. The number of engines and markets tracked comes second, since coverage scales cost almost linearly. Whether earned media is genuinely in scope comes third, and it is the most commonly cut line despite being where most citations live. And whether implementation is included, or handed back to your team, can double or halve the real cost.
Pepper does not publish pricing, so we will not pretend otherwise. What we will say is that the ratio matters more than the total. If you are spending more on being told what to do than on doing it, the shape of the engagement is wrong.
The deliverables worth challenging
This is the section we would most want if we were buying, so it gets the most space.

llms.txt. Google’s official AI search guidance, published 15 May 2026, explicitly mythbusts llms.txt, alongside content chunking, AI-specific rewrites and structured-data over-optimisation. It appears on standard AEO deliverable lists anyway. It costs almost nothing to implement, so it is not worth fighting about, but it should not be presented as foundational and it should certainly not be a headline line item.
Content chunking and AI-specific rewrites. Also on Google’s mythbusted list. If a proposal includes rewriting your existing content into AI-optimised formats as a major workstream, ask what evidence supports it beyond the vendor’s own case studies. Structural improvement to genuinely weak pages is worth doing. Reformatting good pages is expensive theatre.
Branded prompt tracking. A visibility number built mostly from prompts containing your brand name will always look healthy, because you are asking the engine about you. Ask specifically what proportion of the tracked prompt set is unbranded commercial language, and expect most of it to be.
Citation counts with no competitor context. A rising citation count means very little without share of voice. If a competitor grew faster in the same period, you lost ground while your report showed green.
Single-engine coverage. Muck Rack’s data shows ChatGPT includes sources in 96% of responses, Gemini in 82%, and Claude in 55%. Coverage of one engine describes one engine’s behaviour, and a blended average across them describes none.
Guaranteed citations. This one ends the conversation. Google explicitly advises against providers guaranteeing rankings, because third parties cannot access internal ranking systems. Nobody has that access for ChatGPT either.
How the engagement models differ
Three shapes dominate, and they suit genuinely different situations.

An audit or one-off project gives you a diagnosis and a prioritised plan. It suits teams with real internal capacity who need direction rather than hands. The failure mode is well documented: the audit surfaces two hundred fixes, nobody has bandwidth to ship them, and it becomes a backlog rather than a capability.
A monthly retainer buys ongoing work. It suits organisations that need delivery rather than advice, and it is the most common shape. The question to press is what proportion of the retainer is execution versus reporting, because the ratio varies enormously between vendors charging similar amounts.
A platform plus team model gives you the measurement infrastructure and people running it together. This is what Pepper does, so treat the description as interested. Customers log in and run it themselves, setting up a workspace, defining brand profile, competitors and personas, connecting GA4 and Search Console, managing themes and prompts, and building their own agents in the Agent Atlas. Alongside that, a growth team is attached and does the work with them.
It suits teams that want organic run as a continuous function rather than a project. Where it falls short: we do not publish pricing, so budget discovery is a conversation rather than a page. And if you want a self-serve monitoring tool and nothing else, a dedicated tracker will serve you better and cost less.
How to run the procurement
Five moves, in order, that will tell you more than any capabilities deck.
Ask for the prompt set before you ask for anything else. Give them 20 to 30 real commercial questions from your category and ask them to run a mini audit. Then ask for five specific things back: where you appear today, which competitors appear instead, which sources are influencing those answers, why those sources are winning, and what they would change in the next 90 days.
That conversation separates firms that have done this work from firms that have read about it, usually within ten minutes.
Ask what proportion of the scope is earned media. Given that 84% of citations come from earned sources, a proposal weighted almost entirely to on-site work is structurally mismatched to the outcome. There may be good reasons for it in your specific case. Make them say what they are.
Ask how prompts are selected. If a vendor cannot explain a methodology, your visibility number is whatever the prompt list flatters. This single question filters more proposals than any other.
Ask to see raw answers, not a score. You want the actual response text and the cited URLs. A tool that only surfaces a composite number cannot be audited, and neither can the agency using it.
Ask what they would not do. A serious partner prioritises ruthlessly and will tell you which parts of your site are not worth optimising. Anyone who wants to fix everything has not thought about cost.
If you would rather run the diagnostic yourself before briefing anyone, our 7-point AI search audit is the checklist we use on every new account. And if you want it run on your category, book a growth audit and we will show you which sources are deciding your answers today.
What nobody should promise you
Guaranteed citations or positions in AI answers. Nobody controls the ranking systems.
A single AEO score that means something. Composite numbers hide the gap between being mentioned and being cited, which is the most actionable finding available.
Results in one quarter. Technical fixes land faster, but earned authority compounds slowly because it depends on other people publishing about you.
Complete coverage of every engine. Coverage varies by tier and changes monthly. Verify at the tier you are buying, not the tier on the marketing page.
That the service will fix your capacity problem. Measurement is not execution, and a diagnosis nobody can act on is an expensive document.
Frequently asked questions
What do answer engine optimization services include?
Four layers: technical retrievability covering rendering and crawler access, content structure work, earned authority through digital PR and third-party sources, and per-engine measurement. Most proposals cover the first two well and thin out on earned media.
How much do AEO services cost?
Entry-level tooling starts in the low hundreds monthly. Mid-market retainers typically run into the mid four figures to around ten thousand. Enterprise engagements start around ten thousand and rise with markets and languages. One-off audits generally land in the five figures.
Is llms.txt a legitimate AEO deliverable?
Google’s May 2026 guidance explicitly mythbusts llms.txt, along with content chunking and AI-specific rewrites. It costs little to implement, so it is not worth arguing about, but it should not appear as a foundational line item in a proposal.
What is the difference between AEO and GEO services?
Very little in practice. Wikipedia records that no consensus definition separates AEO, GEO, LLMO and AIO in academic literature. Focus on which four layers a proposal actually covers rather than on which acronym the vendor prefers.
How do I know if an AEO agency is any good?
Give them 20 to 30 real commercial prompts and ask for a mini audit showing where you appear, which competitors appear instead, which sources influence those answers and what they would change in 90 days. That reveals more than any deck.
Should AEO be a retainer or a project?
A project suits teams with internal capacity who need direction. A retainer suits teams needing delivery. The question to press on any retainer is what proportion is execution versus reporting, because that ratio varies widely at similar price points.
How long before AEO services show results?
Technical retrievability fixes can show within weeks. Content structure work takes a quarter or so. Earned authority takes two to three quarters, because it depends on other publishers, and no vendor controls that timeline.
Do I need AEO services if I already do SEO?
Possibly not. Google states AI Overviews run on core Search ranking systems, so strong SEO carries over. Buy AEO services specifically when engines mention your brand but do not cite your pages, which signals a citability gap that content alone will not close.
Where to go next
Before you brief anyone, establish which problem you actually have. Run a real prompt set and compare two numbers: how often engines mention your brand, and how often they cite a page from your domain.
If mentions are healthy and citations are not, you have a citability problem and your proposal should be weighted toward earned media. If both are low, you have a visibility problem and content carries more of the load. If your pages are not being retrieved at all, you have a technical problem and no amount of either will help.
That diagnosis should determine the shape of what you buy, and it takes an afternoon.
For the underlying discipline, our guide to answer engine optimization covers the mechanics, and the Visibility, Citability and Retrievability framework covers the three levers and which work moves each. For measurement specifically, what actually matters in AI search measurement covers the metric set worth reporting.
To see where you stand across engines, see where you show up in ChatGPT, Perplexity, Gemini, Claude and AI Overviews. Customers run this themselves in the platform, with a growth team attached. For what compounding looks like, our SalesHood case study documents AI Overview keywords rising from 14 to 97 and rich results growing 291% in five months.
One honest exit. If your category produces very few meaningful monthly prompts, or nobody internally can act on findings, you do not need an AEO retainer yet. Spend the money on earned media directly and revisit in two quarters. We would rather tell you that than sell you a reporting subscription.
Sources
Every study cited here was published in 2026. Pricing ranges reflect published rates across vendors in this category and should be verified at the tier you are buying.
- Google Search Central, AI features and your website, 15 May 2026, including the mythbusting of llms.txt, content chunking, AI-specific rewrites and structured-data over-optimisation, plus the advice against guaranteed rankings.
- Muck Rack, Earned media still drives 84% of AI citations, 7 May 2026. More than 25 million cited links across ChatGPT, Claude and Gemini, 17 industries. Muck Rack sells PR software, so read it as an interested party with a large dataset.
- Suzgun et al., Evaluating Commercial AI Chatbots as News Intermediaries, arXiv, 21 May 2026. 2,100 factual questions across six chatbots, February 2026.
- Wikipedia, Generative engine optimization, on the absence of a consensus academic definition separating AEO, GEO, LLMO and AIO as of early 2026.
- OpenAI crawler documentation and Perplexity bot documentation.
Latest Blogs
TL;DR: This is the most ambitious product we have ever launched. For three editions I have shown you where to find the questions you are losing, the sources AI trusts, and the repetitive work that quietly decides your results. Every one of them still ended the same way, with you going off to do something. […]
Developers stopped reading marketing content years ago and now ask a model instead. That changes what a content partner has to be good at: your documentation is the asset, technical accuracy is the ranking factor, and most agencies cannot write a working code sample. Here is our read on the nine worth considering.
Most SEO guides still teach a pre-AI version of the discipline. This one covers the five pillars in the order that actually matters, what Google’s May 2026 guidance says about AI search, and the three numbers worth measuring once engines start answering for you.