SEO for retail and ecommerce: how to pick an agency that understands product search

The short answer
Retail SEO is a crawl and template problem before it is a content problem, and the reason is arithmetic. A catalogue with 20,000 products and eight filterable attributes can generate millions of URLs, almost none of which should be indexed.
So the agency you want is one that leads with crawl allocation, template architecture and rendering, and treats product copy as the thing you do after those are fixed. Ask about facets first.
Key takeaways
- Google names your exact situation. Crawl budget needs active management above one million pages changing weekly, or ten thousand changing daily. Most mid-sized catalogues clear both.
- “Perceived inventory” is Google’s own term, and it is the retail trap. Google decides how much to crawl partly on how much content it thinks you have, and facets inflate that enormously.
- Two popular fixes do not work. Google states plainly that
noindexdoes not save crawl budget and robots.txt does not reallocate it. - Product pages are templates, not pages. One flaw is a hundred thousand flawed URLs, and page-level auditing will never find it.
- AI shopping surfaces are new and mostly off your site. Review sites, marketplaces and communities feed them more than your product pages do.
A note on where this comes from. We run organic for more than 250 enterprises and track over 10 million prompts across every major engine. What follows is shaped by that, by the client reviews we sit in weekly, and by the retail accounts we inherit where nobody had ever pulled a log file.
What is retail SEO, and what makes it different?
Retail and ecommerce SEO is the practice of making a product catalogue discoverable in search. The discipline is the same as any other SEO. The scale, and what that scale does to crawling, is not.
Three things genuinely differ.
Catalogues generate URLs by default. A content site publishes pages deliberately. A retail site generates them as a side effect of filters, sorts, pagination, sizes, colours and search parameters. Nobody decided to publish those URLs, and Google still has to decide what to do with them.
Google’s crawl guidance names the thresholds. Its large-site docs name three cases. Sites of one million or more pages that change about weekly. Sites of ten thousand or more that change daily. And any site with a large share of URLs sitting in “Discovered, currently not indexed.” A catalogue of moderate size with active pricing clears the second condition easily.
Demand is partly your own doing. Google states that crawl demand depends on perceived inventory, popularity and staleness, and it flags perceived inventory as the one you control. If Google perceives three million URLs where you have twenty thousand products, the crawl spreads across the wrong things.
Our enterprise SEO guide covers what else changes above 10,000 pages, and enterprise SEO audit framework sets out the 47 technical checks we run.
What a retail SEO agency has to cover, at a glance
| Area | What it actually involves | Typical engagement cost | Where most agencies are weak |
|---|---|---|---|
| Crawl allocation | Facets, parameters, pagination, log file analysis | Included in a technical retainer | The single most common gap. Needs log access most agencies never request |
| Template architecture | Category, product, brand and search templates audited as patterns | Included | Often done page by page, which never finds the pattern |
| Rendering | Whether product data exists before JavaScript runs | Engineering time, usually quoted separately | Frequently assumed rather than tested |
| Product and category content | Copy, specifications, unique category text | Per-page or per-batch pricing | Where most agencies start, and where most retainers are spent |
| Reviews and user content | Structured, indexable, kept fresh | Often out of scope | Undersold, and it is real differentiation at scale |
| Off-site and AI shopping | Marketplaces, review sites, communities feeding AI answers | Digital PR line, often absent | The newest gap and the largest |
Costs vary enormously by catalogue size, so treat that column as shape rather than benchmark. What matters is which rows a proposal actually contains.

If you want your own catalogue assessed against these rows before you brief anyone, book a growth audit and we will run the crawl and template analysis with your team.
The facet problem, in numbers
This is the mechanic worth grasp before any sales call, because it is what separates a retail site from a large content site.
Take a catalogue of 20,000 products. Add filters for size, colour, brand, price band, material, availability, rating and sort order. Each mix is a URL. The combinatorial maths runs into the millions, and every one of those URLs is a page Google may discover, request and evaluate.
Almost none of them should be indexed. Many of them should not be crawled at all. And the two fixes most commonly proposed do not do what people think:
noindexdoes not save crawl budget. Google states it will still request the page, then drop it when it sees the directive. The crawl is already spent- robots.txt does not reallocate crawl budget. Google says it will not shift that capacity elsewhere unless you are already hitting your limit
- What does work is not generating the URLs in the first place, handling parameters deliberately, and keeping server responses fast, since Google ties crawl capacity to latency stability and the absence of 5xx errors
An agency that cannot explain this difference will propose noindexing your facets, which is the most common wrong answer in retail SEO and sounds entirely sensible.

What a retail SEO engagement costs
Almost nobody publishes pricing in this category, ours included, and retail varies more than most because catalogue size and template count drive the work rather than page count.
What actually moves the number is not on any rate card. Template count matters more than product count: a 500,000-URL catalogue built from 20 templates is a smaller job than a 50,000-URL one built from 200. Log file access decides whether crawl behaviour gets measured or guessed, and arranging it is usually the longest lead time in the engagement. Engineering capacity decides whether findings ship at all, and it is the variable nobody quotes on.
One useful benchmark exists. WebFX publishes its pricing, with an initial two-month campaign from $5,900 to $12,200 and ongoing monthly work from $2,900 to $9,200. Treat widely quoted hourly and retainer averages with care: the most-cited survey behind them was last updated in August 2024 and is routinely republished as current.
Our cost breakdown of platforms against hiring models the build-versus-buy side, and SEO audit agencies covers what a diagnostic should cost before a retainer starts.
What changed with AI shopping

Product questions increasingly get answered before anyone reaches a product page, and the answer is assembled from sources you mostly do not control.
Muck Rack analysed more than 25 million cited links across ChatGPT, Claude and Gemini in May 2026 and found earned media accounted for 84% of AI citations against 0.3% for paid and advertorial. Muck Rack sells PR software, so read it as an interested party with a large dataset.
Independent academic work points the same way. Huang and colleagues, across 11,000 real queries, found reference and long-form third-party sources disproportionately overrepresented in AI answers. For retail specifically that means review sites, comparison sites, marketplaces and communities, not your product detail pages.
What this changes about agency selection. A retail agency whose entire scope ends at your domain is addressing the part of product search that is shrinking. Ask which review sites, marketplaces and communities are in scope, and who does the outreach.
Our engine-by-engine guide covers the crawler access side, and how to run a GEO audit covers the five-stage diagnostic.
Where each type of agency falls short
The technical specialist. Strong on crawl, logs, rendering and templates. Where it falls short: content at catalogue scale is usually subcontracted or thin, and off-site work is often absent entirely.
The content and merchandising agency. Strong on category copy, product descriptions and merchandising. Where it falls short: they rarely request log files, which means crawl waste goes undiagnosed and the content sits on templates Google is not reaching.
The full-service ecommerce agency. Covers more rows on the table. Where it falls short: depth varies sharply between rows, and the specialists on your account are rarely the ones on the website. Ask who actually works on it.
The AI search specialist. Strong on the newest and fastest-growing gap. Where it falls short: usually no capacity for the technical crawl work that still decides whether your products can be found at all.
And Pepper. We are an agentic organic growth engine. Customers log into the platform themselves. Workspace setup, brand profile, competitors and personas. GA4 and Search Console connected. Themes and prompts managed, GEO analytics read directly, and their own agents built and run in the Agent Atlas. And a growth team is attached to the account doing the work alongside them. Where it falls short: we do not publish pricing, so you cannot size it before a call, and our depth sits in content, authority and AI search rather than large technical migrations. If a replatform is your immediate problem, say so early and we will tell you whether we are the right fit.
Our FMCG and CPG and consumer tech pages cover the verticals closest to retail.
What nobody should promise you
Guaranteed rankings for product terms. Google advises against providers guaranteeing rankings, because third parties cannot access internal ranking systems.
That noindexing facets fixes crawl budget. Google explicitly says it does not. The request still happens.
A content programme as the answer to a crawl problem. Publishing more category copy on a site where Google is spending its crawl on filter mixes changes very little.
That product page optimisation covers AI shopping. Most of what feeds those answers sits on third-party sites.
How to choose an SEO agency for retail and ecommerce
I would run the whole review from one question, asked in the first ten minutes, because the answer sorts the market faster than a skills deck.
“How would you handle our faceted navigation, and what would you need from us to answer that properly?”
A strong answer asks how many facets you have, which mixes generate URLs, whether pagination is crawlable, and whether server log access is available. A weak answer recommends noindexing them, which Google states does not save crawl budget.
Google’s guidance is the anchor throughout. Crawl budget needs managing above one million pages changing weekly. Or ten thousand changing daily. Or wherever a large share of URLs sit in “Discovered, currently not indexed.” Crawl capacity depends on crawl health, meaning server response consistency and the absence of 5xx errors. Crawl demand depends on perceived inventory, popularity and staleness, and perceived inventory is the one you control. A proposal that never mentions any of this is not addressing retail’s actual constraint.
Score the proposal before you sign.

| Area | Weight | What a strong proposal demonstrates |
|---|---|---|
| Crawl and template capability | 35 | Server logs requested as a prerequisite, findings scoped to templates with URL counts, and facet handling explained without prompting |
| Engineering integration | 25 | Developer-ready tickets with acceptance criteria, sized to fit your release cycle, not a PDF and a handover call |
| Content at catalogue scale | 20 | A method for producing unique category and product content at volume, with a quality floor below which a page is not published |
| Off-site and AI shopping | 15 | Named review sites, marketplaces and communities in scope, with someone doing the outreach |
| Revenue-linked measurement | 5 | Reporting by template tied to revenue, not keyword counts or a visibility score |
Those weights sum to 100. Score your own readiness too, because an agency that produces a template remediation plan needs engineering time you may not have booked.
The live test
Send three URLs: your highest-revenue category template, your highest-volume product template, and one URL currently sitting in “Discovered, currently not indexed.”
Ask four things in writing. What is the rendered versus raw HTML difference on these three. Which templates do they belong to and how many URLs share each. What would you need from our systems team in week one. What would change in the next 90 days, and who does it.
Then add the AI half, because product questions increasingly get answered before the click. Give them 30 real commercial questions a shopper would type. Four worked examples: “best running shoes for flat feet under 100 dollars”, “is brand X or brand Y better for sensitive skin”, “which laptop has the best battery life for students”, “what is the difference between these two models”. Ask which sources currently answer those, and how many they could plausibly reach this quarter.
The weak proposal against the strong one
The weaker one opens with a keyword gap analysis, recommends category page copy and product description rewrites, adds schema, and prices per page. It never requests logs, so crawl waste stays invisible, and the content lands on templates Google is under-crawling.
The stronger one opens by asking about facets and logs, maps templates with URL counts, sequences fixes so dependencies resolve in order, and puts content after the templates are reachable. It produces fewer findings and more shipped change.
Red flags you will genuinely hear
“We’ll noindex the faceted URLs to save crawl budget”, which Google says does not work. “We can start without log file access”, on a catalogue of any size. “Here are 4,200 issues”, offered as thoroughness rather than a prioritisation failure. “Guaranteed first page for your head terms.” “We’ll write unique descriptions for all 20,000 products”, quoted without a quality floor. “Core Web Vitals is the priority”, on a site where a third of products are not indexed. And the quiet one: a scope with no line for reviews, marketplaces or third-party sources.
Five questions to ask
- How would you handle our faceted navigation? A good answer asks questions back. A weak one recommends noindex.
- Do you need server logs, and what happens if we cannot get them? A good answer treats logs as a prerequisite and says what confidence is lost without them.
- Will findings be scoped to templates with URL counts? A good answer describes template mapping. A weak one promises a full page list.
- Who writes the tickets, and do they fit our release process? A good answer has seen a sprint board.
- Which third-party sources are in scope, and who does the outreach? A good answer names sites and people. A weak one calls it content amplification.
It all comes down to one principle: in retail, fix what gets crawled before you improve what gets read. An agency that leads with product copy on a site drowning in filter URLs is solving the wrong problem competently.
One closing note that costs us something. Very few providers are equally strong across measurement, execution, earned authority and attribution, ours included, and the good ones will tell you which of the four is their weakest. Our depth sits in content, authority and AI search rather than large technical migrations, so if your catalogue needs a replatform before anything else, a technical specialist is the better first call.
Frequently asked questions
The basics
What makes retail SEO different from regular SEO?
Scale, and what scale does to crawling. Catalogues generate URLs by default through filters, sorts and pagination, so a 20,000-product site can produce millions of URLs nobody decided to publish. Managing what gets crawled becomes the primary constraint.
How do I know if my site has a crawl budget problem?
Check the share of URLs in “Discovered, currently not indexed” in Search Console. Google names that as a trigger at any site size, alongside catalogues above one million pages changing weekly or ten thousand changing daily.
Should I noindex my faceted navigation?
It will not save crawl budget. Google states it still requests the page and then drops it when it sees the directive, so the crawl is already spent. Handling parameters deliberately and not generating the URLs works better.
What should a retail SEO agency ask for first?
Server log files. It is the only source showing what Googlebot actually requested rather than what a crawler could find, and an agency that does not ask for it is inferring your crawl behaviour rather than measuring it.
Choosing and buying
How much does a retail SEO agency cost?
It varies more than most categories, because catalogue size and template count drive the work. Almost nobody publishes pricing, ours included, so plan the review as identically briefed calls rather than a price comparison.
Does AI search matter for ecommerce yet?
Increasingly, and mostly off your own site. Muck Rack found 84% of AI citations come from earned media, and independent research found third-party reference sources overrepresented in answers. For retail that means review sites, marketplaces and communities.
Can one agency cover technical and content at catalogue scale?
Some can, but depth varies sharply between the two. Ask who exactly works on your account rather than which skills appear on the website, and expect one of the two to be stronger.
What is the most common finding on a retail site?
Crawl spent on filter and parameter URLs while revenue templates sit under-crawled. It is invisible in a standard crawl and obvious in server logs, which is why log access matters so much.
Where to go next
Pull your Search Console Page Indexing report. Check one number: the share of URLs in “Discovered, currently not indexed.” Google names it as a crawl budget trigger at any size, and it takes two minutes.
Then count your facets. Multiply the mixes. If the number is far larger than your product count, you have found the constraint before anyone has quoted you.
Our enterprise SEO audit framework sets out the 47 technical checks, and SEO audit agencies covers what a real audit includes and what to pay. For provider selection more broadly, how to choose an organic SEO agency covers the red flags worth walking away from.
For the AI half, how to run a GEO audit covers the five-stage diagnostic and 18 GEO best practices covers what the research says actually gets cited. Our TVS Eurogrip case study shows the work on a consumer brand recovering from a migration.
The honest exit. If your catalogue is under a few thousand products and your facets do not generate crawlable URLs, you do not need a retail specialist. A competent generalist SEO agency will cover you, and you should spend the difference on reviews and third-party presence instead.
Sources
Google’s guidance is quoted from Google’s own docs, checked on 11 September 2026. Every study cited was published in 2026 and checked at the original source.
- Google Search Central, Large site owner’s guide to managing your crawl budget, checked 11 September 2026. Names the three trigger conditions, states crawl capacity depends on crawl health and Google’s resources, states crawl demand depends on perceived inventory, popularity and staleness, and states that
noindexdoes not save crawl budget and robots.txt does not reallocate it. - Google Search Central, AI features and your website, 15 May 2026. States AEO and GEO are part of SEO, and advises against providers guaranteeing rankings.
- WebFX SEO pricing, checked 11 September 2026: initial two-month campaign $5,900 to $12,200; ongoing monthly $2,900 to $9,200; enterprise quoted. One of very few published price lists in the category.
- Ahrefs, SEO pricing survey, 439 service providers, last updated 15 August 2024. Cited only to note that its figures circulate widely as current 2026 rates.
- Muck Rack, Earned media still drives 84% of AI citations, 7 May 2026. More than 25 million cited links across ChatGPT, Claude and Gemini, 17 industries. Muck Rack sells PR software, so read it as an interested party with a large dataset.
- Huang, Goyal, Saha and Chandrasekharan, Answer Bubbles: Information Exposure in AI-Mediated Search, arXiv, 17 March 2026. 11,000 real search queries across five systems. Academic and independent.
- Zhang Kai, He Xinyue and Yao Jingang, From Citation Selection to Citation Absorption, arXiv, 28 April 2026. 602 prompts, 21,143 citations across ChatGPT, Google AI Overview and Perplexity. Academic and independent.
The facet arithmetic and the weightings are Pepper’s own framework, based on the retail accounts we run. They are a working method rather than an industry standard.
Latest Blogs
Your rankings did not change. The page they sit on did. Across 53 brands and 2.43 billion impressions, click-through on queries with an AI Overview fell while click-through on queries without one rose. Here is how to tell which problem you actually have, in twenty minutes.
Most SEO reports are built to look like progress rather than to be read. Twelve metrics worth your attention, eight that are there to fill space, and the four questions that tell you which kind of report you are holding.
One of these terms comes from a peer-reviewed paper with a formal definition and a public benchmark. The other has no canonical source at all. We traced both, and the answer to whether they are the same thing is more useful than a definitions table.
Get your hands on the latest news!
Similar Posts

Digital Marketing
15 mins read
How to improve organic search rankings in 2026

Digital Marketing
15 mins read
What a modern SEO team looks like in 2026

Content
12 mins read