GEO / AI Search

GEO pipeline: does generative engine optimisation actually drive deals?

janvi
•
Posted on 24/09/26•17 min read
GEO pipeline: does generative engine optimisation actually drive deals?

The short answer

Yes, and almost none of it will appear in your pipeline report. The demand-side evidence is strong. 82% of B2B software buyers have sourced recommendations from an AI chatbot, and about half say the biggest effect lands while they are narrowing and comparing. That stage happens inside the answer, so it produces no referral. The attributed-pipeline evidence is weak and contradictory. Published multiples run from worse than organic to five times better.

Key takeaways

  • The demand evidence is good and primary. More than 1,000 B2B software buyers and decision-makers were surveyed, with interviews from more than 50 sales and marketing leaders. 82% had sourced software recommendations from an AI chatbot in the last 24 months.
  • The same research says where it bites. About half of those buyers said AI mattered most when narrowing and comparing options, which is the deciding stage of a B2B purchase.
  • That same research contains no data on whether those buyers convert. The best-sampled study on this subject does not answer the pipeline question, and nobody pretends otherwise.
  • The pipeline claims in circulation are unreliable. Published figures run from 22% better than organic to five times better. The only peer-reviewed study, on e-commerce, found AI traffic converting below every traditional channel except paid social.
  • Almost none of those claims publishes a sample size. One widely repeated figure traces to a vendor’s first-party data with no method. Unsourced multiples circulate on our own site too, and we corrected one yesterday.
  • The lag explains the missing evidence. Citation to traffic runs six to nine weeks. Citation to pipeline runs three to five quarters on enterprise cycles. AI search only became material during 2025, so the first clean cohorts are landing now.
  • Pepper is an agentic organic growth engine and an organic growth partner. Agent Atlas puts the agents in your team’s hands. Pepper’s GEO platform reports Brand Visibility, Domain Prompt Presence and Share of Voice. A growth team works alongside yours. Eight years, more than 250 enterprises, more than 10 million tracked prompts.

A note on where this comes from. I spend most of my time on what changes when a channel’s economics change, and this one has a structural problem rather than a performance problem. The work influences buyers at the exact moment your analytics goes blind. Pepper runs organic for more than 250 enterprises over eight years and tracks more than 10 million prompts across every major engine, and the argument I have most often is not whether GEO works. It is what counts as proof.

Disclosure: Pepper sells GEO services, so a confident yes would suit us. The honest answer is more qualified than that, and this article gives it. Every figure below is traced to a named source with its sample size, or flagged as lacking one. Competitors are named but never linked.

What is GEO pipeline impact, and why is it so hard to see?

GEO pipeline impact is the share of your closed and open pipeline that exists because a generative engine named or cited your brand during the buying process. It is hard to see because the influence and the attribution happen in different places.

A buyer asks an assistant which tools to consider. The assistant names four, including you. The buyer never clicks. Three weeks later they search your brand name, arrive directly, and book a demo. Your analytics records direct or branded search. It is not wrong. It simply cannot see the moment that mattered.

Three properties follow, and they shape everything below.

  • The influence is upstream of the click, so referral data measures the wrong moment.
  • The influence is strongest at the comparison stage, which is exactly where the buyer stays inside the answer rather than visiting four websites.
  • The lag is long. On enterprise cycles, a citation earned this quarter shows up in pipeline several quarters later, which makes short-window tests useless.

Our live GEO ROI model covers how to model the spend once you accept this. This article covers the prior question, which is whether there is anything real underneath it.

The demand evidence is strong

Figure 1: What the best-sampled B2B research does and does not answer. Source: G2 buyer behaviour research, published 22 July 2026.

The strongest B2B evidence available is G2’s 2026 buyer behaviour research, published 22 July 2026. It surveyed more than 1,000 B2B software buyers and decision-makers, paired with interviews from more than 50 B2B sales and marketing leaders.

Its findings on AI use are unambiguous.

  • 82% of buyers sourced software recommendations from an AI chatbot in the last 24 months. Eight in ten, on a sample large enough to take seriously.
  • About half of those buyers said AI had its greatest impact when narrowing and comparing options.

The second finding is the one that matters here, and summaries routinely drop it. Narrowing and comparing is not awareness. It is the stage where a shortlist becomes a preference, which makes it the most commercially valuable moment in a B2B purchase.

And the same report contains no data on whether those buyers convert, purchase or generate revenue. That is not a criticism of the research. It set out to describe behaviour rather than attribution. It is a statement about what the best evidence can and cannot support.

Where it falls short: it is a survey, so it measures what buyers report rather than what they did, and buyers are unreliable narrators of their own process. Read it as strong evidence that AI is present in the buying journey, and as no evidence at all about conversion.

If you would rather find out whether engines name you at the comparison stage than read another survey, book a growth audit and we will run your category’s prompts with you.

The pipeline evidence is not

Here is every published claim we could find about how AI traffic converts relative to organic, gathered on 24 September 2026.

Figure 2: The four claims that state a multiple. The peer-reviewed study is not plotted, because it publishes none.
  • Worse than organic, from the only peer-reviewed study in the set. It covers 973 e-commerce sites, $20 billion in revenue, and 50,000 ChatGPT-referred transactions against 164 million from traditional channels. AI traffic converted below affiliates, organic search, paid search, direct, email and referral. It beat only paid social.
  • 22% better than organic, from a 2026 benchmark report, at 3.49% against organic.
  • About double, from a B2B figure of roughly 4.2%.
  • Three times better, attributed to a vendor’s own first-party data, published with no sample size.
  • Five times better, at 14.2% against 2.8%, a figure we could not locate on the page it was attributed to.

Among the four that state a multiple, the spread runs from 1.22 to 5.1, on the same question, in the same year. Only one of those four publishes a sample size. The fifth position, the peer-reviewed one, cannot be plotted at all, because it reports an ordering rather than a multiple. That is worth sitting with: the only rigorous entry in the set will not give you the number everybody else quotes freely.

This is not a reason to dismiss GEO. It is a reason to distrust every conversion multiple in the category, especially the flattering ones. We repeated an unsourced multiple ourselves and corrected it. The correction cost us a much better-sounding number.

Where it falls short: we gathered what is publicly findable rather than every claim that exists, and a vendor without a published method may well have a rigorous one internally. The point is what a buyer can verify, which is very little.

Why the two bodies of evidence disagree

They are measuring different things, and both are right about the thing they measure.

The demand research measures influence. It asks buyers what shaped their decision, and captures the comparison stage where AI is strongest.

The conversion research measures referrals. It asks what happened to sessions arriving with an AI engine as the referrer, which is the small subset of influenced buyers who clicked.

A product designed to answer without sending a click produces exactly this pattern. High reported influence. Low measurable referral. Unstable conversion figures, computed on a small and unrepresentative tail.

There is also a timing problem, and it is the most underrated fact in this debate.

Figure 3: Why the pipeline evidence does not exist yet. Source: Pepper’s own engagement observations, stated as ranges.

Citation to traffic runs about six to nine weeks. Citation to pipeline runs three to five quarters on enterprise sales cycles. AI search only became material during 2025. So the first cohorts clean enough to prove anything are landing now. Anybody publishing confident B2B pipeline attribution before that had a shorter cycle or a looser definition.

How we weighted the evidence

Four criteria, fixed before we gathered anything. They are our priorities, not measured coefficients.

CriterionWeightWhat it means
Sample size and method published30You can see how many, over what window, and judge how much a figure means without asking anyone
Independent of the seller25The publisher does not sell the thing the number flatters, which rules out most of what circulates
Measures deals rather than sessions25It follows buyers to an outcome a finance team recognises, instead of stopping at a conversion event
B2B rather than e-commerce20The buying committee, the cycle length and the consideration depth all differ enough to change the answer
Figure 4: How we judged each source. Same weighting approach as our GEO agency ranking methodology.

What the weighting produces is uncomfortable. The G2 survey scores well on the first three and is B2B, but it measures influence rather than deals. The peer-reviewed study scores well on all but the last. Nothing available scores well on all four. That is the real state of the evidence, and the reason this article exists.

Where it falls short: weighting published method above everything favours research that is transparent over research that is right. A vendor sitting on excellent unpublished data scores badly here and may be correct anyway.

The GEO pipeline test you can actually run

Since the published evidence cannot answer the question for your business, here is the test that can. It takes two quarters and costs almost nothing.

  1. Freeze a prompt set of 25 buying questions from your own category, and record which brands engines name at the comparison stage. This is the input you are trying to move.
  2. Add one open field to every form, asking what brought them. Store the raw text. This is the only instrument that catches a buyer who never clicked.
  3. Baseline your branded and direct pipeline before you change anything. If GEO works without attribution, this is where it will appear, not in AI referral pipeline.
  4. Ship GEO work against half your category’s prompts and not the other half. A holdout is the difference between a test and a story.
  5. Wait three quarters, not one. On the lag above, a one-quarter read tells you nothing and will be interpreted anyway.
  6. Compare branded pipeline growth between the two prompt groups, and read the self-reported field for mentions of assistants.

What a positive result looks like: branded and direct pipeline grows faster in the covered group, and assistant mentions rise in the self-reported field. What a negative result looks like: no divergence after three quarters. The honest conclusion is then that your category’s buyers are not using engines this way yet.

Where it falls short, and it costs us something: a holdout means deliberately not doing work you believe in for two quarters. That is a real cost, and the reason almost nobody runs one. Without it you have a trend and no counterfactual, and a trend is what everybody already has.

GEO pipeline evidence at a glance

SourceWhat it measuresSampleIndependentWhat it cannot tell youCost to you
G2 buyer behaviour research, 22 Jul 2026Buyer-reported AI use and where it matters1,000+ buyers, 50+ leader interviewsYesWhether those buyers convertedFree to read
Marketing Science, 2026Conversion and revenue per session by channel973 sites, 164M comparison transactionsYes, peer-reviewedAnything about B2B pipelineFree abstract, paywalled full text
Vendor first-party claimsReferral conversion multiplesNot publishedNoAlmost anything, without a methodFree, and worth that
Your own self-reported fieldWhat buyers say brought themYour conversionsYesPrecise share, because recall is imperfectFree. One form field
Your own holdout testWhether covered prompts move pipelineYour categoryYesAnything in under three quartersTwo quarters of foregone work
Prompt trackingWhether engines name you at comparisonYour prompt setYesWhether being named produced a dealFrom about $29 a month

## What this costs to test

  • Three of the six instruments above are free. A form field, a branded pipeline baseline and reading the two published studies.
  • Prompt tracking is the only recurring licence, starting around $29 a month at published entry tiers.
  • The real cost is the holdout, paid in foregone work rather than cash, and it is the one that makes the result mean anything.
  • The three-quarter wait is the expensive part politically, because most marketing leaders are asked to justify spend on a quarterly cycle and this does not fit one.
  • Enterprise platforms that supply the evidence base start around $2,500 a month, covered in our enterprise pricing work.

How Pepper fits

Pepper is an agentic organic growth engine and an organic growth partner, and the honest version of this argument is the one our growth teams have with clients every quarter.

  • Pepper’s GEO platform measures the influence side, which is the side that exists. Brand Visibility for how often engines mention you, Domain Prompt Presence for how often they cite a page from your domain, and Share of Voice for your slice of the category. Those three measure whether you are present at the comparison stage, which the buyer research identifies as the decisive moment. See the platform.
  • Agent Atlas keeps a test like the one above alive across three quarters. Freezing prompt sets, maintaining holdout groups and re-running comparisons on schedule are workflows rather than someone’s reminder. System agents stay fixed, user agents stay editable and versioned, and customers log in and build and run their own inside Atlas.
  • A growth team works alongside yours on the part no platform does, which is agreeing with your revenue operations team what would count as proof before the test starts.
  • Proof rather than adjectives. Acceldata went from 85 to more than 300 top three keywords with 6X organic traffic growth. More in our case studies, and the B2B SaaS practice is where this test was designed.

Where Pepper fits, and where it does not. We are built for considered purchases with long cycles, which is where the influence evidence is strongest and the attribution evidence is weakest. A business that needs attributed pipeline this quarter should not buy GEO from us, and we would say so on a first call. We also publish no pricing, so budget discovery is a conversation rather than a page.

How to choose who to trust on this question

I would judge any partner on one thing before anything else: whether they will tell you what their evidence cannot show.

The framing judgement first, anchored outside my own view. Google’s guidance on optimising for generative AI features states that optimising for generative AI search is still SEO, running on core ranking systems with no separate index, and it advises against providers guaranteeing rankings because no external party has access to those systems. Read that as a filter on anyone promising a pipeline outcome. Then hold the other half, which is that Google documents Google, and the assistants doing most of the comparison work publish nothing equivalent.

The Pepper view on top is narrower. In a market this young, the useful signal is not confidence. It is calibration.

Here is the 100 point scorecard I would run over any partner, including us.

AreaWeightWhat a strong partner demonstrates
Distinguishes influence from attribution30They state plainly which of their numbers measure buyer behaviour and which measure referrals, and never present one as the other
Every figure carries a sample size25Each number in the pitch has a study and a sample behind it, and they say which ones are their own unpublished data
They propose a holdout20The plan includes a group that does not get the work, and they explain what it buys you rather than treating it as optional
They set the review at three quarters15The timeline matches the lag in your sales cycle rather than your reporting cycle, and they say so before you sign
They will describe a negative result10They can tell you what would make them conclude this is not working for you, in advance and in writing

**Then run the live test, on us as readily as on anyone else.** Give any prospective partner 25 buying questions from your own category and a quarter of pipeline data, then hold their read for 90 days before judging it. Engines vary between runs, so a single reading proves nothing. Ask for the provenance of every claim. For example: “which of these numbers measures buyers and which measures sessions”, “what is the sample behind this conversion multiple”, “what would you expect our branded pipeline to do if this works”, “what result would make you tell us to stop”.

Ask them to come back with five things. Which of their figures are influence and which are attribution. The sample behind every number. The holdout design. The quarter at which they expect a readable signal. And the outcome that would count as failure.

A partner who names their failure condition beats one who presents a case study. And one quoting a conversion multiple with no sample has repeated a slide, not read a study.

The weaker way to run this, and it is the common one. Find a flattering multiple. Project it onto expected AI traffic. Buy a programme. Report AI referral pipeline monthly. Watch it come in near zero. Cancel at the end of quarter two, one quarter before the lag would have produced a readable signal.

The stronger sequence. Separate influence from attribution before you start. Baseline branded and direct pipeline. Instrument the invisible part with a form field. Run a holdout across half your prompt set. Read at three quarters. Accept a negative result if you get one. The distinction matters. The first sequence guarantees an inconclusive answer. The second produces a real one, whichever way it falls.

Red flags, each one something a provider actually says.

  • A pipeline guarantee, which Google itself advises against for rankings and which is worse here, because the attribution does not exist to verify it.
  • A conversion multiple with no sample size. The published range runs from below one to five, so any single figure is a choice rather than a finding.
  • Influence statistics presented as attribution. “82% of buyers use AI chatbots” is true and says nothing about your pipeline.
  • A quarterly proof point on a purchase cycle measured in quarters.
  • No holdout offered, and no explanation of what you lose without one.
  • Case studies with no counterfactual. Pipeline grew, and something else in the business may explain all of it.
  • Reluctance to describe failure. A partner who cannot say what a negative result looks like is not running a test.

Five questions worth asking, and what a good answer sounds like.

  1. “Which of your numbers measure buyers and which measure sessions?” A good answer sorts them immediately. A bad answer treats the distinction as pedantic.
  2. “What is the sample behind this multiple?” A good answer names a study. A bad answer names a conference or a dashboard.
  3. “What should our branded pipeline do if this is working?” A good answer predicts a direction and a rough timeframe. A bad answer talks about visibility.
  4. “Will you run a holdout?” A good answer says yes and explains the cost. A bad answer says the data will be clear without one.
  5. “What would make you tell us to stop?” A good answer is specific and uncomfortable. A bad answer does not exist.

If I reduce this to one principle: separate what you can influence from what you can measure, and never let a partner blur them. The blur is where every bad GEO engagement lives.

The honest note that costs us something. Very few providers are strong at measurement discipline, content execution, technical work and earned media at once, and we would not claim uniform strength across all four either. In a market where proof is this hard to get, the temptation to oversell is strongest. That applies to us as much as anyone.

What nobody should promise you

Nobody should promise attributed pipeline from GEO. The attribution to verify it does not exist yet for most B2B businesses. The promise cannot be checked, which is precisely why it gets made.

Nobody should quote a conversion multiple without a sample size. The published range runs from worse than organic to five times better, in the same year, on the same question.

Nobody should present buyer survey findings as evidence of conversion. That 82% figure describes behaviour, not outcomes, and the research that produced it makes no claim otherwise.

Nobody should promise a readable result in one quarter when citation to pipeline runs three to five quarters on enterprise cycles.

Where this stops working, including for us

If your sales cycle is short and your product is simple, the peer-reviewed e-commerce evidence applies to you more than the B2B buyer research does. It is unflattering. Treat AI referrals as discovery traffic.

If you cannot run a holdout, you will not get a causal answer, and you should say so in the report rather than presenting a trend as proof.

If your board needs attributed pipeline within two quarters, you do not need GEO yet and should not buy it. You will cancel it one quarter before it becomes readable, and you will have spent the money for nothing.

Where Pepper fits and does not. We are built for considered purchases with long cycles, which is where the influence is real and the attribution is hardest. For a business needing proof inside a quarter, we are the wrong partner and we would rather say it now.

Where to go next

Decide first whether you are buying influence or attribution. If your organisation can only fund what it can attribute, that is a legitimate constraint. It rules this out for now.

Then baseline your branded pipeline this week, because the test above is worthless without a before. For the adjacent decisions, our model for proving return on AI search investment covers the cost side and break-even, how to build a GEO business case covers the board conversation, the metrics worth tracking covers the influence side, and the enterprise B2B playbook covers the execution. To see whether engines name you at the comparison stage, see where you show up.

Frequently asked questions

Does GEO actually drive pipeline?
The evidence says it influences buying and that most of the influence is not attributable. A survey of over 1,000 B2B buyers found 82% sourced recommendations from AI chatbots, with about half saying the effect was strongest while narrowing options, which produces no click.

Why does AI search influence not show up in pipeline reports?
Because the influence happens at the comparison stage, inside the answer, and the buyer often arrives later through branded or direct search. Your analytics records the last step accurately and cannot see the step that mattered.

How much better does AI traffic convert than organic?
Published claims run from worse than organic to five times better. The only peer-reviewed study, covering 973 e-commerce sites, found AI traffic converting below every traditional channel except paid social. Treat any single multiple with suspicion.

How long before GEO shows up in pipeline?
Citation to traffic runs roughly six to nine weeks. Citation to pipeline runs three to five quarters on enterprise sales cycles. A one-quarter read tells you almost nothing and will be over-interpreted anyway.

How do you prove GEO drives pipeline?
Run a holdout. Ship work against half your category’s prompt set and not the other half, baseline branded and direct pipeline first, add a self-reported field to forms, and compare after three quarters rather than one.

Is AI referral traffic a good measure of GEO success?
No. It counts only buyers who clicked through, which the peer-reviewed evidence puts at under 0.2% of visits. Using it alone is the most common reason a working programme gets cancelled.

Should B2B companies invest in GEO now?
If you sell considered purchases with long cycles and can tolerate a three-quarter proof window, the buyer evidence supports it. If your board requires attributed pipeline within two quarters, wait, because you will cancel before the signal arrives.

What counts as proof that GEO is working?
Divergence in branded and direct pipeline between prompt groups that received work and those that did not, plus a rise in assistant mentions in a self-reported source field. Referral volume alone is not proof.

Sources and further reading

  • G2, 2026 B2B buyer behaviour research, published 22 July 2026. Method: a survey of more than 1,000 B2B software buyers and decision-makers, paired with interviews from more than 50 B2B sales and marketing leaders. Source of the 82% figure and the finding that about half of those buyers said AI mattered most when narrowing and comparing options. Limitations: self-reported behaviour, no geography or fieldwork dates published, and it contains no data on whether AI-sourced buyers convert or generate revenue. We state that absence rather than filling it.
  • Maximilian Kaiser and Christian Schulze, “Frontiers: ChatGPT Referrals to E-Commerce Websites: How Do LLMs Compare Against Traditional Channels?”, Marketing Science, 2026. Method: 973 e-commerce websites, $20 billion combined revenue, August 2024 to July 2025, comparing over 50,000 ChatGPT-referred transactions against 164 million from traditional channels. The only peer-reviewed entry in the set. Limitations: e-commerce only, so it cannot speak to B2B pipeline, which is exactly why this article does not treat it as the answer.
  • The four remaining conversion claims in figure 2 were collated from published pages on 24 September 2026 and are named by their figure rather than linked, per our policy on competitors. None publishes a sample size. One is attributed to a vendor’s own first-party data with no method. One figure could not be located on the page it was attributed to, and is recorded as unverifiable rather than quietly dropped.
  • Pepper, GEO ROI: how to model and prove return on AI search investment. The cost side and break-even calculation this article assumes.
  • Pepper, AEO vs SEO vs GEO. Reports the peer-reviewed conversion finding correctly, and is where we corrected an unsourced multiple we had repeated ourselves.
  • Google Search Central, guide to optimizing for generative AI features, page last updated 10 July 2026. Source of the advice against providers guaranteeing outcomes. Applies to Google Search only.
  • The lag figures and the holdout test design are ours, stated as ranges from our own engagements rather than as measured constants, because we have not published the underlying data and will not present an unpublished number as a finding.

What is not here, and why. No conversion multiple recommended, because the published range spans more than an order of magnitude and picking one would be a choice dressed as a finding. No Pepper client pipeline data, because we have not run the holdout design above across enough accounts to publish a result, and reporting uncontrolled growth as proof is the practice this article criticises. No claim that GEO does not work, which the evidence does not support either. The honest position is that the influence is well evidenced, the attribution is not, and the first clean cohorts are landing now.