How to check if your brand appears in AI search, and whether to believe what you see

The short answer
You can check if your brand appears in AI search in about twenty minutes, free, using nothing but two engines and a notepad. Ask five buyer questions, ask each one twice in different words, and write down who gets named. However, a single check is a much weaker signal than it feels like. Engines do not always search, rewording changes answers, and what gets said about you is not always true. So treat the result as a floor rather than a measurement.
Key takeaways
- Twenty minutes is genuinely enough for a first answer. You do not need a platform to find out whether you have a problem. You need a platform to track whether it is getting better.
- Ask each question twice, reworded. Across 4 benchmarks and 13 models in a 2026 study, meaning-preserving rephrasing changed the answer with mismatch rates above 23%. In other words, one prompt is one sample, not a verdict.
- The engine may not have searched at all. In a July 2026 pilot of 48 buying prompts, ChatGPT searched the web on only 42% of them. Consequently, on the other 58% you were testing the model’s memory, not your visibility.
- When it did search, it searched for something else. The same pilot captured 6.3 sub-queries per searching prompt, and you wrote none of them.
- Appearing is not the same as appearing correctly. A 5,000-prompt benchmark published in April 2026 found factual recall hallucination of 4.2% to 12.7% by model, and citation hallucination of 6.8% to 19.1%.
- Absence is the finding that needs the most care. Above all, do not conclude you are invisible from one run on one engine, because that is exactly the result a bad sample produces.
- Pepper is an agentic organic growth engine and an organic growth partner. Agent Atlas puts the agents in your team’s hands. Pepper’s GEO platform reports Brand Visibility, Domain Prompt Presence and Share of Voice across six engines, including ChatGPT, Perplexity, Gemini and Google AI Overviews. A growth team works alongside yours. Eight years, more than 250 enterprises, more than 10 million tracked prompts.
A note on where this comes from. I run this check with new clients in the first meeting, on a shared screen, because it changes the conversation faster than any deck. Pepper has run organic growth for more than 250 enterprises over eight years and tracks more than 10 million prompts across every major engine. Furthermore, the thing that surprises people is rarely the result itself. It is how much the result moves when you ask the same question a second way.
Disclosure: Pepper sells services that help brands appear in AI answers, so we benefit when you conclude that you are invisible. This article tells you to run the check yourself for nothing. It then tells you why a bad first result is often a bad sample rather than a real problem. Every figure below is traced to a named study with its method and sample size. Competitors are named but never linked.
What is an AI search brand check?
An AI search brand check is a structured test of whether generative engines name your brand when someone asks a question you want to win. Instead of measuring rankings on a results page, you are reading the answer itself.
So the unit of measurement changes. In classic search you had a position, a query and a click. In AI search you have an answer that may name three brands, cite six sources, and never mention a rank at all. Therefore what you are actually recording is presence, prominence and provenance: whether you were named, how favourably, and which pages the engine leaned on.
Two things follow from that, and both matter before you run anything.
- The check is sampled, not measured. You are asking a handful of questions out of an infinite set, on systems that do not answer identically twice. As a result, you get an estimate with error bars, not a number.
- The check has two halves. First, does the brand appear? Second, is what the engine said about it true? Most people only run the first half, and the second half is usually where the fixable problem lives.
How to check if your brand appears in AI search in twenty minutes
Here is the whole method. It runs in one sitting, it costs nothing, and it needs no login and no tool.
1. Pick five questions your buyer would actually ask (3 minutes).
Do not search your brand name. Searching your own name tells you whether the engine has heard of you, which is a different and much easier question. Instead, write the questions someone asks when they do not yet know you exist.
- Category questions: “best [category] tools for [segment]”
- Problem questions: “how do I fix [problem your product solves]”
- Comparison questions: “[known competitor] alternatives”
- Qualifier questions: “[category] for [industry or company size]”
- Decision questions: “is [category] worth it for a [role]”
2. Open a clean session on two engines (2 minutes).
Use ChatGPT and one other, such as Perplexity or Google AI Mode. Log out where you can, or use a temporary chat, because account history and saved memory change what you see. Meanwhile, avoid running this on the browser where you read your own website all day.
3. Ask each question once, then ask it again in different words (10 minutes).
This is the step everyone skips, and it is the step that makes the result worth anything. Rephrase without changing meaning: “best tools for X” becomes “what should I use for X”. Then compare the two answers.
- If both name you, that is a reasonably solid signal.
- If one names you and the other does not, you have variance, not a verdict.
- If neither names you, note it, but hold the conclusion until step 5.
4. Write down three things per answer (3 minutes).
A sheet with five columns is plenty: question, engine, named or not, who else was named, and which sources were cited.
5. Open every source the engine cited (2 minutes).
This is where most of the value is. The cited pages are the pages the engine trusts on your category. Consequently, if a competitor’s comparison page or a stale directory listing is doing the talking, you have found your problem and your fix in the same click.
If you would rather not run this by hand, we will do the first pass with you and hand back the cited-source list for your top prompts. Book a growth audit and bring your five questions.
Four reasons your check may be wrong
Now the part that the rest of the internet leaves out. The method above is sound, but the systems you are testing are noisy in four specific ways, and each one distorts the result differently.

1. The engine may not have searched at all.
An EMGI pilot published July 2026 put 48 SaaS buying prompts through GPT-5.2 with live web search. The model searched the web on only 42% of them. On the remaining 58% it answered from what it already held. So on those runs there was no retrieval event to observe, and nothing you publish this quarter could have changed the answer.
2. When it did search, it searched for something else.
The same pilot captured 127 fan-out queries, an average of 6.3 per searching prompt. You wrote none of them. Therefore, when you test 5 prompts and roughly 2 of them trigger a search, the engine is running something closer to 13 queries of its own devising. Your prompt set is an input, not a map.
3. Rewording flips the answer more often than you would guess.
A 2026 study across 4 benchmarks and 13 models tested meaning-preserving paraphrases. It found mismatch rates above 23%. Its authors conclude that single-prompt correctness is a poor indicator of reliability. That study covered factual question answering and mathematical reasoning rather than brand queries, so treat the exact figure as indicative. Nevertheless, the mechanism transfers, and it is why step 3 exists.
4. Personalisation and session state change what you see, and nobody has published a rate.
Account history, saved memory, location and the model version you happen to get all move the answer. In fact, we looked for a 2026 study that quantifies how much, and we could not find one. So this stays on the list as an unquantified risk rather than a number. That absence is itself worth knowing when a vendor quotes you a precision they cannot have.

Appearing is not the same as appearing correctly
Suppose the check goes well and the engine names you in four answers out of five. Second question: was any of it true?
Presence is string matching, and every monitoring tool on the market sells it. Accuracy is a different job, because detecting it requires your own current source of truth, and no platform holds that. Pepper included.
The scale of the problem is measurable. A benchmark published 23 April 2026 ran 5,000 prompts across five frontier models. Those were GPT-5.5, Claude Opus 4.7, Gemini 3 Pro Deep Think, Grok 4.5 and DeepSeek V4, over three task families, with automated and human review.
- Factual recall hallucination ran 4.2% to 12.7% depending on the model. This is the task family closest to questions about your pricing and features.
- Citation hallucination ran 6.8% to 19.1%. Models invented titles, authors and identifiers.

The caveat matters and we will state it plainly: those were benchmark task families, not brand queries. No 2026 study measures brand-fact accuracy specifically, which is a genuine gap in the evidence rather than something we are glossing over.
In practice, the errors we see most often are stale rather than invented. An engine quotes a price you retired, a tier you renamed, or an integration you dropped, because a third-party page still says so. Those are the most fixable errors in the category. The fix is a correction request to a page you do not own, rather than a change to your own site.
How we weighted what belongs in a twenty-minute check
Plenty of things are worth measuring. Very few are worth measuring in the first twenty minutes. Here is how we decided what made the cut.

| Criterion | Weight | Why it carries that weight |
|---|---|---|
| Runs without a tool or a login | 30% | The check must work before you have budget. Otherwise it is a sales process rather than a diagnostic. |
| Produces an action, not just a number | 25% | Knowing you appear in 40% of answers changes nothing. Knowing which competitor page the engine cited changes your week. |
| Survives the reliability tests | 25% | Anything you cannot reproduce on a second, reworded run does not belong in a twenty-minute conclusion. |
| Works on more than one engine | 20% | Engines overlap little in what they cite, so a single-engine result generalises badly. |
## Reading your result: three outcomes
Your twenty minutes will land you in one of three places, and the right next move is different in each.
Outcome 1: you appear, and what the engine says is accurate.
Good. Do not buy anything yet. Instead, set a cadence, because the useful signal here is change over time rather than today’s snapshot. Monthly is enough for most brands, and weekly mostly measures noise.
Outcome 2: you appear, but the details are wrong or out of date.
This is the most urgent outcome and the cheapest to fix, which is an unusual combination. Work backwards from the cited sources: find the page carrying the stale claim, and correct it at the source. Furthermore, publish a current, unambiguous version of the fact on your own site, because engines resolve conflicts partly by weight of agreement.
Outcome 3: you do not appear at all.
Before you conclude anything, rule out the four distortions above. Re-run on a clean session, reword every prompt, and try a third engine. If you are still absent across engines and phrasings, then the finding is real, and it is a citability problem rather than a monitoring problem. In that case buying a tracking platform first would be spending money to watch a number you already know.
What a twenty-minute check cannot tell you
Being clear about the limits is what makes the rest trustworthy.
- It cannot give you a rate. Five prompts on two engines is a sample of roughly ten answers. That supports “we appear sometimes” and never “we appear 40% of the time”.
- It cannot detect the no-search answers. When the model answers from memory, there is no retrieval event for anyone to observe, including any platform you might buy.
- It cannot tell you why. Absence has several causes: thin coverage, weak third-party corroboration, crawler access problems, or simply a category where the engine defaults to incumbents.
- It cannot measure trend. One sitting is one point. Trend needs a fixed prompt set, a fixed cadence and enough runs to separate movement from variance.
- It cannot tell you what it is worth. Presence and pipeline are different questions, and the link between them is weaker than the category implies.
The check at a glance
| Step | Time | Cost | Tool needed | What it tells you | How reliable |
|---|---|---|---|---|---|
| Write five buyer questions | 3 min | $0 | None | Whether you are testing discovery or vanity | High, if you avoid your brand name |
| Open clean sessions on two engines | 2 min | $0 | None | Removes the most obvious personalisation effect | Partial, session state is not fully controllable |
| Ask each question twice, reworded | 10 min | $0 | None | Whether a result is stable or noise | The single most useful step |
| Log named, competitors, sources | 3 min | $0 | Spreadsheet | Who owns the answer today | High |
| Open every cited source | 2 min | $0 | None | Where the fix actually lives | Highest value per minute |
| Conclude a rate or a trend | Not possible | From ~$29 a month | Platform | Requires fixed prompts and repeated runs | Out of scope for one sitting |
## What this costs
Nothing, in money. About twenty minutes of one person’s time, plus another twenty if you write the results up properly for someone else to read.
The ongoing version costs more. A fixed prompt set, run monthly across several engines and logged consistently, costs roughly two hours a month by hand. That is the point where most teams start looking at a platform. Published entry pricing in this category starts around $29 a month for very limited prompt counts and rises quickly with engines and volume. However, the honest sequence is to run the free check first, because it tells you whether you need the paid one.
How Pepper fits
Pepper is an agentic organic growth engine and an organic growth partner, and that means three things working together rather than one product.
Pepper’s GEO platform is the self-serve workspace. You set up a brand profile, add competitors and personas, connect GA4 and Search Console, and define the themes and prompts you care about. It then reports Brand Visibility, Domain Prompt Presence and Share of Voice across six engines, including ChatGPT, Perplexity, Gemini and Google AI Overviews. Platform Breakdown, New Mentions and Citation Analysis sit underneath. That is the ongoing version of step 4 and step 5 above, run properly and at volume.
Agent Atlas is where your team builds, versions and runs its own agents, with quick runs for a single input and sheet runs for bulk. This is the part that does the boring half of this article at scale. It pulls the cited sources out of a few hundred answers, then tells you which third-party pages keep showing up.
The growth team is attached to the account and does the work alongside yours. That matters most in outcome 2 and outcome 3. Correcting a stale third-party page and earning new corroboration are both slow, unglamorous jobs, and no platform does them for you.
Where it falls short: our platform measures presence, not accuracy. It will tell you that an engine mentioned you and which sources it cited. It will not tell you that the price quoted was eighteen months out of date. That check needs your current source of truth, and no vendor in this category has solved it. We run that part as a manual audit, roughly a day a quarter, and we would rather say so than imply the software covers it.
We have done this for eight years, across more than 250 enterprises, tracking more than 10 million prompts. You can see the shape of the work in the Acceldata case study, and in how we run this for B2B SaaS brands specifically. There are further examples across our case study library.
How to choose what to do after the check
The decision is not which vendor to buy. It is whether to buy anything at all, and in which order. So here are the criteria we apply, weighted, and you should disagree with the weights if your situation differs.
The weighted scorecard
Score each row from 1 to 5, multiply by the weight, and total it out of 100.
| Criterion | Weight | Score 1 means | Score 5 means |
|---|---|---|---|
| Prompts you genuinely need to watch | 25 | Under 10, a spreadsheet is fine | Over 100, manual logging has broken down |
| Engines your buyers actually use | 20 | One engine covers them | Four or more, with real differences between them |
| How fast your category’s answers change | 20 | Stable for months | New comparison pages appear weekly |
| People who need the same report | 15 | Just you | A board pack, monthly |
| Cost of an accuracy error | 20 | Mildly embarrassing | A quoted price you no longer honour |
Under 40, run the free check monthly and spend the money on content instead. From 40 to 70, a light tool plus manual review is proportionate. Above 70, you are past what a spreadsheet can carry, and the reporting overhead alone justifies a platform.
The weaker playbook against the stronger one
The weaker response to a bad check result is to buy a dashboard and watch the number. That feels like action, and it changes nothing, because the number is an output of pages you have not touched. The stronger response works backwards from the sources you found in step 5. Find the three pages the engine keeps citing. Then get your facts corrected, or your brand represented, on each one. One is measurement theatre. The other moves the thing being measured.
Run a live test before you commit
Ask for a trial on your own prompt set, never a demo on theirs. Give the vendor 20 prompts written in your buyers’ language rather than yours. Good examples look like “best invoicing tools for freelance designers”, “alternatives to the market leader for a 50 person team”, or “is this category worth it for a regulated business”. Run them twice, then again at 30 days, and a third time at 90 days. Then check three things. Do the numbers move in ways anyone can explain? Do the cited sources match what you see by hand? And did the platform tell you anything you had not already worked out with a notepad?
Red flags
- A single visibility score with no published method behind it
- A quoted engine count that turns out to include paid add-ons
- Any claim to measure answers where the model never searched
- Prompt-set design done entirely by the vendor, with no view of your buyers
- Reporting on accuracy or sentiment without a stated source of truth
- A monthly refresh sold as real-time
- Guaranteed placement in AI answers, in any wording
Five questions worth asking any vendor
- How many runs per prompt do you average before reporting a change, and what variance threshold has to be crossed first?
- Which of your listed engines are included at my tier, and which are priced as add-ons?
- How do you handle prompts where the model returned no citations at all, and are those runs counted in my visibility score?
- What exactly does your accuracy or sentiment metric compare the answer against, and who maintains that reference?
- Can I export the raw answers, not just the aggregates, so I can audit your numbers myself?
The reducing principle. It comes down to one question: can you already name the pages the engine cites? If you can, you have your roadmap, and a platform would be an expensive way to reprint it. If you cannot, a tool earns its price in the first month by finding them for you. Everything else in this section is a refinement of that one test.
The honest closing note. For plenty of brands the right move after this check is to do nothing for a quarter except fix the cited sources found in step 5. That is unglamorous, it does not need us, and it is more likely to shift the result than any dashboard. Buy monitoring when you have to prove a change to somebody else, or when the manual version has outgrown a spreadsheet. Not before.
What nobody should promise you
- A guaranteed mention. Nobody controls what a generative model says, and anyone who says otherwise is selling a coin flip with a retainer attached.
- A precise visibility percentage. Given no-search answers and run-to-run variance, any single figure carries error bars its owner usually does not publish.
- Full engine coverage. Engine lists move monthly, and the gap between the marketing page and the tier you are on is the most common unpleasant surprise in this category.
- Accuracy monitoring, today. No platform we have tested, including ours, measures whether what the engine said about you is true.
- A fast fix for a real absence. Corroboration takes quarters, because it depends on pages you do not control.
Where this stops working, including for us
This article assumes you are a brand in a category an engine will discuss at all. Consequently, if you sell something so new that the category has no name yet, the check returns nothing useful. The honest read then is a category education problem rather than a visibility problem.
It also assumes a reasonably competitive category. In a market with three players, engines tend to name all three, and the check tells you very little you did not know.
The reliability evidence has a real limit, and we would rather flag it than lean on it. The 23% paraphrase mismatch figure comes from factual and mathematical benchmarks, not brand queries. The 42% search rate comes from a pilot of just 48 prompts. Its own authors call it a pilot rather than a population study. Both point the same way. Neither is a precise number for your category.
And our own limitation stands: Pepper measures presence rather than truth, so the accuracy half of this article is manual work at every vendor, ours included.
Where to go next
Once you have your result, these go deeper on the specific question you now have.
- For the ongoing free method, with the sheet and the monthly cadence, read how to measure AI search visibility without expensive tools.
- To compare detection approaches properly, read how to track brand mentions in AI search.
- For the engine-specific version of step 2, read our ChatGPT brand visibility audit.
- For how many runs a change actually needs before it is real, read AI search tracking across engines.
- To decide whether to buy a platform at all, read why use AI search monitoring tools.
- For a structured version of the whole audit, use the 7-point AI search audit.
- For the three levers behind all of this, read Visibility, Citability and Retrievability.
- For why the engine’s own sub-queries matter, read query fan-out tracking.
Frequently asked questions
How do I check if my brand appears in AI search?
Ask five questions your buyer would ask, on two engines, in a logged-out or temporary session, and ask each one twice in different words. Record whether you were named, who else was named, and which sources were cited. Then open every cited source. It takes about twenty minutes and needs no tool.
Should I just search my own brand name?
No, or at least not only. Searching your brand name tests whether the engine has heard of you, which is a much easier question than whether it recommends you. The valuable test is the category question, asked by someone who does not know you exist yet.
Why do I get a different answer every time I ask?
Because these systems are probabilistic, and because rewording matters more than people expect. A 2026 study across 4 benchmarks and 13 models found meaning-preserving paraphrases produced mismatch rates above 23%. As a result, one run is one sample, so build repetition into the method rather than treating the first answer as the truth.
Does it matter if I am logged in?
Yes. Account history, saved memory and location all influence the answer. Use a temporary chat or a logged-out session where the engine allows it. No published 2026 study quantifies how large this effect is, so we treat it as a real but unmeasured risk.
My brand did not appear at all. Am I invisible?
Not necessarily, and this is the result most worth double-checking. Re-run on a clean session, reword every prompt, and try a third engine first. If you are still absent everywhere, the finding is real, and the problem is citability rather than measurement.
Do I need a paid tool to do this?
Not for a first answer. You need a tool for three reasons only. You have to prove a change to someone else, your prompt set has outgrown a spreadsheet, or several people need the same report. Run the free version first, because it tells you whether the paid one is worth it.
How often should I re-run the check?
Monthly for most brands. Weekly mostly measures run-to-run variance rather than real movement, and quarterly is too slow to catch a competitor’s new comparison page while it is still new.
The engine named us but got our pricing wrong. What do I do?
Treat it as the most fixable problem you have. Find the cited page carrying the stale figure and request a correction at the source, since errors traced to a third-party page are the easiest to resolve. Then publish an unambiguous current version on your own site, because engines weigh agreement across sources.
Sources and further reading
- EMGI, AI search retrieval pilot, published July 2026. 48 SaaS buying prompts across six categories, run through GPT-5.2 with live web search via the DataForSEO LLM API, with 16 identical prompts also run on Gemini, Claude and Perplexity. Source of the 42% search rate and the 6.3 fan-out average. The authors describe it as a pilot rather than a population study, and note that API behaviour may differ from the consumer app.
- Faghih, Cheng, Saha, Pournemat, Gerami and Feizi, “Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy”, arXiv:2607.22554, submitted 18 May 2026. 4 benchmarks, 13 models, meaning-preserving paraphrases across factual question answering and mathematical reasoning. Source of the mismatch rate above 23%. Not a study of brand queries.
- Hallucination benchmark published 23 April 2026. 5,000 prompts across GPT-5.5, Claude Opus 4.7, Gemini 3 Pro Deep Think, Grok 4.5 and DeepSeek V4, three task families, automated plus human review. Source of the 4.2% to 12.7% factual recall range and the 6.8% to 19.1% citation range. Benchmark task families rather than brand queries.
- Faghih et al., Same Question, Different Answers, arXiv preprint, for the paraphrase mismatch finding in full.
- Pepper, our own fan-out range, eight to twelve, 24 June 2026.
- Pepper, the ongoing free measurement method, 26 August 2026, which this article hands off to.
- Pepper, five detection methods compared, 26 August 2026, which states that it does not examine accuracy.
- Pepper, the framework behind the three levers, for where citability sits.
A note on sources. Only studies published in 2026 are cited here. Two older fan-out studies with larger samples fall outside that rule, because they describe a superseded model generation. One 2025 study of pricing accuracy in SaaS buyer queries is excluded on the same rule. It was the most directly relevant finding we saw, which is worth saying out loud.
Latest Blogs
You can check if your brand appears in AI search in about twenty minutes, on two engines, without paying for anything. The method is simple and the result is noisier than most people admit. In one 2026 pilot, ChatGPT searched the web on only 42% of buying prompts, so on the rest there was no retrieval event at all. In another, rewording a question changed the answer more than 23% of the time. This article gives you the check, then gives you the four reasons it can lie, so you can tell a real absence from a bad sample.
A blog traffic drop has five plausible causes and AI search is the most interesting one, which is exactly why people reach for it first. It is also the only one you cannot measure directly, so it should be diagnosed last, by elimination. The first check takes ten minutes and costs nothing: put clicks and impressions on the same chart. In one 2026 study of 53 brands, clicks held flat at around 400,000 while impressions more than doubled, so clickthrough halved without a single visitor being lost. Work through the other four before you conclude anything about AI.
Google publishes one eligibility rule for AI Mode, and in May 2026 it published what AI Mode users actually do. Both matter more than the fan-out arithmetic the category quotes.