How to track brand mentions in AI search: 5 methods

Every method on this page is partial. That is the honest starting point, and it is the thing most guides in this category skip.
AI answers are probabilistic, so the same prompt produces different responses on different runs. Engines differ in whether they cite sources at all. And a mention inside an answer usually leaves no trace in your analytics. So the question is not which method is best. It is which combination covers your blind spots at a cost you can justify.
We run organic for more than 250 enterprises and track over 10 million prompts across every major engine, so we use all five of these, and we have opinions about which ones earn their cost.

Method 1: Manual prompt testing
The cheapest method, and the one most teams skip because it feels unserious.
Write 20 to 30 questions your buyers actually ask, phrased the way they would phrase them rather than as keywords. Then run them across ChatGPT, Perplexity, Gemini and Google AI Overviews on a fixed schedule, and record four things each time: whether your brand appears, where in the answer it sits, how it is described, and which competitors appear instead of you.
What it tells you. Everything the automated tools tell you, at a smaller sample. More usefully, it shows you the actual language engines use about your category, which no dashboard summarises well.
Where it is blind. Sample size and consistency. Because answers vary between runs, a single pass tells you very little, and humans are inconsistent about recording the same thing the same way twice.
Our view. Do this first, before buying anything. It takes an afternoon, it tells you whether you have a problem worth spending money on, and it makes you a far better buyer when you do talk to vendors. Teams that skip straight to tooling usually cannot tell whether the dashboard is answering the right question.
Method 2: First-party citation data
Some of this data now comes directly from the engines, which was not true a year ago.
Bing’s AI Performance report entered public preview for all verified sites on 10 February 2026. It shows cited URLs, triggering queries and citation share, drawn from Microsoft’s own systems rather than inferred from sampling.
What it tells you. Which of your pages get cited and for what, with no sampling error, at no cost.
Where it is blind. Read the scope carefully, because it is narrower than people assume. Microsoft states it covers Copilot, AI-generated summaries in Bing and select partner integrations, that the data represents a sample rather than full citation activity, and that it counts citations rather than visits. It is not a window into ChatGPT.
Our view. Free, first-party and specific, so connect it on day one. Just do not let anyone present it as total AI visibility, because it covers one part of one ecosystem.
Method 3: Crawler and referral log analysis
This method measures whether engines can reach you at all, which is upstream of whether they mention you.
Check your server logs for AI agents specifically. OpenAI runs OAI-SearchBot, ChatGPT-User and GPTBot as three distinct agents with different jobs, and Perplexity runs PerplexityBot and Perplexity-User for scheduled indexing and live fetching. Then segment referral traffic by AI source in GA4.
What it tells you. Whether crawlers are reaching your important pages, and what visitors from AI engines do once they arrive.
Where it is blind. Referral data undercounts badly, because most AI answers resolve without a click. Forrester’s 2026 Buyers’ Journey Survey of 18,000 global business buyers reports companies seeing 10% to 40% traffic declines as research moves into engines. So falling AI referral traffic alongside rising citations is the normal pattern, not a failure.
Our view. This is the most under-used method on the list, and it catches the most expensive problem. A robots.txt line written two years ago by somebody who has since left is a genuinely common cause of invisibility, and no amount of content fixes it.
Method 4: Dedicated AI visibility platforms
The automated version of method one, at a sample size humans cannot match.
Platforms run your prompt set across engines on a schedule and report mentions, citations, cited URLs, competitor share of voice and sentiment. Pricing runs from about $29 a month at the low end to several hundred for real multi-engine coverage.
What it tells you. Trend over time at a defensible sample, plus competitor context you cannot easily assemble by hand.
Where it is blind. Three things worth knowing before you buy. Coverage depends on your tier rather than the marketing page, and entry tiers are frequently ChatGPT only. Prompt selection methodology is rarely disclosed, and a prompt set weighted toward branded queries will always look healthy. And composite visibility scores hide the gap between being mentioned and being cited, which is the finding you can actually act on.
Our view. Worth it once you have more prompts than a person can run consistently. Ask for raw answers and cited URLs rather than a score, because a number you cannot audit is a number you cannot defend in a review.
Method 5: Citation source analysis
The method that tells you what to do, rather than how you are doing.
Instead of tracking whether you appear, read the sources cited in the answers where you do appear, and in the answers where a competitor appears instead. Classify each cited URL as owned, earned or paid. Then find the sources that recur.
What it tells you. Which ten to twenty publications, review platforms, comparison sites and communities actually decide your category. We call that the Citation Core, and most brands cannot name a single member of theirs.
Where it is blind. It is manual, it does not automate well, and it takes judgement rather than analyst time.
Our view. The highest-leverage method here, and the one that changes budgets. Muck Rack analysed more than 25 million cited links across ChatGPT, Claude and Gemini in May 2026, spanning 17 industries, and found earned media accounted for 84% of AI citations against 0.3% for paid and advertorial. If most of what decides your answers is published elsewhere, knowing exactly where is worth more than another visibility percentage.

Why one method is never enough
The engines behave differently enough that a single approach misreads at least one of them.

Figure 3: Citation behaviour varies by engine. Source: Muck Rack, May 2026.
Muck Rack’s data shows ChatGPT includes sources in 96% of responses, Gemini in 82%, and Claude in 55%. For an engine citing barely half the time, citation tracking tells you much less than mention tracking does. A blended average across the three lands near 78% and describes none of them.
This is why we report per engine and separate two numbers that most dashboards merge. Brand Visibility is how often engines name your brand. Domain Prompt Presence is how often they cite a page from your domain. Healthy mentions with flat citations means engines know you and are describing you using somebody else’s page, which sends you to methods three and five rather than to more publishing.
What this looks like when it works
SalesHood is the clearest published example we have, and it is useful because their starting position was a decline rather than a standing start.
They are self-funded, competing against much better-resourced rivals, and their traffic was falling despite content that had historically performed. Elay Cohen, their CMO, described the problem directly: “Traffic was declining, we realized the world of GEO and AEO was happening without us.”
The measurement work came before the content work. Website and content audits, ten technical SEO issues resolved, an EEAT audit, FAQ schema, then AI Overview visibility tracking through Atlas, alongside Reddit, YouTube and podcast strategies.

Keywords appearing in AI Overviews rose from 14 to 97, a 593% increase. Impressions doubled from 200,000 to 417,000. SERP rich results grew 291% across five months, and clicks rose 20% while the industry-wide direction was down. Individual movements show the mechanism: “Sales Enablement Meaning” went from rank 101 to rank 1, and “AI Sales Enablement Platform” from 4 to 1.
Cohen’s summary was blunter than anything we would write for ourselves: “This is more than an agency, more than technology. Move fast and dominate.”
How to combine the five
Run them in this order, because each one tells you whether the next is worth paying for.
Start with method one and method three in the same week. Manual prompt testing tells you whether you have a visibility problem, and log analysis tells you whether it is actually a crawling problem wearing a visibility problem’s clothes. Together they cost nothing but time.
Connect method two immediately, since Bing’s data is free and first-party.
Add method four when your prompt set outgrows what a person can run consistently, which for most teams is somewhere past 30 prompts across four engines.
Then run method five quarterly. It is the slowest and the one that most often changes where the budget goes.
Want the diagnostic run for you before you build any of this? Book a growth audit and we will run your category’s prompt set and show you which sources are winning the answers you want.
What no method can tell you
Whether a mention drove a purchase. Attribution from citation to closed revenue does not close cleanly today, and anyone claiming otherwise is overstating what the data supports.
What your true visibility percentage is. Every number here is a sample against a prompt set you chose, so it is a directional measure of a defined question space rather than an absolute.
What a good number looks like. There is no universal benchmark. Your trend and your distance from the category leader are the only comparisons worth making.
Whether an engine will cite you tomorrow. Model updates re-weight sources without notice, which is why step changes in these metrics are often platform behaviour rather than your performance.
Frequently asked questions
How do I track brand mentions in AI search?
Combine five methods: manual prompt testing, first-party data from Bing’s AI Performance report, crawler and referral log analysis, a dedicated AI visibility platform, and citation source analysis. No single method covers the whole picture.
Can I track AI brand mentions for free?
Partly. Manual prompt testing costs only time, Bing’s AI Performance report is free first-party citation data, and crawler log analysis uses infrastructure you already have. Automated multi-engine tracking at scale requires a paid platform.
How often should I check AI brand mentions?
Monthly for presence and share of voice, quarterly for citation source analysis. Weekly mostly captures noise, because engine outputs vary between runs even when nothing about your content has changed.
Why do AI answers about my brand change every time?
AI answers are probabilistic, so the same prompt produces different responses across runs. This is why repeated sampling matters more than any single check, and why one flattering screenshot proves nothing about your visibility.
What is the difference between a mention and a citation?
A mention names your brand in the answer text. A citation links to a page on your domain as a source. You can have either without the other, and the gap between them is the most useful diagnostic available.
Does Bing’s AI Performance report cover ChatGPT?
No. Microsoft states it covers Copilot, AI-generated summaries in Bing and select partner integrations, on a sampled basis, counting citations rather than visits. It is valuable first-party data, but it is not a view into ChatGPT.
Which engine should I track first?
Whichever your buyers use, then weight by citation behaviour. ChatGPT includes sources in 96% of responses and Claude in 55%, so citation tracking on Claude tells you far less than mention tracking does.
How many prompts do I need to track?
Enough to represent your category honestly, usually 20 to 30 to start and 50 to 200 for an ongoing programme. Refresh frequency matters as much as count, since 50 prompts refreshed daily yields far more signal than 50 refreshed monthly.
Where to go next
Do method one this week, before you evaluate a single vendor. Twenty to thirty real questions, four engines, one afternoon. You will learn whether you have a problem, and you will be able to tell within ten minutes whether a vendor’s demo is answering your question or theirs.
Then check your crawler access, because that is the cheapest fix with the largest downside if it is wrong. Our guide to robots.txt for AI crawlers covers the specific agents, and can AI search bots crawl your website is the check we run on every new account.
For the metric set behind all of this, what actually matters in AI search measurement covers which numbers to report and which to ignore. The Visibility, Citability and Retrievability framework covers which lever each finding points to.
To see it on your own category without building the tracking, see where you show up across ChatGPT, Perplexity, Gemini, Claude and AI Overviews. Customers run this themselves in the platform, with a growth team attached to do the work alongside them.
Sources
Every study cited here was published in 2026. Case study figures are as published by the client on their own case study page.
- Muck Rack, Earned media still drives 84% of AI citations, 7 May 2026. More than 25 million cited links across ChatGPT, Claude and Gemini, 17 industries. Muck Rack sells PR software, so read it as an interested party with a large dataset.
- Forrester, 2026 Buyers’ Journey Survey, January 2026. 18,000 global business buyers. Forrester sells research subscriptions.
- Bing Webmaster Tools, Introducing AI Performance, public preview 10 February 2026.
- OpenAI crawler documentation and Perplexity bot documentation.
- SalesHood case study, verified 25 August 2026. Figures as published, with Elay Cohen, CMO, named on the page.
Latest Blogs
Choosing an organic SEO agency is mostly a filtering exercise. This covers the eleven red flags that should end a conversation, the five questions that surface them in the first call, and what a realistic engagement actually looks like once you sign.
No single method shows you the whole picture of how AI engines mention your brand. Here are the five that exist, what each one genuinely measures, where each one is blind, and how to combine them without paying for capability you already own.
Almost every SEO agency now calls itself an AI SEO agency. A much smaller number changed what they actually do. This compares the firms that rebuilt their model against the ones that rebuilt their homepage, with published results and pricing where they exist.