AI citation analysis software: how the nine leading platforms actually measure citations

The short answer
They do not measure the same thing, and most will not tell you how they measure it. The IAB published a standard on 3 August 2026 that splits this market into Directional and Decision-Grade measurement, and reserves the second tier for methods with stated query volume, sample size and reproducibility. On the reproducibility requirement, our own census of fourteen platforms found zero disclosing run-to-run variance.
Key takeaways
- There is now a standard, and it is seven weeks old. The IAB’s Measuring Visibility in the AI Era, released 3 August 2026, gives the category a shared vocabulary and a quality bar.
- It defines two tiers. Directional measurement “identifies patterns and trends, supporting early signal detection” but is unsuitable for budget decisions. Decision-Grade “meets a higher standard of rigor” on query volume, sample size and reproducibility.
- Most of this category is Directional and sold as if it were Decision-Grade. That is the finding, and it is the IAB’s own distinction rather than ours.
- Reproducibility is where it breaks. Our earlier review of the AI visibility platforms we ranked found none of fourteen publishing run-to-run variance, and nothing has changed since.
- The platforms do not agree on what a citation is. Some count a linked source, some count a brand mention, and some blend the two into one score without saying so.
- So a citation count from one tracker is not comparable to a count from another, including across a tool migration, which is the practical cost most teams discover late.
- Pepper is an agentic organic growth engine and an organic growth partner. Agent Atlas puts the agents in your team’s hands. Pepper’s GEO platform reports Brand Visibility, Domain Prompt Presence and Share of Voice. A growth team works alongside yours. Eight years, more than 250 enterprises, more than 10 million tracked prompts.
A note on where this comes from. I read these platforms the way I read a spec sheet, and the question I ask first is not what the number is but what produced it. Pepper runs organic for more than 250 enterprises over eight years and tracks more than 10 million prompts across every major engine, and we sell a platform in this category, so the standard below applies to us as much as to anyone here.
Disclosure: Pepper sells an AI visibility platform, so this article assesses a category we compete in. We have therefore kept Pepper out of the ranked comparison entirely and scored ourselves in the prose against the same criteria. Every platform claim came from that company’s own public pages, read on 29 September 2026. Where a platform publishes nothing on a criterion, we record that as publishing nothing rather than inferring an answer. Competitors are named but never linked.
What is AI citation analysis software, and what does it actually do?
AI citation analysis software runs a fixed set of prompts across AI engines on a schedule, captures the answers, extracts the sources the engine cited, matches brand and competitor mentions, and stores the result as a time series.
That description is uncontroversial. The disagreement starts one level down, at what counts as a citation.
Our own definition of a citation in AI search draws the line this way: “Being mentioned means the engine said your name. Being cited means the engine trusted your page enough to build the answer on it.” That is a real distinction with different implications, and not every platform draws it.
Three mechanisms sit behind every number these tools produce.
- The prompt set. Which questions get asked, how many, and who chose them.
- The sampling. How often each prompt runs, and how many times, because engines answer the same question differently on different runs.
- The extraction. What the tool records from the answer, and whether a name in the text counts the same as a linked source.
Change any one of those and the number changes. None of them is standardised across vendors, which is why this article exists.
The standard that arrived in August

On 3 August 2026 the IAB released Measuring Visibility in the AI Era, developed with industry partners including Microsoft Clarity. It is the first attempt at a shared vocabulary in this market, and it does two useful things.
First, it organises metrics into four groups, which it calls the 4 P’s.
- Presence. Does the brand appear in an AI response? Tracked through mention rate, citation rate, share of voice and visibility momentum.
- Prominence. Placement, ranking order, and “how substantively publisher content is drawn upon versus superficially cited”.
- Portrayal. Sentiment, framing, hallucination rate and factual inaccuracy rate.
- Persuasion. Recommendation strength and post-citation clickthrough rate.
Second, and more usefully for a buyer, it defines two quality tiers.
- Directional measurement “identifies patterns and trends, supporting early signal detection”, and the framework is explicit that it is not suitable for budget decisions.
- Decision-Grade measurement “meets a higher standard of rigor” across query volume, sample size and reproducibility, and is what the framework says you need before a strategic decision.
That distinction is the most useful thing published about this category all year, because it gives a buyer a question that vendors cannot answer with a feature list: which tier is this, and how do you know?
Where it falls short: the framework is seven weeks old, adoption is voluntary, and the detailed disclosure requirements sit in a downloadable document rather than on the public page, so we have described the tiers and the hierarchy as published rather than quoting requirements we could not verify at source.
If you want to know which tier your current reporting sits in before your next planning cycle, book a growth audit and we will assess it with you.
What the nine platforms actually publish
We checked nine leading platforms against four disclosures a buyer needs before they can interpret any number the platform produces. All read on 29 September 2026.
- What counts as a citation, and whether a mention is distinguished from a cited source.
- Refresh frequency, meaning how often each prompt is re-run.
- Engine coverage, stated explicitly rather than implied.
- Reproducibility, meaning any statement of run-to-run variance.

The pattern is consistent. Engine coverage is published almost everywhere, because it is a selling point. Refresh frequency is published by a minority. A definition of what counts as a citation is rare. Reproducibility is published by nobody.
- Otterly publishes engine coverage, prompt allowances and daily tracking on every tier, which makes it one of the few where a buyer can compute how much data a number rests on.
- Writesonic publishes prompts and answers together, which is effectively a refresh rate, and states engine coverage by tier.
- Profound describes daily updates and says its prompt volume data is built on more than 1.9 billion user prompts. It withdrew its published self-serve pricing earlier this quarter.
- AthenaHQ meters in credits rather than prompts and prices prompt volume off a proprietary model it does not document, which makes the unit hard to interpret.
- Peec AI is frequently described as having the cleanest separation between a brand appearing and a source being used. Its own site publishes no measurement methodology, which we checked directly.
- Scrunch publishes prompt allowances but no refresh frequency, so the allowance cannot be converted into a sampling rate.
- Semrush and Ahrefs Brand Radar both publish engine coverage and refresh cadence, and both are search platforms extending into this category rather than dedicated citation tools.
- Evertune is the one most often cited by others for statistical rigour, which is a reputation rather than a disclosure.
Where it falls short, and it matters: this measures what each platform publishes, not what it does internally. A vendor with an excellent unpublished method scores badly here and may be the best tool in the category. What the check establishes is what a buyer can verify before signing, which is the only thing a buyer has.
Reproducibility is the one that decides everything

The IAB puts reproducibility in the Decision-Grade tier alongside query volume and sample size. It is the requirement this category systematically fails.
Our earlier review of fourteen AI visibility platforms found not one publishing a run-to-run variance figure. Re-checking the nine here produced the same answer.
Why that matters more than any feature. Engines are probabilistic. Ask the same question twice and you can get different sources. Without a variance figure you cannot tell whether a month-on-month move is your work or the weather. Our sampling arithmetic shows how large that noise floor can be, and it is usually larger than the movements people report on.
The practical consequence. By the IAB’s own definition, a platform that publishes no reproducibility statement is offering Directional measurement. The framework says Directional measurement is unsuitable for budget decisions. Most of this category is therefore being sold for a purpose its own industry body says it does not support.
Where it falls short: absence of disclosure is not proof of absence of rigour. Several of these vendors almost certainly measure variance internally. The criticism is of publication, not of competence, and any vendor here could answer it in an afternoon by publishing the number.
How we weighted these platforms
Four criteria, fixed before we read any vendor page. They are our priorities, not measured coefficients.
| Criterion | Weight | What it means |
|---|---|---|
| Reproducibility is stated | 35 | Some published figure for run-to-run variance, so you can tell a real movement from noise |
| A citation is defined | 25 | The platform says in writing what it counts, and whether a name in the text scores the same as a linked source |
| Refresh frequency is published | 25 | How often each prompt re-runs, without which a prompt allowance cannot be converted into a sampling rate |
| Engine coverage is explicit | 15 | The engines are named per tier rather than implied by a logo strip |

Why reproducibility leads. Every other disclosure tells you what the tool does. This one tells you whether its output means anything. It is also the cheapest to publish and the rarest to find, which is a bad combination.
Where it falls short: these criteria reward transparency over capability, deliberately. A platform could publish all four and still extract citations badly. What they test is whether you can interpret the output at all.
AI citation analysis software at a glance
| Platform | Citation defined | Refresh published | Engines explicit | Variance published | Where it falls short |
|---|---|---|---|---|---|
| Otterly | Partly | Yes, daily on all tiers | Yes | No | Engine add-ons can nearly double the bill |
| Writesonic | Partly, via answers per prompt | Yes | Yes, by tier | No | Engine coverage gated to Enterprise |
| Profound | Partly | Yes, daily claimed | Yes | No | Self-serve pricing withdrawn, so scope needs a call |
| Evertune | Not published | Not published | Yes | No | Reputation for rigour is not a disclosure |
| Peec AI | Reported as clean, not published on site | Not published | Yes | No | Nothing verifiable before a demo |
| AthenaHQ | Not published | Not published | Yes, 11 models | No | Credits rather than prompts makes the unit opaque |
| Scrunch | Not published | No | Yes | No | Allowance cannot be converted to a sampling rate |
| Semrush | Partly | Yes, weekly | Yes, four to five | No | A search platform extending into the category |
| Ahrefs Brand Radar | Partly | Yes, monthly | Yes | No | Monthly refresh is coarse for a volatile signal |
| Pepper | Yes, published | Not published | Yes | No | We fail the criterion we weight highest, same as everyone |
**How to read this table against figure 2.** The chart scores strictly, so “Partly” counts as not published. A platform that reports answers per prompt is signalling a refresh rate without defining a citation, and a signal is not a definition. On that strict rule the best score any of the nine reaches is two out of four.
Note the last row. We publish a definition of a citation and we do not publish a variance figure either. On the criterion worth 35 points, this category scores zero across the board, us included.
What this software costs
- Entry tiers start around $29 a month at the cheapest published plans, covered in our comparison of AI SEO tools.
- Mid-market plans run roughly $150 to $500 a month, and the variable that actually changes the bill is engine add-ons rather than prompt count.
- Enterprise platforms are mostly quote-only, and at least one withdrew its published self-serve pricing this quarter.
- The expensive decision is not the licence, it is the migration. Because platforms count differently, switching resets your baseline and you lose the trend you were paying to build.
- Ask for the variance figure before you buy. It costs nothing and it is the single most informative question in the sales process.
How Pepper fits
Pepper is an agentic organic growth engine and an organic growth partner, and we are in this category, so the standard above applies to us.
- We publish a definition of a citation and draw the mention-versus-cited distinction explicitly on our citation analysis page. On that criterion we do better than most of this list.
- We do not publish a run-to-run variance figure either. That is the criterion we weight highest, and we fail it alongside everyone else. It is a fair thing to hold against us.
- Pepper’s GEO platform reports three named metrics rather than one blended score: Brand Visibility for how often engines mention you, Domain Prompt Presence for how often they cite a page from your domain, and Share of Voice for your slice of the category. Keeping mention and citation separate is the whole point, and the gap between the first two is the citability diagnostic. See the platform.
- Agent Atlas keeps the prompt set honest, which matters because the prompt set determines the number more than the platform does. System agents stay fixed, user agents stay editable and versioned, and customers log in and build and run their own inside Atlas.
- Proof rather than adjectives. Acceldata went from 85 to more than 300 top three keywords with 6X organic traffic growth. More in our case studies, and the B2B SaaS practice is where this assessment was built.
Where Pepper fits, and where it does not. We are built for teams who want three separated metrics and a growth team attached, not for a team that wants the cheapest per-prompt tracking. Several tools on this list are cheaper and two of them publish more about their sampling than we do.
How to choose AI citation analysis software
I would ask one question before any demo: what is your run-to-run variance on our prompt set? The answer, or the absence of one, sorts this market faster than a feature matrix.
The framing judgement first, anchored outside my own view. The IAB’s framework is the first neutral standard here, and it says plainly that Directional measurement supports early signal detection but is not suitable for budget decisions. Read that as permission to ask which tier you are buying. Then hold the other half, which is that the framework is voluntary, seven weeks old, and names no compliance mechanism, so nobody is obliged to answer.
The Pepper view on top is narrower. In a category where nobody publishes variance, the vendor who will compute it for you on your own prompt set is demonstrating something the whole market is avoiding.
Here is the 100 point scorecard I would run over any platform, including ours.
| Area | Weight | What a strong platform demonstrates |
|---|---|---|
| They will state variance on your data | 30 | They run your prompt set twice and show you how much the answer moved, before you sign anything |
| Mention and citation are separate metrics | 25 | Two numbers, defined in writing, rather than one blended visibility score with undisclosed weights |
| Refresh frequency is on the pricing page | 20 | You can convert a prompt allowance into a sampling rate without asking a salesperson |
| Raw answers are exportable | 15 | You can see the text the number came from, and take it with you if you leave |
| They place themselves against the IAB tiers | 10 | They will say Directional or Decision-Grade, and justify it, rather than treating the question as unfair |
Then run the live test, on us as readily as on anyone else. Give any platform 25 buying questions from your own category and ask them to run the set twice on the same day. Hold the comparison for 90 days before deciding, because engine behaviour drifts and one reading proves nothing. Ask them for the numbers behind the numbers. For example: “how much did the answer change between the two runs”, “how many times is each prompt run in a month”, “does a brand mention score the same as a linked citation here”, “what happens to our history if we leave”.
Ask them to come back with five things. The variance between two runs. The refresh rate in plain language. The definition of a citation in writing. The export format. And which IAB tier they believe they meet.
A platform that runs your set twice and shows you the difference beats one with a better dashboard. And one that will not define a citation in writing is selling you a number whose meaning it will not commit to.
The weaker way to buy this, and it is the common one. Compare per-prompt prices across three vendors. Pick the one with the most engines. Load a prompt set. Report the visibility score monthly. Discover a year later that the score moves as much on re-runs as it does on work, and that migrating to a better tool would reset the trend.
The stronger sequence. Ask for variance first. Confirm mention and citation are reported separately. Check the refresh rate on the pricing page rather than in a call. Confirm you can export raw answers. Then compare price. The distinction matters because the first sequence optimises for cost per prompt and the second optimises for whether the output can be believed.
Red flags, each one something a platform actually does.
- A single blended visibility score with undisclosed weights, which hides the mention-versus-citation question rather than answering it.
- A prompt allowance with no refresh frequency, which cannot be converted into a sampling rate and so cannot be compared to any other allowance.
- No variance figure, and no willingness to compute one on your data during evaluation.
- Credits instead of prompts, unless the conversion is published.
- Engine coverage shown as a logo strip rather than named per tier.
- No raw answer export, which means you cannot audit the number or take your history with you.
- A guarantee of citations, which Google itself advises against in the equivalent case.
Five questions worth asking, and what a good answer sounds like.
- “What is the run-to-run variance on our prompt set?” A good answer is a number, produced during evaluation. A bad answer is that engines vary.
- “Does a brand mention score the same as a linked citation?” A good answer is no, with two separate metrics. A bad answer is a single score.
- “How many times does each prompt run a month?” A good answer is on the pricing page already. A bad answer requires a call.
- “Can we export the raw answers?” A good answer is yes, in a named format. A bad answer is a dashboard screenshot.
- “Directional or Decision-Grade?” A good answer engages with the IAB tiers. A bad answer has not read them.
If I reduce this to one principle: buy the method, not the dashboard. Every platform here produces a chart, and the chart is the least differentiated thing any of them sells.
The honest note that costs us something. Very few platforms are strong at extraction accuracy, sampling rigour, engine coverage and honest reporting at once, and we would not claim uniform strength across all four either. Sampling rigour is where the whole category is weakest, us included, and it is the one a buyer can force into the open simply by asking.
What nobody should promise you
Nobody should present a citation count as comparable to a competitor tool’s citation count. The platforms define and sample differently, and no standard has been adopted long enough to make them interchangeable.
Nobody should sell Decision-Grade measurement without publishing reproducibility. The IAB put that requirement in writing on 3 August 2026, and the category has had seven weeks to respond.
Nobody should report a single blended visibility score without publishing its weights. A mention and a cited source are different events with different implications, and combining them hides the diagnostic.
Nobody should promise a citation for a fee. Google advises against providers who guarantee rankings, because no external party has access to the ranking systems.
Where this stops working, including for us
If you need a rough signal rather than a defensible number, you do not need most of this software and should not buy the enterprise tier. Directional measurement is genuinely fine and much cheaper. Buy an entry plan, read it as a trend, and keep it out of the board pack.
If you have fewer than about twenty tracked prompts, variance will dominate everything you see, and no platform can fix that with better extraction.
If you are mid-contract with a tool that counts differently from the one you want, the migration cost is the lost baseline rather than the licence, and that usually argues for finishing the year where you are.
Where Pepper falls short. We do not publish a run-to-run variance figure, which is the criterion we weight highest in this article. We publish a citation definition and three separated metrics, and on sampling disclosure two tools on this list currently do better than we do.
Where to go next
Ask your current vendor for the variance on your own prompt set this week. Whatever comes back, or does not, tells you which IAB tier you have been reporting from.
Then check whether mention and citation are separate numbers in your reporting. For the adjacent decisions, our review of fourteen AI visibility platforms ranks the tools themselves, the sampling arithmetic covers how many runs a change needs, which KPIs to use covers what to report upward, and prompt coverage covers the metric underneath it. To see your own citation picture, see where you show up.
Frequently asked questions
What is AI citation analysis software?
Software that runs a fixed prompt set across AI engines on a schedule, captures the answers, extracts which sources were cited, matches brand and competitor mentions, and stores the result as a time series you can trend.
How do AI citation platforms measure citations?
Differently, which is the problem. Some count a linked source, some count a brand name appearing in the text, and some blend both into one score. Only a minority publish which they do.
Is there a standard for measuring AI visibility?
Yes, since 3 August 2026. The IAB’s Measuring Visibility in the AI Era sets out a four-part metric hierarchy and splits measurement into Directional and Decision-Grade tiers. Adoption is voluntary and the framework is new.
What is the difference between a mention and a citation?
A mention means the engine said your name. A citation means the engine used a page from your domain to build the answer. The gap between the two is a diagnostic, and blending them into one score hides it.
Why can’t I compare citation counts between tools?
Because the prompt set, the sampling frequency and the extraction rule all differ, and each changes the number. Migrating platforms resets your baseline, which is usually a larger cost than the licence.
What should I ask a vendor before buying?
Ask for the run-to-run variance on your own prompt set, whether mention and citation are separate metrics, how often each prompt re-runs, whether raw answers export, and which IAB tier they believe they meet.
Do any platforms publish run-to-run variance?
None that we have found. Our review of fourteen platforms found no variance disclosure, and re-checking nine leading tools in September 2026 produced the same result, including for our own platform.
How much does AI citation analysis software cost?
Entry tiers start around $29 a month, mid-market plans run roughly $150 to $500, and enterprise is mostly quote-only. Engine add-ons change the bill more than prompt count does.
Sources and further reading
- IAB, Measuring Visibility in the AI Era, released 3 August 2026, developed with industry partners including Microsoft Clarity. Source of the 4 P’s hierarchy (Presence, Prominence, Portrayal, Persuasion) and of the Directional versus Decision-Grade distinction, including the quoted characterisations of each tier. Limitations: adoption is voluntary, the framework is seven weeks old, and the detailed disclosure requirements sit in a downloadable document rather than on the public page, so this article describes the tiers and hierarchy as published rather than quoting requirements it could not verify at source.
- IAB, the release announcement, 3 August 2026. Source of the definitions of Prominence and of the two measurement tiers.
- Platform public pages, each read on 29 September 2026 and named but not linked, per our policy on competitors. Nine platforms were checked against four disclosures: what counts as a citation, refresh frequency, engine coverage, and any statement of run-to-run variance. Peec AI’s own site was checked directly and publishes no measurement methodology. Scrunch publishes prompt allowances without a refresh frequency. AthenaHQ meters in credits against a proprietary model it does not document. Otterly publishes daily tracking on every tier. No platform published a variance figure. Limitations: this records what each platform publishes, not what it does internally.
- Pepper, the fourteen AI visibility platforms we reviewed. Source of the earlier finding that none of fourteen platforms disclosed run-to-run variance, which this article re-tested.
- Pepper, citation analysis. Source of the mention-versus-citation definition used as the yardstick here, and of our own published position.
- Pepper, AI search tracking and the sampling arithmetic. How many runs a change needs before it clears the noise floor.
- Google Search Central, guide to optimizing for generative AI features, page last updated 10 July 2026. Source of the advice against providers guaranteeing rankings.
Latest Blogs
Healthcare SEO advice has one dominant theme: this is YMYL, so build clinical authority. An analysis of 824,997 health citations shows that advice is correct and badly incomplete. Only 17.3 percent of citations go to authoritative medical sources. Three quarters go somewhere else entirely.
Three questions bring people here and only one of them has a clean answer. Profound no longer publishes prices. The $99 Starter and $399 Growth tiers that most comparison pages still quote, including one of ours, are gone from its pricing page, which now shows a free trial and a custom Enterprise tier. The trial is specified in detail and is genuinely useful: 50 prompts run daily for seven days across three engines, with unlimited seats. On reviews, the counts in circulation range from 140 to over 1,100, so we publish no rating at all.
On 3 August 2026 the IAB published the first industry standard for measuring AI visibility. It splits measurement into two tiers: Directional, which it says supports early signal detection but is unsuitable for budget decisions, and Decision-Grade, which requires rigour on query volume, sample size and reproducibility. We then checked what nine leading citation platforms publish about their own methods. Reproducibility is the requirement almost nobody meets, and our own earlier census of fourteen platforms found not one disclosing run-to-run variance.
Get your hands on the latest news!
Similar Posts

GEO / AI Search
11 mins read
Healthcare SEO services: what actually moves patient acquisition in 2026

GEO / AI Search
15 mins read
Profound pricing, reviews and alternatives: a 2026 buyer’s breakdown

SEO
14 mins read