· · measure AI search visibilityAI visibility metricscitation share

How to Measure AI Search Visibility: The Metrics That Matter

Presence rate, citation share, recommendation rate, sentiment and accuracy: what each metric tells you, how big a sample you need, and how to avoid acting on noise.

How to Measure AI Search Visibility: The Metrics That Matter

Most AI visibility reporting fails in one of two ways: it collapses everything into a single score that cannot be acted on, or it reports movement that is smaller than the measurement error. Useful measurement needs five distinct metrics, measured per query and per engine, and an honest account of how large a change has to be before it means anything.

Measure five things: presence rate, citation share, recommendation rate, sentiment, and accuracy - each per query and per engine. The gap between presence and citation share is the most diagnostic number available: high presence with low citation share means engines will discuss you but source the claim elsewhere, which is an authority problem rather than a content problem. Small query sets carry wide error bars, so treat small movements as noise. CLEO reports all five across six supported engines, up to four active per site.

Why rank is the wrong instrument

In classical search, rank works because the user sees a list. Position four still earns clicks. In a generated answer there is no list. You are in the text or you are not, and a buyer who never sees you cannot choose you later.

That changes what is worth counting. The question is not where you placed but whether you were included, how often, on what basis, and how you were described when you were.

The five metrics

MetricQuestion it answersWhat a low value usually means
Presence rateHow often do you appear in the answer at allThe engine does not consider you part of the category
Citation shareWhat proportion of cited sources are yoursYou are discussed but not sourced; an authority gap
Recommendation rateHow often does the answer actively suggest youYou are listed as an option but not the preferred one
SentimentHow does the answer frame youCompetitor-authored sources are shaping your description
AccuracyIs what it says about you correctStale or absent sources are being filled in by inference

These are not interchangeable, and the pairs that diverge carry the most information.

The presence-to-citation gap is the number to watch

If there is one diagnostic worth extracting from any AI visibility report, it is the distance between presence rate and citation share.

High presence with low citation share is a specific, common, and frequently misread condition. It means the engines know your brand well enough to mention it, but when they need a source to stand behind a claim, they reach for someone else. That is not a content quality problem, and writing more articles will not close it. It is a retrieval authority problem: not enough independent sources point at your pages for the engine to treat them as citable.

The inverse pattern - low presence, high citation share when present - means the opposite. Your pages are trusted when found, but you are absent from most of the category conversation. That is a coverage and discovery problem, and more content genuinely does help.

Reading these two conditions correctly determines whether a team spends the next quarter writing or earning placements. Getting it backwards is expensive.

Sample size, error bars, and not chasing noise

Generated answers vary between runs. Ask the same question twice and you may get different sources. This means a single check is an observation, not a measurement, and a small set of tracked queries carries a wide error band.

The practical rules are simple and widely ignored:

  1. Judge every reported change against the sample it came from. A movement smaller than the error band is not a result, whichever direction it points.
  2. Use repeated checks per query over time rather than one-off snapshots. Streaks of consecutive outcomes are readable; single observations are not.
  3. Keep the query set stable. Changing the questions between periods makes trend comparison meaningless.
  4. Report confidence, not just percentages. A share figure with no stated confidence level cannot tell you whether the work caused the change.
  5. Treat per-query streaks as the actionable unit. A query losing repeatedly is a specific problem with a specific cause, even when the aggregate share is flat.

Small query sets are still worth running. They are diagnostically useful even when they are not statistically strong, because a repeated loss on one query points at something concrete. The error is quoting the aggregate as if it were precise.

Never blend the engines into one number

Engines retrieve from different indexes and weight different things. Google AI Overviews draw overwhelmingly from pages already ranking in the organic top twenty, which ties that surface to conventional SEO. Others retrieve independently, with no correlation to Google rank, and lean more heavily on third-party corroboration. Recency-weighted engines pick up fresh content fastest.

Only 11% of domains are cited by both ChatGPT and Perplexity simultaneously (The Digital Bloom, 2025). A blended average across engines can look stable while one surface collapses and another improves. Segment first, then aggregate, and never act on the aggregate alone.

Measuring the off-site half

Roughly 85% of brand mentions in AI answers originate from third-party pages rather than brand-owned content (AirOps and Kevin Indig). A measurement programme that only watches your own pages is instrumented for the smaller share of the signal.

The corollary metric is which domains are being cited instead of you. That list is the most directly actionable output of any AI visibility report, because it names the specific places where the category conversation is happening without you. For what to do when those sources are not competitor brands at all, see out-competed by arxiv.org in AI answers.

Where CLEO fits

CLEO's GEO engine reports presence, citation share, recommendation rate, sentiment and accuracy per query per engine, across six supported AI engines - ChatGPT, Bing AI Overviews, Google AI Overviews, Perplexity, Gemini, and Claude - with up to four active per site at a time. The GEO Score aggregates eight weighted metrics into a single trend line, with the components visible underneath so the number can be taken apart rather than taken on faith.

Competitor monitoring covers up to 10 competitors with win rate, citation share, average position, and threat tiers, and surfaces which domains are being cited instead of you. The Citation Loop connects Social activity across X, LinkedIn, Reddit, Medium, YouTube, Quora, and Bluesky to citation data, which is how the off-site 85% becomes measurable rather than assumed.

Results are reported with a stated confidence level: for client DisburseCloud, a payment disbursement platform, CLEO took AI citation share from 17% to 67% in 90 days at 95% Wilson confidence. On CLEO's own site, AI Readability moved from 35 to 96, SEO from 40 to 95, and the GEO Score from 14 to 50 in 30 days. A score that is allowed to fall is the point: see why a score you can trust has to be able to go down.

Frequently asked questions

How do you measure AI search visibility?

With presence rate, citation share, recommendation rate, sentiment and accuracy, measured per query and per engine. Rank does not apply, because an answer either includes you or it does not.

What is the difference between presence rate and citation share?

Presence asks whether you appear at all; citation share asks what proportion of the sources cited are yours. A large gap between them means engines will discuss you but source the claim elsewhere, which is an authority problem rather than a content problem.

How many queries do I need?

More than most teams track. Small sets carry wide error bars and cannot support confident claims about overall share, though they remain useful for diagnosing specific repeated losses. Always judge a change against its sample.

Why do the numbers differ so much between engines?

Different indexes and different weightings. Some surfaces are tied to organic rank, others are not, and some strongly favour recency. Blending them into one number hides exactly the variation you need in order to act.

What is the single most useful number?

The gap between presence rate and citation share. It tells you whether your next quarter should be spent producing content or earning third-party corroboration, and those are very different budgets.

Check your score. Start with your baseline. Enter a domain at regencleo.ai/scan - the engine reads what ChatGPT, Google AI Overviews, and Perplexity say in response to your category's most-asked questions, reads where you sit in classical search, and reads what the social surfaces carry. It returns a single page, no login required. Paid GEO plans extend coverage to Claude.

About this article - How to Measure AI Search Visibility: The Metrics That Matter

Presence rate, citation share, recommendation rate, sentiment and accuracy: what each metric tells you, how big a sample you need, and how to avoid acting on noise.

Article details

Published August 1, 2026 by CLEO. Part of The Field Notes - the working journal of the CLEO Presence Engine at regencleo.ai/articles. Topics covered: measure AI search visibility, AI visibility metrics, citation share, GEO measurement.

Published on The Field Notes at regencleo.ai/articles. Learn more about the CLEO Presence Engine at regencleo.ai/engine. Methodology and scoring at regencleo.ai/methodology.