A plain-English guide to the six metrics on the CLEO GEO dashboard: what each one counts, how it is calculated, and what moves it. This is the same methodology the dashboard uses, published in full.
How is the data collected?
Everything on the dashboard comes from one repeatable test run:
- You have a set of tracked queries (real questions your buyers ask).
- Every day, CLEO asks each of those queries to each AI engine on your plan.
- Each answer is stored and analysed for three things:
- Presence: does your brand appear in the answer, or is your site cited as a source?
- Classification: is your brand recommended, merely mentioned, or not mentioned?
- Sentiment: how positively is your brand described, on a 0 to 100 scale.
- Those per-answer facts are rolled up per engine, then averaged across engines to give the numbers you see.
Two rules apply everywhere:
- Presence rule. Your brand counts as “present” only when it is included in the answer text or your domain is cited as a source. A passing string match with no inclusion and no citation does not count.
- Negative answers do not count as wins. If an answer says it cannot find your brand, or actively advises against it, that answer is excluded from the presence-based metrics rather than counted as visibility.
All metrics are on a 0 to 100 scale and are calculated for a single day, then plotted over time.
1. What is AI Presence Rate?
Question it answers: out of every AI answer we tested, how often did your brand actually show up?
AI Presence Rate = (answers where your brand is present / usable answers) x 100
“Usable answers” means all answers for the day across all engines, minus the negative or “brand not found” ones.
Example. 20 queries x 3 engines = 60 answers. 4 are negative or not-found, leaving 56 usable. Your brand appears in 21 of them: 21 / 56 x 100 = 37.5%.
What moves it: publishing content that directly answers the tracked queries, being cited on sources the AI engines trust, and clear entity and schema signals so the engines recognise your brand.
2. What is Grounding Frequency?
Question it answers: how broadly are you visible, engine by engine?
Presence Rate pools every answer together, so a single dominant engine can flatter the picture. Grounding Frequency computes visibility inside each engine first, then averages the engines equally. It is your breadth-of-coverage measure.
Per engine: visibility = (queries where your brand appears / queries tested on that engine) x 100
Overall: Grounding Frequency = average of the per-engine visibility figures
Example.
| Engine | Queries tested | Queries with your brand | Visibility |
|---|---|---|---|
| ChatGPT Search | 20 | 12 | 60% |
| Google AI Overviews | 20 | 4 | 20% |
| Perplexity | 20 | 8 | 40% |
(60 + 20 + 40) / 3 = 40.0%
How to read it: if Grounding Frequency is much lower than AI Presence Rate, your visibility is concentrated in one or two engines. That is an engine-specific gap, not a content-quality problem.
3. What is Recommendation Rate?
Question it answers: when you do appear, are you being endorsed or just listed?
This is the strongest commercial signal on the dashboard. Every answer is classified into one of three labels:
- RECOMMENDED: the AI endorses, suggests or positions your brand as the solution (“I would recommend X”, “X stands out because…”).
- MENTIONED: your brand is named with no endorsement (“competitors include X, Y and Z”).
- NOT_MENTIONED: your brand does not appear at all.
Recommendation Rate = (RECOMMENDED answers / classified answers) x 100
Example. 21 answers include your brand: 6 are RECOMMENDED, 15 are MENTIONED: 6 / 21 x 100 = 28.6%.
How to read it: high presence with a low Recommendation Rate means you are visible but not persuasive. The fix is proof (comparisons, specifics, third-party validation, review signals), not more content volume.
4. What is Share of Voice?
Question it answers: in the sources the AI actually leaned on, how much of the citation space is yours versus your tracked competitors?
This is the one zero-sum metric: it moves relative to your rivals.
Share of Voice = (your citations / (your citations + tracked competitor citations)) x 100
Three deliberate design choices:
- Only tracked competitors count. Wikipedia, news sites, directories and other neutral third parties are not in the denominator. Otherwise every brand would look tiny.
- The top three competitors cap the denominator. Only your three most-cited tracked rivals are counted, so adding more competitors to your watchlist for analysis never mechanically deflates your score.
- No tracked competitors means 100%. With nobody in the denominator, the metric has nothing to compare against.
Example. Across the day’s answers your domain is cited 18 times. Your three most-cited competitors are cited 22, 14 and 9 times (45 total): 18 / (18 + 45) x 100 = 28.6%.
Fair-share adjustment (inside the score only). The headline Share of Voice on the dashboard is the raw figure above. When Share of Voice feeds the overall GEO Score, it is first indexed against the fair share for the field size, so a strong performer in a crowded market is not punished for having more rivals:
fair share = 100 / (competitors + 1)
indexed value = min(100, raw x (competitors + 1))
With three competitors, fair share is 25%. A raw 28.6% is above fair share, so it indexes to 100 (capped). Indexing can only raise the value, never lower it, and it does not change the number displayed on the card.
Your competitor list controls this metric
Share of Voice is the one metric you directly influence from the dashboard. On the Competitors tab you can add, edit and remove the brands CLEO measures you against, and that list is the denominator of the formula above.
| Action | Effect on Share of Voice |
|---|---|
| Add a competitor that gets cited a lot | Goes down, if that domain enters your top three most-cited rivals |
| Add a competitor that is rarely or never cited | No change, it sits outside the top three |
| Remove a competitor from your top three | Goes up, its citations leave the denominator |
| Remove a competitor outside the top three | No change |
| Track no competitors at all | Reads 100%, there is nothing to compare against |
Because only your top three most-cited tracked rivals count, you can safely track up to 15 competitors for analysis and reporting without each extra name dragging the score down.
Changes apply from the next daily run, not retroactively. Each day’s metrics are calculated and stored as a fixed snapshot at run time, so your trend chart stays honest.
Practical guidance: track the rivals you actually compete with for the same buyer. Adding a large publisher or marketplace that happens to be cited often will depress the score without telling you anything useful about your market position.
5. What is Share of Citation?
Question it answers: how often does an AI answer actually cite your website as a source?
Presence is about being named. Citation is about being used as evidence, with a link back to you.
Share of Citation = (answers that cite your domain at least once / total answers) x 100
An answer counts once whether it cites you one time or five times, so a single heavily-linked answer cannot distort the figure.
Example. Of 60 answers, 11 cite your domain: 11 / 60 x 100 = 18.3%.
Alongside the percentage, the dashboard shows a breakdown of the top 10 cited domains across those answers, so you can see exactly which sources the engines are pulling from instead of you.
What moves it: crawler access for AI bots, original data and statistics worth quoting, clean citable page structure, and coverage on the domains that already dominate that breakdown list.
6. How is Sentiment scored?
Question it answers: when the AI talks about you, how favourable is the language?
Each answer that mentions your brand is scored 0 to 100 (0 is the most negative, 50 is neutral, 100 is the most positive). The dashboard figure is the average across all scored answers for the day.
| Score | Reading |
|---|---|
| 60 and above | Positive |
| 40 to 59 | Neutral |
| Below 40 | Negative |
Aspect sentiment. Each answer is also scored separately on three aspects, judged independently and averaged only over the answers that gave a basis to judge them: Quality (reliability, performance, capability), Service (support, responsiveness, experience), and Price (value for money, affordability, positioning). This is why a brand can sit at 72 overall but 48 on Price. That split is usually the most actionable thing on the card.
How do the six metrics roll up into the GEO Score?
The headline GEO Score is a weighted blend of the six metrics, built around a funnel: get found, win the recommendation, reinforce with authority.
| Stage | Metric | Weight |
|---|---|---|
| Found (40%) | AI Presence Rate | 20% |
| Grounding Frequency | 20% | |
| Won (45%) | Recommendation Rate | 25% |
| Share of Voice | 20% | |
| Reinforcing (15%) | Share of Citation | 10% |
| Sentiment | 5% |
GEO Score = sum of (metric value x its weight)
Missing data does not drag you down. If a metric has no data for the day, it is dropped and the remaining weights are re-scaled to total 100%. You are never scored against a blank.
Worked example. Presence 37.5, Grounding 40.0, Recommendation 28.6, Share of Voice 100 (indexed), Share of Citation 18.3, Sentiment 72:
37.5x0.20 + 40.0x0.20 + 28.6x0.25 + 100x0.20 + 18.3x0.10 + 72x0.05
= 7.50 + 8.00 + 7.15 + 20.00 + 1.83 + 3.60
= 48.1
How much should you trust a single day’s score?
The score sits next to a Data Quality badge telling you how much to trust it on any given day. The badge is deliberately kept out of the score itself: it is a confidence signal, not a performance signal. It reads two things:
- Citation Drift: how much your daily visibility swings across the last 30 days. High drift means unstable AI answers, so any single day is a weak read.
- Confidence Interval (margin of error): how tight the estimate is, which depends mainly on how many queries you track. More queries, tighter interval.
| Badge | Condition | Meaning |
|---|---|---|
| High (green) | Drift under 15% and margin of error 5% or less | The score is reliable |
| Moderate (yellow) | Drift 15 to 30%, or margin of error 5 to 10% | Scores may shift; more query volume would sharpen it |
| Low (red) | Drift above 30%, or margin of error above 10% | Not enough stable data for a confident read |
If you are sitting on a Low or Moderate badge, judge trends over weeks rather than reacting to a single day’s number, and consider increasing your tracked query volume.
Quick reference
| Metric | Counts | Divided by | Best used for |
|---|---|---|---|
| AI Presence Rate | Answers where you appear | All usable answers | Overall discoverability |
| Grounding Frequency | Per-engine visibility | Averaged across engines | Breadth across engines |
| Recommendation Rate | RECOMMENDED answers | Classified answers | Endorsement strength |
| Share of Voice | Your citations | You plus top 3 tracked rivals | Competitive position (you control the rival list) |
| Share of Citation | Answers citing your domain | All answers | Source authority |
| Sentiment | Per-answer 0 to 100 scores | Averaged | Tone and perception |
Reading order we recommend: start with Presence and Grounding to see whether you are being found. If those are healthy but Recommendation Rate is not, the gap is persuasion. If Share of Voice is the weak point, the gap is competitive. If Share of Citation lags everything else, the gap is authority and crawlability.
Related reading: the GEO Score reference, A Score Allowed to Fall on why the score uses outcome measures and confidence intervals, and the Presence Score reference for how GEO visibility feeds the wider Presence Index.