Customer Impact

SEO & GEO

New metrics for the AI era: the GEO metrics that actually matter

Copy for AI

Most B2B marketers still measure their visibility with a dashboard from 2018: Google positions, organic traffic and click-through rate. The problem is that those numbers say less and less about what is actually happening. An AI answer appears above position one, fully resolves the user’s question, and there is no click to measure. So you can perfectly well “score well” on your old scorecard while becoming invisible in AI search in practice.

By new GEO metrics I mean a set of measures that no longer track clicks or positions, but instead whether your content gets retrieved, selected, cited and correctly understood by language models such as ChatGPT, Perplexity, Gemini and Google AI Overviews. In this article I explain the full measurement stack and give you the formulas and benchmarks you need to get an honest picture.

Want the broader framework first? Then read the ultimate GEO guide, which explains generative engine optimization (GEO) in full. This article zooms in specifically on the measurement side.

Curious how AI sees you? Run your page through our free GEO check.

Why the old metrics are becoming misleading

There is no shame in admitting it: a large part of our classic SEO metrics has degraded into a misleading proxy.

  • Keyword rankings take no account of AI answers that appear above the first organic position.
  • Organic traffic is increasingly cannibalised by AI answers that satisfy the user without a click through.
  • Click-through rate assumes the click is the goal, while visibility inside the AI answer itself often becomes more important.

The old metrics were not bad: they were right for traditional search. But the game has changed and our scorecards have not. So we have to measure what actually determines visibility in AI search: selection, extraction, salience and presence.

The AI visibility stack: four layers

The most useful way to think about this is as a stack, where each layer builds on the one below it. You cannot be selected if you are not retrieved first, and your position against competitors means little without healthy underlying numbers.

LayerWhat it measuresCore metrics
1. FoundationThe basic mechanics of visibilityRetrieval rate, selection rate, citation frequency
2. QualityHow well you perform when you are visibleCitation coverage, extraction depth, primary source rate
3. PositionYour competitive standingCitation share, competitive index, visibility score
4. HealthSustainability and trendSalience score, freshness index, velocity

I will walk through the most important ones layer by layer.

Layer 1: foundation metrics

These are the base measurements everything else is derived from.

  • Retrieval rate = queries for which your content is retrieved / total number of relevant queries. This is the precondition for everything: if you are not retrieved, you cannot be chosen either. A low retrieval rate (below 30%) usually points to content that is not indexed, does not match search intent, or has too few authority signals. Above 60% you are in good shape.
  • Selection rate = selections / retrievals. As far as I am concerned, this is the core indicator. It measures whether the model actually picks your content after retrieving it from the candidate set. Below 25% is weak, above 50% is strong. Low selection despite good retrieval almost always points to quality or extraction problems, or a relevance mismatch.
  • Citation frequency = citations / total tracked queries. This combines retrieval and selection into a single outcome metric. The rule of thumb: citation frequency is roughly retrieval rate times selection rate. If you are retrieved 50% of the time and cited in 40% of those cases, your citation frequency sits around 20%.

Layer 2: quality metrics

These measure how efficiently you capture value when you are actually visible.

  • Citation coverage = extracted characters / total number of characters on the page. How much of your content investment translates into AI visibility? Below 20% is low and often points to bloated content or a structure in which the extractable part is hard to find.
  • Extraction depth = your extracted words / total number of words in the answer. When you are cited, how much of the answer actually comes from you? Below 10% you are a supporting source, above 25% you are the primary voice with maximum brand presence.
  • Primary source rate = primary citations / total number of citations. How often are you the prominently mentioned, first-cited source rather than a footnote? Primary status drives brand visibility, click-through likelihood and authority association.

Layer 3: position metrics

This is all about your competitive standing. Always compare these numbers with competitors, not with absolute thresholds.

  • Citation share = your citations / total number of citations (all sources). This is the AI equivalent of market share. In competitive markets, 30% or more is a dominant position. Segment this by topic cluster, by platform and by period.
  • Competitive index = your citations / a competitor’s citations. Where citation share measures your overall position, the competitive index tracks a head-to-head duel. Above 1 you are cited more often than that competitor, below 1 you are losing.
  • Visibility score = a weighted summary of several factors. A workable formula weights citation frequency (0.3), extraction depth (0.2), primary rate (0.2) and citation share (0.3), normalising each component to a 0 to 100 scale first.

Layer 4: health metrics

This layer measures sustainability and direction. This is where many teams stay blind.

  • Salience score measures the strength of the associations the model holds with your brand, built from mention probability (the chance the model names you unprompted), recommendation confidence and association clarity. This is the parametric foundation of your visibility: is your brand already anchored in the model, independent of what is retrieved live?
  • Freshness index = weighted average age of your visible content, weighted by extraction frequency. Content that gets extracted often weighs more heavily. Averaging older than 12 months is stale, under 6 months is fresh.
  • Velocity = (citations current period minus citations previous period) / citations previous period. Apply this to every metric. Positive velocity means gaining ground, negative velocity is an early warning before it becomes a crisis.

Platform, query and page level

Aggregated numbers hide variation. Three refinements make your measurement usable for optimisation.

  • Platform level. Calculate citation share and selection rate per platform. You can perfectly well dominate on Google AI Overviews and lag behind in ChatGPT. A platform divergence index (the standard deviation of your platform metrics divided by the average) shows where platform-specific opportunities lie.
  • Query level. Track selection rate per query and your rank per query. That instantly shows you which queries you need to defend (you are number 1), attack (you are losing) or conquer (you are absent).
  • Page and section level. Calculate a page visibility score per page and a section extraction rate (how often a section is extracted per time the page is retrieved). That tells you which content structures actually drive extraction and which ones you need to rework.

Five brand-focused GEO metrics as a complement

The stack above mainly measures behaviour at query level. Alongside it, we use five brand-focused indicators that measure how a model understands and trusts your brand. I have described them extensively in the 5 core indicators of AI visibility, summarised briefly:

  • Prompt Recall Rate (PRR): the percentage of category-relevant prompts in which your brand surfaces spontaneously. The generative equivalent of impression share. Emerging brands sit at 5 to 10%, mature brands at 25 to 40%.
  • Semantic Accuracy Index (SAI): how accurately the model conveys your positioning. Does it call you a “trucking app” or a “fleet intelligence platform”? A low SAI signals narrative drift.
  • Source Diversity Ratio (SDR): the ratio between unique, authoritative sources and the total number of references. The more spread out your mentions, the harder you are to train away.
  • Trust Alignment Score (TAS): a composite that weighs the authority and sentiment of your sources. A neutral Wikipedia plus positive G2 reviews plus factual press gives a high TAS.
  • Model Drift Index (MDI): the speed at which your representation changes across model updates. A high MDI means an unstable story and a risk of reputational decay in the model’s memory.

Want to know how to detect and count citations concretely? I work that out in citation mining: measuring AI citations. The parametric side, how strongly your brand is anchored in the model, is covered in measuring brand salience.

From metrics to dashboard and alerts

Isolated numbers are noise. The value emerges when you place them in context.

  • Trends over snapshots. Use a 30-day rolling average to smooth out volatility and year-on-year comparisons to filter out seasonal effects. Detect break points in the trend and tie them to causes: an algorithm update, a competitor’s move or your own optimisation.
  • Benchmark against the right group. Compare your metrics with the category average, with the leader (the leader gap you see closing or widening over time) and with a realistic peer group of comparable players.
  • Set up alerts. A threshold alert (“citation share below 15%”) and anomaly detection (“selection rate deviates 2 standard deviations from the trend”) catch problems early. Just as valuable are competitor alerts when a rival suddenly gains ground.

An executive dashboard keeps it simple: visibility score, citation share, selection rate, trend and one key insight plus one risk and one action. An operational dashboard goes deeper and shows citation share, selection rate and trend per topic cluster. That way leadership and operators each see what they need without drowning in a mush of numbers.

Frequently asked questions

What is the difference between retrieval rate and selection rate?

Retrieval rate measures how often your content is retrieved for relevant queries, so whether you make it into the candidate set at all. Selection rate measures how often the model then actually picks and uses your content after it has been retrieved. High retrieval but low selection points to a quality or extraction problem: you are found, but not found good enough to be cited.

Which metric is most important to start with?

Start with selection rate and citation share. Selection rate tells you whether your content is strong enough to be chosen, and citation share puts that into competitive context, the AI equivalent of market share. Together they give you an honest picture of your position fastest. Then add velocity to keep an eye on the direction.

Do I need expensive tooling to measure this?

Not to get started. Many of these metrics are measured with discipline rather than infrastructure: build a fixed prompt library of 25 to 50 representative queries, run them through the relevant AI platforms periodically and manually count whether you are mentioned, cited and how prominently. For scale and automation, specialised tooling comes in handy later, but the measurement framework is independent of the tool.

How often should I measure these metrics?

AI answers are probabilistic and change with model updates, so a one-off measurement says little. Measure monthly if you can and work with rolling averages and velocity, so you see trends rather than noise. On top of that, set up alerts for sharp deviations, so you do not only notice three months later that your citation share has collapsed.

Need help?

Want to translate this into execution? See how we approach it with AI visibility.

Free website scan

Enter your website and get an automatic scan within minutes, with concrete technical and SEO improvements. No sales pitch.

Where should we send your report?

We only use your details for your scan. No spam, unsubscribe anytime.