Customer Impact

SEO & GEO

Citation mining: measuring your AI citations systematically

Copy for AI

When I tell clients that an AI model makes millions of choices every day about which sources deserve visibility, they nod. When I then ask which sources get cited in their category, the room goes quiet. Most companies are flying blind. They ask ChatGPT a question now and then, happen to spot a citation and form an opinion from it. That is not a measurement method, that is guesswork.

Citation mining solves that. In this article I explain what it is, which metrics you get out of it and how to set it up yourself. It belongs in the ultimate GEO guide, where I map out the full playing field of generative engine optimization (GEO).

Test your AI visibility: score your page in half a minute with our GEO check.

What citation mining is

Citation mining is the systematic extraction and analysis of the citations AI systems show in their answers. Every time a model cites a source, that is a judgement: this piece of content deserves visibility for this question. Collect and organise thousands of those judgements and a rich information source emerges.

That data tells you:

  • Which domains AI systems trust for which topics.
  • Which specific pages win selection for competing questions.
  • How citation patterns shift over time.
  • Where you stand relative to your competitors.
  • Which content characteristics correlate with selection.

The difference with scattered observations is method. Instead of remembering that you were “cited once by Perplexity”, you structurally measure how often and how prominently you appear, and who you beat.

What the data shows you

When I aggregate citations for a topic, patterns surface that you would otherwise never see.

Domain authority

Count all the citations for a topic and you immediately see who dominates. Often those are the usual suspects (Google, Microsoft, AWS) plus one or two specialised players who have built up authority. If you sit at 2 percent share, you know there is work to do, and exactly who you are up against.

Winners per question

Within a topic, different questions have different winners. Definition questions (“what is X?”) cite encyclopedic sources. Comparison questions (“best X tools 2026”) cite review sites. Implementation questions cite documentation. That nuance decides where you can realistically win. There is no point challenging Wikipedia on a definition question if you score better on the implementation angle.

Content characteristics

By dissecting cited pages, you see what models value: average length, heading density, use of lists, recency of the last update, presence of schema markup. If models mostly cite pages that were recently updated while your article is two years old, you have a concrete, measurable action on your hands.

The metrics that matter

Citation mining only delivers value once you translate it into numbers you can track over time. These are the formulas I use, based on the source material.

MetricFormulaWhat it measures
Citation shareyour citations / total citationsYour share of all citations within a query set
Citation frequencyquestions where cited / total number of questionsHow often you show up at all
Primary citation rateprimary citations / total citationsHow often you are the main source when you are cited
Citation velocity(citations now - citations previous period) / citations previous periodThe speed at which your position is changing
Competitive citation indexyour citations / competitor citationsWhether you are cited more or less often than a rival

A few interpretations I consider important:

  • High frequency, low share means you show up often but rarely prominently. You are present, not dominant.
  • Low frequency, relatively high share means you appear rarely, but when you do, you dominate. A sign to expand your coverage.
  • Citation velocity is your early warning system. A negative velocity betrays a loss of position before it becomes visible in your traffic.
  • A competitive citation index above 1 means you are cited more than the competitor, below 1 that they lead. Track this against several rivals for a complete picture.

These citation metrics go deeper than the 5 core indicators of AI visibility, such as prompt recall rate. Where prompt recall rate measures whether you are mentioned at all, citation mining zooms in on the source attribution itself.

How to measure it: the process step by step

You do not need heavy infrastructure to start. Discipline matters more than tooling. This is the approach I stick to.

Step 1: define your query corpus

Decide which questions you want to track. A good corpus mixes three types:

  • Core questions about your brand and products (“what is [product]?”, “[brand] review”, “[brand] vs. [competitor]”).
  • Category questions where you want visibility (“best [category] software”, “how to tackle [category task]”).
  • Intent questions around the problems you solve (“how do I solve [problem]?”, “[pain point] solutions”).

Aim for 100 to 500 questions, depending on your market. Tip: pull real search queries from Google Search Console and convert them into conversational form. Those are the queries people actually ask.

Step 2: run the queries

Put your questions to the relevant platforms (Google AI Mode, ChatGPT with search, Perplexity, Claude with search) and record for every answer:

  • The full answer text.
  • All citations (URLs).
  • The position of every citation (in the text or in a source list).
  • A timestamp and the platform.

Set your frequency by importance: daily for competitive intelligence, weekly for trends, monthly as a minimum for a baseline.

Step 3: extract the citations

Pull the URLs out of every answer. Normalise them to domain level (all variants of microsoft.com count together), so you can aggregate correctly. Decide upfront whether you count unique or total citations when the same source appears several times.

Step 4: analyse and build a database

For ongoing intelligence, store every citation in a structured way. The essential fields are: a unique ID, the query and its category, the platform, the timestamp, the cited URL and the normalised domain, the position in the answer, the citation type, whether you were the primary source, and the specific passage of text (the grounding snippet) that pointed to your page.

With that database you run share analyses, query-to-domain mappings, trends and gap analyses. A relational database with indexing on your queries is more than enough to start.

From snippets to content optimisation

Beyond URLs, it pays to analyse the text that actually got cited. For every citation you can retrieve the grounding snippet: the exact passage from your page that ended up in the answer. Do that across many citations and you see what models do extract (definitions, statistics, process descriptions, comparisons) and what they ignore (introductory filler, marketing language, vague claims).

That is where citation coverage comes in: the share of a page that actually gets cited, calculated as cited characters divided by total characters. An 800-word page of which 290 words get cited reaches a higher coverage than a hefty 2,400-word page of which only 380 words get picked up. High coverage points to efficient, extractable content. Low coverage betrays bloated text or a structure the model cannot work through.

From intelligence to action

Citation mining generates insight, but insight without action is wasted effort. I always close the measurement loop:

  1. Identify an opportunity in the data: questions where you have zero citations but should be competing, or pages with declining citations.
  2. Take action: create new content, refresh outdated pages, or analyse why a competitor is suddenly winning.
  3. Remeasure the citations after some time.
  4. Attribute the change to your action.
  5. Refine your approach based on the result.

Visually, that measurement loop looks like this, where every round starts at an opportunity and ends at a refinement that sharpens the next round:

MEASUREMENT LOOP From insight to remeasurement repeat & accelerate 01 Opportunity find a gap 02 Action adjust content 03 Remeasure count citations 04 Attribute link the effect 05 Refine adjust approach The loop turns citation mining into an active optimisation engine.

That loop turns citation mining from a passive report into an active optimisation engine. If you want to bundle these insights for stakeholders, read how to create an AI visibility report in which citation data takes centre stage.

Frequently asked questions

What is the difference between citation mining and regular rank tracking?

Rank tracking measures positions in classic search results. Citation mining measures which sources AI systems cite in their generated answers, including the exact passages they extract. It is a separate discipline because AI answers do not work with ten blue links, but with a selection of sources woven into one coherent answer.

How many queries do I need to start?

For a meaningful picture, aim for 100 to 500 questions, depending on the breadth of your market. If you simply want to get going, 25 to 50 carefully chosen prompts already lay a usable baseline. More important than the number is that you measure the same set repeatedly, so you can see trends instead of snapshots.

Do all AI platforms cite in the same way?

No, and that is exactly why you measure per platform. Some systems cite several sources per claim, others prefer one comprehensive source. Some weigh recency heavily, others value source diversity. If you are strong on one platform but weak on another, that is a platform-specific opportunity to adapt your content.

Should I build this myself or buy a tool?

That depends on your resources. Build it yourself if you want deep custom intelligence, have engineering capacity and data ownership is crucial. Buy a tool if time-to-value counts and standard functionality is enough. A common middle road is a hybrid: an external tool for collection and basic analysis, and your own analyses for the competitive intelligence that gives you an edge.

Need help?

Want to translate this into execution? See how we approach it with AI visibility.

Free website scan

Enter your website and get an automatic scan within minutes, with concrete technical and SEO improvements. No sales pitch.

Where should we send your report?

We only use your details for your scan. No spam, unsubscribe anytime.