Customer Impact

SEO & GEO

Grounding snippets: the atomic unit of AI visibility

Copy for AI

Two thirds of your carefully written content will never be used by an AI answer engine. Not because it is bad, but because an AI does not summarise your page: it extracts from it. That changes everything about how you should build content.

A grounding snippet is the specific passage of text an AI system pulls from a source page to work into its answer. Sometimes that is a single sentence, sometimes a paragraph, sometimes a few paragraphs. But it is never your whole page. The snippet is the atomic unit of transfer between your content and the AI answer. Everything outside it (your intro, your navigation, your call to action) simply does not exist for that one query.

In this article I explain what grounding snippets are, why most content does not survive extraction, and how to write passages that do get selected. It is a core piece of the ultimate GEO guide, which covers generative engine optimization (GEO) from foundations to measurement.

How does your page score? Check your GEO readiness with our free GEO tool.

The 32% problem

Here is a number that should change how you look at content: roughly 32%. That is the average share of your page content that actually makes it into AI answers. Anyone who lays thousands of AI answers character by character next to their source pages sees the same pattern every time: broadly a third survives the extraction process.

That is not a bug or an inefficiency. It is how AI search fundamentally works. The model does not distil your argument and does not summarise your reasoning. It picks out the passages that serve the answer it is constructing at that moment. The rest stays invisible.

That means you no longer write pages, you write extractable units. Units that have to function on their own, compete to be selected, and survive the compression that appears as soon as multiple sources contribute to a single answer.

How the extraction process works

To understand why content falls away, it helps to follow the steps an AI system runs through before it picks a snippet.

That process runs in five steps, with each step narrowing the content further until only the passages that genuinely serve the answer remain. The diagram below follows that funnel, from complete page to selected snippet.

EXTRACTION PROCESS From page to snippet 1 Document retrieval AI pulls in candidate pages 2 Content parsing navigation and boilerplate stripped 3 Chunk identification splitting into separate passages 4 Relevance scoring a score per chunk against the query 5 Selection and extraction the snippet that survives On average, about 32% of your content survives this process.
Each step narrows the content down to the passages AI actually uses.
  • Document retrieval. The AI pulls in candidate documents through its search process. Your page is in the selection: it has been found, opened and is being evaluated.
  • Content parsing. The system strips out navigation, ads and boilerplate. What remains is the actual content body.
  • Chunk identification. The content is split into potential chunks: paragraphs, sections under headings, or semantic units.
  • Relevance scoring. Each chunk gets a score on semantic similarity to the query, presence of relevant terms, information density and structural signals such as headings.
  • Selection and extraction. The model chooses which chunks it actually uses. That choice balances relevance, source diversity, the answer’s length limit and redundancy.

The practical lesson: a high relevance score is necessary, but not enough. Your passage also has to beat the alternatives and fit inside the limited space of the answer.

Why most content does not survive

The 32% figure calls for an explanation. Why does so much content not make it?

  • Query specificity. A query serves one need. Your page may cover ten topics, but the question is about one. The other nine are not extracted, not because they are bad, but because they do not serve this answer.
  • Structural overhead. Intros, transitions, conclusions, author bios: useful for the human reader, but ballast for extraction.
  • Information redundancy. Good writing repeats a key point for emphasis. The AI only needs one formulation.
  • Context-dependent content. A passage like “this approach solves, as discussed above, the first problem” is meaningless in isolation. Which approach? Which problem?
  • Competition from other sources. Even relevant content falls away if a competitor phrases the same thing better. The limited space in the answer means not everyone gets in.

The compression effect

One pattern stands out when you analyse extraction across thousands of answers: the more sources an answer uses, the shorter each snippet becomes. That is the compression effect, and it follows a power law.

If an answer draws on one source, the extracted snippet averages around 1,500 characters. At five sources that drops to roughly 1,100 characters per source. At ten sources it compresses to broadly 1,000 characters.

The maths checks out: answers have a length limit, so more sources means less room per source. The strategic implication is significant. You are not only competing to be selected, you are competing for space inside a bounded answer. And the more competitive the query, the less space you get, even when you are selected. That forces information density: every word has to earn its place.

What a strong snippet looks like

Analysing passages that get extracted consistently surfaces recurring characteristics. This is the anatomy of a good grounding snippet.

CharacteristicWhat it means
Standalone meaningThe passage holds up without surrounding context, without “the above” or “as mentioned”.
Explicit entitiesBrands, products and concepts are named, not replaced with “it” or “they”.
Information densityNo filler words (“it is important to note that”), no superfluous qualifiers.
Direct claim”Pages with a clear heading structure are selected more often” beats a hedged, conditional variant.
Structural boundariesThe snippet coincides with a complete paragraph or section, not with a thought cut in half.

A good example of a standalone snippet names its subject explicitly: “Selection Rate measures how often AI systems choose a specific source from the retrieved candidates.” A context-dependent variant (“this metric, which we introduced in the previous section, differs on a few points”) is not extractable.

Contextual collapse and semantic compression

The most common cause of failed extraction is what I call contextual collapse. A human reading your page builds context while scrolling. Pronouns refer back to what was just mentioned, relative terms like “this approach” work because the reader has the context.

When the AI extracts a snippet, that accumulated context does not travel with it. The passage is torn out of its environment and has to function on its own. “The first challenge is the hardest to solve” is clear to a human, but in isolation nobody knows which challenge that is.

The antidote is semantic compression: making sure the meaning sits inside each extractable unit, instead of being spread across the page. Concretely: name the entity explicitly, even if you just mentioned it. “Visibility measurement is the hardest challenge in AI SEO. To tackle visibility measurement, practitioners need new frameworks.” That repetition feels excessive when reading the full page, but the AI reader has no memory of the previous sentence. What looks redundant to humans is essential for extraction.

The role of headings and the first paragraph

Headings play an outsized role in extraction. AI systems use them to recognise content boundaries and topics. A heading is a label for the text beneath it, and it steers the search for relevant passages.

  • Headings as query match. A section headed “What is Selection Rate?” maps perfectly onto that question. This is not keyword stuffing, but headings that mirror how your audience actually asks.
  • Headings as boundary signal. Chunking algorithms use headings to demarcate sections. The content between two headings forms a natural extraction unit, so every section has to stand on its own.
  • The first-paragraph premium. The first paragraph after a heading has the highest chance of being extracted. Front-load your most important information there, and do not waste that space on transitions or context setting.

I develop this logic further in content architecture for AI extraction, where page structure takes centre stage.

The multi-snippet strategy

Most pages should not optimise for one snippet, but for several. See your page as a portfolio of extractable passages: an overview snippet for general queries, detail snippets for deeper questions, a definition snippet for “what is” questions, a procedural snippet for “how do you” questions and a data snippet for evidence-driven questions.

Snippets on the same page do not compete with each other for different queries. They expand the query surface your page can serve. A page with ten strong extractable passages has ten times the extraction surface of a page with one. And for complex questions, several snippets from your page are sometimes extracted together, for example your definition plus your example.

Which snippet fits which query type is largely settled:

  • Definition snippet (“What is X?”): start with the definition, followed by characteristics and a concrete example.
  • Explanatory snippet (“How does X work?”): start with the mechanism, break it into steps.
  • Comparative snippet (“X versus Y”): set out a comparison framework with the differentiating factors.
  • Procedural snippet (“How do you do X?”): numbered steps with a success criterion.
  • Statistical snippet: lead with the key figure, give the source and the benchmark.

How to win that selection battle systematically is something I cover in Selection Rate Optimization.

What this means for your content strategy

All of this forces a few fundamental shifts away from classic writing:

  • Write for extraction first. Make every section valuable on its own first, and only then take care of narrative flow.
  • Embrace strategic redundancy. Name entities explicitly in every snippet, even if you mentioned them in the previous paragraph.
  • Front-load aggressively. Lead with the core information. Make the first sentence of every section extractable.
  • Granular structure. Give every distinct topic its own heading, so extraction algorithms easily recognise discrete units.
  • Explicit over implicit. The AI does not infer nuance, it extracts text. Let every snippet state its meaning directly.

Grounding snippets are the atomic unit of AI visibility. Anyone who understands how they are extracted, what makes them selectable and how to optimise them transforms their content strategy from “written for readers” into “built to be cited”.

Frequently asked questions

What exactly is a grounding snippet?

A grounding snippet is the specific passage of text an AI system pulls from a source page to work into its answer. That can be a single sentence or a few paragraphs, but never your whole page. It is the building block that travels from your content into the AI answer; everything else on the page does not count for that query.

Why does AI only pull about 32% of my content?

Because an AI does not summarise your page, it extracts from it. A query serves one specific need, so only the relevant passages are used. Structural content (intros, transitions, conclusions) and redundant repetition are skipped, and relevant passages sometimes lose out to better passages at competitors.

How do I make a passage extractable?

Make sure the passage holds up without surrounding context. Name entities explicitly instead of using “it” or “this approach”, cut filler words and hedged phrasing, and make a direct claim. Have the passage coincide with a natural structural boundary, such as a complete paragraph under a heading.

Should I optimise each page for one snippet?

No. A strong page offers a portfolio of snippets for different queries: a definition, an explanation, a procedure, a data point. Those snippets do not compete with each other, they expand your query surface. A page with ten extractable passages has ten times the chance of being cited compared to a page with one.

Need help?

Want to translate this into execution? See how we approach AI visibility.

Further reading

  • How the new AI search architecture works (retrieval, generation and RAG)

Free website scan

Enter your website and get an automatic scan within minutes, with concrete technical and SEO improvements. No sales pitch.

Where should we send your report?

We only use your details for your scan. No spam, unsubscribe anytime.