SEO & GEO
The probabilistic nature of AI answers (and what it means for your visibility)
Copy for AI
For twenty years, SEO lived with a reassuring fiction: search was deterministic. You typed a query, you got a result, and that same result came back tomorrow. You could say “we rank third for this keyword”, and that felt concrete. AI search shatters that idea. Ask a question in an AI search engine, ask it again, and you may get a different answer: different sources, different brands, a different structure. That variation is not a bug. It is baked into the way these systems work.
In this article I explain where that randomness comes from, why it changes everything about how you measure visibility, and why “probability” is the only metric that still matters. It is one of the foundations under the ultimate GEO guide, where GEO stands for Generative Engine Optimization.
Measure it yourself: see how ready your page is to be cited by AI with the free GEO check.
What “probabilistic” actually means here
A language model does not pull a ready-made answer out of a database. It generates text token by token, where each token (a word or a piece of a word) is chosen based on a probability distribution over all possible next tokens.
An example. Suppose the model has written “The best approach to…” and needs to pick the next word. It calculates probabilities for every word in its vocabulary: “optimise” might get 15%, “improve” 12%, “strengthen” 8%, and so on across thousands of options. And here is where the randomness enters: the model does not always choose the most likely token. It samples from the distribution. High-probability tokens are chosen more often, but low-probability tokens remain possible.
You can see that same probability distribution below: the most likely word wins most often, but the model regularly reaches for a less likely word. That is exactly where the variation creeps in.
That sampling is steered by a parameter called “temperature”:
- Temperature zero: the model always picks the most likely token. Fully deterministic.
- Higher temperature: the distribution flattens, lower-probability tokens become more likely. More variation, more creativity.
Production systems almost always run at a temperature above zero, and that is a deliberate choice. Deterministic answers feel robotic and repetitive; variation makes the conversation more human. But that same design choice has enormous consequences for measurement: the same question can produce a genuinely different answer, without anything changing in the underlying data.
Why small variation leads to big differences
Variation at the token level stacks up dramatically into variation at the answer level. Take an answer 500 tokens long. Every token is a sampling decision. Even if every decision has a 95% chance of turning out the same as last time, the probability that the full answer is identical is 0.95 to the power of 500. In practice that is zero.
Concretely this means: ask “what are the best AI SEO strategies?” ten times and you get ten meaningfully different answers:
- Different cited sources
- Different brands mentioned
- Different structure and phrasing
- Different emphasis on the various strategies
- Different level of detail per point
These are not small differences in word choice. One answer might mention your brand prominently, the next in passing, and the third might leave you out entirely. All from exactly the same question.
And there is a second source of variation, separate from the sampling. The grounding process (looking up and selecting sources) introduces its own uncertainty. If you want to understand how that fan-out and selection work, read how AI search architecture works. The core of it: the model splits your question into sub-queries, and that split is itself not deterministic.
Three layers of grounding variation
| Source | Why it varies |
|---|---|
| Query fan-out | The model splits the same question differently on Monday than on Tuesday, because the intermediate reasoning samples differently |
| Timing of retrieval | The web changes constantly: new pages, updates, a news moment that temporarily pushes certain sources up |
| Selection threshold | Selection is not a hard yes/no but a probability: a source with a 60% selection chance is sometimes included and sometimes not |
The consequence: your visibility can change without you doing anything. A competitor publishes something new, a freshness signal shifts, and today you appear less often than yesterday.
Why prompt tracking fails
The first generation of AI visibility tools simply applied the old rank-tracking method to AI answers. The recipe:
- Define a set of prompts
- Run those prompts daily against the AI systems
- Note whether the brand appeared in the answer
- Track that over time as a kind of “ranking”
It feels intuitive and it looks like measurement. But it measures the wrong thing.
When you track a prompt daily, you sample one point from a probability distribution. Today the brand appeared, tomorrow it did not. Was that because your visibility changed, or because you happened to draw a different point from a variable distribution? You cannot know. The metric is so noisy that it becomes meaningless.
Worse still: it creates false confidence. “Brand mentioned: yes” in your dashboard does not mean you are visible for that question. Maybe you only appear 30% of the time and today happened to be one of those times. I have seen teams celebrate “improvements” that were nothing but noise, and teams panic over “drops” that were simply regression to the mean.
The sampling problem
“Fine”, you say, “then we run every prompt a hundred times and take the average.” That helps, but it brings new problems:
- The economics do not scale. Running every prompt a hundred times, across hundreds of prompts, across multiple platforms, and repeating that regularly: the API costs add up fast.
- You are still only measuring specific prompts. Users do not ask your tracking questions verbatim. They rephrase, ask follow-ups, use synonyms. The space of possible questions is effectively infinite.
- The prompts you choose introduce bias. Teams naturally pick questions where they expect to score well. Users do not limit themselves to strategically important questions.
From position to probability: the mental shift
This demands a fundamental shift: stop thinking about position, start thinking about probability.
In classic search the question was “what position do we hold for this query?”. In AI search the question becomes “what is the probability that we appear in answers to questions within this topic cluster, and how prominent are we when we do appear?”. That shift touches not only your measurement but also your media choices, because AI search blurs the line between paid and organic and that has consequences for how you divide your budget.
That reframing changes three things at once:
- From rankings to distributions. Instead of tracking whether you ranked, you track a probability distribution. For a topic cluster you appear, for instance, in 45% of relevant answers, with an average prominence of 0.6 when you do appear.
- From keywords to topic clusters. Instead of individual keyword positions, you track visibility across a cluster of semantically related questions. Your cluster visibility is the aggregated appearance rate across all those variants.
- From snapshots to expected value. Instead of celebrating or lamenting individual observations, you calculate expected value across the distribution. You appear in 50% of questions but with lower prominence than before: has your visibility gone up? Only expected-value calculations answer that.
If you want to dig deeper into this measurement method, measuring brand salience will take you further. Brand salience measures how strongly a model associates your brand with relevant topics, independent of any specific question. It is a property of the model’s knowledge and bias, not of a single answer.
Uncertainty is good news (really)
Here is the counter-intuitive insight: the probabilistic nature of AI search is actually good news for anyone who understands it.
In deterministic systems, advantages are fragile. You rank third, a competitor improves their page, and suddenly you rank fourth. In probabilistic systems, advantages are more robust. You have a 45% appearance probability, a competitor improves their content, and maybe your probability drops to 42%. The change is proportional to the improvement, not a binary flip.
Probabilistic systems also reward consistency over tricks. In classic SEO, a clever hack could push you up temporarily. In AI search, durable visibility comes from genuine quality and authority, because those are the factors that steer the probability distribution over time. Anyone who always understood that SEO was about probability and influence feels validated.
How to build a probabilistic measurement system
If probability is the fundamental metric, what does a solid measurement system look like? In broad strokes:
- Define topic clusters. Each cluster is a space of semantically related questions. Make them complete (covering what drives business value), distinct (minimal overlap) and measurable (small enough for statistical significance).
- Sample with diverse phrasings. Per cluster, generate not just keywords but natural questions, rephrasings and different intents. Record for each query whether you appeared, how prominently and which competitors appeared.
- Analyse the distribution. Aggregate into an appearance rate, a prominence distribution and a competitive position.
- Track trends with statistics. Because you are measuring distributions, you need confidence intervals to separate real change from noise. A jump from 45% to 47% may not be significant; from 45% to 55% probably is.
- Segment. Break visibility down by platform, cluster, intent and language. You can be strong on Google but weak in ChatGPT.
Segmenting by platform is not a detail. Models differ sharply in which brands they know and how they answer, which is why cross-model analysis is a fixed part of a mature measurement approach.
Honest uncertainty
Let me be direct here: we are still early in understanding probabilistic AI visibility. The frameworks I describe are based on careful observation, but they have not yet settled. The tools are still being built, the benchmarks do not exist yet.
But the alternative, pretending AI search is deterministic and applying the old tracking methods, does not give real certainty. It gives false certainty: numbers that feel concrete but represent noise. Honest uncertainty beats confident error, and whoever builds a solid probabilistic measurement system now is building a lead that compounds over time.
Frequently asked questions
Why does AI give a different answer to the same question every time?
Because the model generates text token by token by sampling from a probability distribution, steered by a temperature parameter above zero. It does not always choose the most likely word, so small variations stack up into meaningfully different answers. On top of that, looking up and selecting sources introduces extra variation.
Is prompt tracking completely useless then?
For tracking visibility as a ranking metric it is misleading, because every day you sample one random point from a variable distribution. It can have value as a qualitative check (what exactly does the model say, which sources does it cite), but not as a reliable visibility score. For that you need a sufficient sample and statistical analysis.
What is a topic cluster and why does it matter?
A topic cluster is a group of semantically related questions that customers ask around the same subject. Because the space of possible phrasings is practically infinite, you do not measure visibility per keyword but as an aggregated appearance rate across an entire cluster. That gives a far more stable and commercially relevant picture than individual prompts.
How do I get started with probabilistic measurement in practice?
Start by defining a few topic clusters that drive real business value. Per cluster, generate a diverse set of question phrasings, run them with enough repetition against the relevant AI platforms, and aggregate into an appearance rate with prominence and competitive position. Track that with confidence intervals so you can separate real change from noise.
Need help?
Want to translate this into execution? See how we approach this with AI visibility.
Free website scan
Enter your website and get an automatic scan within minutes, with concrete technical and SEO improvements. No sales pitch.
We only use your details for your scan. No spam, unsubscribe anytime.