Data & Tracking
Why your GA4 reports lie: data sampling explained
Copy for AI
Data sampling is what happens when Google Analytics does not process all of your data, but takes a sample and extrapolates your report from it. The problem: as soon as you add segments or filters above a certain threshold, GA4 switches to sampled data and the exact same numbers can hand you opposite conclusions. Short TL;DR: watch the shield icon at the top of your report, because green means full data and yellow means a subset you had better not base a revenue decision on. In this article you will read when sampling kicks in, how to spot it and how to work around it.
At Customer Impact we steer on customers and revenue, not on vanity numbers in a dashboard. And sampled data is exactly that kind of vanity number: it looks convincing, but it can send you in the wrong direction. The honest advice agencies rarely give you, we do give here.
What is data sampling in Google Analytics?
Data sampling means that Google Analytics does not calculate every session individually, but takes a representative sample and scales the outcome up to your full traffic. For a giant dataset that is a logical choice: it keeps your reports fast. The downside is that the outcome becomes an estimate, not an exact count.
In practice you mostly notice it when you dig deeper. A standard report without filters usually runs on full data. But as soon as you build an ad hoc analysis with segments, pick a long period or add complex filters, Google may decide to sample. The same data, processed differently, with different numbers as a result.
If you want to make your tracking and reporting fundamentally reliable, that is exactly the work that falls under our data analytics services: making sure the numbers you steer on are correct.
When does GA4 switch to sampled data?
GA4 does not sample at random. It happens with ad hoc queries that cross an event threshold within the selected period. Important: unlike the old Universal Analytics, GA4 counts in events, not in sessions. According to the documentation on data sampling in GA4, that threshold sits at 10 million events per query on the free Standard version, and at 100 million events by default on Analytics 360 (expandable to a maximum of 1 billion events for more detail).
Important to understand: it is not about your total traffic ever, but about the number of events within the date range and query you request. A busy webshop hits those 10 million events in no time, but a B2B site that pulls a year of data in one go with multiple segments can cross it too.
Sampling mainly shows up in:
- Explorations with many dimensions and segments over a long period.
- Standard reports you add a comparison or filter to that makes the query complex.
- Long date ranges combined with detailed breakdowns.
The standard overview reports you look at daily are usually safe. The risk sits in the moment you want to answer an interesting question, because that is exactly when you run those heavy queries.
How do you spot sampled data in GA4?
Watch the shield icon. At the top of your report in GA4 sits a small shield that indicates data quality. A green shield means your report is based on 100% of the data. A yellow shield means GA4 has applied sampling and you are looking at a subset.
Click that shield and Google shows exactly what percentage of the sessions was used. If you see, for example, that a report runs on 40% of the data, the margin on your numbers is substantial. The lower that percentage, the bigger the chance your conclusion wobbles.
The habit of always glancing at that shield before interpreting a report is one of the simplest yet most valuable reflexes you can pick up. It costs two seconds and saves you from wrong decisions. In a well designed marketing dashboard setup you build that check in by default. If you want to know which channels deliver your best B2B customers, those are precisely the reports where you want to work on full data.
How much can sampled data distort your conclusions?
Far more than you think. In an anonymised client example we saw how the same regex on organic sessions told two completely opposite stories, depending on whether the data was sampled or not.
Applied as a segment, the report showed +13% organic sessions year over year (6,754 versus 5,986). The exact same regex, but applied as a filter on unsampled data, showed -9% year over year (6,012 versus 6,600). So one view says your organic traffic is growing, the other says it is shrinking. Same period, same site, same regex.
Placed side by side, the reversal becomes painfully visible: on sampled data the line climbs, on full data it drops.
Imagine deciding to raise your SEO budget based on that +13%, while reality is a 9% decline. You would be steering on a fiction. This is exactly why at Customer Impact we insist so much on reliable data before a single euro gets reallocated. A dashboard that grows on paper but shrinks in reality costs you real customers.
How do you prevent sampling in your reports?
You cannot fully switch off sampling in the free version, but you can strongly reduce the chance of it. A few practical moves:
- Shrink your date range. Request a shorter period and the sessions stay under the threshold, so GA4 does not need to sample. Pulling three quarters separately is often more reliable than pulling a full year at once.
- Simplify your query. Fewer segments and fewer dimensions at once means a lighter calculation. Build your analysis step by step instead of cramming everything into one exploration.
- Use filters instead of segments where you can. As the client example showed, filters in standard reports sometimes run on full data while segments sample. Test which approach gives you the green shield.
- Export the raw data. For anyone who wants it truly exact, the BigQuery export of GA4 is the gold standard. There you get every event row unsampled and you calculate your own numbers. This is the same layer a solid conversion tracking setup leans on.
The common thread: the lighter and more focused your question, the better your odds of working with full data.
Why does this matter especially for B2B?
You would think sampling is only a problem for sites with enormous volumes. For B2B it is subtler and therefore more treacherous. A B2B site often has relatively little traffic, but traffic that counts heavily: every session can be a potential customer worth thousands of euros.
That is why small distortions make a big difference. A difference between +13% and -9% across a handful of channels can determine whether you scale a campaign up or shut it down. At a webshop with millions of sessions such an error often averages out, but for you there is a concrete decision attached to it that directly touches your roas.
We do not work for webshops or e-commerce, but for B2B companies that steer on lead quality. That means we would rather run three precise reports than one broad report leaning on a 40% sample. A small, fast team that truly understands the data isolates the truth quicker than a big agency that blindly trusts a dashboard. If you want to know whether your own tracking is correct, also read what is tracking as a starting point.
Frequently asked questions about data sampling in GA4
Can I switch off data sampling completely in GA4?
Not fully in the free version. You can avoid sampling by using shorter periods and simpler queries, or by choosing the least sampled option within an exploration. For guaranteed unsampled data you need the BigQuery export or Analytics 360.
What does the yellow shield icon in my report mean?
The yellow shield indicates that GA4 has applied sampling and that your report is based on a subset of the data. Click it to see what percentage of the sessions was used. A green shield means you are looking at 100% of the data.
From how many events does GA4 sample?
According to Google, GA4 counts in events, not in sessions. The threshold for ad hoc queries sits at 10 million events per query on the free Standard version, and at 100 million events by default on Analytics 360 (expandable to 1 billion). It is about the number of events within the requested period and query, not about your total traffic.
Are standard reports in GA4 sampled too?
Most standard overview reports run on full data. The risk of sampling arises as soon as you add comparisons, segments or complex filters, or build an exploration over a long period.
Is sampled data always wrong?
Not necessarily, a sample can sit close to reality. But you do not know in advance how large the deviation is, and as the client example showed, that deviation can be big enough to flip your conclusion entirely. For decisions with budget attached, you are better off working with full data.
Stop steering on numbers you cannot trust
Sampled data looks just as convincing as real data, and that is exactly the danger. If you base your marketing decisions on a 40% sample, you are effectively gambling with your budget. The solution is not a complicated tool, but the right reflexes: look at the shield, simplify your queries and pull in the raw data where needed.
We help B2B companies set up their tracking and reporting so that every decision rests on full, reliable data. No vanity numbers, but real customers and revenue. Book your free intake.
Free website scan
Enter your website and get an automatic scan within minutes, with concrete technical and SEO improvements. No sales pitch.
We only use your details for your scan. No spam, unsubscribe anytime.