Growth & Strategie
Peeking and stopping too early: the biggest of all A/B testing mistakes
Copy for AI
You run an A/B test, take a quick look after two days and see variant B ahead by 18%, nicely green, significant. You push the winner live and move on. A month later, conversion hasn’t gone up at all. What happened? Chances are you fell into the peeking trap: looking too early, stopping too early, and mistaking a lucky streak for a real winner. This is the quietest and most expensive of all A/B testing mistakes. In this article, you’ll read why peeking poisons your decisions and how to fix it.
Let’s be honest up front: this isn’t a story about “just run more tests”. It’s about the discipline to finish a test. That discipline is exactly where a growth programme stands or falls, because every false winner you push live erodes trust in all your following tests.
What A/B test peeking actually is
Peeking means repeatedly checking your experiment’s interim results and stopping the test as soon as the outcome suits you, usually the moment the significance number dips below 5%. It sounds harmless. You’re only looking, right? The problem isn’t in the looking, it’s in the deciding based on that look.
A classic A/B test is built on a single assumption: you fix up front how many visitors or conversions you need, you let the test run exactly that long, and you look at the result exactly once, at the end. That entire statistical machinery, including the famous p-value of 0.05, is calibrated for precisely one look at the end.
The moment you check in every day and allow yourself to stop whenever it suits you, you break that assumption. You give yourself dozens of chances to see “significant”, and by pure chance the meter dips below the threshold now and then, even when there is no real difference between A and B at all.
Why stopping too early inflates your error rate
Picture this: two identical variants, no real difference whatsoever. If you look once at the end, you have roughly a 5% chance of accidentally seeing a “significant” difference anyway. That’s the agreed margin of error, and you can live with it.
But conversion rates fluctuate day to day. One day A is ahead, the next day B. If you look again every day and allow yourself to stop the moment the line turns green, sooner or later you’ll grab a random outlier. The more often you look, the greater the chance such an outlier shows up at least once. Your 5% error rate per look stacks up into a far higher chance across the full run time.
The result is a false-positive winner: a variant that wins in the test but delivers nothing in the real world. And because you stopped the test at the peak of its accidental lead, the uplift also looks wildly overstated. You expect 18% more conversion and you get zero. That’s exactly the pattern many teams see time and again: impressive test results that never show up in the revenue figures.
The insidious part is that peeking is rarely bad faith. It’s curiosity plus time pressure. Everyone wants to know whether the new variant works, and nobody wants to wait two weeks when it already looks “obvious” after three days. But that impatience costs you the reliability of your entire programme.
The classic solution: fix everything up front
The simplest way to avoid peeking is to respect the rules of the classic test. That means agreeing on four things before the test goes live:
- Sample size: calculate up front how many visitors or conversions you need per variant to reliably detect a difference of a given size.
- Run time: let the test run for at least one or two full weeks, so you include weekdays and weekends. Buying behaviour on a Monday differs from a Sunday.
- Stopping rule: agree that you only decide once the sample is complete, not earlier, whatever the interim results do.
- One primary metric: choose up front which number determines the winner, so you don’t go shopping around your dashboard afterwards until you find something significant.
This approach works, but it takes patience and a reasonable volume of traffic. How to calculate that sample size and which thresholds to use, you’ll find in our guide on statistical significance in A/B testing. Want the basics of setting up a test first? Then start with our article on A/B testing in practice.
Sequential testing: allowed to look, without blowing up your error rate
This is where it gets interesting. It isn’t that “looking” is forbidden by definition. It’s forbidden in a test that wasn’t designed for it. Sequential testing flips that logic around: it’s a family of statistical methods built precisely to look at interim results repeatedly without your error rate running away.
The idea behind it is that the bar for “significant” gets stricter the more often you look. In return for the right to decide along the way, you pay a price: a higher bar. Early in the test, the difference has to be very convincing before you’re allowed to stop; later on, the bar gradually becomes more reasonable. That way the method automatically corrects for the fact that you’re looking multiple times.
The big advantage is practical. With a real, strong winner you can stop earlier and push the gain live faster. With variants that barely differ, the method forces you to keep going or to close the test out honestly as inconclusive. So you get speed where you can have it and discipline where you need it.
Most modern testing tools offer some form of sequential testing or a Bayesian variant that handles the same problem. The pitfall is that teams use those tools as though they were a classic test: they check in constantly but still read the numbers through the old lens. Pick one method, understand how it handles interim looks, and stick to it.
The real cause isn’t in the statistics
It’s tempting to see this as a purely technical problem. It isn’t. Peeking is first and foremost an organisational problem. It arises when experimentation is buckshot: a little test here, a little test there, with no agreements on when you decide and why.
That’s why a structured growth process works better than a collection of loose tricks. In a mature approach, testing isn’t an impulsive action but a repeatable cycle with clear stopping rules, shared definitions of success and a rhythm that takes the impatience out. Within such a system, conversion optimisation isn’t a standalone channel but one cog in a growth engine where SEO, content, paid and lead gen are aligned. That’s exactly how we as a growth marketing agency approach experimentation: not as a gamble, but as a steered process that drives leads, pipeline and revenue instead of vanity metrics.
Want to understand the broader logic behind that way of working? Then read our pillar page on what growth marketing looks like as a system. There you’ll see how testing, measuring and deciding connect to the rest of your growth programme.
A practical checklist against peeking
Before you push your next test live, run through these points:
- Have you calculated and written down your sample size and minimum run time up front?
- Have you chosen one primary metric that determines the winner?
- Do you know which method your tool uses: classic with one look at the end, or sequential with multiple looks?
- Do you treat an honest “no difference” outcome as a valid option, or does every test feel like something that simply must produce a winner?
- Do you let every test run for at least a full week, so day-to-day fluctuations cancel each other out?
If you answer yes to these questions, the bulk of your false-positive winners disappears by itself. The results become less spectacular on paper, but far more reliable in practice. And that’s the whole point: decisions that still hold up once they’re live.
Conclusion
Peeking and stopping too early feel like efficiency, but they undermine exactly what you run a testing programme for: reliable decisions. A classic test asks for patience and one look at the end. Sequential testing gives you the freedom to check in, provided you respect the stricter bar that comes with it. In both cases, discipline beats impatience.
Want to embed your experiments in a growth process that drives pipeline and revenue instead of accidental uplifts? Get in touch and we’ll look together at how to make your testing programme more reliable and faster.
Free website scan
Enter your website and get an automatic scan within minutes, with concrete technical and SEO improvements. No sales pitch.
We only use your details for your scan. No spam, unsubscribe anytime.