Growth & Strategie
Statistical power for growth experiments: how to avoid false conclusions
Copy for AI
Most A/B tests prove nothing. Not because the idea was bad, but because the test had too little statistical power to see a real effect. You run it for two weeks, spot a difference of a few percent, declare a winner and roll it out. Three months later the gain has evaporated. TL;DR: power and the minimum detectable effect determine whether an experiment can carry a conclusion at all. If you do not plan them in advance, you steer your entire growth program on chance. In this article you will read what statistical power is, how to choose an MDE and why you have to look at this at the program level, not per individual test.
Honest up front: this is the least sexy part of experimentation and, at the same time, the most important. Whoever skips it builds a testing program that looks busy but delivers no reliable learning.
What statistical power actually means
Statistical power is the chance your test actually detects a genuinely existing effect as significant. Suppose your new variant really does improve conversion. Power is then the probability that your experiment picks up that difference rather than missing it.
A low-power test is like a net with mesh that is too coarse: there are fish in the water, but you do not catch them. The effect is there, your test just does not see it. The result is a false negative: you conclude “no difference” while the variant actually wins. You throw away a good idea.
On the other side sits the risk of hasty conclusions. If you test on too little data and stop as soon as you see a nice number, you mainly pick up noise. On a small scale, random fluctuations look like real gains. That is the false positive: you roll out something that in practice delivers nothing.
Power is therefore about the first error (missing real effects), the significance level about the second (mistaking noise for signal). In practice people often work with power of around 80 percent as a lower bound. Below that means you are structurally leaving winners behind without even knowing which ones.
The minimum detectable effect: your most important choice
Before you start a test, you have to answer one question: how large does the effect have to be at minimum to be worth it? That is your minimum detectable effect, the MDE. It is the smallest difference your test can reliably pick up and that is big enough to base a decision on.
The MDE is not a statistical trick, it is a business choice. An improvement of a fraction of a percent on your main conversion may be real, but delivers too little to justify the implementation and the risk. By deciding in advance which effect matters, you avoid spending weeks chasing margins that change nothing about your pipeline.
There is a hard law here: the smaller the effect you want to be able to detect, the more data you need. A large effect shows up quickly, a small effect demands far more visitors and therefore far more time. Four variables are inextricably linked:
- Baseline conversion: how your current variant performs.
- MDE: the smallest difference you want to be able to see.
- Power: the chance you detect that difference.
- Sample size: the number of visitors per variant.
Fix three and the fourth is fixed. Do you want a small MDE with high power on a low baseline conversion? Then you need a lot of traffic. If you do not have that traffic, you have to increase your MDE or accept that the test stays inconclusive. There is no escaping this calculation.
Calculate it before you start
The mistake most teams make: start a test and check afterwards whether it became significant. That is the world upside down. Calculate in advance how many visitors and conversions you need per variant, and how many calendar days that costs with your traffic.
A simple sample-size calculator is enough. You enter your baseline conversion, desired MDE and power level, and you get the number of visitors per variant. Divide that by your daily traffic and you know how long the test has to run. Sometimes the answer is sobering: a test that would take months is not a test you should run. Better to know that now than after three wasted weeks.
Also watch the duration itself. Always run in full weeks so that day-of-the-week effects and buying cycles are included. B2B traffic on a Tuesday behaves differently than on a Saturday. If you stop halfway through a week, that skews your result. And never stop the moment an interim measurement happens to look significant; that so-called peeking at the data inflates your error rate and is one of the biggest sources of false winners.
Why you look at power at the program level
Here is where the real lesson lies, and it is almost always missed. Power is not a property of a single test. It is a property of your entire experimentation program in relation to the traffic you have.
Suppose: you have a modest amount of traffic on your most important landing page. Then you can only finish a limited number of tests with sufficient power per quarter. That is not a shortcoming, that is your reality. The question is not “can I run this one test?” but “which experiments can I reliably finish with my traffic, and which of those have the largest expected impact?”
That framing changes how you prioritize. On low-traffic pages, small optimizations have no chance: the MDE you can realistically detect is so large that only drastic changes clear the bar. There you do not test a button color, but a completely different approach. On high-traffic pages you can work more finely. You allocate your testing capacity to where power and impact coincide.
This is precisely why experimentation does not stand on its own but is part of a larger whole. If you see growth marketing as the system that orchestrates SEO, content, CRO and paid into one growth engine, then traffic is not a given but a lever. More qualified traffic means more tests with sufficient power, means learning faster, means a faster-running engine. Your experimentation program and your acquisition channels reinforce each other.
At the program level you therefore watch a few things:
- Traffic budget per test: deliberately allocate your visitors across experiments instead of testing everywhere at once and underpowering everything.
- Realistic MDEs per page: match the effect you chase to the traffic the page receives.
- Lead time as a decision factor: a test that takes too long to finish does not belong in your roadmap.
- Learning speed over test count: it is not the number of tests that counts, but the number of reliable conclusions.
When you are better off not running a test
Sometimes the most honest choice is: do not test. If you have too little traffic to reach power on a realistic MDE within a reasonable timeframe, then an A/B test delivers not proof but an illusion of proof. In that case you are better off deciding on the basis of rules of thumb, qualitative research and experience, and preserving your testing capacity for pages where it does work.
That is not a capitulation. It is the difference between a team that thinks it works data-driven and a team that really is. False certainty is more expensive than honest uncertainty, because on false certainty you build wrong decisions.
Want to know how many reliable experiments your traffic can handle and which priorities go with that? Our approach as a growth marketing agency starts exactly here: not with isolated tests, but with a program that fits your numbers and steers on leads, revenue and pipeline. For the broader context, read how to interpret statistical significance in A/B testing and how a growth experiment goes from hypothesis to learning.
Conclusion
Statistical power and the minimum detectable effect are not a statistical footnote, they are the foundation under every reliable experiment. Calculate in advance what you need, choose an MDE that matters for your business, and judge power not per test but at the program level in relation to your traffic. That way you stop steering on noise and start learning what really works.
Want to set up an experimentation program that delivers conclusions instead of false winners? Get in touch and we will look together at your traffic, your goals and the tests that genuinely move you forward.
Free website scan
Enter your website and get an automatic scan within minutes, with concrete technical and SEO improvements. No sales pitch.
We only use your details for your scan. No spam, unsubscribe anytime.