Customer Impact

Growth & Strategie

How to calculate sample size for an A/B test: how much traffic do you need?

Copy for AI

Before you start an A/B test, you should first answer one question: do I have enough traffic to ever pull a verdict out of this? Many teams skip that step, launch a test, and after two weeks stare at a dashboard that says “not significant yet.” By then it is too late. You have lost time on an experiment that was hopeless from the outset. This article shows how to calculate your sample size upfront, so you know whether a test is feasible before you even set it up.

This is the planning side of experimentation. On the page about statistical significance you can read how to interpret a running or finished test. Here everything revolves around what you do before the first visitor sees your variant.

Why calculate your sample size upfront

An A/B test is a bet with your traffic as the stake. You can only put each visitor into an experiment once. If you waste those visitors on a test that is too small to prove anything, you do not get them back. That is why the sample calculation is not an academic exercise, but simply being economical with a scarce resource.

The problem without a calculation comes in two forms. The first: you test too briefly and stop as soon as the dashboard turns green. We call that peeking, and it produces false positives that you leave in production without them actually working. The second: you test endlessly, because the difference you want to measure is so small that you would need months of traffic you will never get. In both cases, a calculation upfront would have saved you from lost time.

A predetermined sample enforces discipline. You agree: this test runs until we have X visitors per variant, and only then do we look at the result. That takes the emotion and the wishful thinking out of the process. For a growth team that steers on pipeline and revenue, that discipline is the difference between learning and gambling.

The three levers that determine your sample

The sample size you need depends on three inputs. Understand these three, and you understand the whole calculation.

1. Your current conversion rate (baseline)

This is the percentage of visitors who already take the desired action: filling out a form, requesting a demo, downloading a quote. Suppose 4 out of every 100 visitors to your landing page submit a request, then your baseline is 4 percent. The lower your baseline, the more traffic you need, because small percentages need more data to estimate reliably.

2. The minimum detectable effect (MDE)

This is the most important and most neglected lever. The MDE is the smallest difference you still consider worth detecting. Do you want to know whether your variant lifts your conversion from 4 to 5 percent? Then your MDE is 1 percentage point. Do you only want to see big jumps, for example from 4 to 6 percent? Then your MDE is 2 percentage points.

Here lies the biggest lever. The smaller the effect you want to be able to see, the more sensitive your test has to be, and the more traffic that costs. A test that has to be able to demonstrate a difference of 1 percentage point demands roughly four times as many visitors as a test that has to see a difference of 2 percentage points. Many teams unconsciously set their MDE too low and then wonder why they need tens of thousands of visitors.

A sensible approach: choose an MDE that is large enough to base a business decision on. An improvement of half a percentage point on a low-traffic page will not change your pipeline. Aim instead for bigger changes that can have a bigger effect.

3. Confidence and power

The third lever consists of two settings that are usually fixed at standard values. Your confidence level (often 95 percent) determines how strict you are on false positives: how often you see a difference that is not there. Your statistical power (often 80 percent) determines how well your test picks up a real difference when it exists. If you make these stricter, for example 99 percent confidence, your required sample rises. For most B2B experiments, 95 percent and 80 percent are a fine starting point.

How the calculation works in practice

You do not need to know the statistical formula by heart. There are free sample size calculators online where you enter your baseline, MDE, confidence, and power, and which work out the number of visitors per variant for you. The value is not in the calculation itself, but in taking the outcome seriously.

The pattern you see every time: a low baseline and a small MDE inflate your required sample enormously. A page that now converts at 2 percent and where you want to prove an improvement of 0.5 percentage points can quickly demand tens of thousands of visitors per variant. If you only have a few hundred visitors per week, that test runs for almost a year. That is not a test, that is waiting.

So turn the reasoning around. First look at the traffic you realistically get on a page in four to six weeks. Then work back to figure out which effect you can detect with that traffic. If it turns out you can only see differences of 3 percentage points or more, you know immediately: subtle tweaks will yield nothing here. Your time is better spent on bigger redesigns or on a page with more traffic.

What if you have too little traffic?

This is the reality for many B2B companies: the traffic is simply too low for classic A/B tests on most pages. That is no reason to stop optimising, but it is a reason to change your approach.

Test bigger changes. With limited traffic you cannot prove a button colour, but you can prove a completely different offer, a radically different headline, or a new page structure. Big changes produce big effects, and big effects can be demonstrated faster with less data.

Concentrate your tests on your busiest pages. One or two pages usually capture the lion’s share of your traffic. There, a reliable test is feasible. On your long tail of pages with few visitors, you are better off steering on reasoned choices and best practices than on statistics.

Or look beyond testing. Sometimes the smartest growth answer is not a better variant of a page, but more qualified traffic going to it. That is exactly where conversion optimisation meets the broader growth engine: SEO, content, paid, and CRO should not run separately, but as one system that steers together on leads and revenue. A test on a page with too little traffic is often a signal that you first need to work on the traffic side.

Sample size belongs in your experiment calendar

Make the sample calculation a fixed step in your experiment process. Before a test goes live, you record: baseline, MDE, required traffic per variant, and estimated runtime. If that does not fit within a reasonable timeframe, the test does not go ahead in this form. This is how you prevent your calendar from filling up with experiments that never reach a verdict.

That discipline is exactly what makes experimentation pay off. A team that tests at random collects a pile of “not significant” and little learning gain. A team that calculates upfront knows which tests are feasible, deliberately chooses bigger effects on busier pages, and thus builds a steady stream of proven improvements. That is the difference between being busy and making progress.

Do you want to embed experimentation in an approach that lets SEO, CRO, content, and paid work together as one whole? As a growth marketing agency we build that growth engine with you, with tests that steer on pipeline and revenue instead of on vanity metrics. Get in touch and we will look together at which experiments are realistic for your traffic.

Free website scan

Enter your website and get an automatic scan within minutes, with concrete technical and SEO improvements. No sales pitch.

Where should we send your report?

We only use your details for your scan. No spam, unsubscribe anytime.