Growth & Strategie
How Long to Run an A/B Test? Plan the Duration Upfront
Copy for AI
How long you run an A/B test decides whether your result is right or not. TL;DR: calculate the duration upfront based on your required sample size and expected effect, then let the test run for at least one to two full business cycles before drawing a conclusion. Stop earlier because a variant looks promising and you are probably measuring noise, not a real difference. In this article you will read how to plan the duration upfront and which mistakes that prevents.
Honest upfront: there is no fixed number that fits every test. Two weeks is a handy rule of thumb, but it is not a law. The right duration depends on how much traffic you get, how big the difference is that you want to prove and what your sales process looks like. Let us make that concrete.
Why you set the duration upfront
The biggest mistake in A/B testing has nothing to do with the tool, but with human behaviour: you look too early, you see variant B is ahead and you stop the test. That is called peeking, and it is the fastest route to the wrong conclusion.
The problem is statistical. Early in a test the numbers swing hard. A variant can be 20 percent ahead after three days and level after two weeks. If you look every day and stop the moment it suits you, sooner or later you will always find a “winner”, even when both variants actually perform equally well. You are then deciding on chance.
The solution is simple in design and hard in discipline: you decide the duration and the required sample before the test starts, and you stick to it. You are free to check in along the way to make sure nothing is broken, but you only draw your conclusion at the agreed endpoint. That turns a test from a gamble into a reliable measurement.
This discipline is exactly what sets growth marketing apart from one-off tricks. Growth becomes predictable when you measure systematically and honestly, not when you interpret optimistically. If you want to understand how testing fits into a bigger system, read our pillar on what growth marketing is as an orchestrating growth engine.
The three factors that determine your duration
Three things together determine how long a test should run. You weigh them against each other upfront.
1. Your sample size
Sample size is the number of visitors or conversions you need per variant to prove a difference reliably. The smaller the effect you want to measure, the more people you need. A difference of 50 percent you can prove with little traffic. A difference of 2 percent demands a far bigger sample.
In practice you use a sample size calculator. You enter three things: your current conversion rate, the smallest effect worth detecting, and your desired confidence level. The calculator gives you the number of visitors per variant. Divide that by your daily traffic and you roughly know how many days your test needs to run.
This is where the reality check kicks in. If you have little traffic and want to measure a small effect, you end up with a duration of months. That is often unworkable. Better to test bolder, more noticeable changes that produce a larger effect, instead of micro-tweaks you will never prove reliably.
2. Full business cycles
Visitors do not behave the same way every day. A B2B audience converts differently on a Monday morning than on a Sunday evening. At the end of the month budgets come into play, at the start of the month other priorities do. If your test only runs a few days, you happen to measure a favourable or unfavourable period.
That is why you always run a test for at least one full week, and preferably two. That way every day of the week carries equal weight. For B2B with a longer orientation phase, two to four weeks is often more realistic, because a prospect rarely converts on their first visit.
Important: also close a test on a full cycle. Stop after ten days and one weekend counts double relative to the working days. Close at seven or fourteen days and every weekday is represented equally often.
3. Your business cycle and sales cycle
For B2B this is the factor that gets forgotten most often. Your conversion on the website is rarely the end of the story. Someone fills in a form, and only weeks later does a quote or a deal come out of it. If you only test on the direct website conversion, you may be measuring more requests of lower quality while actual revenue drops.
So plan your duration partly around your sales process. If you test a change that affects lead quality, your test has to run long enough to see those leads move through the pipeline. A two-week test judged on the number of requests can mislead you if it only becomes clear after six weeks which requests turn into pipeline.
This is exactly why you steer on the right metric. Optimise for revenue, qualified leads and pipeline, not for vanity metrics that look good but bring in no money. A growth marketing agency that builds your growth engine therefore always ties test results back to what happens on the bottom line.
A practical step-by-step plan
Here is how to plan the duration in four steps:
- Decide your target metric. What do you really want to improve? Pick the metric closest to revenue, not the easiest one to measure.
- Calculate your sample size. Enter your current conversion, your minimum relevant effect and your confidence level into a calculator. Note the number of visitors per variant.
- Convert to days and round up to whole cycles. Divide the sample size by your daily traffic. If that works out at twelve days, for example, round up to fourteen so you measure whole weeks.
- Lock the endpoint and do not touch it. Write down the end date and the required sample before you start. Along the way, look only to catch technical errors, not to decide early.
With this approach you avoid the two opposite pitfalls. Too short and you measure noise: you draw conclusions from chance and build on a foundation that is not there. Too long and you waste traffic: every extra week a clear result runs on needlessly is a week you are not working on the next improvement. The optimum sits between the two, and you find it by calculating upfront.
What if you do not have enough traffic?
Many B2B sites simply have too little traffic for classic A/B tests on small effects. That is no reason to stop optimising, but it is a reason to adapt your approach.
In that case, test bigger, clearer changes that produce a larger effect, such as a completely rewritten landing page instead of a different button colour. Bigger effects show up faster. On top of that you can lean on qualitative research, such as session recordings and user interviews, to understand where visitors drop off before you build a variant. That way you do not lose months on a test you will never conclude reliably.
A good test calendar also prevents isolated experiments from running alongside each other without coherence. How to move from isolated experiments to a structural programme, you can read in from isolated A/B tests to a test programme. And if you are still unsure about the basics, what A/B testing exactly is gives you the full foundation.
Conclusion: duration is a choice, not a discovery
You decide the duration of an A/B test upfront, not afterwards. Calculate your sample size, let the test run at least one to two full cycles, take your sales cycle into account and lock the endpoint before you begin. That way every experiment becomes a reliable building block of your growth engine instead of a gamble that happens to pay off or not.
Want to make your experiments part of a system that steers on pipeline and revenue instead of on isolated numbers? Get in touch with us and we will look together at what your test programme should look like for your growth.
Free website scan
Enter your website and get an automatic scan within minutes, with concrete technical and SEO improvements. No sales pitch.
We only use your details for your scan. No spam, unsubscribe anytime.