How much traffic do you need to A/B test? It's the first question every team should ask and the one most skip. The honest answer isn't a round number — it depends on your baseline conversion rate, the size of the change you're trying to detect, and how confident you want to be when you call a winner.
When someone asks how much traffic you need to A/B test, they want a single number. The honest answer is that the number falls out of three inputs: your baseline conversion rate (what the control does today), your minimum detectable effect (the smallest improvement worth catching), and your statistical thresholds (typically 95% significance and 80% power). Change any one of those and the required sample size moves. A high-traffic page testing for a tiny lift can need more visitors than a low-traffic page testing for a large one.
Two things quietly inflate the sample size. The first is a low baseline conversion rate: at 2%, you have far fewer conversions per thousand visitors to work with than at 20%, so you need more traffic to separate signal from noise. The second is a small minimum detectable effect: insisting on catching a 2% relative lift requires dramatically more traffic than being willing to ship only on a 10% lift. Halve the effect you want to detect and the sample size roughly quadruples. Most teams set both of these against themselves — low baseline, ambitious small lift — and then wonder why the test never resolves.
A test that can't reach significance in a reasonable window isn't a small test — it's a non-test. It will produce a number, but the number won't mean anything.
The damage from underestimating traffic isn't that the test errors out. It's that it runs, shows an early swing, and tempts you to call it. Peeking at an underpowered test and stopping on a hopeful day is how teams ship changes that do nothing — or quietly hurt. The fix isn't more discipline at the finish line; it's honest math at the start. Decide the required sample size and duration before launch, and commit to running to it.
Run the math before you build the variant. If your weekly traffic and baseline conversion put a meaningful test within a few weeks, go. If the honest sample size is months away, that's not a reason to lower your standards — it's a reason to do different work. Drive more qualified traffic, ship the obvious UX fixes that don't need a test to justify them, and come back to experimentation when the page can actually carry a test. Flight Path inside Optimize Pilot makes that call for you: it detects when a page is too quiet to test and routes you to growth work until the traffic is there.
Enter what you pay Optimizely, Crayon, Hotjar, and Ahrefs today. See what Optimize Pilot would cost instead — and how many headcount the delta covers.
Enter your baseline conversion rate, minimum detectable effect, and weekly traffic. Get the required sample size per variant and an estimated test duration.
How high-performing CRO teams ship more experiments without sacrificing statistical rigor. Includes the idea-to-ship workflow we see work in practice.
Run your numbers through the sample size calculator before you build a variant. If the test can't finish, Flight Path will tell you what to do instead.