A/B Test Sample Size Calculator

Work out how many visitors a split test needs before you start it, and how long that will take.

Anything you want recorded alongside the result. The calculation uses the fields below.

200 means 2.00%. Whole numbers keep low rates precise.

Relative: 1000 means a 10.00% lift. Absolute: 100 means 1.00 percentage point.

Optional. Leave at zero to skip the duration estimate.

100 runs left today · sign up for more

About this tool

Calculates the visitors per variant a split test needs to detect a given change in conversion rate, and how long that takes at your traffic.

Do this before the test, not after. Most split tests on small sites are not inconclusive — they were never capable of concluding, and running one for a month to find that out is a month spent. Knowing the number up front turns it into a decision: run it, run it longer, or test something with a bigger expected effect.

The result comes with the sizes for neighbouring effects, because the relationship between them is the thing worth internalising. **Halve the change you are looking for and the sample roughly quadruples.** That is why detecting a 1% lift is usually out of reach and a 20% lift is often quick.

Common questions

Why is the number so large?

Because conversion rates are noisy and small differences hide easily inside that noise. Detecting a 5% relative lift on a 2% baseline needs tens of thousands of visitors per arm. This is a property of the measurement, not of the calculator.

What is statistical power?

The chance of detecting a real effect of the size you specified. At 80% power — the usual default — a real effect will be missed one time in five. Raising it to 90% costs roughly a third more traffic.

Should I use absolute or relative effect size?

Relative is usually how the goal is stated ("a 10% lift"), and it is the safer default because it stays meaningful across baselines. A 1 percentage point absolute change is small on a 20% baseline and enormous on a 1% one.

Can I stop early if it looks significant?

No. That is exactly what fixing the sample size prevents. Repeated checking with an early stop inflates false positives to the point where a test with no real difference will usually cross 95% eventually.