A/B Test Sample Size Calculator
Visitors per variant and days to run — worked out before the test starts.
Results
Enter the baseline and the effect to detect.
Add your visitors per day to see how long the test must run.
Visitors needed for other effect sizes
What can a test of a given length detect?
About the A/B Test Sample Size Calculator
Plan an A/B test before it starts: enter your current conversion rate and the smallest improvement worth detecting, and the calculator gives the visitors needed per variant with the standard two-proportion formula — and, from your daily traffic, how many days the test must run. It works for averages too, such as revenue per visitor, from the average and its standard deviation.
Set the confidence level, the power, one- or two-sided testing and the number of variants (with a Bonferroni correction for several), and an unequal traffic split if you use one. Two tables show the trade-off: the visitors needed for other effect sizes, and the smallest effect a test of one to eight weeks could detect at your traffic. The result is a sample size to fix in advance — the number at which you check the result once, instead of stopping the moment it looks significant.
How to use it
- Choose Conversion rate (each visitor converts or not) or Average value (for example revenue per visitor).
- Enter the baseline: the control’s current conversion rate, or its current average and standard deviation per visitor from past data.
- Enter the minimum detectable effect — the smallest change that would make a difference to your decision — as a relative change (+10%) or in percentage points (or units).
- Keep 95% confidence, 80% power and a two-sided test unless you have a reason to change them, and set the number of groups (the control plus each variant).
- Add your visitors per day (and the share of them that enters the test) to see the duration. Round it up to whole weeks, then read the tables to see what a bigger effect or a longer test would mean.
Examples
Baseline 3% · detect +10% relative (3% → 3.3%) · 95% confidence, 80% power, two-sided · 5,000 visitors a day
53,211 visitors per group, 106,422 in total · 22 days → plan 4 full weeks
Baseline 5% → 6% (+20% relative), 95% two-sided (z = 1.959964), 80% power (z = 0.841621)
[1.959964 × √(2 × 0.055 × 0.945) + 0.841621 × √(0.05 × 0.95 + 0.06 × 0.94)]² ÷ 0.01² = 8,157.7 → 8,158 visitors per group
Same 5% → 6% test with 3 variants, Bonferroni α = 0.05 ÷ 3 = 0.0167 per comparison
10,882 per group, 43,528 in total — or 40,143 in total if the control gets 37.6% of the traffic and each variant 20.8%
The control takes part in every comparison, so giving it more traffic than each variant needs fewer visitors overall; the calculator finds the split that needs the fewest.
Average ₹4.20, standard deviation ₹18 · detect +5% (₹0.21) · 95% / 80%
115,331 visitors per group — revenue varies so much from visitor to visitor that small lifts need large tests
Common uses
- Checking before launch whether a test can finish in a reasonable time with your traffic.
- Choosing a realistic minimum detectable effect: the effect table shows what each extra week buys.
- Planning tests with several variants or an unequal split without under-sizing them.
- Explaining to stakeholders why a test has to run for four weeks, not four days.
The formula
For conversion rates the calculator uses the standard formula for comparing two independent proportions without continuity correction (Fleiss, Levin & Paik, Statistical Methods for Rates and Proportions, 3rd ed., Wiley, 2003):
- n per group = [z_α × √(2 p̄ q̄) + z_β × √(p₁q₁ + p₂q₂)]² ÷ (p₂ − p₁)², where p₁ is the baseline rate, p₂ the rate to detect, p̄ = (p₁ + p₂) ÷ 2 and q = 1 − p.
- z_α is z₁₋α/₂ for a two-sided test (1.96 at 95% confidence) or z₁₋α for a one-sided one; z_β is z at the power (0.84 at 80%).
- With an unequal split, r = visitors per variant ÷ visitors in the control: n_control = [z_α √((1 + r) p̄ q̄) + z_β √(r p₁q₁ + p₂q₂)]² ÷ (r (p₂ − p₁)²) with p̄ = (p₁ + r p₂) ÷ (1 + r).
- For averages: n_control = (z_α + z_β)² × σ² × (1 + 1/r) ÷ Δ², with σ the standard deviation per visitor and Δ the difference to detect.
- With several variants each is compared with the control at α ÷ (number of variants) (Bonferroni). The Holm correction in the significance calculator is never stricter than that, so these sizes are enough for it.
This is the same equation R’s power.prop.test() solves: the examples in its documentation — n = 76.7 for 50% → 75% at 90% power, n = 10,451,937 for 50% → 50.1% at α = 0.001 — come out the same here.
For the practice around these numbers — the statistics of online experiments, A/A tests and sample ratio mismatch checks — see Kohavi, Tang & Xu, Trustworthy Online Controlled Experiments (Cambridge University Press, 2020), chapters 17, 19 and 21.
Choosing the minimum detectable effect
The minimum detectable effect (MDE) is the smallest improvement you would act on — not the improvement you hope for. A smaller MDE needs many more visitors: the sample grows with 1 ÷ (p₂ − p₁)², so halving the effect roughly quadruples the visitors needed. The What can a test of a given length detect? table turns the question around: at your traffic, it shows the smallest change one, two, three, four, six and eight weeks could detect. If even eight weeks only detect a +20% change, test a bolder idea, a page with more traffic or a metric earlier in the funnel.
Fix the number in advance — and do not peek
The sample size only works if you analyse the test once, when every group has its visitors. Looking at the result every day and stopping as soon as it turns significant makes false winners far more likely: in Evan Miller’s example in How Not To Run an A/B Test, testing after every observation and stopping at the first result significant at 5% (or after 150 observations) raised the false-positive rate from a nominal 5% to 26.1%. The same article’s rule of thumb, n = 16σ²/δ², is a handy check on the averages mode: at 95% confidence and 80% power the exact factor here is 2 × (1.96 + 0.84)² ≈ 15.7 per group.
Run whole weeks, so that every weekday is in the test — visitors at weekends and on weekdays often behave differently — and keep the test running even if the visitors arrive sooner.
Limitations
- The formulas use the normal approximation. With fewer than about 10 conversions expected per group the sizes are rough, and the version with continuity correction (also in Fleiss et al.) gives somewhat larger numbers.
- The sizes are for a test analysed once at the end. Sequential testing with planned interim looks needs other methods and larger samples.
- For averages, both groups are assumed to have the same standard deviation, taken from past data on the same metric per visitor; revenue with a few very large orders may need more visitors than the formula says.
- The duration assumes steady daily traffic; holidays, campaigns and seasonality change both the traffic and how visitors behave.
- Bonferroni is conservative: with many variants it may ask for slightly more visitors than strictly needed.
Privacy
Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot.
Frequently asked questions
Is the result per variant or for the whole test?
Per group: each variant, and the control, needs that many visitors. The total for all groups is shown next to it. With an unequal split the control and each variant get their own numbers.
What do confidence and power mean?
Confidence (1 − α) protects against false winners: at 95%, a variant that is really no better looks significant only 5% of the time. Power protects against missed winners: at 80%, a real improvement of the size you entered is detected 80% of the time. 95% and 80% are the usual choices.
Why does a low conversion rate need so many visitors?
Because few visitors convert, each group collects few conversions, and small counts are noisy. A +10% lift from a 3% baseline is a change of only 0.3 percentage points — it takes about 53,000 visitors per group to tell it apart from chance at 95% / 80%.
Can I stop the test early if it is already significant?
Not if you want the 95% to mean anything. Stopping at the first significant look inflates false positives several times over. Wait for the planned sample size and whole weeks, then check significance once with the A/B test significance calculator.
Should I always split the traffic 50/50?
With two groups an equal split needs (almost exactly) the fewest visitors in total. With several variants, giving the control a bigger share than each variant can save visitors — the calculator shows that split when it saves at least 2%. If you change the split, enter the planned split in the significance calculator too, so its sample ratio check compares against it.
Should I use a one-sided test to finish sooner?
Only if you decided before the test that you would treat “worse” and “no different” the same way. One-sided tests need about 20% fewer visitors at 95% / 80%, but they cannot detect that a variant is harmful, and the test must then be analysed one-sided as well.