A/B Test Significance Calculator
Is the difference real? p-value, uplift and its interval — conversions or averages.
Results
Enter the numbers for the control and at least one variant.
About the A/B Test Significance Calculator
Find out whether the difference between the versions of an A/B test is real or could just be chance. For conversions — visitors who bought, signed up or clicked — the calculator runs the two-proportion z-test; for averages such as revenue per visitor it runs Welch’s t-test from each group’s size, mean and standard deviation (or from the raw values you paste).
You get the p-value and whether it is significant at your confidence level, the uplift with its confidence interval, the difference in percentage points, an optional Bayesian chance to beat the control and the smallest change the test could reliably detect. With very few conversions it switches to Fisher’s exact test. Compare up to five variants with one control — Holm’s correction keeps the chance of a false winner at your α — and every test is checked for a sample ratio mismatch, the traffic-split problem that silently breaks many tests. Everything runs in your browser.
How to use it
- Choose Conversions (each visitor either converted or not) or Averages (a number per visitor, such as revenue per visitor, with non-buyers counted as 0).
- Enter the control in the first group and each variant below it: visitors and conversions, or visitors, average and standard deviation. Add variant adds up to five variants.
- Keep the confidence level at 95% unless you chose another before the test. Use a two-sided test unless you decided in advance that only an improvement matters.
- If the test was not split equally, enter the planned split in Planned traffic split (for example 20/80) so the sample ratio check compares against it.
- Read the verdict, the uplift and its interval. Copy summary puts the result on the clipboard; Download CSV saves every number.
Examples
Control 480 conversions of 10,000 visitors (4.80%) · New checkout 560 of 10,000 (5.60%)
Uplift +16.7% (95% CI +3.6% to +31.4%) · z = 2.548 · p = 0.011 · chance to beat the control 99.5%
Pooled rate 5.2%, standard error √(0.052 × 0.948 × (1/10,000 + 1/10,000)) = 0.00314, so z = 0.008 ÷ 0.00314 = 2.548.
Control 540 / 12,000 · New headline 610 / 12,100 · Shorter form 598 / 11,950
p = 0.049 and p = 0.067 → Holm-adjusted 0.097 and 0.097 → neither significant at 95%
Holm multiplies the smallest p-value by 2 (two comparisons) and the next one by 1, never letting an adjusted value fall below an earlier one. Judged alone, the headline test would have looked like a winner.
Control 0 conversions of 1,000 visitors · Variant 5 of 1,000
Fisher’s exact test: p = 0.062 → not significant at 95% (the z-test alone would say p = 0.025)
With 5 conversions in all, each group expects only 2.5 — too few for the normal approximation. The exact test asks how often all 5 would land in one group by chance: 2 × (1000 × 999 × 998 × 997 × 996) ÷ (2000 × 1999 × 1998 × 1997 × 1996) = 0.062.
821,588 vs 815,482 visitors on a planned 50/50 split
χ² = 22.77, p = 0.0000018 → sample ratio mismatch: find the cause before reading the result
The example Fabijan et al. (2019) use: a 50.2/49.8 split that looks harmless but is very unlikely by chance with this many visitors.
Control 4,820 visitors, mean 3.85, SD 14.2 · Banner 4,790 visitors, mean 4.31, SD 15.6
Uplift +11.9% (95% CI −4.4% to +28.3%) · t = 1.511 · p = 0.131 → not significant
Revenue varies a lot from visitor to visitor, so an 11.9% lift is not yet distinguishable from chance — the test could only reliably detect changes of about ±22% with these numbers.
Common uses
- Deciding whether a new landing page, checkout or headline really converts better before rolling it out.
- Checking revenue per visitor or average order value, not just conversion rate, for a pricing or shipping test.
- Reviewing a test with several variants without being fooled by the one that won by luck.
- Spotting a broken test early: a traffic split that does not match the plan points to a bug in assignment or tracking.
The formulas
- Conversions — two-proportion z-test: z = (p̂_B − p̂_A) ÷ √(p̂(1 − p̂)(1/n_A + 1/n_B)), with the pooled rate p̂ = (x_A + x_B) ÷ (n_A + n_B) (NIST/SEMATECH e-Handbook of Statistical Methods, §7.3.3). The p-value is 2 × P(Z > |z|) for a two-sided test, P(Z > z) for a one-sided one.
- Difference in percentage points: p̂_B − p̂_A ± z × √(p̂_A(1 − p̂_A)/n_A + p̂_B(1 − p̂_B)/n_B).
- Relative uplift: p̂_B ÷ p̂_A − 1, with the log-ratio interval exp(ln(p̂_B/p̂_A) ± z × √(1/x_B − 1/n_B + 1/x_A − 1/n_A)) − 1 (Katz et al., Biometrics 34), which can never go below −100%.
- Very few conversions: when an expected count — visitors × pooled rate, or visitors × (1 − pooled rate) — is below 5 (Cochran’s rule), the z-test’s normal approximation is unreliable. The p-value then comes from Fisher’s exact test: given the totals, it adds up the probabilities of every possible table that is no more likely than the one observed (R documentation of fisher.test); one-sided, the chance of at least as many conversions in the variant. The difference then gets Newcombe’s hybrid score interval (Newcombe 1998, Statistics in Medicine 17), which stays sensible with zero conversions.
- Averages — Welch’s t-test: t = (m_B − m_A) ÷ √(s_A²/n_A + s_B²/n_B) with the Welch–Satterthwaite degrees of freedom (Welch 1947, Biometrika 34). The uplift interval uses the delta method described by Kohavi, Tang & Xu, Trustworthy Online Controlled Experiments (Cambridge University Press, 2020).
- Several variants: each is compared with the control, and Holm’s step-down method (Holm 1979, Scandinavian Journal of Statistics 6) adjusts the p-values: sort them, multiply the smallest by the number of comparisons m, the next by m − 1 and so on, keeping the running maximum.
- Bayesian chance to beat the control: with a uniform Beta(1, 1) prior, each rate’s posterior is Beta(1 + conversions, 1 + non-conversions), and Pr(p_B > p_A) is computed exactly with the closed form in Evan Miller’s Formulas for Bayesian A/B Testing. For averages it uses normal posteriors with flat priors, an approximation.
Reading the result
The p-value is how often a difference at least this large would appear if the versions performed exactly the same. Below α (0.05 at 95% confidence) the result is called significant. It is not the probability that the variant is better, and a significant result can still be a small one — look at the interval: it shows the range of uplifts that fit the data.
Not significant does not mean “no effect”: it means the test cannot tell. The Detectable (80% power) figure shows the smallest change these visitor numbers could reliably pick up; if a smaller change would still matter to you, the test needs more visitors — the sample size calculator tells you how many.
Sample ratio mismatch (SRM)
When the groups get a noticeably different share of the visitors than planned, something in the experiment is broken: a redirect that loses visitors, a bot filter that only hits one version, a tracking tag that fires late. Fabijan et al., Diagnosing Sample Ratio Mismatch in Online Controlled Experiments (KDD 2019), report that about 6% of experiments at Microsoft had an SRM; a chi-square test detects it.
The calculator runs that chi-square goodness-of-fit test on the visitor counts against your planned split (equal unless you enter one) and flags p < 0.01 — the threshold the SRM Checker browser extension used (SRM Checker FAQ). With an SRM, do not trust the result, whichever way it points, until you have found the cause.
Do not stop at the first significant result
These tests assume you looked at the data once, after a sample size you fixed in advance. Checking every day and stopping as soon as the result turns significant makes false winners far more likely: in Evan Miller’s example in How Not To Run an A/B Test, testing after every observation and stopping at the first result significant at 5% (or after 150 observations) declared a winner 26.1% of the time between two identical versions. Decide the sample size first with the A/B test sample size calculator, and run whole weeks so every weekday is included.
Limitations
- Each visitor must be counted once and belong to one group. For conversions, count visitors who converted (not orders); for several orders per visitor use Averages.
- The z-test and t-test rely on normal approximations. With an expected count below 5 the conversion test switches to Fisher’s exact test (which is conservative: its p-values tend to be a little high), and the calculator warns when a group has fewer than 10 conversions or non-conversions, or fewer than 30 visitors for averages. Revenue data with a few very large orders need large samples.
- The p-values are valid for one look at a sample size fixed in advance — not for a test that was stopped when it first looked significant.
- Holm’s correction covers the comparisons with the control. The confidence intervals are for each comparison on its own and are not adjusted.
- The Bayesian chance uses uninformative priors; it says how likely the variant is to be better at all, not by how much.
Privacy
Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot.
Frequently asked questions
What confidence level should I use?
95% (α = 0.05) is the usual choice. Use a higher level such as 99% when a wrong decision is costly or you test many things at once. Choose the level before you look at the results — changing it afterwards to reach significance defeats the purpose.
My variant has a higher conversion rate. Why is it not significant?
Because a difference that size often appears by chance with the number of visitors you have. The interval shows how uncertain the uplift still is (it includes 0%), and Detectable (80% power) shows the size of change this test could reliably find. Keep the test running to the sample size you planned, or plan the next one with the sample size calculator.
What is the difference between the p-value and the chance to beat the control?
The p-value is a frequentist measure: how surprising the data would be if there were no difference. The Bayesian chance to beat the control is the probability that the variant’s true rate is higher, given the data and a flat prior. A 97% chance to beat the control is not the same as p = 0.03 and is not adjusted for several variants — use the significance verdict to decide, and the chance as extra context.
Why does adding variants make significance harder to reach?
Every extra comparison is another chance for one to win by luck. With four variants at α = 0.05, the chance of at least one false winner can be up to 20%. Holm’s correction raises the bar so that this chance stays at 5% overall, while being less strict than the simpler Bonferroni method.
Should I use a one-sided or a two-sided test?
Two-sided, unless you decided before the test that you only care whether the variant is better and would treat “worse” and “no different” the same way. A one-sided test has more power for improvements but cannot detect harm, and switching to it after seeing the data is not valid.
Where do I get the standard deviation for revenue per visitor?
From your analytics or data warehouse (STDEV.S in a spreadsheet), or paste the raw values — one per visitor, with 0 for visitors who bought nothing — into Work it out from the raw values in the Averages mode, and the calculator fills in the visitors, average and standard deviation for you.
Is my data uploaded?
No. All calculations run in your browser; nothing you enter is sent anywhere.