↑↓ to move ↵ to open Esc to close Browse all tools

A/B Test Significance Calculator

Enter visitors and conversions for both variants to see whether B really beats A, with the uplift, p-value, confidence interval and a sample size planner.

Your split test

Control A
Variant B
Confidence level
Test
Two-sided also tells you when B is worse. Pick before you look at the data.
Relative uplift of B over A
+16.00%

Conversion rate A
Conversion rate B
Absolute uplift
p-value
z-score
95% interval for B − A

Sample size helper

How many visitors each variant needs before you start, to detect the lift you care about. Uses the confidence level and test type above.

%
% relative
Power

Runs on your device, nothing you enter is sent to a server.

What statistical significance tells you

In an A/B test, or split test, you show two versions of a page, email or ad to similar visitors and count how many of each group convert. Variant B will almost never land on exactly the same conversion rate as A, even if the change made no difference at all, because visitors behave randomly. Statistical significance answers one question: is the gap big enough that chance alone is an unlikely explanation?

The p-value is the probability of seeing a difference at least this large if A and B actually performed the same. A small p-value means the result would be surprising under "no real difference", so you conclude that B is really better (or worse). At 95% confidence the p-value has to be below 0.05.

Conversion rate
Conversions ÷ Visitors
Relative uplift
(Rate B − Rate A) ÷ Rate A
z-score
(Rate B − Rate A) ÷ SEpooled

This calculator runs a two-proportion z-test. The pooled standard error is √(p × (1 − p) × (1 ÷ nA + 1 ÷ nB)), where p is the combined conversion rate of both groups. The z-score is turned into a p-value with the normal distribution. The confidence interval for the difference uses each group's own rate.

An example

Variant A had 10,000 visitors and 500 conversions, a 5.00% conversion rate. Variant B had 10,000 visitors and 580 conversions, 5.80%. That is an absolute uplift of 0.80 percentage points and a relative uplift of 16%. The z-score is 2.50 and the two-sided p-value is 0.0123, below 0.05, so B wins with 95% confidence. The 95% confidence interval for the difference runs from about +0.17 to +1.43 points: B is very likely better, but the real lift could be much smaller than the 0.80 points you measured.

Common mistakes

  • Peeking and stopping early. If you check every day and stop the moment the result turns significant, you will crown far more false winners than your confidence level suggests. Fix the sample size first and stick to it.
  • Samples that are too small. With a few hundred visitors only huge differences can reach significance, and the ones that do are often flukes that shrink later. Use the sample size helper before you launch.
  • Running too short. Visitors on a Monday morning behave differently from those on a Saturday night. Run tests for whole weeks, even if you reach the sample size sooner.
  • Testing many things at once. Compare ten variants or ten metrics and one will look significant by luck. Pick one main metric before the test.
  • Treating "not significant" as "no difference". It only means the data cannot tell yet. The confidence interval shows how large or small the true effect could still be.

How many visitors you need

The sample size depends on your baseline conversion rate, the minimum detectable effect (the smallest relative lift worth finding), the confidence level and the power. Power is the chance of detecting a real effect of that size; 80% is the common default. Visitors needed per variant at 95% confidence, two-sided, 80% power:

Baseline rateLift to detectVisitors per variant
2%10% relative80,682
2%20% relative21,109
5%10% relative31,234
5%20% relative8,158
10%10% relative14,751
10%20% relative3,841

Halving the effect you want to detect roughly quadruples the traffic you need, which is why small sites are better off testing bold changes than small tweaks.

How to use the A/B test significance calculator

  1. 1Enter the number of visitors and conversions for variant A (the control) and variant B.
  2. 2Choose the confidence level (90%, 95% or 99%) and a two-sided or one-sided test.
  3. 3Read the conversion rates, uplift, z-score, p-value and the verdict. Before your next test, use the sample size helper to plan how long to run it.

Frequently asked questions

What does statistical significance mean in an A/B test?

It means the difference between the variants is large enough that random chance alone would rarely produce it. At 95% confidence, a difference this big would show up less than 5% of the time if A and B really converted at the same rate. It does not tell you how big the true effect is, which is what the confidence interval is for.

What p-value counts as significant?

The p-value has to be below 1 minus your confidence level: below 0.10 at 90%, below 0.05 at 95% and below 0.01 at 99%. Choose the level before you start the test. 95% is the usual default; use 99% for changes that are expensive to roll out or hard to undo.

Should I use a one-sided or a two-sided test?

Two-sided is the safe default. It checks whether B is different from A in either direction, so it also tells you when B is worse. A one-sided test only asks whether B is better, which makes it reach significance sooner, but it cannot detect a loss and is easy to misuse when chosen after looking at the data.

Can I stop my test as soon as it shows significance?

No. Checking the result every day and stopping at the first significant reading, often called peeking, makes false winners much more likely. Decide the sample size up front with the sample size helper, run the test until you reach it, and ideally cover full weeks so weekday and weekend visitors are both included.

How many visitors do I need for an A/B test?

It depends on your baseline conversion rate and the smallest lift you care about. Small rates and small lifts need a lot of traffic: at a 5% conversion rate, detecting a 10% relative lift with 95% confidence and 80% power takes about 31,234 visitors per variant. Enter your own numbers in the sample size helper.