Short answer

A statistical significance calculator tells you whether the gap between two conversion rates in an A/B test is bigger than chance alone would usually produce. Convertica's free calculator runs a two-proportion z-test on the visitors and conversions for control and variant, and reports the p-value, the lift and a confidence interval for the difference.

  • A result is significant when the p-value is below the significance level you chose before the test started.
  • A p-value is not the probability that the variant is better.
  • Read the confidence interval for the size of the lift, and compare its low end with the smallest lift worth shipping.
  • Use a two-sided test unless you chose one-sided before launch, and do not stop a test the first time it looks significant.
  • Check for sample ratio mismatch before you trust any result.

Check your test result

Free to use, no sign-up. The result updates as you type.

Control (A)
Variant (B)
Test settings

The traffic split you set in your testing tool, used for the sample ratio mismatch check. The result warns when the visitor counts do not fit it (chi-square p below 0.001).

Two-sided p-value

0.0171

Statistically significant at the 95% level: the p-value is below 0.05.

Variant 2.50% vs control 2.00%: +25.0% relative lift, +0.50 percentage points.

95% confidence interval for the difference: +0.09 to +0.91 percentage points.

z = 2.38. 1 minus p is 98.29%, which some tools call "confidence". It is not the probability that the variant is better.

Sample ratio check: 50.0% of visitors are in control against the 50/50 split you set (chi-square p = 1.0000). No mismatch at the 0.001 threshold.

Formula

r1 = c1 / n1    r2 = c2 / n2
q  = (c1 + c2) / (n1 + n2)
SE = sqrt(q(1 - q) * (1/n1 + 1/n2))
z  = (r2 - r1) / SE
p  = 2 * (1 - Phi(|z|))   two-sided
p  = 1 - Phi(z)           one-sided

CI  = (r2 - r1) ± Zc * SEu
SEu = sqrt(r1(1-r1)/n1 + r2(1-r2)/n2)

n is visitors, c conversions, q the pooled rate and Phi the standard normal distribution. Zc is 1.96 for a 95% interval.

Get your free CRO audit

Enter your website and email and the audit starts right away. Watch it check your page live: eight checks, each scored out of 100, and three fixes you can make now.

How to use the statistical significance calculator

  1. Enter visitors and conversions for control (A) and variant (B), from your testing tool or analytics. Use the same date range and the same definitions for both groups.
  2. Choose the significance level you set before the test started. 95% is the default.
  3. Choose the test type. Two-sided is the default and also catches a variant that is worse. One-sided only asks whether B is higher.
  4. Read three things, in order: whether the p-value is below your alpha, the confidence interval for the difference, and any warning under the result.

The calculator assumes each visitor was randomly assigned to one group and counted once, each visitor either converted or did not, and the test ran to a sample size you planned in advance (a fixed-horizon test). It also needs enough data: if a group has fewer than about 5 expected conversions or non-conversions, it shows a warning, because the normal approximation behind the test breaks down.

What does statistical significance mean in an A/B test?

Statistical significance in an A/B test means the gap between control and variant is larger than random variation would usually produce if the change had no effect at all. You set the bar in advance, the significance level, usually 5%. If the p-value falls below it, the result is called statistically significant.

The logic runs backward from how most people think about it. The test starts by assuming the variant makes no difference (the null hypothesis). It then asks how likely a gap this large would be under that assumption. If the answer is "rarely", you reject the assumption and treat the difference as real. Significance is a filter against being fooled by noise, not a measure of how good the change is.

Two kinds of mistake are possible. A false positive (type I error) is calling a winner when the change did nothing; the significance level caps how often that happens. A false negative (type II error) is missing a change that really works; you reduce that with a larger sample, which is what statistical power is about.

What is a p-value, and what is it not?

A p-value is the probability of seeing a difference at least as large as the one in your test if control and variant truly converted at the same rate, assuming the test was run correctly. A small p-value means your data would be unusual under no difference. It is evidence against "no difference", not the chance that your variant wins.

What a p-value is, and the common misreadings
A p-value is notWhy
The probability the variant is betterIt is calculated assuming there is no difference, so it cannot give the probability that there is one. That needs a Bayesian model.
The probability the result is due to chanceIt is the probability of data this extreme if chance were the only thing at work. That is a different conditional.
"Confidence" when written as 1 minus pIn the default example 1 minus p is 98.29%. That is not a 98.29% chance that B wins, even though some tools label it that way.
The size of the effectA tiny lift on a huge sample can have a smaller p-value than a large lift on a small one. Read the confidence interval for size.
Proof the change worksAt a 5% level, about 1 in 20 tests of changes that do nothing will still come out significant.

What is the difference between the significance level and the confidence level?

The significance level (alpha) and the confidence level are the same threshold seen from two sides: the confidence level is 1 minus alpha. A 95% confidence level means alpha is 0.05, so you accept a 5% false positive rate when the change has no effect. That is why "significant at 95%" and "significant at the 5% level" mean the same thing.

The level also sets the critical z value: how many standard errors apart the two rates must be before the result counts as significant. These are the values the calculator uses.

Confidence level, alpha and the critical z value
Confidence levelAlphaCritical z, two-sidedCritical z, one-sided
90%0.11.6451.282
95%0.051.9601.645
99%0.012.5762.326

Pick 95% unless you have a reason not to. Use 99% when a wrong call would be expensive or hard to reverse, such as a pricing or checkout change, and accept that it needs more visitors. Pick the level before the test starts; choosing it after you see the p-value turns the threshold into a formality.

How do you read the confidence interval for the lift?

The confidence interval for the lift is the range of true differences between variant and control that your data are consistent with, at the level you chose. The calculator reports it in percentage points. If the whole interval sits above zero, the variant is significantly better; if it includes zero, the test has not ruled out "no difference".

In the default example the observed difference is 0.0050 (+25.0% relative), and the 95% interval runs from +0.09 to +0.91 percentage points. Divided by the control rate of 2.00%, that is roughly +4.4% to +45.6% in relative terms. Roughly, because the control rate is itself an estimate. A significant result, then, is still compatible with a lift far smaller than the one you saw.

Two habits help. Report the interval, not only the point estimate, when you share a result. And compare its lower end with the smallest lift that would be worth shipping: if even the low end clears that bar, you have a result you can act on with some comfort.

The formula: a two-proportion z-test

The test statistic uses the pooled standard error, which assumes no difference between the groups. The confidence interval uses the unpooled standard error, which uses each group's own rate.

  • r1 and r2 are the conversion rates of control and variant.
  • q is the pooled rate: all conversions divided by all visitors.
  • z is how many standard errors apart the two rates are.
  • p comes from the standard normal distribution: twice the tail area beyond |z| for a two-sided test, the upper tail beyond z for a one-sided test.
  • The confidence interval is the observed difference plus or minus Zc standard errors, where Zc is 1.96 for 95%. For a one-sided test the calculator reports a one-sided lower bound instead.

Because the test and the interval use slightly different standard errors, a result right on the edge can be significant while its interval just touches zero, or the other way round. Away from the edge they agree.

Worked example

Control: 200 conversions from 10,000 visitors. Variant: 250 conversions from 10,000 visitors. Two-sided test at 95%. These are the default values in the calculator above.

  1. Rates: r1 = 200 / 10,000 = 0.0200 (2.00%) and r2 = 250 / 10,000 = 0.0250 (2.50%). Difference = 0.0050, a relative lift of +25.0%.
  2. Pooled rate: q = (200 + 250) / (10,000 + 10,000) = 450 / 20,000 = 0.0225.
  3. Standard error: SE = sqrt(0.0225 × 0.9775 × 0.0002) = sqrt(0.0000043988) = 0.0020973.
  4. z = 0.0050 / 0.0020973 = 2.3840.
  5. p = 2 × (1 - Phi(2.3840)) = 2 × 0.008563 = 0.017126, which rounds to 0.0171.
  6. 0.0171 is below 0.05, so the result is statistically significant at the 95% level.
  7. Confidence interval: SEu = sqrt(0.0200 × 0.9800 / 10,000 + 0.0250 × 0.9750 / 10,000) = sqrt(0.0000043975) = 0.0020970. Margin = 1.959964 × 0.0020970 = 0.0041101. Interval = 0.0050 ± 0.0041101 = 0.0008899 to 0.0091101, or +0.09 to +0.91 percentage points.

So the variant's observed lift is +25.0%, and the data are consistent with a true difference anywhere from about +0.09 to +0.91 percentage points. The real lift could be much smaller than the one observed.

Six results compared: what changes the verdict

The table below runs the same calculator on six sets of inputs at the 95% level. Each row changes one thing from the default example, so you can see what moves a result in or out of significance: the amount of traffic, the test type, and the direction of the difference.

Six A/B test results through the same calculator, 95% level. Intervals in percentage points.
ScenarioTestControlVariantRelative liftp-valueIntervalVerdict
The calculator's default exampleTwo-sided200 / 10,000 (2.00%)250 / 10,000 (2.50%)+25.0%0.0171+0.089 to +0.911Significant
Same data, one-sided testOne-sided200 / 10,000 (2.00%)250 / 10,000 (2.50%)+25.0%0.0086+0.155 or moreSignificant
Same rates, a tenth of the trafficTwo-sided20 / 1,000 (2.00%)25 / 1,000 (2.50%)+25.0%0.4509-0.800 to +1.800Not significant
Very large test, small liftTwo-sided10,000 / 500,000 (2.00%)10,300 / 500,000 (2.06%)+3.0%0.0334+0.005 to +0.115Significant
Variant below control, two-sidedTwo-sided250 / 10,000 (2.50%)200 / 10,000 (2.00%)-20.0%0.0171-0.911 to -0.089Significant
Variant below control, one-sidedOne-sided250 / 10,000 (2.50%)200 / 10,000 (2.00%)-20.0%0.9914-0.845 or moreNot significant
  • Traffic decides more than the lift does. The same +25.0% lift is significant with 10,000 visitors per group (p = 0.0171) and nowhere near it with 1,000 per group (p = 0.4509), where the interval runs from -0.800 to +1.800 points.
  • One-sided halves the p-value when the variant is ahead (0.0171 becomes 0.0086), which is why it is tempting to switch after the fact, and why you should not.
  • A huge test can make a small lift significant. A +3.0% relative lift is significant (p = 0.0334) with 500,000 visitors per group, but the low end of its interval is +0.005 percentage points. Whether that is worth shipping is a business question, not a statistical one.
  • A one-sided test cannot see a loss. The variant below control is significant two-sided (p = 0.0171), but a one-sided test for an increase gives p = 0.9914: it was never looking that way.

Should you use a one-sided or a two-sided test?

Use a two-sided test unless you decided on one-sided before the test began. A two-sided test asks whether the variant differs from control in either direction, so it also flags a variant that hurts conversions. A one-sided test only asks whether the variant is higher, which gives a smaller p-value for a winner and no warning for a loser.

A one-sided test is defensible when a worse result and no result lead to the same action, for example you will keep the control either way, and you wrote that down before launch. It is not defensible as a way to rescue a result that just missed significance two-sided: that doubles your real false positive rate.

What is sample ratio mismatch, and how do you check for it?

Sample ratio mismatch (SRM) means the visitors in each group do not match the split you set, beyond what chance explains. If you set 50/50 and got a clearly uneven split, something in assignment or tracking is broken, and the significance result cannot be trusted, however small its p-value. Check the split before reading the result.

The check is a chi-square goodness-of-fit test on the visitor counts. With two groups it reduces to a z-test on the split. Example: you set 50/50 and recorded 10,000 visitors in control and 10,500 in the variant (48.8% in control).

N        = 10,000 + 10,500 = 20,500
expected = N × 0.5 = 10,250 per group
SD       = sqrt(N × 0.5 × 0.5) = 71.59
z        = (10,000 - 10,250) / 71.59 = -3.49
chi²     = z² = 12.20
p        = 2 × Phi(-|z|) = 0.0005

A p-value of 0.0005 says a split this uneven would be very unlikely from random assignment alone, so this test has a sample ratio mismatch. Because the check is run on every test, many teams use a strict threshold for it, such as 0.001, to avoid false alarms. The calculator above runs this check on your visitor counts every time: set the expected split if your test does not divide traffic 50/50, and it shows a sample ratio mismatch warning under the result when the p-value is below 0.001.

Common causes of SRM:

  • A redirect or a slower-loading variant, so some visitors leave before the variant records them.
  • Bots or internal traffic filtered from one group but not the other.
  • The tracking tag firing on a different event, page or condition in each group.
  • Assignment that depends on something else, such as a cookie banner choice, a logged-in state or a cached page.
  • Changing the traffic split, or pausing a variant, partway through the test.

If you find SRM, fix the cause and rerun the test. Do not try to correct the numbers afterward.

Can you stop an A/B test as soon as it reaches significance?

No, not with a fixed-horizon test like this one. Checking a running test again and again and stopping at the first p-value below 0.05 is called peeking, and it makes false positives far more likely than the 5% you set. The p-value naturally wanders up and down while data come in, and given enough looks it will dip below the line by chance.

Three ways to stop a test honestly:

  1. Fix the sample size in advance with the A/B test sample size calculator, run until you reach it, and read the result once. This is what this calculator assumes.
  2. Run whole weeks. Visitors on a Monday morning rarely behave like visitors on a Saturday night. Ending a test partway through a weekly cycle can bias it toward whichever days happened to be included, even when the sample size is reached.
  3. Use a method built for monitoring. Sequential tests, such as group sequential designs or the always-valid p-values some testing tools offer, are designed to be checked as data arrive. They adjust the threshold for each look. Use their own result, not this calculator, when you test that way.

Looking at a running test to catch a broken variant or a tracking failure is fine. Deciding the winner from those looks is the problem.

What is the difference between statistical and practical significance?

Statistical significance says a difference is probably not noise. Practical significance says the difference is big enough to matter to the business. They are separate questions: a very large test can make a lift too small to be worth the work statistically significant, and a small test can miss a lift that would be worth a lot.

Before the test, write down the smallest lift that would justify shipping the change, given what it costs to build and maintain. After the test, compare the confidence interval with it:

Reading a result against the smallest lift worth shipping
What the interval showsWhat it meansTypical decision
Entirely above the smallest worthwhile liftSignificant and large enough to matterShip it
Above zero, but includes lifts too small to matterProbably real, possibly not worth muchShip if cheap to keep; otherwise weigh the cost
Includes zero and also large liftsInconclusive: the test was too small to tellRerun with a planned, larger sample, or move on
Narrow and close to zeroAny effect is probably too small to matterKeep control and test something bolder
Entirely below zeroThe variant is worseKeep control and note what you learned

Can you test revenue or average order value with a significance calculator?

Not with this one. A statistical significance calculator for conversion rates treats every visitor as converted or not converted. Revenue per visitor and average order value are continuous numbers, usually skewed by a few large orders, so they need a test built for them: Welch's t-test on a large sample, or a bootstrap, often with extreme orders capped first.

It matters because the conversion rate and revenue can move in different directions. A variant that gets more people to buy can also push them toward cheaper items, and a discount banner can raise orders while lowering margin. Decide before the test which metric decides it, and check the others for harm. Convertica's case studies report more than one metric for exactly this reason. The dScryb test, for example, measured paid membership sign-ups as the main metric and revenue per visitor alongside it:

+49.4%

Paid membership sign-ups, 17-day A/B test

Source: dScryb, VWO A/B test, 98% confidence

What happens when you test several variants or metrics?

Every extra comparison is another chance for a false positive. With one comparison at alpha 0.05, a change that does nothing comes out significant 5% of the time. Compare three variants with control, or check ten metrics, and the chance that at least one crosses the line by luck alone is much higher. This is the multiple comparisons problem.

Chance of at least one false positive when nothing really changed, at alpha 0.05, if the comparisons are independent
ComparisonsChance of at least one false positiveBonferroni alpha per comparison
15.0%0.05
29.8%0.025
314.3%0.0167
522.6%0.01
1040.1%0.005
2064.2%0.0025

The figures are 1 minus 0.95 to the power of the number of comparisons. Real metrics are often correlated, so the true figure is usually lower, but the direction holds. Three ways to stay honest:

  • Name one primary metric before the test starts and decide on that. Treat the rest as guardrails or ideas for the next test.
  • Correct for several variants. The Bonferroni correction tests each comparison at alpha divided by the number of comparisons. It is strict but simple. The sample size calculator has a Bonferroni option so you can plan for it; then use the matching level here. With five comparisons, for example, Bonferroni gives 0.01, which is the 99% setting.
  • Treat segment wins with suspicion. "It won on mobile in Canada" after a flat overall result is usually a false positive. Retest the segment on its own if it matters.

Bayesian vs frequentist A/B testing

This calculator is frequentist: it asks how surprising your data would be if there were no difference. Bayesian testing tools start from a prior belief about likely lifts and report the probability that the variant is better, and the expected loss if you pick it. That answer is closer to what most people want to know, but it depends on the prior chosen, and it is not immune to peeking or multiple comparisons just because it is phrased as a probability. With sensible priors and enough data, the two approaches usually agree on which variant to ship. The bigger risks are the ones above: stopping early, broken splits and testing too many things at once.

Checklist: before you call an A/B test

  1. The test reached the sample size you planned before it started.
  2. It ran for whole weeks, covering your normal business cycle.
  3. The visitor split matches the split you set (no sample ratio mismatch).
  4. Both groups were counted the same way: same events, same dates, visitors or sessions in both.
  5. You are reading the primary metric you chose in advance, at the level you chose in advance.
  6. The p-value is below your alpha.
  7. The confidence interval's lower end is above the smallest lift worth shipping, or you have accepted that it might not be.
  8. Guardrail metrics, such as revenue per visitor or refunds, did not get worse.

Two more mistakes to avoid

  • Treating the observed lift as the lift you will get. Significant results from small tests tend to overstate the true effect, because only the lucky high readings clear the bar. Plan on the low end of the interval, not the point estimate.
  • Testing a page that has no clear problem. A test only answers the question you ask it. The most useful tests start from evidence about where visitors get stuck, which is what a CRO audit is for.

Frequently asked questions

What does a statistical significance calculator tell you?

A statistical significance calculator tells you how surprising your A/B test result would be if control and variant really converted at the same rate. This one runs a two-proportion z-test on visitors and conversions and returns the p-value, whether it is below your significance level, the relative and absolute lift, and a confidence interval for the difference.

Is an A/B test significance calculator the same thing?

Yes. An A/B test significance calculator is a statistical significance calculator for two conversion rates, which is what this page is. A sample size calculator is a different tool: it tells you how many visitors to plan for before the test starts. Use the A/B test sample size calculator first and this calculator when the test has finished.

What is a p-value?

The p-value is the probability of seeing a difference at least as large as yours if there were truly no difference between control and variant, assuming the test was run correctly. A small p-value means your result would be unusual under no difference. It does not measure how big or how valuable the difference is.

Is 1 minus the p-value the chance that my variant is better?

No. Many tools label 1 minus p as confidence, but it is not the probability that the variant beats control. In the default example the p-value is 0.0171, and 1 minus p is 98.29%, yet that is not a 98.29% chance that B wins. The p-value is calculated assuming there is no difference, so it cannot give the probability that there is one.

Is a p-value of 0.05 the same as a 95% significance level?

They describe the same threshold from two sides. The significance level, alpha, is 0.05 (5%): the false positive rate you accept when there is no real difference. The confidence level is 1 minus alpha, 95%. People say "significant at 95%" and "significant at the 5% level" for the same test, which is where the 5% or 95% confusion comes from.

Is a p-value of 0.001 significant?

Yes, at any of the usual levels (90%, 95% or 99%), because 0.001 is below 0.10, 0.05 and 0.01. A very small p-value says the data would be very unusual if there were no difference. It still says nothing about whether the lift is large enough to matter: read the confidence interval for that.

What significance level should I use, 0.05 or 0.01?

95% (alpha 0.05) is the usual default for A/B tests. Use 99% (alpha 0.01) when a wrong call would be costly or hard to undo, knowing it needs more visitors to reach. Accept 90% only for low-risk changes. Choose the level before the test starts. Changing it after seeing the result defeats the purpose.

Is a big lift always statistically significant?

No. Significance depends on how much data is behind the lift, not on its size alone. On this page the same +25.0% relative lift has a p-value of 0.0171 with 10,000 visitors per group, and 0.4509 with 1,000 per group, which is not significant. Significance is about the lift relative to the noise in your sample.

My result is not significant. Does that mean the variant has no effect?

No. It means the test did not find enough evidence of a difference. The variant may have a small effect the test was too small to detect. Check whether the test reached the sample size you planned, and look at how wide the confidence interval is: a wide interval that includes useful lifts means the test was inconclusive, not negative.

Should I use a one-sided or a two-sided test?

Two-sided is the safer default because it also detects a variant that is worse than control. A one-sided test only looks for an increase, and its p-value is half the two-sided one when the variant is ahead. It is only valid if you chose it before the test started and would act the same way on a worse result as on no result.

Can I use this calculator for revenue or average order value?

No. This calculator is for conversion rates, where each visitor either converted or did not. Revenue per visitor and average order value are continuous and usually skewed by a few large orders, so they need a different test, such as Welch's t-test or a bootstrap, often with extreme orders capped.

How do you calculate statistical significance by hand?

Work out each group's conversion rate, the pooled rate and the pooled standard error, then divide the difference in rates by the standard error to get z. The p-value is the area of the normal curve beyond z, doubled for a two-sided test. If p is below your alpha, the result is significant. The worked example on this page shows every step.

Find out what is costing you conversions

Enter your website and email and the audit starts right away. Watch it check your page live: eight checks, each scored out of 100, and three fixes you can make now.