Free tool
A/B test sample size calculator
Convertica's free A/B test sample size calculator tells you how many visitors each variant needs before you start a test. Enter your baseline conversion rate and the smallest lift worth detecting, and it returns the sample per variant, the total, and how many days the test should run.
Calculate your sample size
Free to use, no sign-up. The result updates as you type.
Visitors needed per variant
3,841
7,682 visitors in total across 2 variants, control included. At 1,000 visitors a day, that is about 8 days.
Detects a change from 10% to 12% (+20.0% relative, +2.00 percentage points) with 95% significance (two-sided) and 80% power.
Formula
pbar = (p1 + p2) / 2
A = Za * sqrt(2 * pbar * (1 - pbar))
B = Zb * sqrt(p1(1 - p1) + p2(1 - p2))
n = (A + B)^2 / (p2 - p1)^2, rounded up
total = n * variants
Za = z(1 - alpha/2) two-sided, z(1 - alpha) one-sided
Zb = z(power)
p1 is the baseline rate and p2 the rate you want to be able to detect. n is visitors per variant.
Short answer
An A/B test sample size calculator tells you how many visitors each version of a page needs before the result can be trusted. It works from four inputs set before the test starts: your baseline conversion rate, the smallest lift worth detecting, the significance level and the statistical power. Add daily visitors to estimate how long the test must run.
- 95% significance and 80% power are the usual defaults, and raising either needs more visitors.
- Planning for a smaller lift needs far more visitors, because the required sample grows roughly with one over the effect squared.
- Every variant, control included, needs its own full sample.
- Run the test to the planned sample, rounded up to whole weeks, even when the result looks clear early.
- The number is a planning figure, not a promise that the variant will win.
What is A/B test sample size?
A/B test sample size is the number of visitors each version of a page needs to see before the result can be trusted. It is set before the test starts, from four inputs: your baseline conversion rate, the smallest lift worth detecting, the significance level and the statistical power. An A/B test sample size calculator turns those into a visitor count.
Getting it right matters in both directions. Stop too early and a test can crown a winner that was only noise, or miss a real improvement because there was not enough data to see it. Plan a sample far bigger than you need and you spend weeks of traffic on one question. The calculator above finds the minimum that answers your question at the reliability you chose.
How to use the A/B test sample size calculator
To use the A/B test sample size calculator, enter the current conversion rate of the page you will test, the smallest lift you care about, your significance level and power, and the number of variants. Add daily visitors to get a test length. The result updates as you type.
-
Baseline conversion rate
The current rate of the page or step you will test, from your analytics. Use the same conversion and the same visitor count (sessions or users) that you will use to judge the test.
-
Minimum detectable effect
The smallest real change you want the test to be able to detect. Relative means a share of the baseline; absolute means percentage points.
-
Significance level and power
95% significance and 80% power are the defaults. Higher values make the test more reliable and need more visitors.
-
Test type
Two-sided looks for a change in either direction. One-sided looks only for an increase and needs fewer visitors, but it cannot flag a variant that made things worse.
-
Variants
Count the control: a simple A/B test is 2. Each extra variant adds a full sample, and with three or more you can apply a Bonferroni correction.
-
Daily visitors
Optional. How many visitors a day will enter the test across all variants. The calculator divides the total sample by this number to estimate the days.
The calculator assumes an even traffic split, one yes or no conversion outcome per visitor, and a fixed-horizon test: you decide the sample size first and read the result once, at the end.
How do you choose a baseline conversion rate?
Your baseline conversion rate is the rate the control page converts at today, measured the same way the test will be measured. Take it from your analytics for the exact page or funnel step you will test, for the audience that will enter the test, over recent whole weeks without unusual promotions or outages.
- Match the test. If the test runs only on mobile visitors to product pages, the baseline is the mobile product page rate, not the sitewide rate.
- Match the unit. Divide by sessions or by users, whichever your testing tool counts, and keep it the same when you analyze.
- Use whole weeks. A baseline taken from a few weekdays can differ from the rate across a full week.
- Check the tracking first. If analytics and your real orders or leads disagree, fix the measurement before planning a test on it.
The baseline changes the sample a lot. With the relative effect fixed at 20% (95% significance, 80% power, two-sided), a lower baseline needs far more visitors per variant, because the same relative lift is a smaller absolute change:
| Baseline rate | Rate to detect | Visitors per variant |
|---|---|---|
| 1% | 1.2% | 42,693 |
| 2% | 2.4% | 21,109 |
| 5% | 6% | 8,158 |
| 10% (default) | 12% | 3,841 |
| 20% | 24% | 1,683 |
Need the rate first? The conversion rate calculator works it out from conversions and visitors.
What is a minimum detectable effect, relative or absolute?
The minimum detectable effect (MDE) is the smallest true difference between control and variant that a test is designed to detect at your chosen power. A relative MDE is a share of the baseline; an absolute MDE is a difference in percentage points. The same lift can be written either way, and gives the same sample size.
- On a 5% baseline, a 10% relative MDE means detecting 5% to 5.5%: 31,234 visitors per variant.
- An absolute MDE of 0.5 percentage points on the same baseline is the same change, 5% to 5.5%, and needs the same 31,234.
- An absolute MDE of 1 percentage point means 5% to 6%, a 20% relative lift, and needs 8,158.
Mixing the two up is a common source of wrong plans: "1%" can mean a 1% relative lift or one whole percentage point, and the samples differ enormously. The calculator asks which one you mean.
How small an effect should you plan for?
Plan for the smallest lift that would change a decision, and that the change could plausibly produce. Required sample size grows roughly with one over the effect squared. On a 10% baseline, halving the relative MDE from 20% to 10% takes the sample per variant from 3,841 to 14,751, about 3.8 times as many visitors.
| Relative MDE | Change detected | Per variant | Total, 2 variants |
|---|---|---|---|
| 5% | 10% to 10.5% | 57,763 | 115,526 |
| 10% | 10% to 11% | 14,751 | 29,502 |
| 20% (default) | 10% to 12% | 3,841 | 7,682 |
| 30% | 10% to 13% | 1,774 | 3,548 |
| 50% | 10% to 15% | 686 | 1,372 |
Small tweaks to button colors or single words tend to produce small effects, so they need the largest samples. A change to the offer, the page structure or the checkout flow is more likely to produce an effect big enough to measure on modest traffic.
What do significance level and statistical power mean?
The significance level controls false positives: at 95% significance (alpha 0.05), a test with no real difference will still show a "winner" 5% of the time. Statistical power controls false negatives: at 80% power, a real effect of exactly your MDE is detected 80% of the time and missed 20% of the time. Raising either needs more visitors.
A false positive means shipping a change that does nothing, or does harm. A false negative means throwing away a change that worked. Which one costs you more decides where to set the dials. 95% and 80% are conventions, not laws.
| Significance | Power | Visitors per variant |
|---|---|---|
| 90% | 80% | 3,026 |
| 95% (default) | 80% | 3,841 |
| 95% | 90% | 5,142 |
| 99% | 80% | 5,716 |
| 99% | 90% | 7,281 |
Should you use a one-sided or two-sided test?
Use a two-sided test unless you have decided in advance that a variant performing worse would be treated exactly like no change. A two-sided test looks for a difference in either direction; a one-sided test looks only for an increase. One-sided needs fewer visitors because its critical value is lower: z = 1.645 instead of 1.960 at 95%.
In the default example, one-sided needs 3,026 visitors per variant against 3,841 two-sided, 815 fewer. The cost is that a one-sided test cannot tell you a variant hurt conversions, which is often the most useful thing a test finds. Choose the test type before the test starts, never after looking at the data.
How does the number of variants change the sample size?
Every variant, control included, needs its own full sample, so the total grows with each variant you add. Each extra variant is also another comparison with control, and every comparison is another chance of a false positive. A Bonferroni correction divides alpha by the number of comparisons, which protects against that but raises the sample per variant. Multivariate testing, which tests every combination of several changes, multiplies the variants fastest.
With 3 variants and no correction, each comparison runs at alpha 0.05, so the chance of at least one false positive across both comparisons is higher than 0.05 and can be up to 2 × 0.05 = 0.1.
| Variants, incl. control | Comparisons | Total, no correction | Bonferroni alpha | Per variant, Bonferroni | Total, Bonferroni |
|---|---|---|---|---|---|
| 2 | 1 | 7,682 | 0.05 | 3,841 | 7,682 |
| 3 | 2 | 11,523 | 0.025 | 4,652 | 13,956 |
| 4 | 3 | 15,364 | 0.0167 | 5,124 | 20,496 |
| 5 | 4 | 19,205 | 0.0125 | 5,458 | 27,290 |
Without a correction the per-variant sample stays at 3,841; only the total grows. If your traffic is limited, test fewer variants: two strong ideas tested properly beat five tested on too little data.
How long should an A/B test run?
An A/B test should run until every variant reaches the planned sample size, rounded up to whole weeks. Divide the total sample by the visitors who enter the test each day to get the days, then round up to full weeks so weekdays and weekends are counted equally. Never stop before the planned sample, even when the result looks clear.
- Work out the days. Total sample divided by daily visitors in the test, all variants together. The calculator does this when you add daily visitors.
- Round up to whole weeks. Behavior on a Monday morning and a Saturday night can differ. A test that covers three Mondays and two Saturdays weighs them unevenly.
- Cover a business cycle. If customers usually take a week or two to decide, or your sales follow paydays or a monthly billing cycle, run long enough to include a full cycle.
- Avoid unusual periods. A sale, a holiday or a big campaign brings visitors who behave differently. Either avoid them or make sure both variants see them equally for the whole test.
- Do not run forever. Over many weeks, cookies are cleared and returning visitors can see both versions, which blurs the result. If the plan needs months, change the plan (below).
| Daily visitors in the test | Days to reach the sample | Run for |
|---|---|---|
| 250 | 31 | 5 weeks |
| 500 | 16 | 3 weeks |
| 1,000 (default) | 8 | 2 weeks |
| 2,500 | 4 | 1 week |
| 5,000 | 2 | 1 week |
Even when the sample arrives in a couple of days, run at least one full week. A test that reaches its number on a Tuesday and a Wednesday has only measured Tuesday and Wednesday visitors.
For a worked example, a duration table by traffic and the rules for when to stop a test, see the guide on how long to run an A/B test.
What if your traffic is too low for an A/B test?
If your traffic is too low for an A/B test, the calculator will show a test length of months. You can test bigger changes, test where traffic is higher, use fewer variants, or accept a less strict test knowingly. If none of that fits, use research instead of testing to decide what to change.
A worked example shows the trade-off. A page converts at 3% and 500 visitors a day enter the test, split across 2 variants at 95% significance and 80% power, two-sided. The calculator's own formula, solved for the effect, gives the smallest relative lift each test length can detect:
| Test length | Visitors in the test | Smallest relative lift | Change detected |
|---|---|---|---|
| 2 weeks | 7,000 | 41.8% | 3% to 4.25% |
| 4 weeks | 14,000 | 28.8% | 3% to 3.86% |
| 8 weeks | 28,000 | 20% | 3% to 3.6% |
If the change you have in mind is unlikely to lift conversions by that much, the test cannot give you a reliable answer in that time. Options, roughly in the order to try them:
- Test a bigger change. A new offer, a rewritten page or a shorter form is more likely to move the rate enough to detect than a new button color.
- Test where the traffic is. Pick a busier page, or run the change across a whole template (every product page) rather than one page.
- Measure a step with a higher rate. Add to cart or starting a form converts more often than a completed purchase, so it needs a smaller sample. Check that it really leads to the final conversion.
- Use fewer variants. Each variant you drop gives its traffic to the others.
- Relax the settings knowingly. 90% significance or a one-sided test shortens the test, with more wrong calls. Decide this before the test, not after.
- Research instead of testing. User tests, session recordings, on-page polls and customer interviews show what is stopping people without needing a statistical sample. Clear problems, such as a broken form, can simply be fixed.
Convertica's free CRO audit is built for this situation: it checks your page for people and for AI agents and ranks what it finds, so scarce test traffic goes to the changes most likely to matter.
Common A/B test sample size mistakes
The most common A/B test sample size mistakes are stopping a test early because it looks significant, planning for an effect that is unrealistically large, changing the test while it runs, and trusting a test whose traffic split came out uneven. Each one makes the result less reliable than the significance level suggests.
- Peeking and stopping early. Checking every day and stopping the first time the result looks significant inflates false positives far above alpha. Wait for the planned sample. If you need to stop early, use a method designed for it, such as a sequential test.
- Stopping as soon as the number is reached. If the sample arrives mid-week, finish the week.
- An unrealistic effect size. A tiny MDE makes a test impossibly long; an inflated one gives a sample too small to detect the effect that really happens. Plan for a lift you could plausibly get.
- Mixing up relative and absolute. A "2% lift" read as 2 percentage points instead of 2% of the baseline changes the plan completely.
- Mixing up visitor counts. Use the same definition (sessions or users) for the baseline and for the test analysis.
- Adding variants without adding traffic. Each variant needs its own full sample, and more comparisons mean more chances of a false positive.
- Changing the test while it runs. Editing a variant or changing the traffic split mid-test mixes two experiments into one result. If the split changes while conversion rates also change over time, the combined numbers can point the wrong way.
- Ignoring a broken split. If a 50/50 test delivers clearly uneven visitor counts (a sample ratio mismatch), check the setup before you trust the result.
- Using a conversion calculator for revenue. Revenue per visitor and average order value need a formula that uses their standard deviation. This calculator is for conversion rates only.
- Recalculating after the test with the observed effect. Plugging the result back in to see whether the test "had enough power" tells you nothing new. Plan the sample before the test.
The A/B test sample size formula
The A/B test sample size formula used here is the standard formula for comparing two proportions with a z-test, rounded up to a whole visitor. It needs the two conversion rates you want to tell apart and the z-scores for your significance level and power:
n = ( z(1 - alpha/2) * sqrt(2 * pbar * (1 - pbar))
+ z(power) * sqrt(p1(1 - p1) + p2(1 - p2)) )^2
/ (p2 - p1)^2
- p1 is the baseline conversion rate and p2 the rate you want to detect: p1 times (1 + relative MDE), or p1 plus the absolute MDE.
- pbar is the average of p1 and p2.
- z(1 - alpha/2) is the z-score for the significance level: 1.960 at 95% two-sided. A one-sided test uses z(1 - alpha), which is 1.645 at 95%.
- z(power) is the z-score for the power: 0.842 at 80% and 1.282 at 90%.
- Total sample is n times the number of variants, control included. With Bonferroni, alpha is first divided by the number of comparisons (variants minus 1).
Worked example
Baseline 10%, relative minimum detectable effect 20%, 95% significance (two-sided), 80% power, 2 variants and 1,000 visitors a day. These are the default values in the calculator above, and every number below is computed by its code.
- p1 = 0.10 and p2 = 0.10 × 1.20 = 0.12.
- pbar = (0.10 + 0.12) / 2 = 0.11.
- z(0.975) = 1.959964 and z(0.80) = 0.841621.
- A = 1.959964 × sqrt(2 × 0.11 × 0.89) = 1.959964 × sqrt(0.1958) = 1.959964 × 0.442493 = 0.867270.
- B = 0.841621 × sqrt(0.10 × 0.90 + 0.12 × 0.88) = 0.841621 × sqrt(0.1956) = 0.841621 × 0.442267 = 0.372221.
- n = (0.867270 + 0.372221)^2 / (0.12 - 0.10)^2 = 1.536339 / 0.0004 = 3,840.85.
- Round up: 3,841 visitors per variant, 7,682 in total.
- Days: 7,682 / 1,000 = 7.68, rounded up to 8 days, so run the test for 2 full weeks.
Why do sample size calculators give different answers?
Sample size calculators give different answers because they make different choices, not because one is wrong. Before comparing two tools, check that they agree on these points:
- Relative or absolute MDE. Some tools take the effect as a share of the baseline, others in percentage points.
- One-sided or two-sided. Defaults differ between tools, and one-sided needs fewer visitors.
- Per variant or total. Some report visitors per variant, others the total across all variants.
- The statistical method. Tools built for sequential or Bayesian testing engines use different maths from a fixed-horizon z-test, so their numbers are not comparable with this one.
- Formula details. Pooled or unpooled variance, continuity corrections and rounding move the answer by a little.
This calculator shows its formula and its full working, so you can check exactly what it assumes.
How to read the result
The number per variant is the minimum sample that gives the test your chosen power to detect a real effect of your chosen size, at your chosen significance level. It is a planning number, not a promise: reaching it does not mean the variant will win.
- If the real effect is smaller than your minimum detectable effect, the test will often miss it.
- If the real effect is larger, the test is more likely to detect it.
- When the test ends, check the result with the statistical significance calculator, once, on the full sample.
A/B test sample size calculator FAQ
How does an A/B test sample size calculator work?
An A/B test sample size calculator uses the standard formula for comparing two conversion rates. You enter your baseline conversion rate, the minimum detectable effect, the significance level and the power, and it returns how many visitors each variant needs. With the defaults on this page (10% baseline, 20% relative effect, 95% significance, 80% power), that is 3,841 per variant.
How do you calculate sample size for an A/B test?
For an A/B test, or any product experiment with a yes or no outcome, pick your baseline conversion rate, the smallest lift worth detecting, a significance level and a power, then apply the two-proportion formula on this page, or enter the values in the calculator. Multiply the result by the number of variants for the total, and divide the total by daily visitors for the number of days.
What is a minimum detectable effect?
The minimum detectable effect (MDE) is the smallest real change the test is designed to detect. It can be relative (a share of the baseline) or absolute (percentage points). Smaller effects need far more visitors: on a 10% baseline, halving the relative MDE from 20% to 10% raises the sample per variant from 3,841 to 14,751.
Why use 95% significance and 80% power?
They are common conventions, not rules. 95% significance means a 5% chance of calling a winner when there is no real difference. 80% power means an 80% chance of detecting a real effect of your chosen size. Raising either one makes the result more reliable and needs more visitors.
Is a sample of 30, or 100, visitors enough for an A/B test?
Almost never, for conversion rates. The rule of thumb about 30 comes from averages of continuous measurements, not from comparing conversion rates. The sample you need depends on your baseline rate and the effect you want to detect: the default example on this page needs 3,841 visitors per variant.
How long should an A/B test run?
Run an A/B test until each variant reaches the planned sample size, and for whole weeks, so every day of the week is counted equally. In the default example, 7,682 visitors at 1,000 a day takes 8 days, which rounds up to 2 full weeks.
Can I stop the test early if it already looks significant?
Not with this method. The calculation assumes a fixed-horizon test: you set the sample size in advance and analyze once, at the end. Checking repeatedly and stopping as soon as the result crosses the line raises the false positive rate well above the significance level you chose.
Is this the same as an A/B test calculator for results?
No. This A/B test calculator plans a test before it starts. To check whether a finished test is significant, use Convertica's free statistical significance calculator, which runs a two-proportion z-test on your visitors and conversions.
Can I use this calculator for revenue per visitor or average order value?
No. This calculator is for conversion rates, where each visitor either converts or does not. Revenue per visitor and average order value are averages, so their sample size depends on the standard deviation of the metric and needs a different formula.
What if my site does not have enough traffic for the sample size?
Test bolder changes, which have larger effects and need smaller samples. Test on pages or funnel steps with more traffic or a higher baseline rate. Use fewer variants. If a test would still take months, research such as user testing, session recordings, on-page polls and a CRO audit is usually a better use of the time.