Plan a test
A/B test sample size calculator
How many visitors each version needs before the test starts, and how long that takes.
A/B testing
This A/B testing guide walks through the process Convertica uses on client sites: research first, a written hypothesis, a sample size set before launch, a clean setup, an honest read of the result and a framework that keeps good tests coming. It is written for site owners who want fewer false wins.
A/B testing is a controlled experiment that splits visitors at random between the current page and a changed version, to see which converts better. To run one well, find a real problem in your data, write a hypothesis, pick one primary metric, set the sample size before launch, and read the result once the planned sample is reached.
A/B testing is a controlled experiment on a website: visitors are split at random between the current version of a page (the control, A) and a changed version (the variant, B) over the same period, and the version that produces more conversions wins. Randomization is what makes the comparison fair.
Conversions can be orders, leads, sign-ups or clicks to a merchant. Because both versions see the same mix of visitors at the same time, outside factors such as a sale, the weather or a newsletter send hit both groups equally. What is left is the effect of the change, plus random noise, which is why sample size and statistical significance matter.
Conversion testing is the wider name for any controlled test that measures whether a change increases the share of visitors who convert. A/B testing is the most common form. Others are split-URL tests, A/B/n tests and multivariate tests, compared below. All of them answer the same question: did this change cause more conversions?
| Test type | What it compares | Best for | Traffic needed |
|---|---|---|---|
| A/B test (split test) | One control against one variant on the same URL | Most changes: headlines, layouts, offers, forms | Lowest |
| A/B/n test | One control against two or more variants | Choosing between several versions of one idea | Higher: each extra variant needs its own sample |
| Split-URL test | Versions that live at different URLs, with visitors redirected | Redesigns, new templates, new checkout or form flows | Same as A/B, more setup |
| Multivariate test | Every combination of several changed elements | High-traffic pages where element interactions matter | Highest |
| A/A test | Two identical versions | Checking that the tool splits and counts traffic correctly | As for the A/B test it precedes |
Run A/B tests when you are unsure whether a change will help and the page has enough traffic to give an answer. A test replaces opinion with evidence, protects revenue from changes that quietly hurt it, and shows how big a win really is. It is the wrong tool for broken pages and very low traffic.
Not every change needs a test. Fix obvious problems straight away: a broken button, a form error that blocks submission, a checkout that fails on one browser. Test the changes where reasonable people could disagree: a new headline, a different offer, a shorter form, a new page layout.
If a page gets too few conversions for a test to finish in a sensible time, use other evidence instead: session recordings, on-page polls, user testing and a careful before and after comparison. A CRO audit is built for exactly that situation.
To run an A/B test, find a problem in your data, write a hypothesis, pick one primary metric, calculate the sample size, build and check the variant, launch it to a random split, wait for the planned sample, then read the result once and record it. Skipping any step is how false wins get rolled out.
Start from evidence, not a list of ideas. Look for the page or step where many visitors leave, then find out why with recordings, heatmaps, polls and support questions. A test aimed at a known leak is far more likely to matter than a test of a button color.
State what you will change, for whom, what you expect to happen and why, based on what you saw. A written hypothesis keeps the test focused on one idea and tells you what to learn if it loses. The template is below.
Pick the metric that decides the test before it starts: orders, leads, sign-ups or revenue per visitor. Add guardrail metrics you must not hurt, such as refunds or lead quality. A lift in clicks that does not reach orders is not a win.
Enter your baseline conversion rate and the smallest lift worth detecting into the A/B test sample size calculator. It tells you how many visitors each version needs. Divide by daily traffic and round up to whole weeks, so every weekday is counted equally.
Build the change, then test it on real phones and browsers before launch. Check that the variant does not flicker into view, that tracking fires on both versions and that nothing else on the page changed. A broken variant tests your bug, not your idea.
Send visitors to each version at random, usually 50/50, and keep each visitor in the same version on return visits. Avoid launching during a one-off sale or a big campaign unless the test is about that traffic.
Do not stop the test the first time it looks like a winner. Early results swing widely, and checking repeatedly until something looks significant produces false wins. Check daily only for errors and an uneven split, not for the answer.
When the planned sample is reached, check the result once with the statistical significance calculator. Roll out a clear winner, look at segments before deciding on a loser, and write the result in your test log either way.
An A/B test hypothesis is one sentence that names the evidence, the change, the audience, the expected effect and the metric that will decide it. A good hypothesis can lose and still teach you something, because it says why you expected the change to work.
Use this template:
Because we saw [evidence], we believe that [change] for [audience] will [effect]. We will know this is true when [primary metric] changes, measured over [planned sample].
An example for a lead generation form: "Because recordings show mobile visitors stop at the phone number field, and our poll answers mention spam calls, we believe that making the phone number optional for mobile visitors will increase completed form submissions. We will know this is true when submissions per visitor rise, measured over the sample the calculator gives."
| Weak hypothesis | Why it is weak | Stronger version |
|---|---|---|
| "A green button will convert better." | No evidence, no reason, no metric | "Heatmaps show visitors miss the button below the reviews, so moving it above the reviews will increase add to carts." |
| "Let's test a new homepage." | Many changes, nothing to learn if it loses | "Polls say visitors cannot tell what we sell, so a headline naming the product will raise clicks to product pages and orders." |
| "Shorter forms are better." | A general belief, not a finding about this form | "Form analytics show most drop-off at the company size field, so removing it will increase completed demo requests." |
An A/B test needs enough visitors in each version to detect the smallest lift you care about, at the significance level and power you choose. That number depends on your current conversion rate: the lower the rate and the smaller the lift, the more visitors you need. Calculate it before launch, never during.
Four inputs decide the sample size:
Once you know the sample, run the test in whole weeks, so weekend and weekday behavior are both included, and run it until the sample is reached even if the result looks clear sooner. The guide to how long to run an A/B test goes deeper.
To analyze A/B test results, wait for the planned sample, check that traffic was split as intended, then compare the primary metric with a significance test and look at the confidence interval, not just the winner. Check guardrail metrics and the main segments, desktop and mobile above all, before you roll anything out.
| Outcome | What it means | What to do |
|---|---|---|
| Significant winner | The variant beat the control by more than chance would explain at your chosen level | Check guardrails and segments, roll it out, and keep measuring after launch |
| Significant loser | The change hurt the primary metric | Keep the control. Ask what the loss says about your visitors: it is evidence too |
| No significant difference | Any effect is smaller than the test could detect | Do not call it a tie. Keep the simpler version, or test a bolder change |
| Mixed by segment | For example, better on desktop and flat on mobile | Look for a reason in the design, then test a version built for the weaker segment |
In several Convertica client tests, the same change behaved differently on desktop and mobile. A cart test that added shipping and refund policies helped on desktop and made no significant difference on mobile. A sticky call-to-action test went the other way: it helped on mobile and made no significant difference on desktop. If you only read the combined number, you can miss a win, or roll out a change that only works for half your visitors.
A well-run A/B test has a hypothesis based on a visible problem, an equal traffic split, a fixed sample and a result checked against a confidence level. Convertica's test for dScryb, a membership store selling boxed text called scenes, shows each part.
More examples, each with its metric and timeframe, are in the CRO case studies, and these A/B test examples include well-known tests from other companies too.
+49.4%
Paid membership sign-ups, 17-day A/B test
Source: dScryb, paid membership sign-ups, 17-day A/B test, 98% confidence
The most common A/B testing mistakes are stopping a test early, testing with too little traffic, measuring the wrong metric, changing several things without a hypothesis, ignoring segments and trusting an uneven traffic split. Each one produces results that look convincing and then fail to show up in revenue after rollout.
Convertica really goes the extra mile on delivering happy clients and results. Even when initial test didn't prove to be succesful, they'd keep iterating until they found a winner. They push through when needed! Thank you Ayumi & Team.
A/B testing does not hurt SEO when it follows Google's published guidance: never show search engines a different page from visitors, add a canonical link from each variant URL to the original, use temporary 302 redirects for split-URL tests, and end the experiment as soon as you have an answer. Long-running tests can look like an attempt to mislead.
Google sets this out in its page on minimizing A/B testing impact in Google Search. It names cloaking, showing Googlebot one version and people another, as against its spam policies whether or not you are running a test.
Testing tools that load variants with JavaScript cannot sort every visitor. People with scripts or cookies blocked, or on slow connections, usually see the original page and never enter the test, so the original quietly collects extra traffic and its numbers stop being comparable. Convertica's answer, first written up for affiliate sites, uses three pages instead of two:
You compare only the control and the variant, which received equal, properly sorted traffic. Both copies carry a canonical link to the original, as Google recommends. For affiliate sites, give each copy its own tracking ID so merchant reports line up with the test. The full write-up on why affiliate split tests fail walks through the setup.
An A/B testing framework turns one-off tests into a program: research feeds a backlog of hypotheses, a scoring model sets the order, every test meets the same rules for sample size and metrics, and every result goes into a shared log. The framework matters more than any single test, because most of the value comes from what you learn across many.
An A/B testing tool splits visitors between versions of a page, keeps each visitor in the same version, records conversions for each and reports whether the difference is significant. Most tools also include a visual editor, targeting by device or audience, and split-URL and multivariate tests. Google Optimize, once the common free choice, was retired by Google in September 2023.
When you compare tools, check how the variant loads (a visible flicker of the original page can bias the test), whether it can test server-side or only in the browser, how it handles visitors who block scripts, and whether its statistics match how you plan to read results. The guide to A/B testing tools compares the main options.
Some AI assistants now browse sites and take steps for the people they work for, and some crawlers run scripts. Both can land in your test groups. Filter known bots out of test reports and watch for sudden traffic from one source. Also check what a variant looks like in the page HTML: a change that only exists after a script runs may be invisible to an AI agent reading the page. This is a new field and nobody, Convertica included, has years of data on it yet, so treat it as something to check, not a proven effect.
Plan and read your own tests with Convertica's free calculators. Each one shows its formula and a worked example.
Plan a test
How many visitors each version needs before the test starts, and how long that takes.
Read a result
Whether a finished test result is significant, with the confidence interval and a sample ratio check.
Measure
The baseline conversion rate you need before you size a test.
An A/B testing guide explains how to compare two versions of a page with real visitors and decide which one works better. This one covers the full process: research, hypothesis, sample size, setup, reading the result, common mistakes and how to run tests as an ongoing program.
A/B testing shows half of your visitors the current page (A) and half a changed page (B) at the same time, then counts which version gets more conversions. Because both groups arrive in the same period, the difference can be put down to the change rather than to luck or the season, once the sample is big enough.
Find a problem in your data, write a hypothesis, choose one primary metric, calculate the sample size, build and check the variant, launch it to a random 50/50 split, wait for the planned sample without stopping early, check the result for significance, then roll out the winner and record what you learned.
Usually nothing: most people use split testing as another name for A/B testing. Some tools use split-URL testing for the case where each version lives at its own URL and visitors are redirected, which suits bigger changes such as a full page redesign.
It depends on your current conversion rate and the smallest lift you want to detect. A page with few conversions needs a large lift or a long test to reach a reliable answer. The free A/B test sample size calculator gives the number of visitors each version needs before you start.
The most common A/B testing mistakes are stopping a test as soon as it looks like a winner, running tests without enough traffic, measuring clicks instead of orders or leads, changing many things without a hypothesis, ignoring device segments, and not checking that traffic was split evenly.
An A/B testing framework is the repeatable system around individual tests: where test ideas come from, how they are prioritized, the hypothesis and sample size rules every test must meet, who builds and checks each variant, how results are read, and where every result, including the losers, is recorded.
Not if it is set up as Google recommends: do not show search engines a different page from visitors, put a canonical link on variant URLs pointing to the original, use temporary 302 redirects for split-URL tests, and end the test once you have an answer.
A multivariate test changes several elements at once, such as the headline and the image, and tests every combination to see which mix works best and how the elements interact. It needs far more traffic than an A/B test, because the visitors are divided across many more versions.
Not as a standalone testing service. The free CRO audit, an app you can run now, shows what to fix and test. CRO advisory, led personally by Kurt Philip, helps you plan tests and read the results while your own team builds and launches them. With full implementation, Convertica's team builds the fixes from your audit for you.