Short answer

A/B testing is a controlled experiment that splits visitors at random between the current page and a changed version, to see which converts better. To run one well, find a real problem in your data, write a hypothesis, pick one primary metric, set the sample size before launch, and read the result once the planned sample is reached.

  • Fix obvious problems straight away, and save tests for changes you are unsure about.
  • A good hypothesis names the evidence, the change, the audience, the expected effect and the metric that decides it.
  • Run tests in whole weeks, and do not stop the first time a variant looks like a winner.
  • Read the sample ratio first, then the primary metric, revenue per visitor, guardrails and segments.
  • A/B testing does not hurt SEO when you follow Google's guidance on cloaking, canonical links and redirects.

What is A/B testing?

A/B testing is a controlled experiment on a website: visitors are split at random between the current version of a page (the control, A) and a changed version (the variant, B) over the same period, and the version that produces more conversions wins. Randomization is what makes the comparison fair.

Conversions can be orders, leads, sign-ups or clicks to a merchant. Because both versions see the same mix of visitors at the same time, outside factors such as a sale, the weather or a newsletter send hit both groups equally. What is left is the effect of the change, plus random noise, which is why sample size and statistical significance matter.

What is conversion testing?

Conversion testing is the wider name for any controlled test that measures whether a change increases the share of visitors who convert. A/B testing is the most common form. Others are split-URL tests, A/B/n tests and multivariate tests, compared below. All of them answer the same question: did this change cause more conversions?

Types of conversion test, and when each one fits
Test typeWhat it comparesBest forTraffic needed
A/B test (split test)One control against one variant on the same URLMost changes: headlines, layouts, offers, formsLowest
A/B/n testOne control against two or more variantsChoosing between several versions of one ideaHigher: each extra variant needs its own sample
Split-URL testVersions that live at different URLs, with visitors redirectedRedesigns, new templates, new checkout or form flowsSame as A/B, more setup
Multivariate testEvery combination of several changed elementsHigh-traffic pages where element interactions matterHighest
A/A testTwo identical versionsChecking that the tool splits and counts traffic correctlyAs for the A/B test it precedes

Why run A/B tests, and when not to

Run A/B tests when you are unsure whether a change will help and the page has enough traffic to give an answer. A test replaces opinion with evidence, protects revenue from changes that quietly hurt it, and shows how big a win really is. It is the wrong tool for broken pages and very low traffic.

Not every change needs a test. Fix obvious problems straight away: a broken button, a form error that blocks submission, a checkout that fails on one browser. Test the changes where reasonable people could disagree: a new headline, a different offer, a shorter form, a new page layout.

If a page gets too few conversions for a test to finish in a sensible time, use other evidence instead: session recordings, on-page polls, user testing and a careful before and after comparison. A CRO audit is built for exactly that situation.

How to run an A/B test in 8 steps

To run an A/B test, find a problem in your data, write a hypothesis, pick one primary metric, calculate the sample size, build and check the variant, launch it to a random split, wait for the planned sample, then read the result once and record it. Skipping any step is how false wins get rolled out.

  1. Find a real problem

    Start from evidence, not a list of ideas. Look for the page or step where many visitors leave, then find out why with recordings, heatmaps, polls and support questions. A test aimed at a known leak is far more likely to matter than a test of a button color.

  2. Write a hypothesis

    State what you will change, for whom, what you expect to happen and why, based on what you saw. A written hypothesis keeps the test focused on one idea and tells you what to learn if it loses. The template is below.

  3. Choose one primary metric

    Pick the metric that decides the test before it starts: orders, leads, sign-ups or revenue per visitor. Add guardrail metrics you must not hurt, such as refunds or lead quality. A lift in clicks that does not reach orders is not a win.

  4. Calculate the sample size and duration

    Enter your baseline conversion rate and the smallest lift worth detecting into the A/B test sample size calculator. It tells you how many visitors each version needs. Divide by daily traffic and round up to whole weeks, so every weekday is counted equally.

  5. Build the variant and check it

    Build the change, then test it on real phones and browsers before launch. Check that the variant does not flicker into view, that tracking fires on both versions and that nothing else on the page changed. A broken variant tests your bug, not your idea.

  6. Launch to a random split

    Send visitors to each version at random, usually 50/50, and keep each visitor in the same version on return visits. Avoid launching during a one-off sale or a big campaign unless the test is about that traffic.

  7. Wait for the planned sample

    Do not stop the test the first time it looks like a winner. Early results swing widely, and checking repeatedly until something looks significant produces false wins. Check daily only for errors and an uneven split, not for the answer.

  8. Read the result, roll out and record

    When the planned sample is reached, check the result once with the statistical significance calculator. Roll out a clear winner, look at segments before deciding on a loser, and write the result in your test log either way.

How to write an A/B test hypothesis

An A/B test hypothesis is one sentence that names the evidence, the change, the audience, the expected effect and the metric that will decide it. A good hypothesis can lose and still teach you something, because it says why you expected the change to work.

Use this template:

Because we saw [evidence], we believe that [change] for [audience] will [effect]. We will know this is true when [primary metric] changes, measured over [planned sample].

An example for a lead generation form: "Because recordings show mobile visitors stop at the phone number field, and our poll answers mention spam calls, we believe that making the phone number optional for mobile visitors will increase completed form submissions. We will know this is true when submissions per visitor rise, measured over the sample the calculator gives."

Weak and strong A/B test hypotheses
Weak hypothesisWhy it is weakStronger version
"A green button will convert better."No evidence, no reason, no metric"Heatmaps show visitors miss the button below the reviews, so moving it above the reviews will increase add to carts."
"Let's test a new homepage."Many changes, nothing to learn if it loses"Polls say visitors cannot tell what we sell, so a headline naming the product will raise clicks to product pages and orders."
"Shorter forms are better."A general belief, not a finding about this form"Form analytics show most drop-off at the company size field, so removing it will increase completed demo requests."

Not sure what to test first?

The free CRO audit app checks your page for people and for AI agents and ranks what it finds, so your test traffic goes to the changes most likely to matter. Watch it run live on your page, and get the report by email.

How much traffic and time does an A/B test need?

An A/B test needs enough visitors in each version to detect the smallest lift you care about, at the significance level and power you choose. That number depends on your current conversion rate: the lower the rate and the smaller the lift, the more visitors you need. Calculate it before launch, never during.

Four inputs decide the sample size:

  • Baseline conversion rate: how the page converts today. The conversion rate calculator works it out.
  • Minimum detectable effect: the smallest lift that would be worth acting on.
  • Significance level: how much risk of a false win you accept. A 95% confidence level is the common default.
  • Statistical power: the chance of catching a real effect of that size. 80% is the common default.

Once you know the sample, run the test in whole weeks, so weekend and weekday behavior are both included, and run it until the sample is reached even if the result looks clear sooner. The guide to how long to run an A/B test goes deeper.

How to analyze A/B test results

To analyze A/B test results, wait for the planned sample, check that traffic was split as intended, then compare the primary metric with a significance test and look at the confidence interval, not just the winner. Check guardrail metrics and the main segments, desktop and mobile above all, before you roll anything out.

The A/B testing metrics to read, in order

  1. Sample ratio. If you set a 50/50 split and one version got noticeably more visitors, something is wrong with the setup. Fix it before trusting any result. The significance calculator includes a sample ratio check.
  2. Primary metric. The one metric you chose before launch, with its p-value and confidence interval.
  3. Revenue per visitor. For stores, a variant can raise orders and lower order value, or the reverse. Revenue per visitor shows the net effect.
  4. Guardrails. Refunds, lead quality, unsubscribes or support contacts: anything the change could hurt.
  5. Segments. Device, new versus returning visitors and traffic source. Treat segment results as ideas for the next test, not as wins, unless the test was planned and sized for that segment.
What each A/B test outcome means, and what to do next
OutcomeWhat it meansWhat to do
Significant winnerThe variant beat the control by more than chance would explain at your chosen levelCheck guardrails and segments, roll it out, and keep measuring after launch
Significant loserThe change hurt the primary metricKeep the control. Ask what the loss says about your visitors: it is evidence too
No significant differenceAny effect is smaller than the test could detectDo not call it a tie. Keep the simpler version, or test a bolder change
Mixed by segmentFor example, better on desktop and flat on mobileLook for a reason in the design, then test a version built for the weaker segment

Why desktop and mobile results often differ

In several Convertica client tests, the same change behaved differently on desktop and mobile. A cart test that added shipping and refund policies helped on desktop and made no significant difference on mobile. A sticky call-to-action test went the other way: it helped on mobile and made no significant difference on desktop. If you only read the combined number, you can miss a win, or roll out a change that only works for half your visitors.

What a well-run A/B test looks like

A well-run A/B test has a hypothesis based on a visible problem, an equal traffic split, a fixed sample and a result checked against a confidence level. Convertica's test for dScryb, a membership store selling boxed text called scenes, shows each part.

  • The problem: search results showed content that many visitors' accounts could not open, with no clear next step, and clicking it led to an error message.
  • The change: call-to-action buttons on the search results, the access level shown for each result, and visitors sent straight to the store page instead of the error.
  • The setup: the original case study says it ran "on desktop and mobile using VWO (our preferred testing software), with traffic split equally between the control and the variant."
  • The sample: 20,000 visitors during the test period. The test ran for 17 days and reached a 98% confidence level.
  • The metric that decided it: paid membership sign-ups, with call-to-action clicks and revenue per visitor read alongside.

Read the dScryb case study

More examples, each with its metric and timeframe, are in the CRO case studies, and these A/B test examples include well-known tests from other companies too.

+49.4%

Paid membership sign-ups, 17-day A/B test

Source: dScryb, paid membership sign-ups, 17-day A/B test, 98% confidence

Common A/B testing mistakes to avoid

The most common A/B testing mistakes are stopping a test early, testing with too little traffic, measuring the wrong metric, changing several things without a hypothesis, ignoring segments and trusting an uneven traffic split. Each one produces results that look convincing and then fail to show up in revenue after rollout.

  • Stopping when it looks good. Peeking at a running test and stopping at the first significant reading inflates false wins. Decide the sample first.
  • Too little traffic. An underpowered test can only detect huge effects, so most real effects come back as "no difference".
  • The wrong metric. Clicks, scroll depth and time on page are easy to move. Orders, leads and revenue per visitor are what pay.
  • No hypothesis. A redesign that changes ten things can win or lose without telling you why.
  • Ignoring segments. A result averaged across devices can hide a change that helps one group and hurts another.
  • An uneven split. If one version gets more visitors than planned, the data is suspect. Find the cause before reading the result.
  • Overlapping changes. Launching a sale, a new campaign or another test on the same page during the test muddies the answer.
  • Not checking the variant. A layout bug on one browser or a flicker of the original page can decide the test on its own.
  • Calling a flat result a tie. "No significant difference" means the test could not see an effect that small, not that there was none.
  • Forgetting the losers. Unrecorded losing tests get run again a year later. Keep a log.

Convertica really goes the extra mile on delivering happy clients and results. Even when initial test didn't prove to be succesful, they'd keep iterating until they found a winner. They push through when needed! Thank you Ayumi & Team.

Niels ZeeTrustpilot review

Does A/B testing affect SEO?

A/B testing does not hurt SEO when it follows Google's published guidance: never show search engines a different page from visitors, add a canonical link from each variant URL to the original, use temporary 302 redirects for split-URL tests, and end the experiment as soon as you have an answer. Long-running tests can look like an attempt to mislead.

Google sets this out in its page on minimizing A/B testing impact in Google Search. It names cloaking, showing Googlebot one version and people another, as against its spam policies whether or not you are running a test.

Convertica's triple setup for split-URL tests

Testing tools that load variants with JavaScript cannot sort every visitor. People with scripts or cookies blocked, or on slow connections, usually see the original page and never enter the test, so the original quietly collects extra traffic and its numbers stop being comparable. Convertica's answer, first written up for affiliate sites, uses three pages instead of two:

  1. Original: the ranking page, which keeps a small share of traffic (10% in Convertica's setup) plus every visitor the tool cannot sort.
  2. Control: an exact copy of the original at its own URL, with 45% of traffic.
  3. Variant: a copy with your change, with the other 45%.

You compare only the control and the variant, which received equal, properly sorted traffic. Both copies carry a canonical link to the original, as Google recommends. For affiliate sites, give each copy its own tracking ID so merchant reports line up with the test. The full write-up on why affiliate split tests fail walks through the setup.

How to build an A/B testing framework

An A/B testing framework turns one-off tests into a program: research feeds a backlog of hypotheses, a scoring model sets the order, every test meets the same rules for sample size and metrics, and every result goes into a shared log. The framework matters more than any single test, because most of the value comes from what you learn across many.

  1. Research continuously. Analytics, recordings, polls, support tickets and reviews feed new ideas every month, not only at the start.
  2. Keep one backlog. Every idea is written as a hypothesis with its evidence, so ideas from the boss and from the data compete on the same terms.
  3. Prioritize with a score. Score each idea for impact, confidence and ease (ICE) or potential, importance and ease (PIE). The CRO audit guide compares the two.
  4. Set test rules. One primary metric, a sample size calculated before launch, whole weeks, no stopping early, and a check of every variant on real devices.
  5. Agree who does what. Someone owns the hypothesis, someone builds, someone checks the build, someone reads the result. In Convertica's CRO advisory, Kurt Philip directs the program and your own team builds and launches the tests.
  6. Log everything. Hypothesis, screenshots, dates, sample, result and what you will try next, for winners and losers alike.
  7. Re-test big wins. A large lift on a small sample is worth confirming before you build more tests on top of it.

What does an A/B testing tool do?

An A/B testing tool splits visitors between versions of a page, keeps each visitor in the same version, records conversions for each and reports whether the difference is significant. Most tools also include a visual editor, targeting by device or audience, and split-URL and multivariate tests. Google Optimize, once the common free choice, was retired by Google in September 2023.

When you compare tools, check how the variant loads (a visible flicker of the original page can bias the test), whether it can test server-side or only in the browser, how it handles visitors who block scripts, and whether its statistics match how you plan to read results. The guide to A/B testing tools compares the main options.

A/B testing when AI agents visit your site

Some AI assistants now browse sites and take steps for the people they work for, and some crawlers run scripts. Both can land in your test groups. Filter known bots out of test reports and watch for sudden traffic from one source. Also check what a variant looks like in the page HTML: a change that only exists after a script runs may be invisible to an AI agent reading the page. This is a new field and nobody, Convertica included, has years of data on it yet, so treat it as something to check, not a proven effect.

Free calculators for planning and reading tests

Plan and read your own tests with Convertica's free calculators. Each one shows its formula and a worked example.

Plan a test

A/B test sample size calculator

How many visitors each version needs before the test starts, and how long that takes.

Read a result

Statistical significance calculator

Whether a finished test result is significant, with the confidence interval and a sample ratio check.

See all free calculators

A/B testing guide: common questions

What is an A/B testing guide for?

An A/B testing guide explains how to compare two versions of a page with real visitors and decide which one works better. This one covers the full process: research, hypothesis, sample size, setup, reading the result, common mistakes and how to run tests as an ongoing program.

What is A/B testing in simple terms?

A/B testing shows half of your visitors the current page (A) and half a changed page (B) at the same time, then counts which version gets more conversions. Because both groups arrive in the same period, the difference can be put down to the change rather than to luck or the season, once the sample is big enough.

How do you do an A/B test?

Find a problem in your data, write a hypothesis, choose one primary metric, calculate the sample size, build and check the variant, launch it to a random 50/50 split, wait for the planned sample without stopping early, check the result for significance, then roll out the winner and record what you learned.

What is the difference between A/B testing and split testing?

Usually nothing: most people use split testing as another name for A/B testing. Some tools use split-URL testing for the case where each version lives at its own URL and visitors are redirected, which suits bigger changes such as a full page redesign.

How much traffic do you need for an A/B test?

It depends on your current conversion rate and the smallest lift you want to detect. A page with few conversions needs a large lift or a long test to reach a reliable answer. The free A/B test sample size calculator gives the number of visitors each version needs before you start.

What are the most common A/B testing mistakes?

The most common A/B testing mistakes are stopping a test as soon as it looks like a winner, running tests without enough traffic, measuring clicks instead of orders or leads, changing many things without a hypothesis, ignoring device segments, and not checking that traffic was split evenly.

What is an A/B testing framework?

An A/B testing framework is the repeatable system around individual tests: where test ideas come from, how they are prioritized, the hypothesis and sample size rules every test must meet, who builds and checks each variant, how results are read, and where every result, including the losers, is recorded.

Can A/B testing hurt SEO?

Not if it is set up as Google recommends: do not show search engines a different page from visitors, put a canonical link on variant URLs pointing to the original, use temporary 302 redirects for split-URL tests, and end the test once you have an answer.

What is a multivariate test?

A multivariate test changes several elements at once, such as the headline and the image, and tests every combination to see which mix works best and how the elements interact. It needs far more traffic than an A/B test, because the visitors are divided across many more versions.

Does Convertica run A/B tests for clients?

Not as a standalone testing service. The free CRO audit, an app you can run now, shows what to fix and test. CRO advisory, led personally by Kurt Philip, helps you plan tests and read the results while your own team builds and launches them. With full implementation, Convertica's team builds the fixes from your audit for you.

Find out what is costing you conversions

Enter your website and email and the audit starts right away. Watch it check your page live: eight checks, each scored out of 100, and three fixes you can make now.