On-site AI
AI shopping assistants
When a store should add an AI assistant, the risks, and how to test it.
AI and conversion
AI A/B testing means using AI to help plan, build, run or read split tests. It can save hours of work. It cannot create the visitors a test needs, and it cannot turn a guess into evidence. This guide shows where AI for A/B testing helps, where it misleads, and how to use it without fooling yourself.
AI A/B testing means using AI to find test ideas in your research, write or design variants, shift traffic between versions or summarize results. It speeds up the wordy parts of testing, but it cannot replace the visitors a test needs: real people still see different versions, and the statistics still decide whether a difference is real.
AI A/B testing is the use of artificial intelligence in any part of a split test: finding test ideas in your research, writing or designing the variants, moving traffic between versions while the test runs, or reading the results. The comparison itself does not change. Real visitors still see different versions, and the numbers still decide.
The phrase covers several different jobs, and tools that sell "AI testing" usually mean only one or two of them. Before you buy anything or change how you test, it helps to know which job you are talking about, because each one helps in a different way and fails in a different way.
| Job | What AI does | Where it helps | What to watch |
|---|---|---|---|
| Research | Sorts survey answers, reviews, chat logs and session notes into themes | Turns a pile of feedback into a short list of problems in an afternoon | It can invent a theme that is not in the data. Spot-check the quotes it groups |
| Hypotheses | Drafts "if we change X, Y will improve, because Z" statements | Gets past the blank page and suggests angles you missed | Generic best practice dressed up as insight. Every hypothesis needs evidence from your site |
| Variants | Writes headlines and copy, suggests layouts, builds mockups | Many options quickly, so you test bolder ideas | Bland copy, claims you cannot back up, off-brand tone |
| Traffic allocation | Shifts visitors toward the version that is doing better (a multi-armed bandit) | Short promotions where earning during the test matters more than learning | Less reliable answers about how big the difference is |
| Analysis | Summarizes results and searches segments for differences | Faster write-ups and a second pair of eyes on the data | Chance findings in small segments, and summaries that misstate numbers |
| Prediction | Estimates a winner with simulated users or models of past tests | Ranking ideas before you spend traffic on them | A prediction is not a result. It never replaces the test |
AI is good at the slow, wordy parts of A/B testing: reading large amounts of customer feedback, drafting hypotheses, producing variant copy and mockups, checking a test setup for obvious mistakes, and writing up results. These tasks used to eat most of a testing program's time, so the gain is real even though the statistics stay the same.
AI cannot create the visitors a test needs, cannot make a small difference easier to detect, and cannot turn a prediction into proof. A test's reliability depends on traffic, conversion rate and the size of the effect you want to measure. No model changes that arithmetic, whatever the tool's marketing says.
The number of visitors a test needs grows fast as the change you want to detect gets smaller: halving the lift you want to detect roughly quadruples the sample. That is a property of the statistics, not of the software. If your site cannot reach the sample a test needs in a few weeks, an AI tool will not change that. Use the A/B test sample size calculator before you start.
Ask a tool to search every segment of a test (device, country, new or returning, traffic source) and it will find a segment where a variant "won". Look at enough slices and some will differ by chance alone. Treat a segment result as a new hypothesis to test, not as a finding.
A test tells you which version did better. It does not tell you why, and neither does a model guessing after the fact. The "why" comes from the research you did before the test: recordings, surveys and the questions customers ask.
Language models write fluent summaries that can misreport a conversion rate, swap the control and the variant, or describe a result as significant when it is not. Check every number in an AI-written report against your testing tool, or a statistical significance calculator.
A classic A/B test keeps a fixed split, such as 50/50, until it reaches a sample size set in advance. An AI-driven multi-armed bandit moves traffic toward the version that is ahead while the test runs. A bandit earns more during the test. A fixed test tells you more reliably whether, and by how much, one version beats the other.
| Method | How traffic is split | Best for | Trade-off |
|---|---|---|---|
| Classic A/B test | Fixed split until a planned sample size is reached | Lasting changes: page layouts, pricing presentation, checkout steps | Half your visitors see the weaker version for the whole test |
| Multi-armed bandit | Shifts toward the version that is ahead as data comes in | Short campaigns, such as a sale headline that only runs for a week | Weaker evidence on the size of the difference; early luck can lock in a loser |
| Contextual bandit | Picks a version per visitor based on what is known about them | Personalization on high-traffic sites with many visitor types | Needs a lot of traffic and still needs a holdout group to prove it helps |
If you are deciding something you will live with for a year, such as a new product page template, use a fixed test. If you are choosing between three banner messages for a weekend sale, a bandit is a reasonable choice. A contextual bandit is personalization in disguise, and it should be measured the same way: see AI ecommerce personalization.
To use AI for A/B testing well, feed it your own research rather than asking for generic ideas, rank its suggestions yourself, edit every variant by hand, plan the sample size before launch, choose a fixed test or a bandit on purpose, let the test run to its planned end, and check every number in the write-up.
Give the AI your material: survey answers, reviews, support tickets, chat logs, notes from session recordings and your analytics funnel. Ask it to group the problems it finds and to quote the source for each one. Ideas drawn from your own customers beat a list of best practices every time.
Ask for hypotheses in the form "if we change X for these visitors, Y will improve, because of evidence Z". Then score them on impact, confidence and ease. The scoring is where your knowledge of the business matters, so do not hand it over.
Use AI to produce several versions of a headline, offer or layout, then pick and rewrite. Remove anything you cannot back up: invented urgency, made-up review counts, guarantees you do not offer. A variant that wins with a false claim is a liability, not a result.
Decide the smallest lift worth detecting, then calculate how many visitors each version needs and how long that takes at your traffic. If the answer is several months, test a bigger change or fix the problem without a test.
Use a fixed split for lasting decisions and a bandit only for short, disposable choices. Write the choice and the stopping rule down before the test starts, so nobody changes them halfway through.
Do not stop early because a dashboard shows a winner. While the test runs, ask AI to help check that the split matches what you set, that bots are filtered, and that conversions in the testing tool match your analytics and your orders.
Read the outcome in your testing tool or a significance calculator. Let AI draft the summary and the next hypotheses, then check every figure against the source and record the result, including tests that lost.
To A/B test an AI feature such as a chat assistant, a recommendation widget or personalized content, show it to a random share of visitors and hold it back from the rest. Compare revenue per visitor and conversion rate between the two groups over full weeks, not how many people used the feature.
This matters because AI features are usually sold on numbers that cannot show whether they helped: chats started, recommendations clicked, "assisted revenue". Many of the visitors who use a feature were going to buy anyway. Only a holdout group, a share of visitors who never see the feature, shows what the feature added.
Agent readiness
Some AI assistants now visit websites on a person's behalf. They can affect a test in two ways: as visitors who never buy, and as readers who may never see your variant.
PERSON VIEW (VARIANT B)
Product page, variant B
Linen duvet cover that stays cool all night
Order by 2pm for next-day delivery.
Illustrative example
AGENT VIEW (PAGE HTML)
Illustrative example: a client-side testing script swaps the headline and adds a delivery note after the page loads. A person in variant B sees the new version. An agent reading the page HTML sees the original.
When an AI agent or crawler loads a page inside a test, the testing tool may assign it to a variant and count it as a visitor who did not convert. A little of this is noise. A lot of it, landing unevenly, can blur a result. Check what your testing tool filters, and compare test traffic with your analytics for the same pages and dates.
Many testing tools change the page with a script after it loads. People see the variant. An agent that reads the page HTML may only ever see the control. That is fine during a test, but a winner that stays in the testing tool, instead of being built into the site, is invisible to agents for good.
The most common mistakes with AI in A/B testing are trusting predicted winners, stopping tests early because a tool says so, testing generic AI copy that nobody asked for, believing segment "wins" found after the fact, and judging AI features by engagement instead of revenue per visitor in a holdout test.
AI A/B testing is worth it for a small site only in the parts that do not need traffic: research, hypotheses, copy drafts and write-ups. Bandits, contextual personalization and automatic segment discovery need far more visitors than most small sites get. With limited traffic, fix clear problems directly and save tests for big changes.
A useful rule: if your planned test needs more than a couple of months to reach its sample size, it is not a test you can run. Either test a bolder change, which needs fewer visitors to detect, or make the change because the evidence is strong and watch your numbers afterwards. A CRO audit is a good way to find the changes that do not need a test at all.
Audit and advisory
Convertica has run conversion research and A/B tests for clients since 2017, and its CRO case studies show the results with metric, timeframe and source. Today it offers a free CRO audit, then CRO advisory or full implementation. The free CRO audit, an app you can run now, checks your page for people and for AI agents and gives you three fixes you can make now. CRO advisory, led personally by founder Kurt Philip with Convertica's CRO team, helps you decide what to test, how to read the results and where AI tools are worth using. In advisory your own team builds the changes; with full implementation, Convertica's team builds the fixes from your audit. Both are priced after your audit.
In this series
On-site AI
When a store should add an AI assistant, the risks, and how to test it.
Personalization
What to personalize, what to skip, and how to measure it with a holdout group.
Free tool
Check whether a finished test result is real before you act on it.
AI A/B testing is the use of artificial intelligence in any part of a split test: drafting hypotheses from your research, writing or designing variants, shifting traffic toward better variants as the test runs, or summarizing results. The test itself still compares real visitors, so it still needs enough traffic to give a reliable answer.
No. AI can predict which version might win, but a prediction is not a result. Only real visitors choosing between real versions show what works on your site. Use AI to get to better test ideas faster, then let the test decide.
AI makes the work around a test faster: research, variant drafts and write-ups. It does not shorten the time a fixed A/B test needs to reach its planned sample size, because that depends on your traffic, your conversion rate and the size of the change you want to detect.
A multi-armed bandit is a test that moves more traffic to the better-performing version while the test is still running, instead of keeping a fixed split. It earns more during the test but tells you less about how big the difference really is. It suits short promotions more than lasting design decisions.
Some tools claim to predict winners with simulated users or models trained on past tests. Treat those predictions as a way to rank ideas, not as evidence. Your visitors, products and prices are specific to you, and only a live test measures them.
Show the feature to a random share of visitors and hold it back from the rest, then compare revenue per visitor and conversion rate between the two groups over full weeks. Do not judge it by how many people used the feature, because many of them would have bought anyway.
They can. Bots and AI agents that load your pages may be counted as visitors in a variant without ever buying, and agents that read page HTML may not see changes made by a client-side testing script. Check your testing tool's bot filtering, and build winning changes into the site itself.
The research and drafting help is worth it at any size. Automated traffic allocation and segment discovery need far more traffic than most small sites have. On a small site, fix clear problems directly and test only the few big changes your traffic can measure.