What A/B testing is for
Websites live in noisy conditions. Traffic mix changes daily, campaigns start and stop, seasons shift demand. A before-and-after comparison cannot separate the effect of your change from all of that. An A/B test can, because both versions run at the same time, to randomly assigned visitors, under the same conditions.
That makes A/B testing the closest thing marketing has to proof. It is also easy to do badly. A poorly designed test produces a confident-looking answer that is wrong, which is worse than no answer, because people act on it.
An A/B test is a question asked of your customers; write the question down before you hear the answer.
The anatomy of a sound test
Fig. 01 · Process
Tap to explore
Six steps to a trustworthy test
The hypothesis should come from research, as described in the CRO process. The primary metric should be as close to business value as your traffic allows: purchases or qualified leads rather than clicks. Guardrail metrics, such as average order value, refund rate or lead quality, make sure a 'win' on the primary metric is not quietly costing you elsewhere.
Sample size: decide before you start
The number of visitors you need depends on three things: your baseline conversion rate, the smallest effect you would care to detect, and how much statistical certainty you want. Low baselines and small effects need very large samples. This is arithmetic, not a matter of opinion, and it is why many low-traffic pages simply cannot support tests of minor changes.
A common rule of thumb for roughly 80% power at a 5% significance level approximates the visitors needed per variant as 16 × p × (1 − p) ÷ d², where p is the baseline rate and d is the absolute difference you want to detect. Use a proper calculator for real decisions; the tool below gives a quick sense of scale.
Calculator
Rough sample size per variant
An approximation for a two-variant test (80% power, 5% significance). Treat it as a sense check, not a substitute for your testing tool's calculator.
Visitors needed per variant
51,733.33
Approximation: 16p(1-p)/d², where d is the absolute difference
= 16 * (base / 100) * (1 - base / 100) / ((base / 100 * mde / 100) * (base / 100 * mde / 100))
Days to reach sample (two variants)
25.87
Round up to whole weeks to cover weekday and weekend behaviour
= 2 * 16 * (base / 100) * (1 - base / 100) / ((base / 100 * mde / 100) * (base / 100 * mde / 100)) / daily
Defaults are illustrations. Use your own numbers. Nothing you enter leaves this page.
How long to run a test
Run until the pre-calculated sample is reached and for at least one full business cycle, usually one or two complete weeks, so that weekday and weekend behaviour, pay-day effects and campaign rhythms are represented. Stopping on a Tuesday because the dashboard turned green is the single most common way to manufacture a false winner.
There is also a ceiling. Tests that run for many weeks suffer from cookie deletion, returning visitors seeing both versions and changing external conditions. If a test needs months to conclude, the effect you are looking for is probably too small to matter for that page, and a bolder change is the better use of traffic.
Reading the result
Statistical significance tells you how surprising the observed difference would be if there were really no difference. It does not tell you the size of the effect with precision, and it does not tell you the result will hold forever. Look at the confidence interval, not just the headline lift: a result of '+8%, with a plausible range from +1% to +15%' carries different implications from a precise +8%.
- Clear win on primary metric, guardrails stable. Ship it, record the learning, consider what the result says about customers.
- No significant difference. Not a failure. The change did not matter enough to detect, which is useful to know. Ship whichever version is simpler.
- Clear loss. Valuable. Your hypothesis about the customer was wrong; work out what that implies.
- Win on primary, harm on a guardrail. Do not ship until you understand the trade-off in business terms.
A/B, multivariate or split URL?
Fig. 02 · Comparison
Tap to explore
Choosing a test type
Split-URL testing is an A/B test in which each variant lives at a different URL. It suits large structural changes, such as a new landing page template, that would be awkward to build as an in-page variation. Make sure search engines see the canonical version and that redirects are temporary.
The traps that produce false winners
Myth vs reality
A/B testing traps
Client-side or server-side testing
Client-side tools change the page in the visitor's browser after it loads. They are quick to set up and let marketers build variants without engineering, which is why most programmes start there. The cost is performance and fragility: scripts add weight, variants can flicker, and complex changes break when the underlying page is updated.
Server-side testing decides the variant before the page is sent. It is faster for visitors, works for logic such as pricing, search ranking or checkout flows, and is cleaner for single-page applications. It needs developer involvement for every test. Mature programmes usually run both: client-side for quick message and layout tests, server-side for anything structural or performance-sensitive.
What to do with inconclusive results
Many tests end with no significant difference, and teams often treat this as wasted effort. It is better read as information. Either the change did not affect the decision, or its effect was smaller than the test could detect. Both tell you to look elsewhere for the barrier.
Resist the urge to extend an inconclusive test indefinitely in the hope it will tip over. Instead, revisit the research behind the hypothesis, consider whether the change was too timid to matter, and record the outcome so nobody repeats it. A programme where every test 'wins' is usually a programme where the analysis is too generous.
What to test first
Test where the decision happens and where research points. That usually means the value proposition and headline on high-traffic entry pages, proof and reassurance near the call to action, pricing presentation, and friction in forms and checkout. Leave colour and microcopy tweaks for pages with abundant traffic and nothing bigger to fix.
Finally, consider privacy and consent. Testing tools set identifiers and process behavioural data, so they belong inside your consent framework. In India, factor in the DPDP Act; elsewhere, the equivalent local rules. A test that breaks consent obligations is not a win at any lift.
Consent also affects the sample. If only visitors who accept tracking enter the test, your results describe that group, which may behave differently from those who decline. That is usually acceptable, but it should be stated when results are reported, so nobody assumes the lift applies to every visitor.
Key takeaways
- 01A/B testing randomises visitors between versions so the effect of a change can be isolated from everything else.
- 02Calculate sample size and duration before launch, and run for full business cycles.
- 03Judge on one primary metric close to business value, with guardrails to catch hidden harm.
- 04Stopping at the first significant moment and over-slicing segments are the commonest sources of false winners.
- 05On low-traffic pages, test bold changes or rely on research rather than chasing tiny effects.
Frequently asked
- What is A/B testing in marketing?
- A/B testing is an experiment that shows different versions of a page, email, ad or experience to randomly assigned groups and compares their performance on a chosen metric. Because the groups are randomised and run at the same time, differences in outcome can be attributed to the change rather than to traffic or seasonal variation.
- How long should an A/B test run?
- Until it reaches the sample size calculated before launch and covers at least one full business cycle, usually one or two complete weeks. Ending early because the result looks significant is a common cause of false winners. If the required duration runs to months, test a bolder change instead.
- What sample size do I need for an A/B test?
- It depends on your baseline conversion rate, the smallest effect you want to detect and the statistical power and significance you choose. Lower baselines and smaller effects need much larger samples. Use your testing platform's calculator before launch, and accept that some pages lack the traffic to test small changes.
- What is statistical significance in A/B testing?
- Significance indicates how unlikely the observed difference would be if the variants truly performed the same. A conventional threshold is 95% confidence. It does not measure the size or permanence of the effect, so read it alongside the confidence interval and guardrail metrics before deciding.
- Can you A/B test with low traffic?
- You can, but only for larger effects. Test bold changes to high-traffic steps, use a conversion event closer to the page if the final conversion is rare, and lean more on qualitative research. Sometimes the honest answer is to make a research-backed change and monitor it carefully rather than run an underpowered test.
Published by Fabulous.Media, a network of specialist marketing agencies. Updated 9 October 2026. Platform features change often; check current official documentation before acting on platform-specific detail.





