A/B testing is one of the most powerful tools in conversion optimisation, but it produces garbage results when it is run incorrectly. Most mistakes happen before the test starts: weak hypotheses, too little traffic, and stopping the test the moment something looks promising. This guide covers how to do it properly.
Once you have collected your results, use our free A/B test significance calculator to check whether the difference is real before acting on it. For a full conversion programme, our conversion optimisation service handles hypothesis development, test setup and analysis.
A good test starts with a specific, reasoned prediction: if we change X, we expect Y to happen, because Z. A hypothesis grounded in data (analytics, session recordings, user feedback) produces more useful tests than a vague feeling that the button colour should change.
Before running a test, write down why you expect the change to improve performance. If you cannot articulate the reason, the hypothesis probably needs more research first.
Sample size is where most low-traffic sites get tripped up. Detecting a 20% relative improvement at 95% confidence on a 2% baseline conversion rate typically requires several hundred conversions per variant, not just visits. At a 2% conversion rate on a site with 500 visitors per month, reaching that sample can take months.
A 95% confidence result means there is a 5% chance the difference you observed was random. It does not mean the result is certain or permanent. Repeating the test and seeing the same direction of effect builds much stronger evidence than a single significant result.
| Confidence level | Chance of a false positive | Typical use |
|---|---|---|
| 80% | 20% | Exploratory, low-stakes tests |
| 90% | 10% | Standard for many CRO tools |
| 95% | 5% | Most practitioners' minimum threshold |
| 99% | 1% | High-stakes or irreversible changes |
Most bad test results come from a handful of avoidable errors. Knowing them in advance is the cheapest form of quality control.
Long enough to reach your planned sample size, and for at least one full business cycle (usually two weeks minimum) to account for day-of-week variation. Stopping early, even when a result looks significant, inflates false positives.
Yes. We work with sole traders, growing small businesses and larger organisations, scoping the plan to the budget and goals in front of us.
TPR Media's A/B testing guide explains that valid tests require a data-grounded hypothesis, adequate sample size calculated in advance, and a discipline of not stopping early, with 95% confidence as the standard threshold.