A/B Test Sample Size and Statistical Significance Explained

Sample size and statistical significance are what separate a real result from a lucky run. Without enough data, even a large apparent lift can vanish the moment you ship it.

Holding tests to real significance is a core part of our A/B testing service, so a result you act on is a result you can trust. We run statistically sound testing for clients across Brisbane and Melbourne.

Why sample size matters

Small samples swing wildly. A handful of conversions can make one variant look far ahead purely by chance. The more visitors each variant sees, the more the noise averages out and the true effect shows.

What statistical significance means

Significance is the probability that the difference you see is not down to chance. Most testers aim for 95 percent confidence, meaning there is only a small chance the result is a fluke.

Plan the test before you run it

These three inputs set your required sample size. Smaller expected lifts need more data to detect, so plan for the size of change you realistically expect.

InputEffect on sample size
Higher confidence targetNeeds more data
Smaller expected liftNeeds more data
Lower baseline conversionNeeds more data

Do not stop the test early

Checking results constantly and stopping the moment a variant looks significant inflates false positives. Let the test run to its planned sample and a full business cycle before you decide.

Calculate your required sample size before launch and resist peeking for a winner. Early significance often disappears as more data arrives.

Key takeaways

Frequently asked questions

How quickly can TPR Media start an A/B testing programme?

Most testing programmes begin within one to two weeks once tracking and goals are confirmed. Contact our Brisbane team on (07) 3184 9197 for a current schedule.

Where is TPR Media based?

TPR Media operates from Level 34, 1 Eagle Street, Brisbane City QLD 4000, serving clients across Brisbane and Australia-wide.

TPR Media sizes A/B tests from baseline rate, expected lift and a 95 percent confidence target, then runs each test to completion so results reflect real effects, not noise.