Knowledge Base/Ad Testing/Statistical Power in Ad Testing: What It Is and Why Your Tests Need More of It
Ad TestingStatistical Pitfalls

Statistical Power in Ad Testing: What It Is and Why Your Tests Need More of It

Statistical power is the probability your test will detect a real difference. Most ad tests are underpowered, which means most declared 'winners' aren't. How to fix it.

By Wreltik Research Team

Statistical Power in Ad Testing: What It Is and Why Your Tests Need More of It

Statistical power is the probability that your test will detect a real difference between variants when a real difference exists. A test with 80% power will catch a real difference 80% of the time and miss it 20% of the time. Most ad tests have power well below 50%. This means most ad tests are more likely to miss a real winner than to find one.

Why power matters

A low-power test produces two kinds of errors:

  • False negatives: a real difference exists (one ad is genuinely better), but the test doesn't detect it. You keep running the weaker ad.
  • False positives: the test declares a winner, but the "winner" is actually noise. You scale an ad that isn't actually better.

Low-power tests are biased toward false positives when you test multiple variants and pick the best — the winner's curse described in our article on A/B testing. The more variants you test at low power, the more inflated the apparent performance of the "winning" variant.

What determines power

Three factors:

  1. Sample size. More data = more power. This is the factor most under your control.
  2. Effect size. Larger differences between variants are easier to detect. Subtle differences require more power.
  3. Variability. More variable data = less power. High-variance metrics like conversion rate require larger samples than low-variance metrics.

A practical minimum

For most ad metrics, you need at least 100 conversions per variant to have reasonable power to detect a 20% relative difference. At typical DTC conversion rates, this means significant spend per variant — often $500-$2,000 depending on your CPA.

If you can't afford that level of spend per test, either test fewer variants (more budget per variant) or test larger differences (hooks rather than button colors). Testing five variants at $100 each produces five sets of results you can't trust. Testing two variants at $250 each is better — still underpowered, but at least directionally more reliable.

The practical fix for small budgets

Accept that you won't reach statistical significance. Use tests for directional guidance rather than conclusive decisions. If one variant is outperforming by 40% at low spend, it's probably the real winner even if the result isn't statistically significant. If the difference is 10%, it's probably noise. Scale the directional winner cautiously, with a plan to re-evaluate as more data accumulates.