Knowledge Base/Ad Testing/Why Your A/B Test Results Are Lying to You
Ad TestingStatistical Pitfalls

Why Your A/B Test Results Are Lying to You

A/B testing is the gold standard for ad creative — until it isn't. Common pitfalls that make test results misleading and how to design tests you can actually trust.

By Wreltik Research Team

Why Your A/B Test Results Are Lying to You

The dirty secret of ad creative testing is that most "winners" aren't. They're noise that looked like signal because the test was designed to find a winner, not to find the truth.

The low-budget trap

You test two ads at $50 each. Ad A gets a 2.1% CTR. Ad B gets 1.6%. You declare A the winner and scale it.

At $50 spend, the difference between 2.1% and 1.6% is probably a handful of clicks. One extra click — from someone who clicked by accident, or who would have converted anyway — flips the result. You're not measuring ad performance. You're measuring luck.

At minimum spend, you need large differences to be confident. A 2x difference in conversion rate at $50 spend: probably real. A 30% difference: could go either way. A 10% difference: you're guessing.

The fix is boring but necessary: either spend more on the test, or test fewer things so each test gets enough budget to mean something.

The "three ads, one winner" problem

Testing three ads, picking the winner, and reporting its numbers as the expected performance. This is called the winner's curse.

If you test enough variants, random chance guarantees that one of them will look good — even if all of them are equally mediocre. The smaller your sample and the more variants you test, the more inflated your winner's numbers will be.

A practical rule: if you test N variants, discount the winner's performance by roughly 10–20% for every doubling of N. Four variants? The winner is probably 10–20% worse than it looks. Eight variants? 20–40% worse. This isn't precise, but it's directionally honest. Most people do the opposite — they take the winner's numbers at face value and project them forward.

Creative fatigue isn't factored in

The test ran for three days and your winner crushed it. So you scale it to your full budget.

Week one: still good. Week two: performance starts sliding. Week three: the "winner" is delivering worse results than the ad it beat. What happened?

The test measured response to a novel stimulus. Novelty itself drives attention — the brain is wired to notice new things. Once the ad has been seen by your target audience a few times, the novelty premium evaporates. The creative underneath matters, but the boost from newness is gone.

Testing over longer windows helps. So does holding back a portion of your audience from the initial test so you can measure the winner against a fresh group later. Neither solves the problem completely, but both reduce the size of the lie.

Platform effects get ignored

Your test ran on Meta. The winner gets moved to TikTok. It flops. Not because it's bad — because the platforms are different environments with different user expectations and different optimization algorithms.

Creative testing needs to happen on the platform where the ad will run. Cross-platform extrapolation is guesswork dressed up as strategy.

What honest testing looks like

  1. Define what you're measuring before you run the test. Not "which ad is better" — better at what? CTR? Conversion rate? ROAS? Pick one and stick to it. If you pick the metric after seeing results, you're cherry-picking.

  2. Run until you have enough data, not until you're impatient. If your budget means you need three weeks to get statistically meaningful results, wait three weeks. Ending early because you're curious is the most expensive shortcut in testing.

  3. Hold out a validation set. Scale the winner to 30% of your audience first. If it holds for two weeks, take it to 70%. This costs a few days. It's cheaper than scaling a false winner to 100% of your budget.

  4. Track whether your winners actually win. Most teams never go back and check whether ads that "won" the test actually outperformed at scale. If you don't measure the accuracy of your testing process, you'll never improve it.