Skip to Main Content

Email A/B Testing Is an ROI Problem, Not a Statistics Problem: What Underpowered Tests Cost the P&L

Email ab testing often fails before it starts, underpowered tests produce false wins that cost more than they return. Here’s what that means for your P&L.

A bearded man wearing a black shirt and wireless earbuds sits in a brightly lit, modern airport terminal.
Robert Haydock
CEO, Zembula

If your email team runs A/B tests, there’s a high chance they’re losing money, not because the tests are poorly designed, but because they’re structurally underpowered before you even hit send. Email ab testing at typical retail send sizes rarely has the statistical power to detect meaningful revenue lifts. That means most ‘winners’ are noise, not insight. And when you scale noise across your entire list, you’re not optimizing. You’re gambling, with QA time, creative resources, and real revenue on the line.

The problem isn’t with the idea of testing. It’s with the instrument. Most email teams use campaign-level A/B testing as their primary method for proving what works. But at a 2.5% conversion baseline, detecting a 10% lift with 95% confidence and 80% power requires about 60,000 recipients per variant. Most tests run with 1,000 to 5,000 per group. That’s not just underpowered, it’s closer to a coin flip.

And yet, that coin flip becomes a slide in the next executive presentation. The ‘win’ gets rolled out company-wide. The team celebrates. But finance sees no corresponding lift in email revenue. Why? Because the test was never designed to measure what matters: true incrementality.

The Math Behind Underpowered A/B Tests

Let’s be clear: statistical significance is not the same as business significance. A test can be ‘significant’ and still be meaningless in practice. The real issue is power, the probability that your test will detect a real effect if one exists.

For email ab testing, power is crushed by low conversion rates. Here’s what it takes to detect a 10% relative lift at 95% confidence and 80% power:

  • Open rate (20% baseline): ~6,500 recipients per variant
  • Click rate (7% baseline): ~21,000 per variant
  • Conversion rate (2.5% baseline): ~60,000 per variant
  • Purchase rate (0.5% baseline): ~330,000 per variant

Most ESPs recommend 1,000-2,000 recipients per variant. At a 2.5% conversion baseline, that yields about 4% power, meaning you only have a 4% chance of detecting a real 10% lift. But the false-positive rate is still 5%. So, when you do get a ‘winner, ’ it’s more likely to be noise than signal.

As Peep Laja from CXL puts it: “A statistically significant result is only evidence if the sample size is large enough; with small baseline conversion rates, detecting small relative lifts requires very large samples most marketers never calculate before launching.”

The P&L Cost of False Positives

So what happens when you act on a false positive?

First, you roll out a change company-wide that doesn’t actually work. That’s not just wasted effort, it’s a real cost. You’ve spent engineering time, QA cycles, and creative labor to deploy something that won’t move the needle.

Second, you’ve blocked that same team from working on something that would move the needle. Testing calendars fill up fast. Every underpowered test is an opportunity cost.

Third, you’ve contaminated your list. If you re-split for the next test, some recipients have already seen the ‘winner’, but others haven’t. That re-split contamination erases any chance of clean longitudinal learning.

And if you’re testing multiple modules? The problem explodes. Testing N modules independently requires 2^N variants in a traditional ESP setup. One email becomes 16, 32, even 64 variants. Most teams can’t scale that. So they don’t test, or they test in silos that never interact.

The result: a testing program that feels active but delivers no measurable ROI.

Why the CFO Doesn’t Trust Email A/B Testing

Ask any finance leader why they don’t rely on email test results. They’ll tell you: the wins never show up in the P&L.

There’s a reason. Most email ab testing is optimized for speed, not accuracy. Teams peek at results mid-test, calling winners early. But as Evan Miller explains in How Not To Run an A/B Test: “Checking an in-flight A/B test for significance at multiple interim points inflates the real false-positive rate well beyond the nominal 5%.”

Then there’s revenue attribution. A single email might generate $100 in sales, but was that because of the email, or because the customer was already planning to buy? Last-touch attribution can’t answer that. And if one customer spends $500 while others spend $50, that outlier swings revenue-per-send wildly, making it hard to trust short-term lifts.

Worse, most ESPs default to a 14-day window. At a 2.5% conversion rate, 1,000 per variant and 14 days means you’re not just underpowered, you’re structurally incapable of detecting revenue effects.

No wonder the CFO isn’t impressed.

The Alternative: Channel-Level Holdout Testing

There’s a better way: the longitudinal holdout test.

Instead of testing one email at a time, assign subscribers to exposed or control at their first open. Keep them locked. Then measure the full UTM-scoped email channel revenue over time, calculated as transactions times blended AOV.

Why does this work?

  • It eliminates last-touch bias by measuring full channel impact.
  • It smooths out order-size outliers by using blended AOV.
  • It resolves in about four weeks to 95%+ significance.
  • It’s re-runnable whenever leadership asks, “Are we sure this is working?”

As PostHog notes: “To detect a 5% lift at 95% confidence with 80% power on a 1% baseline conversion rate, you need approximately 31,000 users per group.” For a 10% holdout, that’s 620,000 total active users. But once set up, it runs continuously, no new splits, no new builds.

This isn’t just statistically sound. It’s finance-ready. When the CFO asks, “What’s the ROI of your email program?” you can point to a number they trust.

Run Both, Confuse Neither

Some will ask: can’t we just do both?

Yes, but only if you keep the purposes separate.

Use variant-level RPM and CTC attribution to tune what’s live every day. Measure revenue per message and conversion-to-checkout across every module. This tells you what’s working right now, without setting up a test.

Use the holdout to prove the program’s overall impact. This is the only number that answers, “Would we make less money if we turned email off?”

Conflate them, and you get false confidence. Measure a single email’s lift versus its own baseline, and you’ll see big numbers, but that’s consistent with zero channel incrementality. Revenue shifts freely between sends under last-touch. A ‘staggering’ single-email lift can just mean you stole sales from next week.

Keep them separate. One optimizes. One proves.

The Audit Before the Board Deck

Before you put any past email test ‘win’ into a forecast, ask three questions:

  1. What was the baseline conversion rate? If it was below 3%, the test was likely underpowered unless you had 50K+ per variant.
  2. Was the sample size fixed in advance? If you peeked or extended the test, the false-positive rate is inflated.
  3. Was the result measured at the channel level? If it’s based on last-touch revenue from one email, it doesn’t prove incrementality.

If the answer to any of these is ‘no, ’ treat the win as interesting, not proven.

Compare your historical test performance against our 2025 email performance benchmark report. The average personalized-block CTC is 13.6%. The typical daily batch email converts at 2.5%. Use these as sanity checks, because if your ‘win’ beats the benchmark by a huge margin, it’s more likely noise than breakthrough.

Key Takeaways

  • Most email ab testing is underpowered, especially for revenue goals, meaning ‘winners’ are often noise.
  • At a 2.5% conversion baseline, you need ~60,000 recipients per variant to detect a 10% lift with confidence.
  • False positives carry real P&L costs: wasted builds, rollout of ineffective changes, and lost opportunity.
  • Traditional A/B testing fails finance because it can’t prove true incrementality.
  • The solution is a longitudinal, person-locked holdout measuring UTM-scoped email revenue over time.
  • Use variant-level RPM/CTC for daily optimization, but only the holdout to prove program-level impact.
  • Before trusting any past test win, audit it for power, fixed sample size, and channel-level measurement.

For more on how to run tests that finance trusts, see our guides on email holdout testing, the executive case for longitudinal testing, and why email is a performance marketing channel. You can also explore how measuring ROI in email shifts the conversation from opens to revenue.

A bearded man wearing a black shirt and wireless earbuds sits in a brightly lit, modern airport terminal.
Robert Haydock
CEO, Zembula

Robert Haydock co-founded Zembula with the mission to give retail performance marketers measurements through image personalization so they can grow revenue from owned channels.

Grow your business and total sales

Book a Demo
Full Width CTA Graphic