Email A/B Tests Rarely Give the CEO a Real Answer. A Smart Banners Holdout Does.
“Variant B won” is not an answer a CEO can fund. The sample-size math shows why most email A/B tests read out noise, and how a channel-level Smart Banners holdout proves real program growth in about four weeks.
This week, somewhere, a marketing team is showing a slide that claims “variant B won by 22 percent,” and a CEO is being asked to approve the next quarter’s email budget based on that finding. I lead a company embedded in retail email programs, so I’ll say it directly: most of those reports cannot answer the real question behind the budget ask. The math on sample size explains why. That is why we built a channel-level holdout into Smart Banners instead of adding another A/B testing dashboard.
The pressure is real and growing. Litmus’s research on email ROI shows 63 percent of marketing leaders now face greater CFO scrutiny over email spending, up from 52 percent, while 21 percent of marketers admit they cannot report their program’s actual return. Email budgets are now audited like ad budgets. The measurement must hold up just as well, or better.
The question the CEO is actually asking
“Variant B won” answers a narrow question: which of two emails performed better on one send. The CEO is asking a strategic question: if we invest more in this program, will we see new revenue that would not have happened otherwise? These are not the same questions, and only the second one justifies increasing spend.
The context sharpens the issue. Upcounting’s analysis of ecommerce ad accounts found average ROAS at 2.87 in 2025, with returns down in 13 of 14 industries, while acquisition costs keep rising. Every CMO is searching for a channel where the next dollar works harder, and email has the structural advantage: an owned audience with first-party identity and no auction inflating costs. I made that case in full in Email Is a Performance Marketing Channel, and the Math Proves It.
But performance budgets require performance-grade proof. Paid media brings lift studies and incrementality tests to budget meetings. Email often brings a subject line test. If email wants ad-budget treatment, it needs ad-grade validation, and campaign-level A/B testing is not that tool.
The sample-size math that kills most email A/B tests
Assume a 2.5 percent baseline. That is a realistic benchmark for the share of recipients who complete a meaningful action, and you can review how this baseline varies by use case in our latest email performance benchmark report.
Now run the power calculation. Detecting a 10 percent relative lift on that baseline (from 2.5 percent to 2.75 percent) at 95 percent confidence and 80 percent power requires about 61,000 recipients per variant. That means over 120,000 randomized recipients in a single campaign to reliably spot a lift worth discussing in a finance meeting. You can verify this using Evan Miller’s sample size calculator; the result changes little under reasonable assumptions.
Most retail brands do not have a 120,000-person test cell for one campaign. So teams test with what they have. A test with 1,000 recipients per variant at a 2.5 percent baseline has only about 5 percent statistical power, meaning a real 10 percent lift would be missed 19 times out of 20. Worse, when an underpowered test does reach statistical significance, the effect size is often inflated. These exaggerated results become the ones that end up on slides. Peep Laja’s team at CXL has been clear about this for years: significance is not meaningful unless the sample size is sufficient.
Now consider open rates. Detecting the same 10 percent relative lift on a 20 percent open rate baseline takes about 6,500 recipients per variant, less than a ninth of the sample needed for conversion. This gap explains why email teams default to subject line testing. Subject line tests deliver results by Friday. Revenue tests rarely deliver at all. This is not laziness; teams naturally test what their list size allows, and ESP tools reinforce the habit by treating subject line and conversion tests as the same feature with identical defaults. The issue is the CEO does not fund open rates.
Valid statistics, wrong instrument
Assume you have the volume. A brand sending to 500,000 people per campaign can power a conversion test. Campaign-level A/B testing still fails to answer the CEO’s question in three ways, and no increase in sample size fixes these issues.
The operational tax. Each A/B test requires two full email builds and two QA passes. Test three modules independently and you are managing eight versions of one email, since variants grow as 2 to the power of N. Teams respond by reducing test frequency, which slows learning.
Re-randomization. The ESP reshuffles the audience each send. Tuesday’s treatment group becomes Thursday’s control group, so no longitudinal signal accumulates. You can never see how the program affects the same people over time. That compounding effect is what the budget should buy, but the method cannot detect it.
Per-email lift is not incrementality. Under last-touch attribution, revenue shifts freely between triggered and batch campaigns. Improve one campaign dramatically and it may claim conversions the abandoned-cart flow would have captured two hours later. A large single-email lift can happen with zero net growth for the channel. I detailed this flaw in Email A/B Testing Is Statistically Broken. Here’s the ROI Case for Holdout Measurement.
The Smart Banners holdout: a channel-level answer in about four weeks
The solution is not a larger list. It is a broader test. Stop comparing one campaign to another and ask whether the entire program grows email revenue. This question has a measurement design a CFO can audit, and it is the test we run for Smart Banners programs:
- Person-locked assignment. Each subscriber is assigned to treatment or control at first open and remains in that group for the duration. No reshuffling, so the longitudinal effect is preserved.
- A 50/50 split. Traditional 10 percent holdouts limit their own power; PostHog’s holdout testing guide estimates the user volume needed to detect a small holdout at low conversion rates in the hundreds of thousands. A balanced split maximizes insight from every subscriber.
- Channel-level scope. Revenue is tracked by UTM and pooled across broadcast, triggered, and transactional emails. Revenue shifting between flows cancels out because both sides are included in the test.
- Transactions times blended AOV. Comparing raw revenue lets a single $900 order skew results. Counting transactions and applying a single blended average order value removes order-size noise, so the test resolves faster.
- Roughly four weeks to an answer. Every send during the test period contributes to the same two cohorts, so events accumulate in a way no single campaign can match. Lists too small for a one-send conversion test still reach 95 percent significance or better.
- Re-runnable. When leadership asks the same question next year, run the test again.
Not every program should run this test immediately; volume and timing still matter, and I covered the criteria in Email Holdout Testing: When It’s Worth Running and When It’s a Waste of Time.
The collapsed-pixel control only Smart Banners can run
The key detail that keeps this design clean is what the control group receives: the same email, but the Smart Banner slot appears as a transparent 1×1 pixel. One build, one send, one placement. The only difference between groups is the module. Treatment versus literal nothing, for the same audience at the same time.
Paid media cannot create this control. Ghost ads and PSA tests use approximations of nothing, and geo holdouts trade individual precision for regional averages. Since Apple’s App Tracking Transparency launched, Ruler Analytics estimates ad platforms see only 40 to 60 percent of conversions, so even the treatment side is partially blind. Email has none of these limits. Delivery logs are person-level, you know exactly who saw the module, and conversions link to first-party identity that does not decay when a privacy prompt appears. I explained the mechanics in The Collapsed-Pixel Holdout: Why Smart Banners Run the Incrementality Test Paid Media Structurally Cannot, and for a practical look at what these modules do, start with The Ultimate Guide to Smart Banners.
Two numbers, two jobs
Attribution still has value. It just has a different role. Block-level RPM and click-to-conversion, measured over a 7-day click window, guide weekly optimizations: which module earns its place, which offer deserves the spotlight. Inside Smart Banners, this happens through a Multi-Arm Bandit or fixed split, using one email build and no extra QA. No explosion of variants, no second round of testing.
The holdout serves a separate purpose: proving to finance that the program expanded email revenue. Across the Zembula platform, personalized Smart Banner and Smart Kicker content achieves click-to-conversion rates far above the 2.5 percent whole-email benchmark. We are so confident in the channel-level outcome that our contracts include a guaranteed 10x ROAS floor for Smart Banners. Compare that to 2025’s average ecommerce ROAS of 2.87, and you have the kind of comparison that belongs in a budget review. The latest performance benchmarks break down these results by use case.
Two numbers, two purposes. Attribution informs weekly adjustments. The holdout justifies budget increases. Confusing these roles leads teams to defend statistical noise in front of decision-makers who could approve growth. If a test result lands on your desk this quarter, ask one thing first: what was it measured against? “The other variant on one send” is at best a tactical insight. “A person-locked holdout across the entire channel for a month” is credible evidence. That is the answer Smart Banners are designed to deliver.
Key takeaways
- To detect a 10 percent relative lift from a 2.5 percent conversion baseline requires approximately 61,000 recipients per variant, totaling over 120,000. Most typical test groups have only about 5 percent statistical power, so most declared wins are statistical noise.
- The same 10 percent lift on a 20 percent open rate baseline requires roughly 6,500 recipients per variant, about one-ninth of the sample needed for conversion. This explains why subject line tests deliver results while revenue tests often do not.
- Even a fully powered A/B test at the campaign level cannot prove overall program growth, because audiences are re-randomized every send and last-touch attribution allows revenue to shift between flows.
- A channel-level Smart Banners holdout, with assignment locked at first open, a 50/50 split, UTM-tagged revenue across all email types, and measurement using transaction count times blended AOV, reaches 95 percent significance or better in about four weeks and can be repeated as needed.
- The collapsed-pixel control, where the Smart Banner renders as a 1×1 transparent pixel for the control group, creates a true test against nothing, a control setup that paid media cannot replicate.
- Use block-level RPM and click-to-conversion to optimize live content. Use the holdout to prove channel growth. Bring the second number to budget discussions.
Grow your business and total sales



