Email A/B Testing Plan Template
Four documents and four sheets that work out whether your list can detect the effect you are testing for, before you test it.
Free download · No account needed
Detectable Effect by Metric · the calculator, inverted
One send can see 3.4% on opens and 51.5% on orders
Fellgate Outdoors · 38,400 subscribers, one campaign a week, 19,200 per arm
| Sample per arm | Where it comes from | Open 41.20% | Click 2.84% | Order 0.386% |
|---|---|---|---|---|
| 19,200 | one send, 50/50 split | 3.4% | 17.4% | 51.5% |
| 38,400 | two sends pooled | 2.4% | 12.2% | 35.2% |
| 76,800 | four sends, one month | 1.7% | 8.5% | 24.3% |
| 3,840 | a 20% test holdout | 7.7% | 40.9% | 132.1% |
| 500 | the flow significance floor | 21.3% | 131.7% | none |
| 50 | the campaign significance floor | 65.2% | none | none |
Nothing in a subject line moves order rate by half. So the test that gets run is the one whose answer does not matter.
The same floors, counted in events
500 recipients per arm is 206.00 opens, 14.20 clicks and 1.93 orders. 50 recipients per arm is 20.60 opens, 1.42 clicks and 0.19 orders. A threshold that can be met by two orders will call a coin flip a winner, and it is the threshold most programmes are trusting.
Why a low baseline is punishing
The required sample divides by the squared absolute difference, and a 10% lift on 0.386% is a difference of 0.0386 points. Detecting it takes 425,277 per arm, roughly 188 times what the same relative lift costs on a 41.20% open rate.
Revenue per recipient is worse, not better
It is a mean rather than a rate, so the order-value spread adds variance. At $106.70 average order value with a $48.00 spread, a 10% lift needs 487,409 per arm against the order rate's 425,277. It is also the metric every testing guide names.
95% confidence, 80% power, two-proportion form with each arm's own variance. The revenue row uses the two-sample form on a mean.
Fellgate Outdoors ran fourteen email tests last year. The sending platform declared a winner on twelve of them, every one configured correctly and read against the platform's own published significance rule. Recompute the sample each of those twelve actually had, and four could detect the lift they reported. The other eight reported a lift smaller than the smallest lift their own sample could resolve, which means they are consistent with nothing happening at all. Two thirds of the year's winners were noise.
The check that would have caught every one of them takes three numbers and about a minute. Fellgate mails 38,400 subscribers weekly, so a 50/50 split gives 19,200 per arm. At that sample one send can detect a 3.4% relative lift in open rate, a 17.4% lift in click rate and a 51.5% lift in order rate. Nothing anybody has ever written into a subject line moves order rate by half. Detecting a tenth of that takes 22.15 sends.
That gap is why almost every email test in existence is about subject lines. Open rate is cheap to test because 41.20% of subscribers do it, and order rate costs roughly 188 times the sample for the same relative lift, because the formula divides by a squared absolute difference. Pair this with the lifecycle email programme pack and the win back email template, or take the files blank from the template library.
What's in the pack
Sample Size Calculator
Six lift columns against your own baselines, per arm, with the number of weekly sends each cell costs written underneath it.
Detectable Effect by Metric
The same formula inverted, which is the version you use in practice: the smallest lift each realistic sample can actually see.
Test Register
One row per test with events per arm and the detection floor as planning fields, plus last year's tests re-audited against the sample they had.
Result Log
Finished tests including the null ones, with a Peeks column and the simulated false positive rate that makes it necessary.
Power and Sample Size Method
The formula, why a squared denominator makes small effects expensive, and the five ways the calculation comes out wrong.
Test Protocol
Two gates before anybody builds a variant, the eight-step sequence, and the four things a test is not allowed to be.
Hypothesis Format
A five-slot form whose mechanism clause is what makes a null result informative rather than quietly abandoned.
Reading a Result
The three mistakes that each produce a confident number, and what a properly powered null result is genuinely worth.
How to use it
- 1
Open in River, or take it blank
Give River your send history and let it compute the floors, or download the four documents and four sheets and run the formula yourself.
- 2
Send twenty sends and any past tests
Recipients, opens, clicks and orders per send. If you have a record of tests already run, that is what the audit needs.
- 3
Read the detection floors first
Four baselines, the sample one send gives you per arm, and the smallest lift each metric can resolve at it.
- 4
Then size, or refuse, each test
Every register row gets its events per arm and detection floor before a variant exists, and tests that cannot resolve do not launch.
Frequently asked questions
Is this template free?
Yes. Download the Word documents and CSV sheets with no account, no card and no email gate. Edit with AI is the optional half: River computes your baselines, the detection floors and the audit of tests you have already run. Every pack sits in the template library.
How big does my list need to be for A/B testing?
It depends entirely on the metric, which is why the question has no single answer. On a 0.386% order rate, detecting a 10% lift needs 425,277 per arm. On a 41.20% open rate the same relative lift needs 2,266. Same list, a factor of 188 apart.
What is a minimum detectable effect?
The smallest lift your sample could find if it were really there. It is the sample size formula solved for the effect instead of the sample, and it is the number that belongs in the plan. A lift below it is a null result, not a small win.
My platform said the test was significant. Is that not enough?
It is a published threshold rather than a power calculation, and it does not claim otherwise. Klaviyo's rule is 500 recipients per variation for flows with a 90% win probability. At a 0.386% order rate, 500 recipients per arm holds 1.93 orders.
How long should an email A/B test run?
Until the planned sample accrues, and not to a calendar. For a low-baseline metric that usually means pooling the same assignment across consecutive sends: four weekly sends give 76,800 per arm and take the order-rate floor from 51.5% to 24.3%.
Can I check the results while the test is running?
Looking is free, deciding is not. Simulated across 40,000 tests with identical arms, seven daily reads call something significant 16.84% of the time against a nominal 5.00%. Fix the number of looks in advance and record it, which is what the Peeks column is for.
What is the campaign threshold, and why does it differ?
Klaviyo documents that a campaign needs 50 recipients per variation and a 90% win probability, against 500 for a flow. At Fellgate's rates, 50 recipients per arm contains 1.42 clicks and 0.19 orders, so it resolves nothing below open rate.
Find out what your list can actually detect
Take the Word documents and CSV sheets blank, or open this exact pack in River and send it twenty sends of history.
Edit with AI