River
Y CombinatorBacked by Y Combinator
FREE TEMPLATE

Email A/B Testing Plan Template

Four documents and four sheets that work out whether your list can detect the effect you are testing for, before you test it.

Free download  ·  No account needed

Detectable Effect by Metric  ·  the calculator, inverted

One send can see 3.4% on opens and 51.5% on orders

Fellgate Outdoors  ·  38,400 subscribers, one campaign a week, 19,200 per arm

Sample per armWhere it comes fromOpen 41.20%Click 2.84%Order 0.386%
19,200one send, 50/50 split3.4%17.4%51.5%
38,400two sends pooled2.4%12.2%35.2%
76,800four sends, one month1.7%8.5%24.3%
3,840a 20% test holdout7.7%40.9%132.1%
500the flow significance floor21.3%131.7%none
50the campaign significance floor65.2%nonenone

Nothing in a subject line moves order rate by half. So the test that gets run is the one whose answer does not matter.

The same floors, counted in events

500 recipients per arm is 206.00 opens, 14.20 clicks and 1.93 orders. 50 recipients per arm is 20.60 opens, 1.42 clicks and 0.19 orders. A threshold that can be met by two orders will call a coin flip a winner, and it is the threshold most programmes are trusting.

Why a low baseline is punishing

The required sample divides by the squared absolute difference, and a 10% lift on 0.386% is a difference of 0.0386 points. Detecting it takes 425,277 per arm, roughly 188 times what the same relative lift costs on a 41.20% open rate.

Revenue per recipient is worse, not better

It is a mean rather than a rate, so the order-value spread adds variance. At $106.70 average order value with a $48.00 spread, a 10% lift needs 487,409 per arm against the order rate's 425,277. It is also the metric every testing guide names.

95% confidence, 80% power, two-proportion form with each arm's own variance. The revenue row uses the two-sample form on a mean.

Fellgate Outdoors ran fourteen email tests last year. The sending platform declared a winner on twelve of them, every one configured correctly and read against the platform's own published significance rule. Recompute the sample each of those twelve actually had, and four could detect the lift they reported. The other eight reported a lift smaller than the smallest lift their own sample could resolve, which means they are consistent with nothing happening at all. Two thirds of the year's winners were noise.

The check that would have caught every one of them takes three numbers and about a minute. Fellgate mails 38,400 subscribers weekly, so a 50/50 split gives 19,200 per arm. At that sample one send can detect a 3.4% relative lift in open rate, a 17.4% lift in click rate and a 51.5% lift in order rate. Nothing anybody has ever written into a subject line moves order rate by half. Detecting a tenth of that takes 22.15 sends.

That gap is why almost every email test in existence is about subject lines. Open rate is cheap to test because 41.20% of subscribers do it, and order rate costs roughly 188 times the sample for the same relative lift, because the formula divides by a squared absolute difference. Pair this with the lifecycle email programme pack and the win back email template, or take the files blank from the template library.

What each test needs, what last year's tests actually had, and what happens when you watch

The Sample Size Calculator, the register audit it produces, and the simulated cost of reading a test daily.

Sample Size Calculator

Subscribers needed per arm at 95% confidence and 80% power, with the weekly sends that takes underneath. Baselines are medians across Fellgate's last twenty comparable sends, not pooled totals.

MetricBaseline+5%+10%+15%+20%+30%+50%
Open rate41.200%9,0212,2661,01056925288
sends needed0.470.120.050.030.010.00
Click rate2.840%220,02256,30325,59814,7226,8282,662
sends needed11.462.931.330.770.360.14
Order rate0.386%1,660,778425,277193,491111,35851,73020,232
sends needed86.5022.1510.085.802.691.05
MetricMeanSDCoeff of variation+10% per armsends+30% per armsends
Revenue per recipient$0.4119$7.257417.62487,40925.3954,1572.82

The cell that decides the programme is order rate at +10%: 425,277 per arm, or 22.15 weekly sends. Nobody runs a subject-line test for five months, which is why almost nobody has run one that could resolve. Read the row across instead: a 30% change needs 2.69 sends and a 50% change needs 1.05.

Revenue per recipient needs more sample than the order rate it is built from. 487,409 per arm against 425,277, because the order-value spread adds variance a proportion test does not carry. A revenue-per-recipient winner from a single send is a report of which arm caught the larger orders.

The denominator is squared, so halving the effect you want to detect quadruples the sample. That is why aiming for a small improvement is an expensive ambition rather than a modest one.

Test Register, re-audited

Fellgate's twelve months of tests, each re-read against the sample it actually had. Powered means the reported lift is at or above the smallest lift that sample could detect.

TestMetricSplitPer armEvents per armReported liftMDEPoweredPlatform called it
T01 Subject: question vs statementOpen100%19,2007,910.46.1%3.4%yeswon
T02 Subject: first name tokenOpen100%19,2007,910.42.4%3.4%nowon
T03 Preheader lengthOpen30%5,7602,373.13.8%6.3%nowon
T04 Send time 6am vs 10amOpen100%19,2007,910.49.3%3.4%yeswon
T05 Hero image vs product gridClick100%19,200545.311.2%17.4%nowon
T06 CTA copy: Shop vs See the rangeClick100%19,200545.37.4%17.4%nowon
T07 One CTA vs threeClick100%19,200545.318.6%17.4%yeswon
T08 Button colourClick20%3,840109.15.5%40.9%noinconclusive
T09 Free shipping bannerOrder100%19,20074.121.4%51.5%nowon
T10 Discount 10% vs free shippingOrder100%19,20074.133.1%51.5%nowon
T11 Review quotes in templateOrder100%19,20074.19.8%51.5%nowon
T12 Urgency line in subjectOrder30%5,76022.215.2%103.1%nowon
T13 Plain text vs designedClick100%19,200545.340.2%17.4%yeswon
T14 Category nav in headerOrder20%3,84014.88.7%132.1%noinconclusive
14 run4 powered12 won

Eight of the twelve declared winners reported a lift below their own detection floor. Two thirds of the year's winners are consistent with no effect. This is not carelessness: every one ran to the platform's published significance rule, and nothing in that workflow asks whether the sample could see the effect.

All four survivors are open-rate or click-rate tests at a full 50/50 split. Four order-rate tests were run and none of them was resolvable. So the year's surviving evidence is entirely about the metrics with the weakest link to revenue, which is a consequence of the arithmetic rather than a choice anybody made.

Spent instead on order-rate tests of a 30% change, each needing 51,730 per arm or 2.69 sends, 19 of them fit in the same 52 sends and every one resolves.

Why the Peeks column exists

40,000 simulated A/B tests in which the two arms are genuinely identical, accruing 2,742 subscribers per arm per day for seven days. Click rate, 2.840% in both arms.

Reading conditionCalled significant at 95%NominalWhat it means
Read once, at the planned sample5.14%5.00%The procedure works. One in twenty, as advertised.
Read every day for seven days16.84%5.00%About one in six, with nothing happening at all.
Inflation3.28xThe cost of watching a dashboard that updates.

Watching a test produces a winner roughly one time in six when there is no difference to find. That is the default behaviour of every platform dashboard that updates while a test is live. The fix is not vigilance: it is a planned sample, a fixed number of looks, and a Peeks column that records whether both held.

An underpowered test that does reach significance overstates the effect. Only the large draws clear the threshold, so the observed lift is selected for being large. A result sitting at the detection floor should be read for direction and not for size, which is why Result Log's adopt decisions say so explicitly.

Two phrases the protocol refuses. Trending towards significance describes an unplanned look. Directionally positive describes a lift below the detection floor. Both convert a null into a finding.

What's in the pack

01

Sample Size Calculator

Six lift columns against your own baselines, per arm, with the number of weekly sends each cell costs written underneath it.

02

Detectable Effect by Metric

The same formula inverted, which is the version you use in practice: the smallest lift each realistic sample can actually see.

03

Test Register

One row per test with events per arm and the detection floor as planning fields, plus last year's tests re-audited against the sample they had.

04

Result Log

Finished tests including the null ones, with a Peeks column and the simulated false positive rate that makes it necessary.

05

Power and Sample Size Method

The formula, why a squared denominator makes small effects expensive, and the five ways the calculation comes out wrong.

06

Test Protocol

Two gates before anybody builds a variant, the eight-step sequence, and the four things a test is not allowed to be.

07

Hypothesis Format

A five-slot form whose mechanism clause is what makes a null result informative rather than quietly abandoned.

08

Reading a Result

The three mistakes that each produce a confident number, and what a properly powered null result is genuinely worth.

How to use it

  1. 1

    Open in River, or take it blank

    Give River your send history and let it compute the floors, or download the four documents and four sheets and run the formula yourself.

  2. 2

    Send twenty sends and any past tests

    Recipients, opens, clicks and orders per send. If you have a record of tests already run, that is what the audit needs.

  3. 3

    Read the detection floors first

    Four baselines, the sample one send gives you per arm, and the smallest lift each metric can resolve at it.

  4. 4

    Then size, or refuse, each test

    Every register row gets its events per arm and detection floor before a variant exists, and tests that cannot resolve do not launch.

Frequently asked questions

Is this template free?

Yes. Download the Word documents and CSV sheets with no account, no card and no email gate. Edit with AI is the optional half: River computes your baselines, the detection floors and the audit of tests you have already run. Every pack sits in the template library.

How big does my list need to be for A/B testing?

It depends entirely on the metric, which is why the question has no single answer. On a 0.386% order rate, detecting a 10% lift needs 425,277 per arm. On a 41.20% open rate the same relative lift needs 2,266. Same list, a factor of 188 apart.

What is a minimum detectable effect?

The smallest lift your sample could find if it were really there. It is the sample size formula solved for the effect instead of the sample, and it is the number that belongs in the plan. A lift below it is a null result, not a small win.

My platform said the test was significant. Is that not enough?

It is a published threshold rather than a power calculation, and it does not claim otherwise. Klaviyo's rule is 500 recipients per variation for flows with a 90% win probability. At a 0.386% order rate, 500 recipients per arm holds 1.93 orders.

How long should an email A/B test run?

Until the planned sample accrues, and not to a calendar. For a low-baseline metric that usually means pooling the same assignment across consecutive sends: four weekly sends give 76,800 per arm and take the order-rate floor from 51.5% to 24.3%.

Can I check the results while the test is running?

Looking is free, deciding is not. Simulated across 40,000 tests with identical arms, seven daily reads call something significant 16.84% of the time against a nominal 5.00%. Fix the number of looks in advance and record it, which is what the Peeks column is for.

What is the campaign threshold, and why does it differ?

Klaviyo documents that a campaign needs 50 recipients per variation and a 90% win probability, against 500 for a flow. At Fellgate's rates, 50 recipients per arm contains 1.42 clicks and 0.19 orders, so it resolves nothing below open rate.

Find out what your list can actually detect

Take the Word documents and CSV sheets blank, or open this exact pack in River and send it twenty sends of history.

Edit with AI