River
Y CombinatorBacked by Y Combinator
FREE TEMPLATE

Beta Program Plan Template

Writing exit criteria before the beta is the easy half. This template checks whether the panel you recruited can produce the evidence they ask for.

Free download  ·  No account needed

Every beta template already tells you to write exit criteria before you start, and that advice is right and not sufficient. A criterion written on day zero can still be one your panel could never satisfy. Ten Enterprise verdicts reads like a reasonable bar until you notice that two Enterprise accounts will ever return feedback. At that point the beta has no way to end on evidence, so it ends on the day somebody loses patience with it, which is the failure the criteria were written to prevent in the first place.

So this pack costs each criterion against the people who can actually produce its evidence. Every row carries the source, the threshold, and the count of participants who could satisfy it if they were perfectly cooperative. Threshold minus pool sorts the gate into three piles: reachable, short by a countable amount, and unsatisfiable by any panel your customer base could supply. At the fictional Trellwood that came out four, four and one, and the whole exercise took an afternoon before any invitation went out.

Two conversions follow, and both are arithmetic rather than judgement. The gap divided by the rate evidence is arriving turns a shortfall into a date, which ends the argument about extending. The threshold divided by the invitation-to-evidence rate gives the invitation count the criteria actually needed, which at Trellwood was 379 against the 240 that went out. Send the scope, whatever anybody has said about the bar, and your customer counts by segment. What comes back first is the count of criteria the programme could never have closed, then a launch brief worth writing.

Nine criteria against a panel of 240 invitations

The gate with its pools attached, the funnel underneath it, and every feedback item pointed at the criterion it bears on.

Exit Criteria Tracking

Nine criteria, costed against the panel

Illustrative, for a fictional field-service scheduling product called Trellwood. 240 invitations, 78 accepted, 51 enabled, 34 used it, 19 used it enough to count, 12 returned structured feedback.

CriterionSourceNeedPoolWeeks to closeVerdict
E-1 No P1 open past five days, two weeks runningdefect tracker0 open340Reachable
E-2 A full week of routing with no manual fallbacktelemetry30196.4Short by 11
E-3 Route time down 8% on the account’s own baselinetelemetry20190.6Short by 1
E-4 Positive verdict from ten Enterprise accountsfeedback register101neverImpossible
E-5 Parity on the field mobile apptelemetry by platform891.2One spare
E-6 Three trades in the sample, five accounts eachparticipant register54pool-cappedShort by 1
E-7 Documentation rated sufficientfeedback register15122.8Short by 3
E-8 No rise in support contact ratesupport dataall340Reachable
E-9 Two accounts willing to be namedfeedback register2120Already met

E-4 is not short, it is impossible

One Enterprise account of 40 invited returned feedback, a rate of 2.5%. Ten verdicts at that rate needs 400 Enterprise invitations, and 74 Enterprise accounts exist. No extension closes it and no recruitment closes it. It gets rewritten as two named interviews, before the beta opens rather than at the gate.

Participant Register

A funnel, not a list of names

The last column is what makes it worth maintaining. A drop-out stops being one fewer name and becomes a hole in a specific criterion.

AccountTierTradeStageWeeks usedEvidence it can produce
Thurlow FacilitiesMid-marketHVACjudged10E-1, E-2, E-3, E-5, E-7, E-8, E-9
Bracken HeatingSMBHVACjudged9E-1, E-2, E-3, E-5, E-8, E-9
Marsden PlumbingSMBPlumbingjudged7E-1, E-2, E-3, E-7, E-8
Sowerby ElectricalSMBElectricaljudged6E-1, E-2, E-3, E-7, E-8
Ivywell ElectricalSMBElectricalactive5E-1, E-7, E-8
Nettleford GroupEnterpriseHVACactive3E-1, E-7, E-8
Kingsdale PlumbingSMBPlumbingenabled, never used0none
Crossfield ElectricalSMBElectricalnever enabled0none, and it is the cheapest route to E-6

The number every status report prints

51 participants. Seventeen of those enabled the feature and never opened it, and 19 have used it enough for their experience to count as evidence. Every pool in the gate is drawn from the 19, and no weekly report distinguished them from the 51.

Feedback Register

Every item names the criterion it bears on

The one field that keeps the gate current without anybody re-reading the register, and that stops the exit memo becoming a summary of mood.

FromWhat they saidBears onTypeDisposition
Calder MechanicalSplit shifts get one route, so the afternoon crew inherits the morning orderE-2defect P1Fixed 9 Feb
Bracken HeatingNine per cent off drive time in week one against their own December baselineE-3measurementCounted, above the 8% bar
Ivywell ElectricalOptimiser ignores that a part has to be collected before the second jobE-2design gap P2Open
Sowerby ElectricalSame parts-collection problem, arrived independentlyE-2design gap P2Merged. Two of four electrical accounts
Ellerby ServicesTheir dispatcher overrides the route most mornings and could not say whyE-2signalNeeds a session
Nettleford GroupSetup docs assume one depot and they run elevenE-7gap P3Backlog. The only Enterprise voice on record
Thurlow FacilitiesWilling to be named as a reference on the GA announcementE-9commitmentE-9 met, confirmed in writing

Silence is a row too

Twelve items came from eight of the 34 active accounts. Seven accounts have used the feature enough to count and have never returned anything, which is the entire three-account gap on E-7 and the reason that criterion is short.

What is in the pack

01

Exit Criteria Tracking

One row per criterion with its evidence source, threshold, the count of participants who could satisfy it, the gap, the weekly rate that evidence is arriving at, and the number of weeks that gap represents.

02

Participant Register

A funnel, not a list. Invited, accepted, enabled, used once, used enough to count, returned feedback, with the criteria each account can serve. Declines and never-enabled accounts stay in, because they are the informative rows.

03

Feedback Register

Every item logged against the criterion it bears on, or explicitly against none. Defects carry a severity and whether they reproduced, measurements carry the number and the baseline, confirmations get counted.

04

Programme Design

Scope, cohort, invitation count and duration, all derived from the criteria rather than chosen. Includes the wind-down sequence, written before it is needed and therefore fairer to the participants.

05

Exit Criteria and Participant Agreement

The gate in full, and the document you send participants. The agreement states the time cost in minutes per week and which criterion that participant is carrying, which is the best retention mechanism available.

06

Feedback Guide, plus a weekly sweep

How to ask so answers land as evidence, and an automation that recomputes every gap, flags a pool that has fallen below its threshold, and names the accounts that quietly stopped. A drop-out worth understanding needs a call rather than a form.

How it works

  1. 1

    Send what exists

    The feature scope, anything anybody has said about what finished looks like, your customer counts by segment, and the invitation list if there is one. Contradictory statements are useful rather than a problem.

  2. 2

    Each criterion gets a source and a pool

    River attaches the system the answer comes from, then counts the participants who could physically produce that evidence, applying every segment restriction the criterion carries in series rather than once.

  3. 3

    Threshold minus pool

    The gate sorts into reachable, short, and unsatisfiable by any panel you could recruit. Short criteria get a week count and a recruitment count. Impossible ones get a proposed rewrite before the invitations go out.

  4. 4

    Then it runs weekly

    Current counts refresh from each register, gaps recompute against the recent rate as well as the whole-programme rate, and accounts that stopped surface with the criteria they were carrying named beside them.

Frequently asked questions

Every beta template says define exit criteria first. What is different?

They all stop at writing the criteria down. None of them checks whether the panel can produce the evidence those criteria ask for. That is one subtraction per criterion, threshold minus the pool that can serve it, and it is available before a single invitation goes out. At Trellwood it retired one criterion and resized four.

How do you count the pool for a criterion?

Start at the narrowest funnel stage the criterion actually needs, then apply its segment restrictions in series. A criterion needing sustained use draws on accounts using it enough to judge, not accounts enabled, and those were 19 and 51. A criterion needing an Enterprise verdict draws on the intersection of two small numbers.

How many people should we invite?

Take the binding criterion's threshold and divide by your observed invitation-to-evidence rate. Trellwood needed 30 accounts using it enough to judge, its rate was 7.9%, so the criteria required 379 invitations and 240 went out. That is 32% of the whole customer base, which is a decision somebody senior should make on purpose.

Does this work for a mobile app beta?

Yes, and the platform constraints become criteria with thresholds already set. Apple caps a TestFlight programme at 10,000 external testers, and newer Google Play personal accounts must run a closed test with 12 testers opted in continuously for 14 days. Opted-in continuously is a pool, not an invitation count.

Our beta is already running and going long. Too late?

This is the better time, because the funnel rates are real rather than assumed. You get the gap on every open criterion, the weeks each one needs at the rate evidence is actually arriving, and the ones no extension can close. Trellwood's binding criterion was 6.4 weeks out, on a beta planned for six weeks and running eleven.

What do we do about a criterion nobody can satisfy?

Rewrite it, and record why. Ask what decision the criterion was protecting, then find evidence that answers the same question and is obtainable. Ten Enterprise verdicts protecting a segment decision becomes two named interviews. Deleting it quietly is the one option that costs you at the gate.

How does this fit with the rest of a launch?

It sits just before it. The gate passing is what makes a launch brief and the release notes honest, and the known gaps the memo names are what support and documentation need. If the feature is going out gradually afterwards, adoption gets measured separately.

Find out which of your criteria the panel can satisfy

Send the criteria and your customer counts by segment. The first thing back is the count of criteria no achievable panel could ever close.

Cost my exit criteria