River
Y CombinatorBacked by Y Combinator

Product & DesignFree

Product Experiment Design and Brief

The brief calculates the smallest effect your real traffic can detect, so an experiment that cannot answer the question never gets run.

Start here

River designs the test before you run it. The hypothesis, the baseline numbers and your actual traffic to the surface being changed go in. The brief comes back with the metric that decides it, the smallest effect that traffic can detect, and the runtime that reaches it. Where the answer is that this experiment cannot resolve the effect you care about inside a quarter, the brief says so on the first line rather than in a footnote three weeks later.

Most ideas do not work, which is the whole reason to test them. Kohavi and colleagues reported that only one third of the ideas tested at Microsoft improved the metric they were designed to improve. Run a test with too little traffic and you lose the ability to tell the third that worked from the two thirds that did not. Every brief template on page one asks for a hypothesis. None of them checks whether the hypothesis is answerable.

Built for the product manager about to spend a sprint on a test, and for the growth team running three at once on overlapping traffic. Reach for it before the ticket is written, not after the results come in. Reading the result once it is finished is the experiment readout. Where users are dropping out in the first place is funnel and retention analysis. What the change becomes in the plan is the roadmap template.

A growth team reviewing conversion baselines and traffic volumes before designing a test
Built for the hour before the experiment ticket gets written, when the sizing question is still cheap to ask.

The number that decides whether a test is worth running

The minimum detectable effect is the smallest change a test can reliably find. Evan Miller's calculator states it exactly: the smallest effect that will be detected a chosen percentage of the time, with everything smaller indistinguishable from no change at all. It falls out of three numbers you already have: the baseline rate, the traffic reaching the surface, and how long you are willing to wait. Compute it first and half the arguments about experiment design disappear before they start.

A checkout step converting at 3.1 percent, with 1,200 visitors a week reaching each variant. At the conventional thresholds, two weeks of traffic can only detect a lift of about 1.6 points, which is a 50 percent relative improvement. Nobody moves a checkout 50 percent by rewording a button. The half-point you actually expect needs roughly seventeen weeks. That is a ten-minute calculation, and it produces a very different sprint plan.

The other half of a brief is stating in advance what result changes what you do. A test with no decision attached produces a number and an argument. Write down now what you will ship if the variant wins, what you will ship if it loses, and what you will do if the interval straddles zero, which on most tests is the likeliest outcome. The brief carries all three, so the readout has something to be checked against rather than interpreted.

How it works

  1. State the hypothesis

    What you believe will change, for whom, and the mechanism you think produces the change.

  2. Hand over the numbers

    Baseline conversion on the surface being changed, and how much traffic actually reaches it.

  3. Size the test

    The smallest effect that traffic can detect, and the weeks needed for the one you expect.

  4. Commit the decision

    What ships on a win, on a loss, and on a result that lands on zero.

What you get

  • The minimum detectable effect your real traffic supports, not a rule of thumb
  • Runtime per arm to reach the effect you actually expect, in weeks
  • One primary metric named up front, with the guardrails that would veto a win
  • What you will do on a win, a loss, and an inconclusive result
  • A plain verdict on whether this test can answer the question at all
  • The cheaper alternatives when the traffic will not support a clean test

Common questions

What if the answer is that I do not have the traffic?

Then you have saved a sprint, which is the most valuable output this produces. The brief says how much traffic you would need and how long it would take. Then it gives the alternatives: test a bigger change, move the test earlier in the funnel where volume is higher, or accept a qualitative answer from a handful of sessions.

Can I just run it until it looks significant?

No, and the brief fixes the runtime in advance precisely to stop that. Checking daily and stopping the first time a p-value dips below the threshold inflates the false positive rate well past the 5 percent you think you are running at. If you need to peek, the brief says so and sizes the test for a sequential method instead.

How many metrics should a test have?

One that decides it, and a short list of guardrails that can veto a win. Twenty metrics on a dashboard means something crosses the threshold by chance on almost every test, which is how a losing variant gets shipped on the strength of a segment. The brief names the decider before any data exists.

Does it work for tests that are not A/B?

Yes. A holdout, a staged rollout by cohort, a before-and-after with a control group, and a switchback all get the same treatment. What would have to be true, what you can measure, and whether the design can distinguish your effect from the noise. The sizing changes; the question the brief asks does not.

We run several tests at once. Does that matter?

It matters when they touch the same users or the same surface. The brief asks what else is running and flags the overlaps, because two experiments changing the same page make each other's results hard to attribute. Where they are genuinely independent it says so, and where they are not it proposes an order.

What happens after the test finishes?

The brief becomes the thing the result is checked against, which is what stops a post-hoc story replacing the original question. Hand both to the experiment readout and it reports the result against the decision you committed to, including when the honest answer is inconclusive.

Product Experiment Design and Brief

Fill in the form and your workspace opens with the work already underway.