River
Y CombinatorBacked by Y Combinator

Product & DesignFree

A/B Test Results Readout Template

The confidence interval, the effect the test was actually powered to find, and a count of every comparison you made along the way.

Start here

River's experiment readout reads the numbers rather than asking you to summarize them. Variant counts, exposures, guardrails, dates and whatever segments you cut go in. What comes out is the effect with its confidence interval, a statement of the smallest effect the test was actually built to detect, the guardrails checked, and a decision with the reasoning attached. Every figure shows its arithmetic, so a disputed number is settled by reading the line rather than rerunning the query.

The templates that rank for this query are well made, and they all end in the same place. One offers three choices for the result: control wins, variant A wins, or inconclusive. Three words to carry everything the interval knows. The standard advice for a null is worse. A widely used readout skill tells you to look for subgroups with significant effects, which is significance hunting offered as a remedy. None of them reads your numbers; every one waits for you to have done the thinking already.

Built for the product manager writing up a test the team hoped would win, the growth engineer running several a week, and the analyst who has to defend a number in a room. Reach for it when the result is not a clean win, which is most of the time. If the experiment has not been designed yet, the experiment design and brief comes first, and a run of readouts eventually becomes the twelve-month roadmap narrative.

Not significant is not the same as no difference

A readout format that only works for winners is broken most of the time it is used. Microsoft published the count from its own experimentation platform: only one third of the ideas tested improved the metric they were designed to improve. The remaining two thirds split between flat and actively negative. If two out of every three readouts you write are non-wins, the null case is not the edge case. It is most of the job.

Here is the shape of the problem. A checkout test runs to 62,000 users per arm and converts 4.00% against 4.17%, a relative lift of 4.2% at 87% significance. The template says inconclusive, so the change gets reverted. The 95% interval, though, runs from a 1.3% loss to a 9.7% gain, and at 80% power this test could only ever have detected a lift of 7.8% or more. It was never built to find the effect it found.

Then comes the segment hunt. Optimizely documents the catch in its own product, that when you segment results, the false discovery rate control is not maintained, and it tells you to use segments for exploration rather than decisions. The American Statistical Association puts the general form plainly: statistical significance does not measure the size of an effect or the importance of a result. An honest readout says how many places you looked, not only the one that paid.

How it works

  1. Paste the table

    Exposures and conversions per arm, the dates, the guardrails, and any segments you cut while looking.

  2. Recompute it

    River derives the effect, the interval and the detectable effect at the sample size the test actually reached.

  3. Read the result

    A readout stating what the test established, what it ruled out, and what it was never able to see.

  4. Rewrite for the room

    Ask in chat for the leadership version or the engineering one. Both come from the same numbers.

What you get

  • The effect with its confidence interval, rather than a label chosen from three
  • The smallest effect the test was actually powered to detect, stated up front
  • Guardrail metrics reported whether or not they moved, and never quietly dropped
  • Every comparison counted, including the segments and variants that came to nothing
  • A decision written with its reasoning and the cost of being wrong each way
  • A sheet holding the arithmetic behind each figure, so anyone can recheck it

Common questions

Our test hit 87% significance. Can we ship it?

That is a decision rather than a statistic, and the readout gives you what the decision needs. It reports the interval, the effect the test could detect, and the cost of being wrong in each direction. Shipping at 87% is defensible when the downside is small and the upside is large. It just has to be said out loud.

What is the difference between inconclusive and no effect?

Inconclusive means the test could not tell the two apart. No effect means the interval is narrow and sits close to zero. A test whose interval runs from a 1% loss to a 10% gain has established almost nothing, and calling that no effect reads your own blind spot as a finding.

We found a significant lift in one segment. Is that real?

Possibly, and the segment alone cannot tell you. Significance across segments is not corrected for how many segments you examined, so the more places you look, the likelier one lights up by chance. The readout records how many cuts were made, which turns the finding into a hypothesis worth its own test. The same caution applies to a satisfaction score cut by segment, where who answers moves too.

Does it need our raw event data?

No. Exposures and conversions per arm cover the primary metric, plus the same figures for any guardrail or segment you want included. Paste the summary table your platform already shows you. Given dates and daily counts it also checks whether the result held steady or rested on a single day. Pasting the block rather than the file is how a regression output gets written up too.

Can it handle a Bayesian result instead of p-values?

Yes. Give it the posterior and the credible interval and it reports in those terms, because the underlying job is unchanged. What matters is stating the range of effects the data supports and the range it excludes, whichever framework produced them. It will not quietly convert one into the other.

Who is the readout written for?

You choose, and the register changes with the answer. The engineering version keeps the interval and the power calculation. The leadership version leads with the decision and the money, carrying the caveat rather than burying it. Both come from one set of numbers, so they cannot disagree. For the wider pack, use the executive summary generator.

We run tests weekly. Does this help across a programme?

It does, because the readouts become comparable. Same fields, same statements, same treatment of nulls, so a quarter of experiments reads as one record instead of thirty documents in thirty voices. Pinning down the metrics themselves helps too, which is what a KPI dashboard description is for.

A/B Test Results Readout Template

Fill in the form and your workspace opens with the work already underway.