River
Y CombinatorBacked by Y Combinator

Research & PolicyFree

Sample Size Justification and Power Note

Send the design and what you can realistically recruit, get the required sample across a range of effect sizes rather than one.

Start here

River takes the design, the outcome type and the effect you are willing to assume, then runs the calculation as a grid rather than a point. One column is the required sample at each plausible effect size. The other direction matters more: given the sample you can actually recruit, what is the smallest effect you could detect at eighty per cent power. Both numbers get compared against the smallest effect that would mean anything, and the note says whether the study is worth running.

Search this and you get calculators. Enter alpha, enter power, enter your effect, receive one integer. The regulators' own harmonised statistical guideline asks for the opposite. It says it is important to investigate the sensitivity of the sample size estimate to a variety of deviations from these assumptions, facilitated by providing a range of sample sizes for a reasonable range of deviations. The grid is the expected form. The single integer is the deviation from practice.

Written for the applicant filling in a power section, the author answering a reviewer who disputes the assumed effect, and the student whose supervisor asked for a justification rather than a number. It uses the design your code will actually implement, which a methods section written from the code that ran already documents. After collection, the results section built from your output block reports what the sample delivered, and a cleaning log recording every decision explains any gap between recruited and analysed.

The funded design had 39 per cent power

A two-arm trial with a continuous primary outcome, two-sided alpha at 0.05 and eighty per cent power. The grant assumed a standardised difference of 0.50, which gives 63 per arm and 126 in total. That is the number in the application. Run the grid and the same design needs 129 per arm at 0.35, 175 at 0.30 and 252 at 0.25. The funded sample is right for one assumption and wrong for every neighbouring one.

Turn it round. At 126 total the power against 0.35 is 50.2 per cent, against 0.30 it is 39.1 per cent, and against 0.25 it is 28.9 per cent. If the literature's smaller trials cluster between 0.20 and 0.35, this study is more likely to miss a real effect than to find it. That is the sentence a reviewer is reaching for when they dispute your assumed effect, and it is better to write it yourself.

Now the feasibility side. Say the sites can deliver 240 analysable participants, 120 per arm. The smallest effect detectable there at eighty per cent power is 0.36, which is above the 0.25 the clinical team called the minimum worth having. Reaching 0.25 needs 504, or 2.10 times the ceiling. And with 18 per cent attrition, analysing 240 means recruiting 294, so a calculation run on the recruited figure is 54 people short before anyone drops out. The study as designed cannot answer its own question.

How it works

  1. State the design

    The arms, the outcome type, the test, and the analysis the sample has to support.

  2. Build the grid

    Required sample computed across the whole range of effects the literature makes plausible.

  3. Invert it

    The smallest detectable effect at the sample you can actually recruit and analyse.

  4. Write the verdict

    Whether the detectable effect is small enough to matter, and what it would take.

What you get

  • The required sample at every plausible effect size, not just the one you assumed
  • The smallest effect your feasible sample can detect, which is the reviewer's real question
  • Power against the effects the literature actually reports, stated for each one
  • The recruitment target inflated for attrition, since the calculation governs the analysed sample
  • Every assumption named with its source, so a reviewer disputes the source not your judgement
  • A verdict on whether the feasible sample can detect an effect worth detecting

Common questions

A reviewer says my assumed effect is too optimistic. What now?

Send them the grid. Their objection is about one cell, and a grid answers it without conceding the study. The honest version states the power at their effect as well as yours, then says what sample their assumption would require. If your feasible sample cannot reach it, that is worth knowing before you argue rather than after.

Can I run this after the data is collected?

Yes, and it is a different question then. Post hoc power computed from the observed effect is circular and the run will not produce it. What it will produce is the minimum detectable effect at your realised sample, which tells a reader what the study could never have found. That is a legitimate limitation paragraph.

Where should the assumed effect come from?

Published data or earlier trials, and the note records which. An effect taken from a small published study is usually inflated, so the grid deliberately includes cells below it. Where the only precedent is your own pilot, say so and treat the pilot estimate as a ceiling rather than a point, because pilots are not sized to estimate effects.

Does it handle designs other than two groups?

It handles what you describe: paired comparisons, more than two arms with a stated multiplicity correction, cluster designs with an inflation factor, and regression with a target coefficient. Each has different assumptions and the note lists them separately. For a cluster design the intracluster correlation is the assumption the grid varies, because that is where the sample actually goes.

Why inflate for attrition rather than just recruiting more?

Because the calculation governs the analysed sample, not the recruited one, and the two are routinely confused in applications. Regulatory guidance is explicit that the figure refers to the subjects required for the primary analysis. Recruiting your calculated number and losing a fifth of them leaves you underpowered with a calculation that looks correct on paper.

What if the answer is that my study is not worth running?

Then the note says it plainly, with the arithmetic, and that is the most valuable outcome available. It also lists what would change the answer: a more sensitive outcome measure, a longer follow-up, another site, or a narrower question the sample can carry. Knowing this before collection is worth considerably more than knowing it after.

What do I get back?

A Sheet holding the grid: required sample at each effect size, power at each effect for your feasible sample, and the recruitment target after attrition. A Doc writing the justification with every assumption named and sourced, the sensitivity to each, and the verdict on whether the detectable effect matters.

Sample Size Justification and Power Note

Fill in the form and your workspace opens with the work already underway.