Robustness Checks Template for Research
Two documents and two sheets that test a headline finding against reasonable alternatives and log every result, including the one that breaks it.
Free download · No account needed
Specification Register
Alderwood Summer Bridge Program
| Spec | Estimate | 95% CI | p | Result |
|---|---|---|---|---|
| A — Headline (n=676) | 0.28 | [0.10, 0.46] | .0019 | Baseline |
| F — Survey completers (n=390) | 0.09 | [-0.16, 0.34] | .4887 | Breaks |
Restricting to the 390 of 842 students who answered an optional follow-up survey drops the estimate by two-thirds and erases its significance. The row stays in the register rather than the footnotes.
A finding survives review because of what a researcher chose to test, not just what the data showed. Simmons, Nelson, and Simonsohn's 2011 analysis of researcher degrees of freedom names four common ones: which covariate to include, when data collection stops, which conditions to compare, and which dependent variable to report. Stacking all four can push a false-positive rate to 61 percent. A check chosen after seeing that it flatters the result is that same flexibility under a different name. A plotted grid of alternative models cannot say which kind it is looking at.
This space writes the reasoning before it writes a result. Setup holds one standing rule and a Specification Rationale naming the threat each planned check answers, filed before a single alternative model runs. A specification-curve package like specr computes that whole grid and warns against using it to somehow arrive at a better estimate. This space sits one layer above the grid: it tests the model already fitted in a methods section, not a substitute for the analysis or for the record of what produced it.
In the worked example, a community college's matched-comparison estimate of a summer bridge program's effect on first-term GPA is 0.28 points. It holds at 0.24 restricted to Pell-eligible students and at 0.22 with the largest contributing high school dropped. It falls to 0.09 and loses significance once the sample is limited to the 390 of 842 students who completed an optional follow-up survey. The comparison table condenses all six specifications for an appendix; formatting it for a target journal is a separate, later pass, and the analytic sample underneath every specification is unchanged throughout.
What's in the pack
Specification Register sheet
Every planned check's estimate, standard error, confidence interval, exact p-value and sample size against the headline's own row, classified as holds, attenuates or breaks rather than left for a reader to judge.
Comparison Table sheet
The register condensed to the form that pastes into a paper's appendix or a reviewer response: one row per specification in the same order, headline first.
Specification Rationale
The threat each planned check answers and why it is real for this specific design, written before a single alternative model runs, so a check chosen for a reason reads differently from one chosen after seeing the result.
Robustness Narrative
The pattern read in prose once every row is in: what held, what attenuated and by how much, and what the one that broke actually means for the claim, written last rather than alongside the checks themselves.
A specification earns its row by the threat it answers, not by what it finds
The standing space rule every prompt reads first. A grid of alternative models has no column for why any one of them was worth computing; this rule is what fills that in before a single check runs.
How to use it
- 1
Open in River, or download it
Open the pack in River and the agent writes the rationale and register with you from your own headline result, or download the two documents and two CSV sheets instantly, filled in with the worked example.
- 2
Send the headline result and the design
The estimator, the sample, the coefficient with its standard error and N, and anything about the data collection or the sample's structure that could plausibly threaten it, before any alternative model gets named.
- 3
Name each threat before running its check
Specification Rationale gets one entry per planned check, naming what changes and why that threat is real for this design, filed before that check's result exists anywhere in the space.
- 4
Log every result, then read the pattern
Specification Register and Comparison Table take every planned check's real result, including the one that breaks the finding, and Robustness Narrative reads what the full pattern does and does not establish.
Frequently asked questions
Is this template free?
Yes. Download the whole pack as a Word document and CSV sheets with no signup and no credit card, filled in for the worked example. Edit with AI is a separate, optional path that has the agent write the rationale and the register from your own headline result. The template library holds the rest of the packs.
What format are the downloaded files?
CSV for Specification Register and Comparison Table, and a Word document for Specification Rationale and Robustness Narrative, all in one zip. They open in Excel, Numbers, Sheets, Word, Pages and Google Docs with nothing to convert. Inside River the same content opens as live Docs and Sheets.
Doesn't a specification-curve tool already do this?
A tool like specr computes a whole grid of alternative models and plots how far an estimate moves, and its own documentation warns against using that grid to arrive at a better estimate. This space adds what a grid has no column for: the threat each check answers, named before it runs, and the result logged whatever it shows.
What do holds, attenuates and breaks actually mean?
Holds means the estimate and its significance both survive a check. Attenuates means the direction survives but the magnitude or significance weakens, the way a clustering correction can leave a point estimate unchanged while widening its interval past conventional significance. Breaks means the finding does not survive at all.
Does this run my analysis or choose my headline model for me?
No. It starts from a headline result you already have, the estimator, the sample, and the coefficient with its standard error, and tests that result against alternative choices. The methods section it came from and the code that produced it are a separate, earlier step.
What happens if a robustness check breaks my finding?
The row stays in Specification Register and Comparison Table rather than getting dropped or footnoted out. Robustness Narrative states plainly what a broken check means for the claim, usually that it holds for a narrower population than the headline sample suggested, not that the whole finding is invalid.
Find out whether your finding survives the alternatives
Send the headline result and the design behind it. The rationale, the register and the narrative come back with every planned check logged, including the one that breaks it.
Edit with AI