Data Journalism Analysis Template
Every data story template asks you to state your caveats. Write down the choices instead, run all of them, and publish the distribution.
Free download · No account needed
A finding is not a number. It is one number out of a grid of defensible ones, and the grid is computable. Write down every choice the analysis has to make, write the alternative a reasonable analyst would defend for each, and run the whole cross product before you write a sentence. Doing it in that order matters, because once the headline figure is known each alternative stops being assessed on its merits and starts being graded against it.
This pack ships filled in for one story at The Cordell Review, an invented paper, working the food service inspection records of Marleigh State Department of Health, an invented agency. Under the specification the reporter chose, restaurants in the poorest ZIP quartile were ordered closed 2.31 times as often as those in the richest. Seven analyst choices with two or three defensible alternatives each make 648 specifications. Of those, 144 reach 2.0, 553 show some disparity, and 95 reverse the finding.
The median specification gives 1.48, not 2.31, and that gap is invisible from the one number. So is the reason: the denominator carries the result, because the state inspects a poorest-quartile restaurant 11.8 times over seven years against 19.2 in the richest. Switch to a per-facility rate and adjust for chain ownership, both defensible, and the finding lands at 0.97. That is the first thing a press office does, which is a good reason to do it first.
What's in the pack
Specification Register
Every choice the analysis makes, every alternative a competent critic would defend, which one is in use, and why the others are arguable. Including the choices inherited from the cleaning pass, because a reader cannot tell the difference.
Sensitivity Check, run as a cross product
All 648 specifications, not one variation at a time. Changing one thing while holding the rest fixed misses the case that matters: two defensible choices that are each survivable alone and not survivable together.
The distribution, as four counts
How many specifications support the claim as written, how many weaken it, how many reverse it, and the median next to the one you chose. No verdict, because a count survives an argument and the word robust does not.
The load-bearing choice, ranked
Each choice measured by how far the median moves across its own alternatives. The top one is the paragraph the story owes the reader. The bottom one is the argument nobody needs to have again.
The specification a press office will build
The two most damaging defensible swaps from different choices, applied together and computed. Being the person who worked it out is worth more than being the person who gets shown it.
Analysis Workings
One line per quantity with its source and derivation, so a reader can recompute the headline figure without asking. Numerators, denominators, every rate ratio, and the frequency ratio that explains the gap between them.
Findings Memo
The claim the grid actually supports, next to the two versions it does not: the flat assertion only a minority of specifications reach, and the hedge that throws the finding away. Plus what the finding is not.
Methodology Note, for publication
Written for a rival newsroom, an academic and the subject's press office. The counts, the register, the distribution, the load-bearing choice, specific limitations, and a corrections promise the grid makes keepable.
Expert Review Request
A bounded ask with the workings attached and three questions, aimed at somebody who could break the finding. Sent before publication, with a real deadline, and never asking whether it looks right.
How to use it
- 1
Start from a file with a change log
Open the pack and send the cleaned dataset, or download the blank register and sheets. If the cleaning is not documented, the dataset cleaning pack is the pass before this one.
- 2
Write the grid before the number
Every choice, every defensible alternative, including the ones the cleaning already decided for you. Report the grid size before running anything, because knowing it changes how the finding gets written.
- 3
Run all of it, report counts
The full cross product, then the four band counts, the median against your chosen specification, and the range. Then what the reversals have in common.
- 4
Name the choice, then send it out
Publish the load-bearing choice and its mechanism yourself, compute the hostile specification, and put the grid in front of two reviewers who would disagree with each other.
Frequently asked questions
Is this template free?
Yes. Download the register, workings, sensitivity sheet and the three documents as Word and CSV files, no signup and no credit card. "Edit with AI" is the optional path for anyone who wants the agent to build the grid from a cleaned dataset and run it.
Is this not just a robustness check with extra steps?
A robustness check varies one thing and reports that the result held. This runs the cross product and reports the distribution, which catches the case a one-at-a-time check cannot: two defensible swaps that each leave the finding standing and together take it to 0.97.
Where does the method come from?
Specification curve analysis, from the Nature Human Behaviour paper that named it. Its argument is that analytical decisions which are defensible, arbitrary and motivated bias results towards the author's own narrative, and that none of that variability shows up in a standard error.
Does a grid this size not just make every finding look weak?
It makes conditional findings look conditional, which they are. On the worked file 553 of 648 specifications point the same way, and that is a stronger sentence than the unqualified one. What the grid kills is the flat claim that only its own top quartile supports.
How do I pick which alternatives belong in the grid?
One test: would somebody who understood the data and did not want the story to be true defend it in front of an editor? Not whether it is better. Padding a grid with alternatives nobody would argue for is worse than having no grid, because the curve then reads as reassurance.
The period and the geography feel like nitpicks. Do they matter?
Sometimes, and the ranking tells you which. Starting a series at the first period you happen to hold is a documented way to manufacture a trend, and administrative data is routinely grouped to the wrong geography for the question. Here they are worth 0.28 and 0.31.
Where does this sit in the rest of the reporting?
The pass before it is cleaning: the dataset cleaning pack produces the working copy this grid inherits choices from. Downstream, the chart selection brief picks the honest chart for whatever the grid supports, and the fact check pack checks each published figure against these sheets. The investigation evidence pack holds what no dataset can establish.
Run the grid before somebody else does
Download the specification register, workings and sensitivity sheet as Word and CSV files, or open this exact pack in River and send it your cleaned dataset.
Edit with AI