River
Y CombinatorBacked by Y Combinator
FREE TEMPLATE

Data Journalism Analysis Template

Every data story template asks you to state your caveats. Write down the choices instead, run all of them, and publish the distribution.

Free download  ·  No account needed

A finding is not a number. It is one number out of a grid of defensible ones, and the grid is computable. Write down every choice the analysis has to make, write the alternative a reasonable analyst would defend for each, and run the whole cross product before you write a sentence. Doing it in that order matters, because once the headline figure is known each alternative stops being assessed on its merits and starts being graded against it.

This pack ships filled in for one story at The Cordell Review, an invented paper, working the food service inspection records of Marleigh State Department of Health, an invented agency. Under the specification the reporter chose, restaurants in the poorest ZIP quartile were ordered closed 2.31 times as often as those in the richest. Seven analyst choices with two or three defensible alternatives each make 648 specifications. Of those, 144 reach 2.0, 553 show some disparity, and 95 reverse the finding.

The median specification gives 1.48, not 2.31, and that gap is invisible from the one number. So is the reason: the denominator carries the result, because the state inspects a poorest-quartile restaurant 11.8 times over seven years against 19.2 in the richest. Switch to a per-facility rate and adjust for chain ownership, both defensible, and the finding lands at 0.97. That is the first thing a press office does, which is a good reason to do it first.

One finding, 648 specifications, one load-bearing choice

The Sensitivity Check, the Specification Register, and the counts in Analysis Workings behind the headline figure.

Sensitivity Check

The Cordell Review, an invented paper. Marleigh State Department of Health food service inspections 2018 to 2024, 63,155 rows after cleaning. Closure orders, poorest ZIP quartile against richest.

 SpecificationsShare
2.0 or more, the headline as written14422%
1.5 up to 2.017126%
Above 1.0 but under 1.523837%
1.0 or below, reverses the finding9515%
Grid648100%
MeasureValue
The specification the reporter chose2.31
Median of all 6481.48
Lowest0.62
Highest3.96
Showing some disparity553 of 648

What the 95 reversals have in common

Choice and alternativePresent in
Chain or independent: adjusted89 of 95
Inspection types: all types included69 of 95
Denominator: per 1,000 facilities56 of 95

Never the word robust. The counts say more and cannot be argued with. A finding that holds in 553 of 648 defensible specifications and reverses in 95 is real and conditional, which is a more useful sentence for an editor than either verdict. The chain adjustment sitting in 89 of the 95 reversals is not a robustness statistic, it is a fact about the story: the disparity is substantially about which restaurants are chains.

Which choice carries the result

Each choice ranked by how far the median moves across its own alternatives.

ChoiceMedian spreadLowest medianHighest median
Denominator0.781.24 per facility2.02 per inspection
Inspection types included0.641.211.85
Chain or independent0.571.211.79
Income measure0.311.311.62
Period0.281.351.63
Scale change at CB-2020.10.221.381.60
27 orphan county records0.011.481.49

One swap at a time, everything else as chosen

SwapFindingChange
As chosen2.31 
Per 1,000 facilities1.42-0.89
Per 1,000 facility-years1.55-0.76
Adjust for chain ownership1.57-0.74
All inspection types1.82-0.49
Routine plus complaint2.80+0.49

The specification a press office will build

Both swaps togetherFinding
Per 1,000 facilities, chain adjusted0.97

Two defensible swaps, each survivable alone, not survivable together. Changing one thing at a time misses exactly this case, which is why the whole cross product gets run. The bottom row of the ranking earns its place too: whether to keep 27 records under an unexplained county code was argued about for an hour during cleaning and moves the median by 0.01.

Analysis Workings

One line per quantity, so the headline figure can be recomputed without asking.

 QuantityPoorest quartileRichest quartile
W-01Closure orders, routine inspections402200
W-02Routine inspections11,84213,609
W-03Per 1,000 inspections33.914.7
W-04Distinct facilities1,004709
W-05Per 1,000 facilities400.4282.1
W-08Routine inspections per facility11.819.2
 Rate ratio, poorest against richestValue
W-09Per 1,000 routine inspections, as published2.31
W-10Per 1,000 facilities1.42
W-11Per 1,000 facility-years1.55
W-12Inspection frequency, poorest against richest0.61

W-12 equals W-10 divided by W-09, which is the whole explanation: the state inspects a poorest-quartile restaurant 11.8 times over seven years against 19.2 in the richest, so a per-inspection rate and a per-facility rate answer two different questions and the gap between them is exactly that frequency difference. Both numbers are true. The sentence that publishes names which one it is using.

What's in the pack

01

Specification Register

Every choice the analysis makes, every alternative a competent critic would defend, which one is in use, and why the others are arguable. Including the choices inherited from the cleaning pass, because a reader cannot tell the difference.

02

Sensitivity Check, run as a cross product

All 648 specifications, not one variation at a time. Changing one thing while holding the rest fixed misses the case that matters: two defensible choices that are each survivable alone and not survivable together.

03

The distribution, as four counts

How many specifications support the claim as written, how many weaken it, how many reverse it, and the median next to the one you chose. No verdict, because a count survives an argument and the word robust does not.

04

The load-bearing choice, ranked

Each choice measured by how far the median moves across its own alternatives. The top one is the paragraph the story owes the reader. The bottom one is the argument nobody needs to have again.

05

The specification a press office will build

The two most damaging defensible swaps from different choices, applied together and computed. Being the person who worked it out is worth more than being the person who gets shown it.

06

Analysis Workings

One line per quantity with its source and derivation, so a reader can recompute the headline figure without asking. Numerators, denominators, every rate ratio, and the frequency ratio that explains the gap between them.

07

Findings Memo

The claim the grid actually supports, next to the two versions it does not: the flat assertion only a minority of specifications reach, and the hedge that throws the finding away. Plus what the finding is not.

08

Methodology Note, for publication

Written for a rival newsroom, an academic and the subject's press office. The counts, the register, the distribution, the load-bearing choice, specific limitations, and a corrections promise the grid makes keepable.

09

Expert Review Request

A bounded ask with the workings attached and three questions, aimed at somebody who could break the finding. Sent before publication, with a real deadline, and never asking whether it looks right.

How to use it

  1. 1

    Start from a file with a change log

    Open the pack and send the cleaned dataset, or download the blank register and sheets. If the cleaning is not documented, the dataset cleaning pack is the pass before this one.

  2. 2

    Write the grid before the number

    Every choice, every defensible alternative, including the ones the cleaning already decided for you. Report the grid size before running anything, because knowing it changes how the finding gets written.

  3. 3

    Run all of it, report counts

    The full cross product, then the four band counts, the median against your chosen specification, and the range. Then what the reversals have in common.

  4. 4

    Name the choice, then send it out

    Publish the load-bearing choice and its mechanism yourself, compute the hostile specification, and put the grid in front of two reviewers who would disagree with each other.

Frequently asked questions

Is this template free?

Yes. Download the register, workings, sensitivity sheet and the three documents as Word and CSV files, no signup and no credit card. "Edit with AI" is the optional path for anyone who wants the agent to build the grid from a cleaned dataset and run it.

Is this not just a robustness check with extra steps?

A robustness check varies one thing and reports that the result held. This runs the cross product and reports the distribution, which catches the case a one-at-a-time check cannot: two defensible swaps that each leave the finding standing and together take it to 0.97.

Where does the method come from?

Specification curve analysis, from the Nature Human Behaviour paper that named it. Its argument is that analytical decisions which are defensible, arbitrary and motivated bias results towards the author's own narrative, and that none of that variability shows up in a standard error.

Does a grid this size not just make every finding look weak?

It makes conditional findings look conditional, which they are. On the worked file 553 of 648 specifications point the same way, and that is a stronger sentence than the unqualified one. What the grid kills is the flat claim that only its own top quartile supports.

How do I pick which alternatives belong in the grid?

One test: would somebody who understood the data and did not want the story to be true defend it in front of an editor? Not whether it is better. Padding a grid with alternatives nobody would argue for is worse than having no grid, because the curve then reads as reassurance.

The period and the geography feel like nitpicks. Do they matter?

Sometimes, and the ranking tells you which. Starting a series at the first period you happen to hold is a documented way to manufacture a trend, and administrative data is routinely grouped to the wrong geography for the question. Here they are worth 0.28 and 0.31.

Where does this sit in the rest of the reporting?

The pass before it is cleaning: the dataset cleaning pack produces the working copy this grid inherits choices from. Downstream, the chart selection brief picks the honest chart for whatever the grid supports, and the fact check pack checks each published figure against these sheets. The investigation evidence pack holds what no dataset can establish.

Run the grid before somebody else does

Download the specification register, workings and sensitivity sheet as Word and CSV files, or open this exact pack in River and send it your cleaned dataset.

Edit with AI