River
Y CombinatorBacked by Y Combinator
FREE TEMPLATE

Nonprofit Program Data Collection Template

Three documents and three sheets that design one collection backwards from every funder's indicator definitions, so a single set of fields answers every report.

Free download  ·  No account needed

Every nonprofit data collection template runs forwards. Start at the logic model, list your indicators, build a form per programme, then add another spreadsheet when a new funder arrives. Even the most rigorous version of that, a per-indicator reference sheet, specifies one indicator at a time and never asks what any two of them have in common. So the collection grows with the reporting obligation instead of with the facts sitting underneath it, and fragmentation is the predictable result rather than a failure of discipline.

This pack runs the other way. It takes each executed agreement, quotes the indicator definitions, splits them into the clauses that each have to be true, and names the facts those clauses require. In the worked example six awards ask Cedar Line Community Health for 22 indicators. Those definitions contain 58 separate testable clauses, and the clauses name 18 facts. Eight are already clean, so the whole distance between what six funders were promised and what the organisation can produce is ten fields.

Then every gap gets two properties nothing else computes. How many indicators it blocks, which puts a household key at the top at four, all four on one state contract, derivable from an address already on every intake form. And whether this cohort's value still exists, because two of the ten are gone: nobody records in June why somebody stopped attending in November. The indicators you committed to sit upstream of this and the impact report sits downstream.

Six awards, 22 indicators, 18 facts underneath them

Every funder definition broken into clauses, the union of the facts they name, and each gap ranked by what it blocks and whether it can still be recovered.

Funder Indicator Crosswalk

Illustrative, for a fictional organisation called Cedar Line Community Health. Diabetes education programme, Q3 FY2026, six awards with a live reporting obligation. This runs the opposite direction from a reporting crosswalk: from each funder's own wording to the facts that have to already exist for the figure to be possible.

AwardThe indicator, as the agreement words itClausesFacts it needsMissingProducible
A-01Unduplicated individuals with a billable encounter3D-01 D-03 D-04noneyes
A-01Race and ethnicity on the federal minimum categories3D-01 D-07D-07no
A-01Household income against the poverty guideline3D-01 D-09D-09no
A-02Households receiving service2D-01 D-02 D-03D-02no
A-02Households by county2D-02 D-10D-02no
A-03Individuals with three or more service contacts3D-01 D-03noneyes
A-03Reason for leaving among those who did not complete3D-01 D-05 D-06 D-17D-17no, and gone
A-05Service encounters delivered3D-03noneyes
A-06Clinical measure six months after completion3D-01 D-06 D-14 D-18D-14 D-18no, and gone
All six awards22 indicators, nine of them shown above5818 distinct facts10 gaps7 of 22
The compression, which is the reason to work backwardsCount
Awards with a live reporting obligation6
Indicators those awards require22
Testable clauses inside those definitions58
Distinct facts the clauses name18, or 3.2 clauses per fact

A per-indicator reference sheet, which is the most rigorous thing in general use, specifies one of these rows at a time. It never asks what any two of them have in common, so it cannot tell you that a person key sits under 17 of the 22 and a household key under four. The union is where the collection stops growing with the reporting obligation. Nobody has to build 22 collection processes. They have to hold 18 facts, and eight of the 18 are already clean. Note the two rows marked gone: those are not late, and no work before the deadline recovers them.

Unified Data Structure

The union of the facts, one row each, carrying the grain, the occasion it is captured on, how many funder indicators it blocks, and whether this cohort's value can still be obtained. Ten of the 18 are gaps. They block 15 of the 22 indicators.

FactGrainOccasionBlocksRecoverable nowTierHours
D-02 Household keyHouseholdIntake, and any address change4From the address already held114
D-07 Race and ethnicityPersonIntake, self-reported2By asking, with recall error29
D-09 Income and household sizeHouseholdIntake, and any change2By asking23
D-12 Knowledge check scorePerson per occasionEntry and exit2On paper, needs keying in112
D-18 Post-exit contact routePersonIntake, as consent2Gone for this cohort334
D-11 Language of deliveryContactEvery session1On the educator's roster14
D-14 Clinical measurePerson per occasionIntake, exit, 6 mo, 24 mo1Only where this clinic ordered it320
D-15 Monitor issuedPersonThe day it is handed over1Paper supply log, needs joining16
D-16 Monitoring frequencyPerson per occasionSession six1By asking24
D-17 Reason for stoppingEnrolmentThe moment somebody stops1Gone for this cohort25
Eight more factsPerson key, contact row, billable flag, enrolment, attendance, date of birth, county, named cliniciancleanAlready collected at the right grainn/a0
Six funder figures, one table, four columnsFigureCost per unitFacts it filters on
Internal dashboard: anyone with at least one contact1,000$412D-01 D-03
A-01: unduplicated individuals with a billable encounter812$507D-01 D-03 D-04
A-02: households receiving service612$673D-01 D-02 D-03
A-03: individuals with three or more contacts447$922D-01 D-03
A-04: new clients, first contact in the period289$1,426D-01 D-03
A-05: service encounters delivered3,146$131D-03
A-06: participants completing all six sessions214$1,925D-01 D-05 D-06
Spread on the six person-based figures4.7x4.7xSame $412,000 of spend under every one

The grain rule is what makes that second table possible. A-05 counts encounters and A-01 counts people, so the row is a contact with a person key on it, because a table of person totals cannot be decomposed back into 3,146 encounters. A-02 counts households and A-01 counts individuals, so the row is a person with a household key on it. Aggregation runs one way. Every one of those six figures is a filter on four columns, and the only one that does not derive today is A-02's, because the household key does not exist yet. 188 of the 1,000 had a contact and no billable encounter, which is exactly the distance between the internal figure and A-01's, and neither number is an error.

Data Quality Check

Every row is a query that returns rows, run on the occasion rather than in the reporting week. A check that cannot return a row is a reminder, and reminders do not fail. The result column is a count, not a verdict.

CheckGrainWhat a failure costsRun whenResult
Person key present on every contactContactEncounters cannot roll up to people, so five of six figures dieNightly0 of 3,146
No person key shared by two peoplePersonThe unduplicated count is wrong in the flattering directionWeekly4 to review
Household key on every personPersonA-02 has no producible indicator at allWeekly1,000 of 1,000
Household size present wherever income isHouseholdNo poverty ratio, so two indicators stay blankOn intake612 of 612
Entry score present before an exit scorePersonThe change cannot be computed and the exit score reads as a levelOn exit63 of 214
Same instrument at both endsPersonThe change is an instrument artefact rather than a resultOn exit0 of 151
Exit reason on every early exitEnrolmentA-03 loses an indicator, permanently, for this cohortOn exit117 of 117
Post-exit contact consent recordedPersonThe 6 and 24-month occasions expire while the programme runsOn enrolment214 of 214
Closing the ten gaps, ranked by what has to be true firstFactsHoursCost of the timeProducible after
Where it stands todayn/a0$07 of 22
Tier 1: nobody is asked a new question436$1,51214 of 22
Tier 2: one form change, no new instrument421$88220 of 22
Tier 3: needs a party outside the organisation254$2,26822 of 22
What one fragmented fact costs instead136.1 a year$1,516 a yearand it never improves

Tier 1 is the whole finding. 36 hours of keying and joining, $1,512 at a blended $42 an hour, and it takes this year's cohort from seven producible indicators to 14 without asking anybody who has already left the programme for anything. It is the only tier that works backwards. Tier 2 reaches 20 of 22 and can do nothing at all for the cohort in the building, which is why the two are separated. The bottom row is the comparison that gets tier 1 approved: race and ethnicity sits on three forms with three category sets, so 812 records are recoded by hand for every quarterly report, 36.1 hours a year, on a fact that a nine-hour form change fixes outright.

What is in the pack

01

Funder Indicator Crosswalk, run backwards

From each agreement's own wording to the facts that have to exist before the figure is possible, with the clause count on every definition.

02

Unified Data Structure

The union of those facts, one row each, so a fact three funders need is one field rather than three that will eventually disagree.

03

The grain rule, applied per field

Each field set to the finest grain any single funder asked for, because aggregation runs one way and the coarse question destroys the fine answer.

04

Gaps ranked by what they unblock

Not by how hard they look. The largest blocker in the worked example is also the cheapest thing on the list, and nobody guesses that.

05

A recoverability column no other template has

Already held and needs joining, still obtainable by asking, or gone because the occasion passed. Only the third one changes what you do today.

06

Data Quality Check as queries

Each check returns rows and runs on its occasion, so intake problems surface at intake rather than in the week the report is due.

How it works

  1. 1

    Send the agreements, not a summary of them

    Executed awards with their reporting guidance and any portal template. A paraphrase drifts toward the fact you already collect, which is exactly the fact the gap analysis needs to ignore.

  2. 2

    Decompose every indicator into clauses

    Each definition split into the separate things that have to be true, then the facts those clauses name. Do all of them before drawing a field, because the saving is in the union.

  3. 3

    Set the grain and take the union

    One field per fact, at the finest grain anybody asked for, with the occasion it is captured on. Funder wordings that map onto an existing field become mappings rather than new fields.

  4. 4

    Rank the gaps, then tier the fixes

    By indicators unblocked, then split by what has to be true before the work can start. The figures you have already filed tell you which gaps were being papered over by hand.

Frequently asked questions

How is this different from a logic model?

A logic model decides which indicators should exist at all, working forwards from the programme. This works backwards from agreements already signed and decides what has to be captured, on which occasion, for those indicators to be answerable. Different direction, different artifact, and you want both.

Which gap do we close first?

The one blocking the most indicators, which is rarely the one that feels urgent. In the worked example a household key blocks four indicators, all four on one state contract, and it derives from an address already on every intake form. The gap that feels most pressing blocks two and needs an outside party.

What if a follow-up window has already closed?

Say so, in writing, before the funder finds it. Name the indicator, the reason it is unavailable for this period, and the change that makes the next cohort's version available. An estimate presented as a measured value is worse than the blank it replaces, and it turns a data gap into a finding.

Should we collect detailed race and ethnicity if no funder asks?

Yes, because the aggregation only runs one way. The revised federal standard requires collection beyond the minimum categories as a default, with the detail aggregating up into them, on one combined self-reported question. A form asking five categories can never produce the detail.

Can a spreadsheet do this, or do we need a database?

A spreadsheet is fine, and the pack ships three. What matters is that one fact has one field at one grain with its occasion named. The gift and grant records you already reconcile prove the point: the tool was never the problem, the second field holding the same fact was.

How long do we have to keep the underlying records?

Federal award records must be retained three years from submission of the final financial report, longer while litigation, a claim or an audit finding is open, and that covers the records rather than only the reports. So corrections get recorded as new dated rows, never as edits in place.

Does collecting more detail create a privacy problem?

It creates a duty you already have. A recipient must take reasonable measures to safeguard protected personally identifiable information and anything the agency designates sensitive. That covers the paper forms in the drawer and every per-funder export, which the audit preparation will ask about.

Find out how few facts you actually need

Send the executed agreements and your current intake forms. The first thing back is the clause count, the union of the facts, and which gaps are already gone.

Design my collection