Program Evaluation Plan Template
Three documents and three sheets that pair every finding with a specific alternative explanation and a number that bounds it.
Free download · No account needed
Most program evaluation plan templates get you to a clean question and a data collection checklist, then stop. The part that decides whether a funder's evaluator trusts the result, naming exactly what else could have produced each finding and how much, gets left as a line about limitations at the end. This pack builds the evaluation around that line instead. Every finding in the Findings Report carries a paired row in the Limitation Register: a named mechanism and, wherever the data allows it, a number bounding how much of the finding it could explain.
In the worked example, Cedar Line Community Health evaluates the support calls between sessions in its six-session diabetes course. Of 331 enrollees, 214 complete and hold a paired knowledge-check score, averaging a 4.6-point gain of 20. Split by calls received, the gain rises from 3.1 points to 4.4 to 6.0, a pattern that looks like the calls are working. There is no comparison group, though, and under the What Works Clearinghouse's own group design standards, a single group measured before and after cannot establish what would have happened otherwise.
The gap is not proof, and the same design that produced it can check it: staff assign calls informally to whoever seems to be struggling, so the highest-call group started 4.1 points behind the lowest-call group before session one. By post-test that gap had closed by 70.7%, which is what regression to the mean predicts on its own whenever a group is selected for an extreme baseline. The Limitation Register states both numbers next to the finding, which is what turns outcome data from a well-run collection into something a skeptical reader can check.
What is in the pack
Evaluation Plan, dated before the data
States the question, the design, and what it can and cannot establish, before the Outcome Analysis sheet has a number in it.
Findings Report with the limitation attached
Every finding carries its alternative explanation in the same paragraph, not a separate section a reader can skip.
Recommendation Note, kept separate from the claim
States what to do next without needing the finding to be proven causal first, and names what it would take to test it properly.
Outcome Analysis
The completion funnel and the overall pre and post knowledge-check gain, with the exclusion rate stated next to it.
Participation and Dose Analysis
Gain by dose band, plus the baseline-gap check that tells a real pattern apart from regression to the mean.
Limitation Register, quantified
Every finding paired with a specific alternative explanation and, wherever the data allows, a number that bounds it.
How it works
- 1
Send the outcome data and say what comparison exists
Participant-level records, and whether a comparison group, a matched cohort, or none was used. This is not the reporting counterpart that turns counts into a document for funders; it asks whether the programme caused the change.
- 2
Build the funnel and the dose bands
Enrolled, completed, measured, on the outcome instrument and threshold your logic model already commits to, then split by whatever varies across participants that the programme does not force to be equal.
- 3
Check the baseline gap before crediting the pattern
Compare the highest- and lowest-dose bands at baseline, size the gap against the data's own spread, and compute how much of it closed by post-test before the pattern gets called a result.
- 4
Write the register, then the narrative
Every finding gets a paired alternative explanation and, wherever the data allows it, a number. The recommendation stays separate from the causal claim, so a low-cost action can proceed either way.
Frequently asked questions
Do we need a comparison group to evaluate a programme?
Not to detect a pattern, but you do to attribute it. A single group measured before and after, the design in the worked example here, can show that knowledge-check scores rose. It cannot show the programme caused the rise, and the Limitation Register says so rather than implying otherwise.
How do we know a dose-response pattern isn't just regression to the mean?
Check whether the dose groups started equal. In the worked example the highest-dose group's baseline was 4.1 points below the lowest-dose group's, and 70.7% of that gap closed by post-test, with no programme effect required to produce it. Run that check before publishing any pattern where the dose was not randomly assigned.
What do we do about people who never finished the programme?
Count them and say what is known about them, on the same page as every finding. In the worked example 117 of 331 enrollees, 35.3%, never reached a post-test and sit outside every result below. State the exclusion rate next to the headline figure rather than letting a completers-only sample pass as the full picture.
Can we still report a finding if we cannot prove it caused the result?
Yes, if the recommendation and the causal claim stay separate. A low-cost, plausibly-helpful action can be worth continuing on a pattern alone. State the pattern, name the alternative explanation in the same sentence, and let a board decide with both in view rather than one.
Does a funder ever require this level of rigor, or just a report?
Most ask for a report of counts and cost, not a causal evaluation, and conflating the two is how attribution language ends up in a document that cannot support it. Save this space for the specific question a design change or a renewal decision actually needs answered.
Should the Limitation Register go to funders, or stay internal?
Send it, or at minimum summarise it in the narrative. A funder's evaluator who finds an unstated confound discounts the whole report. One who is told the pattern and the confound together in the same sentence has no reason to discount anything else in it.
Does this tell us whether we can afford to run the programme?
No. This evaluates whether the programme works, not what it costs. A program budget pack allocates overhead into the full cost and checks it against what your own grants actually pay, which is a separate question from whether the outcomes are real.
See what your data can actually establish
Send the outcome data and say what comparison exists. The plan, the dose-response check and the limitation register come back paired.
Build my evaluation