River
Y CombinatorBacked by Y Combinator
FREE TEMPLATE

Clinical Quality Improvement Project Pack

Three documents and three sheets that freeze the baseline median before any change is tested, then check every new point against that same frozen line.

Free download  ·  No account needed

Run Chart Data, screening view

Six straight points below the frozen median

Illustrative weeks for a fictional four-provider primary care practice, Wrenfield Family Health, testing a two-touch reminder workflow against its weekly no-show rate.

WeekNo-Shows / 240Ratevs. Frozen Median (13.75%)Run
162711.25%Below1
172510.42%Below2
182410.00%Below3
19229.17%Below4
20239.58%Below5
21218.75%Below6 (shift confirmed)

Six consecutive weeks below a median a 12-week baseline froze before this reminder workflow started is a shift, per the run chart's own rule. Weeks 13 through 15 do not appear here: the second of those three came back above the median, so cycle one alone had not yet cleared the threshold.

Full 24-week chart, the other three rules, and the arithmetic: below.

Most quality improvement templates hand a team a grid to plot data on and a PDSA worksheet, then leave the actual proof to a glance at whether the line looks better afterward. Per a peer-reviewed primer on interpreting QI data, it is "crucial to have baseline data before initiating an improvement project," since without it there is no way to confirm a problem existed, let alone that anything afterward improved it. An eyeballed before-and-after average skips that requirement and calls the difference an improvement regardless.

This pack freezes that baseline first: at least 12 consecutive points, collected before any change, with their median locked as the line every later point is checked against. Per Robert Lloyd's chapter on run charts, a shift of 6 or more consecutive points on one side of that line, or a trend of 5 or more moving the same direction, is what counts as a real signal. Wrenfield Family Health, a fictional four-provider practice, froze its weekly no-show-rate baseline at 33 no-shows, 13.75 percent of 240 scheduled appointments, before testing anything.

Cycle one (a call alone) did not hold: two weeks fell below 33 no-shows, then a third came back above it, so no rule fired. Cycle two added a text reminder, and weeks 16 through 21 ran six straight below the frozen median, confirming a shift worth 494 recovered visit-slots a year. A plain before-and-after average across the same 24 weeks would show a similar-looking drop, from 13.68 to 10.28 percent, with no way to tell whether that gap is the reminder working or the ordinary variation the practice already had.

Every document in the pack

The full run chart against the frozen median, which of the four rules actually fired, and why the median-basis result and the naive average agree on the drop but not on whether it is real.

Run Chart Data, all 24 weeks

Twelve baseline weeks, then twelve tested against their median

Baseline, weeks 1-12 (median frozen here: 13.75%)Post-intervention, below the frozen medianPost-intervention, above the frozen medianWeek 21: shift confirmed, Rule 1
13.75%
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24

Weeks 1-12 set the frozen median at 33 no-shows, 13.75 percent of 240 scheduled appointments. Week 15, the only post-intervention bar above the line, is why cycle one (the call alone) did not yet confirm anything. Weeks 16 through 21, six straight below the line, is what confirms the shift, the sixth week after a text reminder was added to the existing call.

The four rules, applied to this project

Only these four decide whether a run of points is real

RuleThresholdIn this project
Shift6+ consecutive points on one side of the frozen medianFired at week 21 (weeks 16-21)
Trend5+ consecutive points each higher or lower than the lastNot the rule that fired here
RunsToo many or too few crossings of the median for the point count3 runs across the 12 post-intervention weeks
Astronomical pointA single point dramatically different from the restNot present in the tested weeks; illustrated instead by a pre-baseline storm week below

A storm three weeks before baseline week 1 closed local roads for two days and pushed that week's no-shows to 61 of 240, 25.42 percent, 28 more than the frozen median. It was excluded from the baseline entirely rather than averaged into it, which is what an astronomical point means: dramatic enough to investigate on its own, not dramatic enough to move the line everything else gets checked against.

Rule definitions and the 6-point and 5-point thresholds: Robert Lloyd's chapter on run charts, cited above.

Median-basis result vs. the naive average

Two ways to read the same 24 weeks

 Frozen median methodPlain before/after average
Baseline (weeks 1-12)33 no-shows, 13.75%32.83 no-shows, 13.68%
Post (weeks 13-24)23.5 no-shows, 9.79%24.67 no-shows, 10.28%
Change9.5 fewer/week, 28.8% relative drop8.17 fewer/week, 24.9% relative drop
Confirmed signal?Yes: Rule 1 fired at week 21No way to tell from an average alone

Both columns show a real drop; that was never the disagreement. The average cannot say whether 8 to 9 fewer no-shows a week is the reminder workflow working or the kind of week-to-week swing Wrenfield's schedule already had before anyone changed anything. Six consecutive weeks below a line fixed in advance is what answers that, at week 21 specifically, not an average of two arbitrary windows.

Annualized at 52 weeks and Wrenfield's blended $135 average reimbursement per completed visit, the frozen-median basis is roughly 494 recovered visit-slots, about $66,690 a year.

What's in the pack

01

Project Charter

The aim statement, the measure's operational definition precise enough that two different counters agree on the same number, and the baseline collection plan, filled in before any intervention gets named.

02

Baseline and Follow-up Measurement

Every measurement period's raw count and rate in one place, split cleanly into the frozen baseline window and everything measured afterward, so the median governing every later rule check is always traceable back to the exact points it came from. The same respect for a measure's exact definition governs which six measures a group reports under a MIPS reporting checklist.

03

PDSA Cycle Log

One row per cycle: what was tested, at what scale, the prediction made before results came in, and what actually happened. Each row closes with the adapt, adopt or abandon decision, so a later reader can see exactly which test was running when a rule fired.

04

Run Chart Data

Every point checked against the frozen median with the running consecutive-point count and whichever of the four rules fired, if any, so the chart records a specific, falsifiable signal instead of a line a reader has to interpret alone.

05

Intervention Description

What changed, on what scale, who owned it, and the prediction made in advance, recorded per cycle rather than once for the whole project. A second cycle that adds a new element is a different test from the one before it, not an edit to it.

06

Results Report

Names the rule that fired, the point it fired on, and the PDSA cycle running at the time. States the size of the change in the aim statement's own units, and reports plainly when no rule fired within the planned time frame rather than calling a trend a result. A confirmed change's own median becomes the next frozen baseline, the same version discipline a clinical registry's specification applies to its own required fields.

How to use it

  1. 1

    Open the pack, or download it

    Open the pack in River and describe the process you want to improve, or download the blank Word and CSV files and run the project on your own.

  2. 2

    Send your baseline, or start collecting it

    If fewer than 12 consecutive pre-intervention points exist, the agent sets the measurement frequency and says plainly how many periods away a usable baseline is. If they exist, it computes and freezes the median immediately.

  3. 3

    Log each PDSA cycle as you test it

    After every new measurement period, the agent checks the point against the frozen median and reports the running consecutive-point count and whether a rule just fired, rather than leaving that read to a glance at the chart.

  4. 4

    Report the result once a rule fires or the timeline ends

    Either outcome writes the Results Report. A confirmed shift's own median becomes the frozen baseline for whatever the team targets next.

Frequently asked questions

Is this template free?

Yes. Download the three documents and three sheets as Word and CSV files with no signup and no card. "Edit with AI" is a separate, optional path for teams that want the agent to run the project with them. The rest of the library is at the template library.

What format are the downloaded files?

Word documents (.docx) for the Project Charter, the Intervention Description and the Results Report, and CSV (.csv) for the three sheets, zipped into one download. They open natively in Word, Pages, Google Docs, Excel, Numbers and Sheets.

Why freeze the baseline median instead of recalculating it as new data comes in?

Because a moving median chases the data it is supposed to judge. If it recalculates every time a point is added, a genuinely working intervention pulls the average down with it, and the line a later point gets compared against keeps sliding toward whatever that point is about to show. Freezing it first is what lets six consecutive points below the line mean something specific.

Does a confirmed shift mean a payer contract improves too?

No, and it answers a different question. This pack confirms whether the underlying rate itself changed in a statistically defensible way. Whether that change moves a specific payer's reconciliation payment depends on that contract's own measure definitions and thresholds, which is what a value-based contract performance review reads once a statement actually arrives.

What happens if no rule ever fires?

The Results Report says so directly instead of describing an encouraging-looking trend as a win. The team can extend data collection, test a different change, or close the project without a confirmed result. All three are legitimate outcomes; the only outcome this pack will not produce is a report that calls a slope a signal.

Does this pack help pick which measure to improve?

No, it assumes the measure is already chosen and turns it into a baseline, a test, and a rule-based read on the result. Picking which measure is worth a project's time, such as a gap a care gap and population health report already flags, happens before this pack's first document gets filled in.

Freeze the baseline before you test the change

Download the blank pack as Word and CSV files, or open this exact pack in River and let the agent freeze your baseline median and check every new point against it.

Edit with AI