Maturity Model Assessment Template
Every level is an observable criterion, the next level is a list of things that must exist, and the re-score names why each level moved.
Free download · No account needed
A maturity model is a measuring instrument, and almost every published one is a description instead. It tells a client what level 3 looks like and never what would prove they are at it, so the assessment turns into a negotiation about adjectives. Here every level is one or more observable criteria with the evidence that satisfies each: a named document, a dated record, a standing forum with minutes. Two assessors reading the same evidence reach the same verdict, and where they would not, the criterion gets rewritten.
The instrument gets used twice, which is the part nothing published treats as a measurement problem. A level differs between the baseline and the re-score for four reasons and only one is improvement: the practice changed, the evidence improved and the baseline was wrong, the criterion was revised, or the population being measured changed. Movement Register types every difference and restates the baseline for the last three, which is the discipline an auditor applies across periods.
What the next level requires is written as artifacts, dated records and standing forums, each with an owner and an effort, so it can be resourced rather than aspired to. Priority Sequence groups those things by the artifact rather than by the dimension, because one master data set routinely unblocks three dimensions and a radar chart cannot show an overlap. Scoring a client once against a rubric you reuse is the diagnostic assessment pack; valuing what the assessment found is operational diagnostic.
What's in the pack
Level Definitions
The instrument: one row per dimension per level, each level an observable criterion with the evidence that satisfies it, and the whole sheet carrying a version and a date. Versioning is what makes the second measurement mean anything.
Score by Dimension
The baseline: the level, the single criterion capping it, the evidence behind the verdict, the target agreed with the client, and criteria met and unmet counts for anything at level 1. A capping criterion worth writing up formally becomes a finding in the findings and recommendations pack.
Movement Register
Written only at the re-score. One row per changed dimension carrying baseline as issued, baseline restated, re-score, and which of the four causes applies, with the evidence for the cause as well as for the level.
Gap Analysis
Every unmet criterion between here and the target, resolved into the artifact, record or standing forum that satisfies it, with an owner, an effort and a note where the same thing appears against another dimension. Where the gap is exposure rather than capability, the risk and controls assessment sizes it against the policy that would respond.
Priority Sequence
The order, grouped by the thing rather than by the dimension, ranked on how many dimensions and levels each one carries. It states where it disagrees with the radar chart, because the client is looking at the radar.
What This Model Measures
The document that travels with every score: what a level is, why there is no overall number, and why the top level is not the target. The clearest published statement of that last point is NIST's, that its Tiers do not represent maturity levels.
The Re-score Protocol
Fix the instrument version, the re-score date and the population in scope before the work starts. The population is the one that changes without anybody telling you, and it is why a divested site reads as a decline.
What Moved and Why
The sponsor document. It leads with the restated comparison, shows the as-issued one beside it, and explains every difference. A phantom regression explained beforehand costs nothing; the same one found in the meeting costs the engagement.
How to use it
- 1
Open in River, or take it blank
Install the pack and hand River your dimensions and any rubric you already use, or download the Word documents and CSV sheets and build it yourself.
- 2
Build the instrument first
Every level becomes observable criteria with the evidence that satisfies each, then the sheet gets a version and a date. Nothing is scored until that text exists.
- 3
Score the baseline
Criterion by criterion from level 2 upward, stopping at the first unmet one. Anything landing at level 1 also records how many of level 2's criteria are met.
- 4
Sequence, then re-score
The gap becomes things with owners, ranked on shared artifacts. At the agreed date the frozen instrument runs again and every difference gets a cause.
Frequently asked questions
Is this template free?
Yes. The download is Word documents and CSV sheets, no account and no card. Edit with AI is the optional half: River turns your level table into observable criteria, scores the baseline against them, and runs the re-score later. The other packs are in the template library.
What format are the downloaded files?
Word (.docx) for the five documents and CSV (.csv) for the five sheets, zipped together. Excel, Numbers and Google Sheets open the sheets straight off the download. Nothing to convert, and the instrument stays editable in whatever you already use.
How is this different from a five-level maturity matrix?
A matrix describes each level. This measures against it, twice. Every level carries the evidence that would prove it, the gap to the next one is a list of things with owners, and the re-score types every difference so a tightened criterion never reads as a client regression.
Why does a dimension at level 1 need extra columns?
Because level 1 carries no distance information. The published models define it as the designation for an organisation that achieved none of the other levels, so one unmet criterion and fifteen print the same number. The sheet records criteria met and unmet instead.
Can I change a criterion mid-engagement?
Yes, and it is often necessary. Changing one creates a new instrument version, and the space rule then obliges a re-score of the baseline on the new text before any comparison goes out. In the worked example that turned an apparent one-level regression into no change at all.
Does it produce one overall maturity score?
No, deliberately. Dimensions score independently and an average hides the single dimension capping the others, which is the one thing the profile exists to surface. The deliverable is a level per dimension with the criterion blocking each one.
What if the client reorganises during the engagement?
That is a population change and the protocol handles it. The baseline is restated on the current population rather than the score falling, so a divested site reads as a change in what is measured. Naming the sites in scope up front is what makes that restatement uncontroversial, unless the reorganisation itself is the finding, which is an organizational design assessment instead.
Build a model that survives its second use
Take the Word documents and CSV sheets blank, or open this exact pack in River and let it turn your level table into criteria you can check.
Build my model