River
Y CombinatorBacked by Y Combinator
FREE TEMPLATE

Maturity Model Assessment Template

Every level is an observable criterion, the next level is a list of things that must exist, and the re-score names why each level moved.

Free download  ·  No account needed

A maturity model is a measuring instrument, and almost every published one is a description instead. It tells a client what level 3 looks like and never what would prove they are at it, so the assessment turns into a negotiation about adjectives. Here every level is one or more observable criteria with the evidence that satisfies each: a named document, a dated record, a standing forum with minutes. Two assessors reading the same evidence reach the same verdict, and where they would not, the criterion gets rewritten.

The instrument gets used twice, which is the part nothing published treats as a measurement problem. A level differs between the baseline and the re-score for four reasons and only one is improvement: the practice changed, the evidence improved and the baseline was wrong, the criterion was revised, or the population being measured changed. Movement Register types every difference and restates the baseline for the last three, which is the discipline an auditor applies across periods.

What the next level requires is written as artifacts, dated records and standing forums, each with an owner and an effort, so it can be resourced rather than aspired to. Priority Sequence groups those things by the artifact rather than by the dimension, because one master data set routinely unblocks three dimensions and a radar chart cannot show an overlap. Scoring a client once against a rubric you reuse is the diagnostic assessment pack; valuing what the assessment found is operational diagnostic.

Six dimensions, two measurements, and the three differences that were not movement

Both comparisons ship, always. The restated one is stronger than the naive one here, which is the usual result once the phantom regressions come out.

Level Definitions

Instrument v1.0 for a fictional contract packaging business, Ferndale Packaging.

DimensionLvlObservable criterionEvidence that satisfies itVer
Demand planning2Each site produces a forecast on a stated cadence and a named person owns itForecast files with dates, plus a named owner per site1.0
Demand planning3One forecasting procedure at all sites, every site override recorded against itThe procedure as one document, plus a dated override log1.0
Supplier management2Supplier performance recorded against agreed measures and reviewed on a stated cadenceThe performance record, plus minutes with attendance and carried actions1.0
Quality management3A documented inspection procedure names a sampling basis and is used at all sitesThe procedure showing the sampling basis, plus completed records per site1.1
Data and systems2One item master with a named owner and a documented rule for adding an itemThe single master, the named owner, the written add rule1.0

The Quality management row is at v1.1 on purpose. Version 1.0 said documented inspection procedure and two assessors read it differently, so the text was tightened in month 2 to require a stated sampling basis. That obliges a re-score of the baseline on the new text before anything is compared.

Levels are cumulative, so the number is never the finding. A dimension sits at 3 only if it meets every level 1, 2 and 3 criterion, which means every score ships with the lowest unmet criterion beside it. That criterion is the reason for the number and the shortest way out of it.

A criterion whose evidence is that the team says so is not finished. Rewrite it until it names something that either exists or does not.

Movement Register

Baseline at month 0, re-score at month 7, both on instrument v1.1.

DimensionBaseline as issuedBaseline restatedRe-scoreNaiveRealCause
Demand planning223+1+1Practice change
Supplier management112+1+1Practice change
Production scheduling22200Three of four criteria now met
Quality management322-10Criterion revision
Data and systems122+10Evidence correction
Continuous improvement211-10Population change
Total111012+1+23 of 6 were not movement

The naive comparison is worse and less defensible. It reads three dimensions up, two down, net one level. Against the restated baseline it reads two up, none down, net two, because a tightened criterion, a record nobody supplied at baseline and a divested site are not capability changes.

Data and systems gained nothing. The add rule was always there, inside a system administration guide dated fourteen months before the baseline. The baseline was wrong, so the baseline moves rather than the client.

Production scheduling is the most improved dimension in the model and its level did not change: one of four criteria met at baseline, three now, four weeks of work outstanding.

Priority Sequence

Ordered on shared artifacts, with the radar chart’s answer shown beside it.

#The thing that has to existDimensions it unblocksLevelsWksRadar would rank it
1One item master, deduplicated, with an owner and a written add ruleData and systems, Demand planning, Production scheduling395 of 8
2Supplier performance record plus a standing monthly reviewSupplier management141 of 8
3One scheduling procedure at all three sitesProduction scheduling166 of 8
4One forecasting procedure with an override logDemand planning157 of 8
5Scheduling deviation recordProduction scheduling048 of 8
6Improvement selection criterion and decision recordContinuous improvement132 of 8
7Capability measurement plan and collection mechanismQuality management0264 of 8

The radar says start with supplier management, the shortest spoke. The sequence starts with the item master, the fifth spoke, because the same artifact sits behind unmet criteria in three dimensions and carries three levels. A chart of spokes cannot show an overlap, so it always recommends the wrong thing first.

Rank 6 is held rather than cheap. Three weeks of work, one level, and the site carrying that dimension’s evidence is in a divestment process, so the artifact would leave with it.

What's in the pack

01

Level Definitions

The instrument: one row per dimension per level, each level an observable criterion with the evidence that satisfies it, and the whole sheet carrying a version and a date. Versioning is what makes the second measurement mean anything.

02

Score by Dimension

The baseline: the level, the single criterion capping it, the evidence behind the verdict, the target agreed with the client, and criteria met and unmet counts for anything at level 1. A capping criterion worth writing up formally becomes a finding in the findings and recommendations pack.

03

Movement Register

Written only at the re-score. One row per changed dimension carrying baseline as issued, baseline restated, re-score, and which of the four causes applies, with the evidence for the cause as well as for the level.

04

Gap Analysis

Every unmet criterion between here and the target, resolved into the artifact, record or standing forum that satisfies it, with an owner, an effort and a note where the same thing appears against another dimension. Where the gap is exposure rather than capability, the risk and controls assessment sizes it against the policy that would respond.

05

Priority Sequence

The order, grouped by the thing rather than by the dimension, ranked on how many dimensions and levels each one carries. It states where it disagrees with the radar chart, because the client is looking at the radar.

06

What This Model Measures

The document that travels with every score: what a level is, why there is no overall number, and why the top level is not the target. The clearest published statement of that last point is NIST's, that its Tiers do not represent maturity levels.

07

The Re-score Protocol

Fix the instrument version, the re-score date and the population in scope before the work starts. The population is the one that changes without anybody telling you, and it is why a divested site reads as a decline.

08

What Moved and Why

The sponsor document. It leads with the restated comparison, shows the as-issued one beside it, and explains every difference. A phantom regression explained beforehand costs nothing; the same one found in the meeting costs the engagement.

How to use it

  1. 1

    Open in River, or take it blank

    Install the pack and hand River your dimensions and any rubric you already use, or download the Word documents and CSV sheets and build it yourself.

  2. 2

    Build the instrument first

    Every level becomes observable criteria with the evidence that satisfies each, then the sheet gets a version and a date. Nothing is scored until that text exists.

  3. 3

    Score the baseline

    Criterion by criterion from level 2 upward, stopping at the first unmet one. Anything landing at level 1 also records how many of level 2's criteria are met.

  4. 4

    Sequence, then re-score

    The gap becomes things with owners, ranked on shared artifacts. At the agreed date the frozen instrument runs again and every difference gets a cause.

Frequently asked questions

Is this template free?

Yes. The download is Word documents and CSV sheets, no account and no card. Edit with AI is the optional half: River turns your level table into observable criteria, scores the baseline against them, and runs the re-score later. The other packs are in the template library.

What format are the downloaded files?

Word (.docx) for the five documents and CSV (.csv) for the five sheets, zipped together. Excel, Numbers and Google Sheets open the sheets straight off the download. Nothing to convert, and the instrument stays editable in whatever you already use.

How is this different from a five-level maturity matrix?

A matrix describes each level. This measures against it, twice. Every level carries the evidence that would prove it, the gap to the next one is a list of things with owners, and the re-score types every difference so a tightened criterion never reads as a client regression.

Why does a dimension at level 1 need extra columns?

Because level 1 carries no distance information. The published models define it as the designation for an organisation that achieved none of the other levels, so one unmet criterion and fifteen print the same number. The sheet records criteria met and unmet instead.

Can I change a criterion mid-engagement?

Yes, and it is often necessary. Changing one creates a new instrument version, and the space rule then obliges a re-score of the baseline on the new text before any comparison goes out. In the worked example that turned an apparent one-level regression into no change at all.

Does it produce one overall maturity score?

No, deliberately. Dimensions score independently and an average hides the single dimension capping the others, which is the one thing the profile exists to surface. The deliverable is a level per dimension with the criterion blocking each one.

What if the client reorganises during the engagement?

That is a population change and the protocol handles it. The baseline is restated on the current population rather than the score falling, so a divested site reads as a change in what is measured. Naming the sites in scope up front is what makes that restatement uncontroversial, unless the reorganisation itself is the finding, which is an organizational design assessment instead.

Build a model that survives its second use

Take the Word documents and CSV sheets blank, or open this exact pack in River and let it turn your level table into criteria you can check.

Build my model