River
Y CombinatorBacked by Y Combinator
FREE TEMPLATE

Consulting Assessment and Scoring Rubric

Four documents and four sheets built on a rubric that keeps its scale between clients, so your third assessment is worth more than your first.

Free download  ·  No account needed

Scoring Rubric  ·  D1 Demand planning  ·  rubric v3

One dimension, four anchors, no adjectives

LevelAnchorWeakest evidence that satisfies it
1A forecast exists for the next period and one named person produced itArtifact
2It is produced on a stated cadence from a stated input set, and the last three existArtifact
3Forecast error is calculated each period against actuals and reviewed by a named forumRecord
4A method change in the last four periods traces to a recorded error analysisRecord

Every anchor passes three tests. Somebody who was not in the room could go and look. The verdict is met or not met. And two different assessors reading the same evidence reach the same answer. That third test is the one the adjectives fail.

Three forecasts, not one. One well made forecast is a person. Three consecutive ones with the same structure is a process, and the difference is the whole content of level 2.

Level 3 will not take a statement. A planning manager describing a monthly accuracy review is evidence that a review is believed to happen. The error calculation and the forum's minutes are evidence that it does.

Level 0 is simply that the level 1 anchor is not met. It is not the same cell as not assessed, and the report says which.

Search this and you get five columns of adjectives. Basic, developing, defined, managed, optimised. Two competent assessors put the same client two levels apart and neither one is wrong, because there was never anything there to be wrong about. The template that ranks is either a consultancy grading its own practice or a slide deck of level descriptions with no rule about what evidence a level requires. Neither answers the two questions that decide whether an assessment holds up.

An anchor replaces the adjective with a sentence somebody else could check. Level 2 on demand planning is not developing, it is that the forecast comes out on a stated cadence from a stated input set and the last three exist. Three, not one, because one good forecast is a person and three is a process. Every anchor is observable, binary, and gives two different assessors the same verdict. That third property is the one the adjectives fail.

Levels are cumulative and dimensions score independently, which is how published models actually work. C2M2 requires every practice at a level and below it, and CMMC Level 2 requires a met result on all 110 requirements with no partial credit. So the deliverable is a profile with a named blocking anchor per dimension, never one number. Score it from what the client's data room and an interview programme produced, or take the Word documents and CSV sheets blank.

The dimension that meets level 3 and scores 1

Every level in Level by Dimension traces to a row in the Evidence Log. Nothing in the profile was typed in from an impression.

Scoring Rubric

Extract for a fictional scientific instruments manufacturer, Coldharbour. Rubric v3.

DimLvlAnchorMin evidenceWt
D21Every order is in one system of record and can be found by referenceObservation1.0
D22Exception handling is documented and the document matches what people doObservation1.0
D23Order cycle time is measured end to end and reported on a cadenceRecord1.0
D32Schedule changes are logged with a reason and a decision makerArtifact1.25
D51Somebody can name the system of record for each core data objectStatement0.75
D62Each measure has a written definition and it is used consistentlyStatement1.25
D63Targets are set from the organisation's own history rather than from ambitionRecord1.25

The word and in D2 level 2 is doing the work. A well written procedure that nobody follows evidences level 1. The anchor requires the document and a walkthrough that agrees with it, which is twenty minutes at somebody's desk and the cheapest real check in the pack.

D5 level 1 is the one anchor a statement satisfies. The anchor is about whether anybody in the organisation knows, so asking three people independently is the correct instrument rather than a concession.

Weights inform the conversation and are never summed. There is no total in this pack, no weighted average and no percentage of anchors met.

Almost nothing scores at D6 level 3. When something does, it is usually the strongest single finding in the assessment.

Evidence Log

Insufficient evidence is a filled row. The reason is frequently the finding.

RefDimLvlKindWhat it isEnough?
E-004D21ObservationWatched three order refs located live in the ERP from a paper noteYes
E-005D22ArtifactException handling procedure dated 2023No
E-006D23RecordERP cycle time report, six periods, definition attachedYes
E-003D13StatementPlanning manager describes a monthly accuracy reviewNo
E-008D32ArtifactSchedule change log, reason column presentNo
E-010D51StatementThree people independently named the same system of recordYes
E-011D62StatementTwo consumers described on-time delivery differentlyNo

E-005 is a document that contradicts practice. The procedure exists, is well written, and diverges from the observed walkthrough at two steps. The anchor is not met and the divergence is worth more than the score.

E-008 is a column that exists and is empty. A reason field on all 52 rows and a reason written in 11 of them. A self-assessment scores this level 2 because the log exists; the anchor asks what is in it.

E-011 is a nine point gap on the same underlying data. One consumer counts from order date, the other from confirmed date, and everything built on top of that measure inherits the ambiguity.

E-003 comes down because a level 3 anchor never rests on a statement. The review may well happen; nothing seen so far shows that it does.

Level by Dimension

The blocking anchor is the deliverable: the reason for the score and the way out of it.

DimLevelAbove the blockBlocking anchorEffortTarget
D1 Demand planning1L3 on a statement onlyL2: the last three forecasts exist and only two doLow3
D2 Order management1L3 fully metL2: the procedure and the observed practice diverge at two stepsLow3
D3 Production scheduling1NoneL2: 41 of 52 change rows carry no reasonLow2
D4 Supplier management2NoneAt targetMedium2
D6 Performance mgmt1NoneL2: two consumers define on-time delivery differentlyMedium3
D7 Quality managementNot assessedNoneOut of scope this phase3

D2 is the row worth the whole engagement. Level 3 is met in full and level 2 is not, so the dimension scores 1. Cycle time is being measured end to end over a process that varies by operator, which means somebody built the measurement before the process was stable. A blended score reports this as a 2 and loses it entirely.

Three of these clear for no budget. Populate the reason field. Reinstate the monthly forecast. Publish one definition. Handing those back to the client for free is what makes the priced half credible.

D4 is at target and is reported as at target. Six dimensions below the top level is not six problems, and a report that says it is has manufactured two arguments it did not need.

D7 is not assessed, not zero. They sit in the same place on a chart and say opposite things.

Benchmark Comparison

Only clients scored on the same rubric version belong in this sheet.

DimThis clientC-04C-07C-09MedianVersionComparable
D112312v3Yes
D212222v3Yes
D421232v3Yes
D611111v3Yes
D7Not assessed22Not assessed2v3No
D8Not assessed3222v2No

Four clients in a row failing the same anchor is a statement about the anchor. Nobody in this book has cleared D6 level 2. That is worth rereading before the fifth assessment, and it is only visible because all four were scored on one version.

D8 is excluded on version, not on data. C-04 and C-07 were scored on v2, where the level 3 anchor did not require a documented cover arrangement. Re-score their evidence against v3 or leave the row out. Comparing across versions reprices the whole back catalogue and nobody downstream can see it happened.

A blank is not a zero. D7 was out of scope for this client and for C-09, so the median is computed over two clients and the sheet says so.

With fewer than three prior clients on one version, leave this sheet out of the report. A sample of one is an anecdote with a chart on it.

What's in the pack

01

Scoring Rubric

One row per dimension per level, each carrying an observable anchor and the weakest evidence kind that satisfies it.

02

Evidence Log

Four evidence kinds in strength order, a source reference that resolves, and insufficient recorded as a filled row rather than a blank.

03

Level by Dimension

The highest level fully met, what is met above the block, and the single blocking anchor with what would clear it.

04

Benchmark Comparison

Your own prior clients on the same rubric version, with cross-version rows excluded rather than quietly averaged in.

05

Rubric Anchors and Scoring Guide

The three tests an anchor has to pass, why levels are cumulative, and the rule that a level 3 never rests on a statement.

06

Assessment Questionnaire

Requests that name a document, an export at a stated grain or a walkthrough, including where the written procedure and the real one diverge.

07

Score Report

Six sections, scope before any level appears, and the blocking anchors led by the cheap ones. No total anywhere in it.

08

Scope Note for Follow-on Work

Blocking anchors grouped by what clears them, the cheap half handed back free, and a re-score offer against the same rubric version. Running that re-score so its movement holds up is the maturity model pack.

How to use it

  1. 1

    Open in River, or take it blank

    Open the pack in River and send whatever you already assess with, or download the Word documents and CSV sheets and build the rubric yourself.

  2. 2

    Turn adjectives into anchors

    Six to nine dimensions, four levels each. Every anchor names a countable artifact, a stated cadence, a named forum or a date inside a window.

  3. 3

    Score from evidence, not impressions

    Work up the ladder and keep testing past the block. Every level gets a log row saying what it rests on, including the rows where the evidence did not reach.

  4. 4

    Report the profile and the block

    A level and a target per dimension, one blocking anchor each, and the cheap ones first. Then group them into a follow-on that is two workstreams rather than six.

Frequently asked questions

Is this template free?

Yes. The zip is Word documents and CSV sheets, no account and no card. Edit with AI is the optional half: the agent turns whatever you assess with into anchors, then scores a real client from the evidence you send. The rest sit in the template library.

What format are the downloaded files?

Word (.docx) for the four documents and CSV (.csv) for the four sheets, zipped together. Excel, Numbers and Google Sheets open the rubric, the evidence log, the profile and the comparison straight off the download.

Why is there no overall maturity score?

Because dimensions score independently, and a single blended figure cannot say where to start or be compared to anything without a weighting nobody will agree to. You get a level per dimension plus the one anchor capping each, which fits on the same page and is actionable. For money rather than a level, a diagnostic that prices each finding is the companion.

Can I reuse the rubric on the next client?

That is the point of it. Every score records the rubric version it was made against, so client three is graded on the same scale as client one and the median of your own book becomes a benchmark no template can sell you.

What happens when I tighten an anchor?

The version bumps, and you get two honest options: re-score the affected prior clients from their own evidence logs, or leave them out of the comparison and say why. Comparing a v2 score against v3 reprices your back catalogue invisibly.

The client already did a self-assessment. Is that usable?

As an artifact about what they believe, yes, and the gap between their rating and the anchored score is often the most interesting page in the report. It is not evidence for a level. Levels come from documents, system records and walkthroughs. The same gap appears when a client self-rates change readiness; the change readiness assessment scores that from its last few initiatives instead.

Two systems report the same measure differently. What then?

The definition anchor is not met and the divergence is the finding. Where you need the difference explained row by row rather than just flagged, reconciling two exports that disagree is the run for it.

Score a client on anchors

Take the Word documents and CSV sheets blank, or send River the spreadsheet you assess with today and get the anchors back to argue with.

Edit with AI