River
Y CombinatorBacked by Y Combinator
FREE TEMPLATE

Performance Calibration Meeting Template

Last cycle's calibration ran five hours and twenty minutes, walked the list alphabetically, and never reached the last third of the company.

Free download  ·  No account needed

This pack reads every submitted review before the session and computes the agenda from them. A row reaches the room only if it trips a named trigger: an away-from-middle rating with no observable outcome, a two-level move, a contradiction with something already on file, or a rating written against goals retired mid-cycle. That turned 186 reviews into 31 rows. The other 155 get ratified in one act by a named signer. Reviews that arrive specific trigger less, which is what review writing is for.

A meeting template cannot do this, because it has no access to your reviews. The sharper finding is that 20 of those 31 rows belonged to a pattern in one manager's ratings rather than to the people being rated. One manager had nine of eleven reports above the middle with two naming a checkable outcome, against a cohort median of 71 percent. Discussing those nine individually produces nine conversations that each end with somebody saying she is good. Discussed as one, it takes fifteen minutes.

Written for people leads who have run a calibration that ran out of time in the middle of the alphabet. The second mechanism is the sentence: no rating moves without the words the manager will say to the employee, naming the work it turned on. If nobody can write it, the submitted rating stands. The competencies being rated should be the ones the job description called must-have. Thirteen ratings changed in the worked example and the distribution moved by one row, which is the number that answers whether calibration is a quota exercise.

Why the outlier is usually the manager

Spread and evidence density are read together or not at all. Two managers with the same spread can be doing opposite things, and only the evidence separates them. The 22 percent row and the 86 percent row are the two ends of this cycle.

Distribution by Manager

Illustrative sample from a 24-manager view. Marloe Logistics, fictional warehouse software company, H2 cycle, 186 reviews.

ManagerReviewedAbove middleEvidence densityMoved 2+Pattern
Priya Raman1182%22%0P1 Loose scale
Owen Whitlock140%86%1P2 Strict scale
Dominic Searle944%64%1P3 Bad inputs
Anneke Vos757%79%0None
Rosalind Achebe862%81%1None
Theo Vandersteen757%66%1None
Marcus Chen667%72%0None
Company18652%71%7

Submitted: 17 Outstanding, 79 Exceeds, 82 Meets, 8 Below. Density range across 24 managers is 22 to 86 percent, and that range is what calibration is actually for.

Priya has the highest spread and the lowest density. Owen has the lowest spread and the highest, so his reviews are specific and his bar simply sits above the written scale. Nobody escalates a strict manager, which is why that one only ever shows up here.

Outlier Register

31 of 186 rows tripped a trigger, on 38 hits. Evidence text is quoted as written, because a paraphrase always reads more convincing than the original.

EmployeeManagerRatingTriggerEvidence as writtenOutcome
Connor DuffyPriya RamanExceedsEvidence gap“No concerns at all. Continues to exceed.”Moved to Meets
Adaeze NwosuPriya RamanExceedsEvidence gap“One of the strongest people on the team.”Held. Cut the validation pass from 5 days to 2.
Steph NowakOwen WhitlockMeetsContradicted on file“Handles a heavy queue well. Meets across the board.”Moved up. Named by 2 other teams.
Callum ReidyDominic SearleExceedsRetired goals“Delivered against the Q3 partner-API objective ahead of schedule.”Moved to Meets. Goal retired 2025-10-02.
Yusuf BarreGregor LindqvistOutstandingGap + moved 2“Outstanding half. Closed the largest deal in company history.”Moved to Exceeds. 4 of 5 competencies blank.
Suki NakamuraEmmeline ShawBelowGap + moved 2“Coverage gaps and 2 missed handovers in December.”Rating held. Plan deferred, no prior record.
Aoife BrennanRosalind AchebeOutstandingGap + moved 2“Grew the mid-market segment 34 percent.”Held. The outcome is the role.

19 evidence gaps, 7 two-level moves, 6 contradictions on file, 6 rated against retired goals. Seven rows tripped two triggers.

A rating being high or low is not a trigger. 17 Outstanding and 8 Below ratings were submitted; 11 of those 25 reached this register, on evidence rather than position. Making the ends a trigger puts 25 rows on the agenda and rebuilds the volume problem.

Adjustment Log

Every change carries the sentence the manager will say. Written on the row in session, before the room moves on.

EmployeeFromToThe sentence the manager will say
Ruben OyelaranExceedsMeets“I rated you above the middle and when I looked at what I had written next to it I had not named anything you actually did. That is on my writing rather than your work.”
Steph NowakMeetsExceeds“Two other teams named you for getting the Ridgeway go-live unblocked over that weekend and I had not counted it, because it did not happen inside my queue. It counts.”
Callum ReidyExceedsMeets“The goal I rated you against was retired in October when the roadmap changed. Against the goals in force this lands at the middle. That is a records problem, not a judgement about you.”
Nadia EllisExceedsOutstanding“I mentioned the buddying in one line and then did not count it. Both of those hires cleared their day-30 milestone and the records show it.”
Suki NakamuraBelowBelowRating held. Improvement plan DEFERRED: it would have been the first written notice. A rating and its consequence are separate decisions.
OutstandingExceedsMeetsBelowTotal
Submitted1779828186
Final1778838186
Shift0-1+100

13 ratings changed, 6 up and 7 down, and the distribution moved by one row. That is the strongest evidence available that the session calibrated raters rather than filling quotas. Session ran 2 hours 43 against a 3-hour booking.

What is in the pack

01

Distribution by Manager

Spread per manager and team with evidence density beside it, because neither column means anything alone. Plus a group cut, computed only where the population passes a stated reporting threshold.

02

Outlier Register

Every triggered row with the evidence quoted verbatim, and pattern rows grouped under the manager they belong to so twenty of them resolve in three conversations.

03

Session Agenda

A running order with minutes on every item and a total against your booked time. Patterns first, individual cases hardest-first, ratification block last and never omitted.

04

Adjustment Log

Filled in live, on the row. Separate columns for what happened to the rating and what happened to the consequence attached to it, because those are two decisions.

05

Calibration Protocol

Two pages the room agrees before item two: what counts as evidence, how a rating changes, and the fact that nobody present has an allocation to spend.

06

Facilitator Guide

Scripts rather than principles. What to say when a manager hears re-substantiation as a downgrade, when nobody argues for the person at the bottom, and when the clock beats you.

07

Two rules every prompt reads first

A distribution is a place to look and never a quota. No adjustment without a sentence the manager can say out loud to the person it concerns.

How it works

  1. 1

    Send the submitted reviews

    All of them, including the ones you think are uncontroversial. An export, a folder, pasted text. Last cycle's ratings and the goal record if you have them.

  2. 2

    Get the distribution with density beside it

    Spread per manager next to the share of away-from-middle ratings that name something checkable. The two columns together are what separates a lenient rater from a strong team.

  3. 3

    See which rows belong to a pattern

    Rows are grouped by whether the finding is about the rater or the rated. In the worked example three patterns absorbed 20 of 31 triggered rows.

  4. 4

    Run a session that fits

    An agenda with a time total, a log written live with the sentence on every change, and per-manager briefs out before anybody delivers an outcome.

Frequently asked questions

Why is my team not on the agenda?

Because every away-from-middle rating you wrote named something checkable, nothing on file contradicts any of them, the goals were in force, and nobody moved two levels. Those 155 rows are ratified in one act by a named signer, which is a recorded decision rather than the meeting running out of time. Running the cycle that way is how the agenda gets shorter.

Do you force a distribution?

No, and the worked example is the proof. Thirteen ratings changed and the spread moved by one row. A distribution tells you where to look. Two managers with identical spreads can be doing opposite things, and only the evidence behind the ratings separates them.

How do you tell a lenient manager from a strong team?

Evidence density, never the histogram. Nine high ratings that each name a shipped deliverable or a date met are a strong team. Nine that read like character references are a fact about the manager. The action for the first is nothing at all. The same test runs on a hiring loop.

Is a rating of four out of five evidence?

No. A number is a conclusion the manager reached, and the register asks what it was reached from. Evidence is something an outsider could go and check: the migration that went live on the committed date, the training week she ran, the deadline missed and the dependency behind it.

What happens to the low ratings nobody argues about?

They get a defensibility check as a document review rather than meeting time: a prior written record dated before the cycle closed, a named outcome, nothing inconsistent on file. Seven of eight passed in the worked example. The one that failed kept its rating and lost its improvement plan.

Why does every change need a sentence?

Because the employee never sees the meeting, only a manager delivering a rating. A manager who cannot explain a change either says it went to calibration or invents a reason. EEOC guidance recommends employees can have appraisals reviewed and corrected (Section 15), and an unexplainable rating cannot survive that.

Do you look at ratings by demographic group?

Yes, at company level, because ratings feeding promotion and pay form part of a selection decision and impact is expected to be disclosable by group (29 CFR 1607.4). Only two of six cuts passed the threshold here, and saying so beats four meaningless ratios. Onboarding evidence starts in the onboarding pack.

Find out how many rows actually need the room

Send the submitted reviews and last cycle's ratings. The first thing back is the count of rows that tripped a trigger, how many of those belong to a manager rather than to a person, and what the agenda comes to in minutes.

Compute my agenda