River
Y CombinatorBacked by Y Combinator
FREE TEMPLATE

Sales Forecast Process Template

Every rep's commit multiplied by what their own closed history says it has been worth, and refused where it cannot be.

Free download  ·  No account needed

Six quarters of commit against what landed. The mean is not the finding. The spread is.

RepSix quarterly ratiosMeanSpreadVerdict
H. Vasquez1.01 · 0.99 · 0.99 · 1.02 · 1.00 · 1.011.0030.011×1.00
A. Redgate0.90 · 0.87 · 0.88 · 0.90 · 0.86 · 0.890.8830.015×0.88
R. Castellan0.71 · 0.68 · 0.71 · 0.68 · 0.74 · 0.660.6970.026×0.70
K. Oyelaran1.23 · 1.25 · 1.24 · 1.19 · 1.31 · 1.221.2400.037×1.24
M. Thibault1.35 · 0.67 · 1.38 · 0.64 · 1.31 · 0.711.0100.338refused
T. Bergström— · — · — · — · 0.93 · 1.081.0052 qtrs×0.91

M. Thibault's mean of 1.010 reads as a rep who needs no correction at all, and a multiplier there would publish false precision on a book that has swung from plus 38% to minus 36%. That commit gets set by inspecting all 14 deals in it instead. Six survived.

Every forecast process guide is an agenda plus a set of category definitions. Time-box the call, ask for evidence rather than confidence, demote anything undefended. All correct, all necessary, and none of it accumulates. Run that process for eight quarters and the ninth starts exactly where the first did, because nothing in it remembers what any rep's commit has been worth. The roll-up is still the arithmetic sum of what six people typed.

Calderwood Systems measured every rep two ways over six quarters. Bias is whether a rep over- or under-calls, and a multiplier corrects it: one rep has landed 69.7% of commit with a dispersion of 0.026, which makes them the most useful rep on the board once multiplied. Dispersion is how much the error moves, and nothing corrects it. One rep's mean is 1.010, reading as no correction needed. The six quarters ran 1.35, 0.67, 1.38, 0.64, 1.31 and 0.71.

So a multiplier is issued only where at least three quarters exist and the dispersion is tight enough to mean something, and the rest get their commit set by inspecting every deal in it. A $6,590,000 submission becomes $5,694,000, and a reported $860,000 gap to quota becomes $1,756,000, which is 2.04 times larger. It pairs with the hygiene rules that decide which deals belong in a submission and with two revenue views that agree, and it reads a category field that may be updating itself from the deal stage.

What the calibration is worth, and where the revenue was actually sitting

A fictional B2B software company: six reps, a $7,450,000 quarterly quota, six quarters of submissions against outcomes. Every figure on these sheets reconciles against the others.

Forecast vs Actual History

Commit as submitted at each period's lock, against total closed-won for that rep in the period. Landed means the rep's whole number, not the portion of the commit list that closed, because the board is given the whole number.

QuarterCommitted at the lockClosed wonNaive errorCalibrated using only prior quartersCalibrated errorReps a multiplier could be issued to
Q3 2024$5,070,000$5,109,000−0.8%$5,070,000−0.8%0
Q4 2024$5,350,000$4,564,000+17.2%$5,350,000+17.2%0
Q1 2025$5,460,000$5,494,000−0.6%$5,460,000−0.6%0
Q2 2025$5,420,000$4,660,000+16.3%$5,069,865+8.8%4
Q3 2025$6,570,000$6,530,000+0.6%$6,135,303−6.0%4
Q4 2025$6,570,000$5,874,000+11.8%$6,194,432+5.5%4
Mean absolute error, Q2 to Q4 20259.6%6.8%

A 29.5% reduction in mean absolute error over the three quarters where the method was live. The first three quarters carry no calibration because there was nothing to calibrate from, which is the honest cost of the method: three quarters of accumulation before it does anything at all.

And Q3 2025 got worse. A quarter the naive sum happened to call within 0.6% came out 6.0% low once calibrated. A calibration trades a large one-directional error for a smaller two-sided one, and any single quarter can land on the wrong side of that trade. The claim is about the mean absolute error over several quarters, stated as exactly that.

Forecast Roll-up

Q1 2026. A multiplier is issued only at three or more quarters of history with a dispersion at or below 0.15. Everything else gets the team default or a number set by inspecting the deals.

RepQuotaCommit submittedBest case submittedQuartersMean ratioDispersionMultiplierCalibrated commitCorrectionBasis
A. Redgate$1,450,000$1,310,000$1,940,00060.8830.0150.88$1,152,800−$157,200Own history
H. Vasquez$1,150,000$1,040,000$1,460,00061.0030.0111.00$1,040,000$0Own history. A 1.00 earned, not assumed
K. Oyelaran$1,050,000$760,000$1,320,00061.2400.0371.24$942,400+$182,400Own history. Sandbags consistently
M. Thibault$1,300,000$1,180,000$2,010,00061.0100.338None$764,000−$416,000Set by inspecting 14 commit deals. 6 survived
R. Castellan$1,500,000$1,420,000$2,180,00060.6970.0260.70$994,000−$426,000Own history. Largest and steadiest correction
T. Bergström$1,000,000$880,000$1,290,00021.0050.0750.91$800,800−$79,200Team default. Two quarters is not a history
All reps$7,450,000$6,590,000$10,200,000$5,694,000−$896,000Four on own history, one default, one inspection

The naive sum reads 88.5% of quota, an $860,000 gap that produces no action. The calibrated roll-up reads 76.4%, a $1,756,000 gap that is 2.04 times larger. That ratio is the most useful number in the space, because a shortfall of twice the reported size changes what leadership does about it.

The corrections run both ways, which is what establishes the method is measuring rather than discounting. K. Oyelaran's multiplier of 1.24 adds $182,400 back. Best case is reported at face value and never calibrated, because the multipliers were fitted on commit and applying one to best case is a category error.

Category Movement

Q4 2025, read as a transition matrix from the submission lock rather than as a net movement figure. Where each deal sat when the number went in, against where it ended.

Category at the lockDealsValue at the lockWonLostStill openWin rate by countValue landedShare of its own valueShare of all landed valueTarget close rate
Commit95$6,570,00061122264.2%$4,320,00065.8%73.5%Above 90%
Best case100$5,400,00026245026.0%$1,038,00019.2%17.7%40 to 60%
Pipeline273$12,900,00011442184.0%$412,0003.2%7.0%Below 20%
Omitted123$3,700,000334862.4%$104,0002.8%1.8%Near zero
All categories591$28,570,00010111437617.1%$5,874,00020.6%100.0%

Headline forecast accuracy: $5,874,000 landed against a $6,570,000 commit, which reads 89.4% and looks like an ordinary shortfall. Underneath it, two large failures with opposite signs. $2,250,000 of committed value did not land, which is 34.2% of the commit. And $1,554,000 of what landed had never been in commit at all, which is 26.5% of everything that landed. They partly cancel, so no single accuracy percentage can show either one.

Commit delivering well under its target while best case also sits under its band is one diagnosis rather than two problems: commit is overflowing downward, so the definition is being applied loosely and the arithmetic is fine. And a quarter of the revenue arriving from outside the number is a forecasting failure however close the total came.

What the pack does that a forecast agenda does not

01

Separates bias from dispersion, and only corrects one

Bias is whether a rep over- or under-calls and a multiplier fixes it completely. Dispersion is how much their error moves and no multiplier touches it. Every accuracy report blends the two into one percentage.

02

Refuses to issue a multiplier that cannot work

Under three quarters of history gets the team default. A dispersion above 0.15 gets no multiplier at any history length, because a mean of 1.010 on ratios from 0.64 to 1.38 publishes false precision.

03

Sets the refused reps' numbers by inspecting the deals

Every deal in that commit against the four conditions, asked one at a time with an answer each. Steady submissions earn a cheap process and unpredictable ones earn an expensive one, and the way out is three steady quarters.

04

Backtests itself and reports the quarters it made worse

Each closed quarter calibrated from only the quarters before it. Mean absolute error fell from 9.6% to 6.8% across three live quarters, and one of those three came out further off than the naive sum.

05

Reads category movement as a transition matrix

Where the value that landed was sitting when the number went in, rather than a net count of deals that moved. That is the only view that finds the 26.5% of landed revenue which never appeared in the commit.

06

Reports the gap to quota twice

Naive and calibrated, with the ratio between them. A reported $860,000 shortfall that is really $1,756,000 is the difference between a quarter nobody acts on and one somebody does. Whether next quarter's pipeline can even cover that revised number, by segment and by deadline, is what a coverage and capacity check runs next.

07

Withdraws a multiplier when its basis breaks

A territory change, a segment change, a quota move of more than about a fifth, or a change to the category definitions resets a rep to untested and back onto the team default.

How the pack runs

  1. 1

    Send four to six quarters of submissions

    Commit and best case per rep as submitted at each lock, plus total closed-won per rep in each period. Board decks and submission emails work fine where the platform did not retain the locked values.

  2. 2

    Measure each rep twice and backtest the method

    Bias and dispersion separately, then the whole calibration replayed over closed quarters using only prior history, with the quarters it got wrong reported rather than dropped.

  3. 3

    Issue the multipliers and the refusals

    Fitted multipliers where the gates allow one, the team default where the history is too short, and a deal-level inspection list against the four commit conditions for everybody else.

  4. 4

    Run the call and assemble the number

    The agenda comes from the exceptions the calibration found, so measured reps are not walked through their books. Five lines go upward, including the method's own backtested error.

Frequently asked questions

Is this just telling reps their numbers are wrong?

Two of the six reps got corrections upward. A rep who consistently lands 124% of commit is being under-read, and their multiplier added $182,400 back. The method measures rather than discounts, and the row that proves it is the one that adds value.

Why not just tell the over-callers to submit lower?

That moves the bias without touching the dispersion. A rep instructed to be conservative becomes conservative by an amount nobody has measured, which is how the bias got there. A multiplier is a correction applied where it can be measured instead of one asked for where it cannot.

One rep averaged 1.010. Why is that a problem?

Because the six quarters behind it were 1.35, 0.67, 1.38, 0.64, 1.31 and 0.71. The mean says no correction is needed, which is exactly wrong. A multiplier is a single scalar: it shifts the centre of a distribution and leaves the spread untouched.

How many quarters before this beats just adding up the submissions?

Three, and it does nothing at all before that. On the worked example the first three of six quarters carry no calibration because there was nothing to calibrate from. That is the honest cost, and any page claiming a forecast method works from day one is describing an agenda.

Does our forecast category field already track this?

Check how it is maintained first. HubSpot documents that with the automate-forecast-categories setting on, the category updates itself when a deal moves to a new stage, so the field is a projection of the stage and carries no rep judgment. Most teams do not know which mode they are in.

Our forecast never matches the CRM number anyway.

Different problem, and worth settling first. Two totals that disagree an hour before the call need a deal-by-deal reconciliation rather than a calibration. This pack assumes one agreed pipeline number and calibrates the judgment layered on top of it.

What if our commit definition changes?

Every multiplier fitted before the change is withdrawn, and the affected reps go back on the team default until three fresh quarters accumulate. It is the rule people most want to skip. Getting the underlying definitions computable first is what stops it firing every quarter.

Find out what your reps' commits have been worth

Send four to six quarters of submissions and what landed in each, and get every rep measured on bias and dispersion separately.

Get the template