River
Y CombinatorBacked by Y Combinator
FREE TEMPLATE

Lead Scoring Model Template

Four documents and four sheets that fit every attribute weight to your own conversion history, then validate the score by conversion within band.

Free download  ·  No account needed

Assigned Points Against Fitted Weights

Duffield, 12,400 leads, 405 closed-won, base rate 3.27%

One joint logistic fit across every candidate attribute on 18 months of leads with outcomes. Joint, so two signals that fire for the same person do not both take full credit. Illustrative, for a fictional warehouse management software company.

AttributeLeadsWonRateLiftOldFitted
Integration docs visit91611212.23%3.74x0+25
Demo requested, under 30d7017911.27%3.45x+30+21
200 to 999 employees3,7302175.82%1.78x+5+17
Industry: distribution or 3PL6,4712914.50%1.38x+8+17
More than one warehouse4,9392114.27%1.31x0+10
Whitepaper or ebook download3,505912.60%0.79x+10-6
1,000 or more employees2,075331.59%0.49x+10-8
Title: C-level1,313161.22%0.37x+15-13
Demo requested, over 30d762354.59%1.41x+30dropped
Opened 3+ emails in 7 days4,9701462.94%0.90x+5dropped

Then the validation, which is the part no points table has

Conversion by score band on the same leads. Top-decile lift goes from 2.84x to 4.44x. The old model's top band held 40 of 12,400 leads, so it was never a group anyone could staff against, and its middle band held 51% of the book at the base rate. Zero out its single largest weight and its lift falls to 1.56x against a 1.00x floor for scoring at random, so sixteen of its seventeen attributes were carrying almost none of the ranking between them.

And the threshold gets priced. Of the 405 closed-won deals, 262 scored below the old MQL line of 45 and went to the nurture stream. Hold the volume handed to sales constant and the fitted model's line at 51 passes the same 2,048 leads while holding 230 of those deals instead of 143.

Every lead scoring template is a points table somebody assigned. Adobe's own shipped behaviour scoring program is a fair example of the genre: a contact form worth +30, a webinar worth +20, a PDF download worth +5, no activity worth -10. Its authoring guide tells you to set those numbers relative to the importance of each action. Starting out wrong is unavoidable. The problem is that nothing in the table ever tells you it is wrong, so the tuning never converges.

So this pack fits the weights instead. The worked example runs on 12,400 leads and 405 closed-won deals from a warehouse management software company. Three attributes flip sign there: C-level titles convert at 1.22% against 3.51% for everything else, employers over a thousand people at 1.59%, gated content downloads at 2.60%. Five of twenty-one candidate terms come out entirely for lack of evidence. Two attributes nobody had ever scored carry real weight, and one is now the largest term in the model.

Then the threshold gets priced in the only currency that matters, which is the closed-won deals it held back. On the same history, 262 of the 405 wins scored below the old MQL line and were routed to nurture. Holding the volume handed to sales constant, the fitted model's line carries 230 of them instead of 143. Pair it with lead routing rules, which decides which owner a qualified lead reaches, and the SLA that governs what happens next.

Weights with conversions behind them, a validation, and a priced threshold

Attribute Weights with the old points next to the fitted ones and every dropped term keeping its reason, the band validation on both models, and the threshold priced in closed-won deals.

Attribute Weights

Illustrative, for a fictional warehouse management software company called Duffield. Twenty-one candidate terms, sixteen scored. The five without a weight stay in the sheet with the reason, because that row is what the next recalibration re-tests against.

TermLeadsWonRateLiftOldFittedWaldChange
Integration docs visit91611212.23%3.74x0+2510.3never scored before
Demo requested, under 30d7017911.27%3.45x+30+217.6freshness window kept
200 to 999 employees3,7302175.82%1.78x+5+178.4raised 12 points
Industry: distribution or 3PL6,4712914.50%1.38x+8+178.0raised 9 points
Title: VP or Director2,5371295.08%1.56x+10+145.9carried over
Known ERP in the stack2,9631454.89%1.50x+5+115.2carried over
Pricing page visit, under 30d1,433916.35%1.94x+10+114.5freshness window kept
ROI calculator used692385.49%1.68x+8+113.4carried over
More than one warehouse4,9392114.27%1.31x0+105.1never scored before
Title: Manager3,0801254.06%1.24x+5+93.9carried over
Returned 3+ times in 14 days1,891874.60%1.41x+10+93.6carried over
Region: US or Canada9,4773253.43%1.05x+5+62.4carried over
Whitepaper or ebook download3,505912.60%0.79x+10-6-2.8sign flip
Personal email domain1,425342.39%0.73x-15-7-2.1kept negative, lighter
1,000 or more employees2,075331.59%0.49x+10-8-2.2sign flip
Title: C-level1,313161.22%0.37x+15-13-2.7sign flip
Demo requested, over 30d762354.59%1.41x+300-not distinguishable from zero
Pricing page visit, over 30d1,654663.99%1.22x+100-not distinguishable from zero
Webinar registration2,132673.14%0.96x+50-not distinguishable from zero
Opened 3+ emails in 7 days4,9701462.94%0.90x+50-not distinguishable from zero
Analyst report request13264.55%1.39x+50-under the 12-deal floor

Two of the five dropped terms are the stale halves of split signals. A demo request converts at 11.27% inside thirty days and 4.59% outside it, and the outside half fails the significance rule, so an old demo request now scores zero rather than a discounted weight. Integration documentation is the counterexample and the more useful half of the finding: 12.58% inside thirty days against 11.88% outside, so the uniform decay curve every scoring guide recommends would have quietly discounted the strongest signal in the model for being three months old.

Conversion by Score Band

Same leads, same bands, both models. This is the only artifact that decides whether the new model ships, and ordering by itself is not the test: most points tables do rank in roughly the right direction.

BandInherited points tableFitted model
LeadsOf bookWonRateLeadsOf bookWonRate
0 to 193,26226.3%491.50%3,60729.1%200.55%
20 to 396,33251.1%1812.86%4,73438.2%731.54%
40 to 592,17417.5%1155.29%3,01624.3%1414.68%
60 to 795924.8%549.12%8236.6%10512.76%
80 to 100400.3%615.00%2201.8%6630.00%
All leads12,400100%4053.27%12,400100%4053.27%
Lift at top decile---2.84x---4.44x
Lift with the largest weight zeroed---1.56x---4.02x

Both models rank in the right order, so the two tests that separate them are about the shape. The inherited table's top band holds 40 of 12,400 leads, which is not a group anybody can staff against, and its 20 to 39 band holds 51% of the book at 2.86% against a 3.27% base, which is the base rate with a label on it. Then zero out its single largest weight, the demo request at 30 points, and top-decile lift falls from 2.84x to 1.56x. The same test on the fitted model leaves 4.02x, so its ranking is spread across attributes rather than resting on one.

The threshold, priced in closed-won deals

Duffield routes anything at or above the line to an account executive the same day and everything below it into the nurture stream. Every scoring guide discusses where to set the line as a matter of sales capacity. None of them crosses the score against the outcome to ask what the line already cost, which is one join on the export.

Inherited table at 45Fitted model at 51
Leads at or above the line1,975 (15.9%)2,048 (16.5%)
Closed-won inside the line143 of 405 (35.3%)230 of 405 (56.8%)
Closed-won below the line262175

The new line was not chosen, it was solved for. The old line passed 1,975 leads, so the new one is the score at that rank on the fitted model. Same work for sales, 87 more closed-won deals inside the threshold, 21.5 points more of all wins. A recalibration that quietly hands sales forty percent more leads is a staffing proposal wearing a model's clothes.

The swap sales needs to be told aboutLeadsWonRate
Passed by both models782--
Passed by the old model only1,193302.51%
Passed by the new model only1,2661179.24%

Sixty percent of the routed book changes hands, and a rep who has worked the top of the old list for two years will notice. Ranked by the points each attribute gave the average lead in the dropped group against what it gives now: demo requested -13.8, whitepaper download -8.4, C-level title -8.1, employer over a thousand people -6.8. The largest contributor is not the demo request. It is the demo request older than thirty days, which the old table scored at a flat 30 and which now scores zero, and 762 of those were sitting in the same same-day queue as a form filled that morning.

What is in the pack

01

Attribute Weights

Every term with its lead count, closed-won count, conversion rate, lift, old points, fitted points and Wald statistic

02

Score Distribution

Decile by decile, with the cumulative share of all wins at or above each line and the route that decile gets

03

Conversion by Score Band

The validation, both models on the same leads and bands, ending in lift at the top decile

04

Model Change Log

One row per change with the evidence for it in the next column, plus open items dated to the next recalibration

05

Scoring Methodology

Why the fit is joint, the two evidence rules, what the fit disagreed with the old table about and by how much

06

Signal Definitions

Every scored term, the field it reads from, and the four kinds of change that void its weight until a refit

07

Threshold and Rollout Note

The threshold priced in won deals, the volume-matched replacement, and the shadow month before routing changes

08

Recalibration Procedure

The refit schedule, five triggers that pull one forward, and what gets re-tested rather than just refitted

How it works

  1. 1

    Send the leads with outcomes on them

    One row per lead, with the closed-won flag on the same row as the attribute values, plus the date each behavioural signal fired. A year works, eighteen months is better. Bring the current points table too, because the comparison column is the most persuasive thing in the finished sheet.

  2. 2

    Fit every weight at once

    One joint model across every candidate attribute, including the ones nobody scores. Joint matters: 61% of the leads that read integration documentation also visited pricing, so a table scoring both at full value pays twice for one person's afternoon.

  3. 3

    Drop what the outcomes cannot support

    Two rules, written down before the fit runs. A term needs twelve closed-won deals behind it, and a coefficient inside plus or minus 1.96 standard errors scores zero rather than a small number. Then refit without the dropped terms and publish the refit.

  4. 4

    Validate, then price the line

    Conversion by band on both models, the share of the book in each band, and lift with each model's largest weight zeroed. Then the closed-won deals below the current threshold, and the volume-matched line that recovers most of them.

Frequently asked questions

What do I need before this is useful?

A lead export where the outcome and the attribute values sit on the same row, covering at least a year. The closed-won count matters more than the lead count: forty deals cannot support twenty weights, and the honest output there is a five-term model with a note about volume rather than twenty coefficients with nothing behind them.

Our CRM already has predictive lead scoring

Then use it, and read what it needs. Salesforce documents that Einstein Lead Scoring uses a global model trained on other customers' anonymised data until your org has 1,000 leads and 120 conversions in six months. Below that line you are scoring your leads with somebody else's history, and the vendor says so.

Why fit jointly instead of one attribute at a time?

Because the signals travel together and a per-attribute table double counts them. On the worked example 61% of the leads that read integration documentation also visited pricing and 40% also requested a demo. Scoring all three at full value ranks the people who clicked around most at the top, rather than the people most likely to buy.

Some of the fitted weights are negative on attributes we like

Publish them. A C-level title converting at 1.22% against 3.51% for everything else is a result, and suppressing it is the same act as assigning the points by opinion. If the real reason is that enterprise deals are slow rather than bad, that is a cycle-length finding and it belongs in its own field, not in a fudged coefficient.

Should behavioural points decay after 30 days?

Test it per signal instead of applying one curve. Two of nine behavioural signals earned a freshness window on the worked example. Integration documentation did not: 12.58% inside thirty days against 11.88% outside it. A uniform half-life would have discounted the strongest term in the model for being three months old.

How do we pick the new threshold?

Do not pick it. Solve for it by holding the lead volume handed to sales constant, then report the share of the routed book that changes hands. On the worked example that share is 60%, which is what the conversation with sales is actually about. It pairs with lifecycle stage definitions, since the threshold is a stage boundary.

Does this cover routing, hygiene or the handover agreement?

No, and the boundary is deliberate. Which owner a scored lead reaches is a routing question, the score depends on deduplicated records to be computable at all, and what the two teams owe each other around the handoff belongs in the SLA. An account ownership overlap is a separate, earlier problem: which single owner an account should have at all.

Find out what your points table is worth

Send the lead export with outcomes on it. The first thing back is conversion by score band on your current model, next to the fitted one.

Fit my weights to my own history