River
Y CombinatorBacked by Y Combinator
FREE TEMPLATE

Product OKR Examples Template

Every OKR template gives you a form. This one goes looking for the query behind each number, and finds the ones nobody can read.

Free download  ·  No account needed

A key result is a claim that a specific number will move by a specific amount. It can be empty in exactly two ways and every OKR template in circulation checks neither: the number cannot be read anywhere, or the move being asked for has never once happened. Both checks require going and looking at data before the period starts, which is inconvenient, because the appeal of an OKR form is that it can be filled in during a workshop.

So the register carries two columns no template has. Data Source holds the actual query, the table and column or the event and property, and it is not satisfiable by writing "analytics". Best Move On Record holds the largest change that metric has made in a single period in its own history, which turns "is this ambitious or delusional" from a matter of temperament into a ratio. In the worked period, nine proposed key results produced five with a query, three with none, and one deliverable in disguise.

Of the five that could be checked, all five asked for a bigger move than the metric had ever made, at a median of four times. One survived, because a new matching model was in staging that had not existed when the record was set, and that is exactly the answer the check exists to find. Capacity for the period comes from the quarterly planning pack, and where the effort actually went last time comes from the portfolio review.

Nine proposed key results, checked before anybody committed to them

Key Result Register, the gaps it produces, and the weekly pace that comes out of it.

Key Result Register

Illustrative, for a fictional expense management product called Quillon. One objective: make submitting an expense something people finish on the day it happens.

Proposed key resultQuery behind itBaselineMove askedBest everOverVerdict
Median submit lag to 3.0 daysexpenses.submitted_at8.4 days5.4 days1.14.9xTarget implausible
Mobile submission share to 55%event: platform31%24 pts64.0xTarget implausible
Receipt auto-match to 85%matcher job logs62%23 pts92.6xStretch, mechanism exists
Policy rule configured to 70%policy_rules table44%26 pts73.7xTarget implausible
Cost per report to $2.10finance, monthly only$2.90$0.80$0.155.3xWrong grain
Submission satisfaction 4.5 of 5no survey existsnone---No instrumentation
Cannot-submit tickets down 40%tag created 5 wks ago5 weeks---Incomplete baseline
Time to first approval to 7 days3 conflicting definitions11, 14 or 19---Contested definition
Ship the mobile receipt scannernot a metric----Not a key result

Five have a query. Four do not, and they fail four different ways that cost four different things: engineering, a backfill, one decision, and being on the wrong list entirely. Of the five that could be checked, every single target exceeded the metric's own record, at a median of 4.0x. Auto-match survived at 2.6x because the new matching model was in staging and had not existed when the previous record was set. None of the other four had an answer like that.

Instrumentation Gap List

A separate sheet because it has a different owner and a different deadline. It goes to engineering in week one, sized.

GapBlocksKindWorkDueWhat happens if it slips
No in-product survey anywhereKR-04No instrumentation2 eng-weeksWeek 2Ungradeable at the review, and argued about instead of read
Cannot-submit tag has 5 weeks of historyKR-05Incomplete baseline1 eng-weekWeek 2A 40% fall claimed against five weeks that included a release
New company means three different thingsKR-08Contested definition0 eng. One decisionWeek 1Cheapest row here and the one still open in week nine
Platform property missing on 4% of eventsKR-02Partial coverage2 daysWeek 3Mobile share understated by up to 4 points against a 14-point target
Match failures logged with no reason codeKR-03Diagnostic3 daysWeek 4The number moves and nobody can say why, so next period's target is another guess
Cost per report arrives monthly, not weeklyKR-09Wrong grain4 eng-weeksdeferredNot blocking. KR-09 became a monitored metric rather than a key result

Three engineer-weeks and one decision closes every blocking gap. That total is what makes the list actionable, because three engineer-weeks is a conversation somebody can have in week one and "we need better instrumentation" is not. The free row is the trap: it costs no engineering, which is exactly why nobody schedules it.

Progress Tracker

Actual against required pace, every week. On track means at or ahead of pace. It does not mean anybody feels good about it.

WeekSubmit lagMobile shareAuto-matchPolicy rulesSatisfactionCall
Baseline8.4 days31%62%44%no readingCommitted
Week 38.2 need 7.833% need 34.366% need 67.445% need 47.3-Policy rules behind
Week 48.135%69%46%3.9 first readTarget reset to 4.2
Week 67.640%74%47%4.0At risk, no initiative
Week 77.442%76%51%4.0Recovering
Week 116.348%83%58%4.2On pace
Week 135.9 / 6.050% / 45%84% / 85%60% / 58%4.3 / 4.24 of 5 met

Every target came from the metric's own history, at a median of 2.3x its record rather than 4.0x. The pace column caught policy rule configuration in week three, when it needed 1.1 points a week and had averaged 0.3 with nothing assigned to it, ten weeks before a status meeting would have found it. The original nine, committed as written, would have finished nought for five and left four with no number at all.

What is in the pack

01

Key Result Register

One row per key result carrying the query that produces the number, the baseline read from that query, and the largest move the metric has ever made in a period.

02

Four verdicts, not a checkbox

Measurable, no instrumentation, incomplete baseline, contested definition. Plus a fourth for the deliverable in disguise. Each one costs something different to fix.

03

Instrumentation Gap List

Every gap sized in engineer-weeks with an owner and a date inside week three, and the consequence at the review written next to it so the work gets scheduled.

04

The history check

Every target expressed as a multiple of the metric's own record. Above 3x needs a named mechanism that was not present when that record was set, or it gets reset.

05

Measurement Plan

Window, population, contested definitions and contaminated weeks, all settled before the period rather than negotiated during the review that depends on them.

06

Progress Tracker

Required pace alongside actual, weekly. Two consecutive readings behind pace gets a named initiative or an explicit note that nothing is being done.

How it works

  1. 1

    Send the drafts

    Key results however rough, plus anything that says what you can measure today. A taxonomy, a dashboard list, or just your tools.

  2. 2

    Find every query

    Each key result gets a table and column, an event and property, or a named report. Whatever has none becomes a sized gap.

  3. 3

    Check every target

    Against the biggest move that metric has ever made. Above three times the record needs a mechanism, written down next to it.

  4. 4

    Commit what survives

    Usually three to five rather than nine, each with a query, a baseline, a defensible target and a weekly pace.

Frequently asked questions

What do I need to have before this is useful?

Draft key results and some idea of what you can measure. An event taxonomy, a dashboard list, warehouse table names or just the tools you use. An objective alone is a better starting point than most, because the key results get proposed from what the data supports. Where the objective itself is unsettled, the tree that hangs work off an outcome comes first.

Is this saying ambitious targets are bad?

No. The check exists to find the ambitious targets that have something behind them. A target at three times a metric's record with a capability shipping this period is the best kind of goal, and it survives. One at four times with nothing new behind it is a number that arrived in a room.

What counts as a data source?

A table and a column, an event and a property, a named dashboard tile, or a report somebody runs by hand. One sentence you can write down. "Analytics" is not a source, and neither is "we track that", because in week thirteen somebody has to reproduce the number and it may not be you.

Why separate the gaps into their own sheet?

Because they have a different owner and a different deadline. Gaps go to engineering in week one, sized, with the review consequence written next to each. The cheapest one is always a contested definition, costing no engineering at all, which is precisely why it is the row still open in week nine.

Does it grade the OKRs at the end?

No. No scoring, no 0.7 convention, no rollup. Grading is an organisational ritual and this is a measurement method. A key result either reached its target or it did not, the tracker says which, and anything whose gap never closed is reported as ungradeable rather than estimated.

Where does the definition of a key result come from?

The canonical framing puts it plainly: an objective is what is to be achieved, no more and no less, and the key results are the measurable part. This pack takes the word measurable literally and asks what produces the number, the same way the roadmap pack asks what produced the date.

Why does the order matter so much?

Because instrumentation is downstream of the metric. Amplitude's own taxonomy playbook builds a tracking plan as objectives, then key metrics, then the events that produce them. A key result written after the period starts is already too late to be instrumented inside it.

Find out which of your key results has no number

Send the drafts and whatever says what you can measure. The first thing back is the query behind each one, or the fact that there is not one.

Check my key results