River
Y CombinatorBacked by Y Combinator
FREE TEMPLATE

Incident Escalation Matrix Template

Five sheets and four documents built so a responder woken at two in the morning classifies the incident by lookup, not judgement.

Free download  ·  No account needed

Severity Definitions

Kesterline, field-service software, 6 on the on-call rota

Written from10 incidents in the 12 months to 31 October
Indexed onSymptoms visible at minute zero, not impact
Reclassification rate50%, which is why this is version 3

Each level carries the incidents that landed in it

LevelWhat the responder can see at minute zeroReal incidents that landed hereIf you cannot tell
SEV1Login, dispatch or card capture failing, not slow. Or writes failing while reads look fine.14 Mar auth out 51 min. 2 Jun card tokens rejected 22 min. 9 Sep failover stuck 68 min.Call it SEV1
SEV2One named feature failing, rest of the product fine. Or one account above 500 seats fully down.21 Jan invoice PDFs down 4 h. 8 Apr largest account locked out 37 min.Call it SEV2, say so
SEV3Visibly wrong, fewer than 50 accounts, and support has a workaround to hand out.29 May export capped at 1000 rows. 12 Aug crash on one Android version.No workaround means SEV2
SEV4One report, nothing broken for anybody else, no revenue path affected.6 Mar a customer's own firewall blocking our webhooks.No ambiguity here

Why the incidents column is the whole sheet

Nobody woken by a pager compares the thing in front of them against a definition. They ask whether this is as bad as the last bad one. A level carrying three dated precedents answers that; four adjectives do not.

The 9 September entry is the reason SEV1 names a symptom rather than an impact. Reads were healthy, so the dashboard stayed green and it ran as a SEV2 for thirty-one minutes.

Every escalation matrix template hands you the same four adjectives. Critical, major, minor, cosmetic, with acknowledge times somebody picked because they sounded proportionate. Adjectives do not survive two in the morning. Nobody woken by a pager compares what they are looking at against a written definition; they ask whether this is as bad as the last bad one. So every level here is written around the real incidents that landed in it, with dates and durations, and classification becomes a lookup.

The definitions everybody ships are written from the postmortem's point of view: data loss, revenue impact, all customers affected. None of that is knowable while the pager is going off. Federal incident handling guidance concedes it, asking handlers to weigh the current impact and the likely future impact if the incident is not immediately contained. So the index is the symptom: what the alert says, whether the thing fails or is merely slow, whether support has a workaround to hand out.

Then the ladder nobody ships. Every template escalates when nobody is making progress, and none of them escalates when nobody answered at all, which is the failure that actually happens after midnight. It has different intervals and it starts by duplicating the call. The same real-history logic sets response targets for slower internal requests too, in the internal service SLA pack. Send River the incident history, or take the sheets blank, alongside the map of the process the incident interrupts, the matrix saying who owns each step and the runbook built from the same history.

Five reclassifications, two escalations that never fired

The log records what each incident was called at minute zero as well as what it turned out to be, which is the only measurement that says whether the definitions are runnable.

Boundary Cases

Illustrative rows for a fictional field-service software company, Kesterline. Pairs of incidents that felt the same and were treated differently.

LineIncident AIncident BWhat actually separated themSentence added
SEV1 / SEV214 Mar, nobody could log in17 Jul, calendar taking 90 s to loadA cannot be waited out. B could, and 80% of users did.Slow with everything completing is never SEV1
SEV1 / SEV22 Jun, all card captures rejected8 Apr, our largest account locked outCard capture is contractually SEV1 at any volume. One big account is not, even at 900 seats.Name the flows that are SEV1 at any volume
SEV2 / SEV321 Jan, invoice PDFs down 4 h29 May, export capped at 1000 rowsNo workaround for A. Support could tell B's callers to filter and export twice.A workaround support can hand out is the line
SEV2 / SEV311 Oct, dispatch down for 34 accounts3 Feb, wrong timezone on job sheetsBoth small. A was a core flow with no workaround, B was cosmetic on a non-core flow.Core flow plus no workaround beats the count
SEV3 / SEV412 Aug, crash on one Android version6 Mar, customer's own firewallOurs was broken in A. Nothing of ours was broken in B.A fault entirely on their side is SEV4

Classification never fails in the middle of a level. Nobody looks at a total outage and wonders whether it is a SEV3. It fails at the lines, and a definition describes a centre, so four more paragraphs of definition fix nothing.

Every row ends in a sentence, and the sentence is the output. A boundary case with no rule added to the definitions is an anecdote. Two of these five turned out to be the same missing rule.

The 9 September failover has no pair, because nothing separated it from a SEV1 except that the symptom was invisible from outside. That got its own line on the SEV1 row.

Escalation Path by Severity

Two ladders in one grid. The column says which trigger each row belongs to.

SevElapsedLadderTriggerActionAuthority added
SEV10bothAlert fires or a human declaresPage primary. Open the channel.none
SEV15 minno answerNot acknowledgedCall the primary. Do not page again.none, duplicate first
SEV110 minno answerStill nothingPage secondary and manager togetherPull anybody off other work
SEV120 minno answerNobody at allCall the VP on the printed cardWake anybody in the company
SEV115 minno progressAcknowledged, no causeManager joins. Second engineer arrives unasked.Rollback, second engineer
SEV130 minno progressNo resolution pathVP joins. Status page up, 30 min cadence.Credits and external comms
SEV130 mincontractualCard capture affectedNotify the named customer contactobligation, not a decision
SEV215 minno answerOwning team silentPage the product-wide primarynone
SEV260 minno progressNo cause identifiedManager joins and decides if this is a SEV1Can reclassify upward
SEV33 daysno progressNo fix scheduledReclassify or accept with a named ownerOwns the decision either way

The no-answer ladder starts by duplicating the contact, on a voice call. A page that failed once will usually fail twice for the same reason, so repeating the channel spends the interval that matters most.

Each no-progress step adds authority, not headcount. A ladder that adds observers is a meeting. And the sheet says outright that the clock fires it, because escalating feels like a complaint about the responder and that is why it so often does not happen.

The no-answer ladder has no final row. An unacknowledged top-severity incident with nobody reachable needs no exit condition: you keep calling.

Contact Roster

Roles on the ladder, names here, with a fallback and a verification date on every row.

RoleChannelNo answer inThenWorks with no company systemsVerified
Primary on-callPage then voice5 minSecondaryYes, mobile on the card2 Nov
Secondary on-callPage then voice5 minEngineering managerYes, mobile on the card2 Nov
Engineering managerVoice call10 minVP EngineeringYes2 Nov
VP EngineeringVoice call15 minCEOYes2 Nov
CEOVoice callend of chainNothing. Keep calling.Yes2 Nov
Database on-callVendor line, account number on the card15 minAccount manager's mobileYes, deliberately not in our systems14 Oct
Named customer contactDirect line, out-of-hours mobilecontractualEscalate inside their organisationYes, in the printed pack14 Oct
The paging app itselfWeb and mobilen/aVoice calls from the printed cardYes, once the card exists2 Nov
Status pageWeb consolen/aSecond admin with independent credentialsYes, by design2 Nov

The no-answer column is what makes this a roster rather than a directory. A directory tells you who to call. A roster tells you what to do when they do not pick up, which is the state the document exists for.

The paging app gets a row because it is a dependency. If it is down or the account is locked, the whole matrix is a list of names nobody is being told about. Same reason the vendor's account number lives on a card: it is otherwise stored inside the system you are ringing about.

Verified is a date somebody actually reached that person on that channel. Not the date the row was written. An unverified row is indistinguishable from a wrong number until the night it matters.

Escalation Log

Two severity columns per incident, which is the experiment. Almost nothing records the first one.

DateIncidentCalled firstFinallyAck inEscalatedShould it have
14 MarAuth returning 500s for everybodySEV1SEV12 minYes, 15 minYes
9 SepFailover stuck, writes down, reads fineSEV2SEV14 minYes, 31 minSooner
2 JunAll card tokens rejectedSEV2SEV13 minYes, 9 minImmediately
17 JulCalendar taking 90 seconds to loadSEV1SEV21 minYes, 0 minNo, and that is fine
8 AprLargest account entirely locked outSEV2SEV26 minYes, 45 min no-answerYes
21 JanInvoice PDF generation downSEV2SEV211 minNo, held 3 h 40Yes at 60 min
11 OctDispatch failing for 34 accountsSEV3SEV218 minNoYes at 60 min
19 FebRegional outage, 03:12 onsetSEV1SEV127 minYes, 20 min no-answerYes
6 MarCustomer's own firewall blocking usSEV3SEV440 minNoNo
totals10 incidents, 12 months5 reclassifiedworst 27 min2 should have fired and did not

A fifty percent reclassification rate is not a functioning matrix. That number is why the policy is on version 3, and it is unobtainable from any template that never records what an incident was called at minute zero. Two of the five came from one missing rule, now a single sentence.

The should-it-have column is worth more than the response times. On 21 January one engineer was nearly there for two hours, and nothing anywhere produced an alert, a missed target or a red cell. It exists only because somebody wrote it down against themselves.

17 July is an over-escalation and it is logged without criticism. Over-calling cost one message and woke three people. Under-calling on 9 September cost thirty-one minutes of the wrong response.

What's in the pack

01

Severity Definitions

Every level names the real past incidents that landed in it, with dates and durations, so classification is a lookup against precedent. The test is whether somebody three weeks into the job can run it at three in the morning, and they are disproportionately the person on call.

02

Indexed on minute zero, not on impact

Data loss, revenue effect and blast radius are postmortem findings. The columns hold what the responder can actually see: what the alert says, whether the thing fails or is merely slow, whether support has a workaround to hand out. Each answerable in thirty seconds by somebody half awake.

03

Boundary Cases

Pairs of near-identical incidents that landed on opposite sides of a line, with the fact that actually separated them and the one sentence added to the definition as a result. Classification never fails in the middle of a level, and a definition describes a centre.

04

The no-answer ladder

Nobody answered is a different trigger from nobody is making progress, and it is the one that fails after midnight. Federal guidance prescribes it outright: duplicate the contact, then escalate and keep repeating until somebody responds, naming failed phones and personal emergencies as the causes.

05

Escalation Path by Severity

Both ladders in one grid with a column saying which trigger each row belongs to. Every no-progress step states the authority it adds rather than the person it adds, because a ladder that adds observers is a meeting.

06

Contact Roster

A no-answer path and a verification date on every row, plus rows for the paging app and the status page, since the tool that pages people is itself a dependency. Verified means a date somebody reached that person on that channel.

07

Escalation Log

Two severity columns per incident, so the reclassification rate becomes visible. The filled example runs at fifty percent, and it also records the two escalations that should have fired and did not, which is the real failure mode and the one nothing else logs.

08

After-hours Procedure and Escalation Standard

One page written to be read on a phone by somebody who has been asleep, plus the four rules behind it. Once the matrix holds, reconstructing the timeline afterwards has real severity data to work with. A handoff that's merely incomplete, not broken, is what a handoff agreement covers instead.

How to use it

  1. 1

    Open in River, or take it blank

    Install the pack in River and hand it your history, or download the four documents and five CSV sheets and fill them yourself.

  2. 2

    Send the incident history

    A ticket export, an incident channel, a postmortem folder, a status page history. Twelve months is plenty and six is workable.

  3. 3

    Name the external clocks

    The shortest notification window anybody outside the company has put on you. That is the one severity threshold you do not have to invent.

  4. 4

    Log the first classification

    From the next incident onward, record what it was called at minute zero as well as what it turned out to be.

Frequently asked questions

Is this free, and what do I get?

Free, with no signup gate on the download. Four documents and five spreadsheets. The AI half is optional: send an incident history and River writes each severity level around the incidents that actually landed in it. More in the template library.

Why not just use the standard four severity levels?

Because critical, major, minor and cosmetic are adjectives, and adjectives are unrunnable when somebody is half awake. A level carrying three dated precedents answers the question a responder is actually asking, which is whether this is as bad as the last bad one.

How many severity levels should we have?

Whatever the response behaviour in your history supports. Sort past incidents by who got woken, how fast somebody acknowledged and whether the status page went up. Those behaviours cluster, and the clusters are your real levels. It is frequently not four.

Where do the acknowledge and resolve times come from?

Externally set clocks wherever they exist, because they are documented somewhere a reader can check and not negotiable at two in the morning. Contractual notification windows, credit tiers in a customer agreement, regulatory reporting boundaries. Where nothing external exists, your own history at the ninetieth percentile.

What makes a threshold actually usable during an incident?

A number rather than a judgement. Regulators write them in exactly the right shape: the communications outage rules put the boundary at thirty minutes affecting at least 900,000 user minutes, notified within 120 minutes of discovery. A product of scale and duration, and a clock starting at discovery.

What is the reclassification rate for?

It is the only real measure of whether the definitions work. Record what each incident was called first as well as finally, then count. Under one in ten and the definitions are runnable. Half means they are not, and the boundary cases sheet says which line is broken.

Do we need the escalation timers automated?

Automate them once the definitions stop moving. A timer firing on a level that gets reclassified half the time just pages people about the wrong thing faster. Get the reclassification rate down first, then wire the intervals into whatever already pages you.

Write levels somebody can run

Take the documents and CSV sheets blank, or install this pack in River and send it twelve months of incident history.

Edit with AI