Incident Escalation Matrix Template
Five sheets and four documents built so a responder woken at two in the morning classifies the incident by lookup, not judgement.
Free download · No account needed
Severity Definitions
Kesterline, field-service software, 6 on the on-call rota
| Written from | 10 incidents in the 12 months to 31 October |
| Indexed on | Symptoms visible at minute zero, not impact |
| Reclassification rate | 50%, which is why this is version 3 |
Each level carries the incidents that landed in it
| Level | What the responder can see at minute zero | Real incidents that landed here | If you cannot tell |
|---|---|---|---|
| SEV1 | Login, dispatch or card capture failing, not slow. Or writes failing while reads look fine. | 14 Mar auth out 51 min. 2 Jun card tokens rejected 22 min. 9 Sep failover stuck 68 min. | Call it SEV1 |
| SEV2 | One named feature failing, rest of the product fine. Or one account above 500 seats fully down. | 21 Jan invoice PDFs down 4 h. 8 Apr largest account locked out 37 min. | Call it SEV2, say so |
| SEV3 | Visibly wrong, fewer than 50 accounts, and support has a workaround to hand out. | 29 May export capped at 1000 rows. 12 Aug crash on one Android version. | No workaround means SEV2 |
| SEV4 | One report, nothing broken for anybody else, no revenue path affected. | 6 Mar a customer's own firewall blocking our webhooks. | No ambiguity here |
Why the incidents column is the whole sheet
Nobody woken by a pager compares the thing in front of them against a definition. They ask whether this is as bad as the last bad one. A level carrying three dated precedents answers that; four adjectives do not.
The 9 September entry is the reason SEV1 names a symptom rather than an impact. Reads were healthy, so the dashboard stayed green and it ran as a SEV2 for thirty-one minutes.
Every escalation matrix template hands you the same four adjectives. Critical, major, minor, cosmetic, with acknowledge times somebody picked because they sounded proportionate. Adjectives do not survive two in the morning. Nobody woken by a pager compares what they are looking at against a written definition; they ask whether this is as bad as the last bad one. So every level here is written around the real incidents that landed in it, with dates and durations, and classification becomes a lookup.
The definitions everybody ships are written from the postmortem's point of view: data loss, revenue impact, all customers affected. None of that is knowable while the pager is going off. Federal incident handling guidance concedes it, asking handlers to weigh the current impact and the likely future impact if the incident is not immediately contained. So the index is the symptom: what the alert says, whether the thing fails or is merely slow, whether support has a workaround to hand out.
Then the ladder nobody ships. Every template escalates when nobody is making progress, and none of them escalates when nobody answered at all, which is the failure that actually happens after midnight. It has different intervals and it starts by duplicating the call. The same real-history logic sets response targets for slower internal requests too, in the internal service SLA pack. Send River the incident history, or take the sheets blank, alongside the map of the process the incident interrupts, the matrix saying who owns each step and the runbook built from the same history.
What's in the pack
Severity Definitions
Every level names the real past incidents that landed in it, with dates and durations, so classification is a lookup against precedent. The test is whether somebody three weeks into the job can run it at three in the morning, and they are disproportionately the person on call.
Indexed on minute zero, not on impact
Data loss, revenue effect and blast radius are postmortem findings. The columns hold what the responder can actually see: what the alert says, whether the thing fails or is merely slow, whether support has a workaround to hand out. Each answerable in thirty seconds by somebody half awake.
Boundary Cases
Pairs of near-identical incidents that landed on opposite sides of a line, with the fact that actually separated them and the one sentence added to the definition as a result. Classification never fails in the middle of a level, and a definition describes a centre.
The no-answer ladder
Nobody answered is a different trigger from nobody is making progress, and it is the one that fails after midnight. Federal guidance prescribes it outright: duplicate the contact, then escalate and keep repeating until somebody responds, naming failed phones and personal emergencies as the causes.
Escalation Path by Severity
Both ladders in one grid with a column saying which trigger each row belongs to. Every no-progress step states the authority it adds rather than the person it adds, because a ladder that adds observers is a meeting.
Contact Roster
A no-answer path and a verification date on every row, plus rows for the paging app and the status page, since the tool that pages people is itself a dependency. Verified means a date somebody reached that person on that channel.
Escalation Log
Two severity columns per incident, so the reclassification rate becomes visible. The filled example runs at fifty percent, and it also records the two escalations that should have fired and did not, which is the real failure mode and the one nothing else logs.
After-hours Procedure and Escalation Standard
One page written to be read on a phone by somebody who has been asleep, plus the four rules behind it. Once the matrix holds, reconstructing the timeline afterwards has real severity data to work with. A handoff that's merely incomplete, not broken, is what a handoff agreement covers instead.
How to use it
- 1
Open in River, or take it blank
Install the pack in River and hand it your history, or download the four documents and five CSV sheets and fill them yourself.
- 2
Send the incident history
A ticket export, an incident channel, a postmortem folder, a status page history. Twelve months is plenty and six is workable.
- 3
Name the external clocks
The shortest notification window anybody outside the company has put on you. That is the one severity threshold you do not have to invent.
- 4
Log the first classification
From the next incident onward, record what it was called at minute zero as well as what it turned out to be.
Frequently asked questions
Is this free, and what do I get?
Free, with no signup gate on the download. Four documents and five spreadsheets. The AI half is optional: send an incident history and River writes each severity level around the incidents that actually landed in it. More in the template library.
Why not just use the standard four severity levels?
Because critical, major, minor and cosmetic are adjectives, and adjectives are unrunnable when somebody is half awake. A level carrying three dated precedents answers the question a responder is actually asking, which is whether this is as bad as the last bad one.
How many severity levels should we have?
Whatever the response behaviour in your history supports. Sort past incidents by who got woken, how fast somebody acknowledged and whether the status page went up. Those behaviours cluster, and the clusters are your real levels. It is frequently not four.
Where do the acknowledge and resolve times come from?
Externally set clocks wherever they exist, because they are documented somewhere a reader can check and not negotiable at two in the morning. Contractual notification windows, credit tiers in a customer agreement, regulatory reporting boundaries. Where nothing external exists, your own history at the ninetieth percentile.
What makes a threshold actually usable during an incident?
A number rather than a judgement. Regulators write them in exactly the right shape: the communications outage rules put the boundary at thirty minutes affecting at least 900,000 user minutes, notified within 120 minutes of discovery. A product of scale and duration, and a clock starting at discovery.
What is the reclassification rate for?
It is the only real measure of whether the definitions work. Record what each incident was called first as well as finally, then count. Under one in ten and the definitions are runnable. Half means they are not, and the boundary cases sheet says which line is broken.
Do we need the escalation timers automated?
Automate them once the definitions stop moving. A timer firing on a level that gets reclassified half the time just pages people about the wrong thing faster. Get the reclassification rate down first, then wire the intervals into whatever already pages you.
Write levels somebody can run
Take the documents and CSV sheets blank, or install this pack in River and send it twelve months of incident history.
Edit with AI