River
Y CombinatorBacked by Y Combinator
FREE TEMPLATE

Incident Status Page Update Template

Six documents and four sheets that record what you promised beside what you posted, so the silence between updates becomes a number.

Free download  ·  No account needed

Every status page template solves wording. There is no shortage of advice about avoiding jargon, not blaming a vendor, and never committing to a fix time, and all of it is correct. None of it is what went wrong during your last incident. What went wrong is that an update went out at 15:33 and the next one went out at 16:44. In between, a customer refreshing the page had no way to tell a team mid-failover from a team that had gone home.

Seventy-one minutes of silence is the event people remember, and it is not a wording problem. Nobody remembers to post while firefighting, which is exactly when posting matters most and exactly when nobody has attention to spare. So this pack treats the next update as a dated obligation instead of a good habit. Every update states the time of the next one, that time is the commitment, and the Notification Log records what was promised beside what actually arrived.

That log carries four columns and only one of them exists anywhere today. Posted at, which a status page history already gives you. Promised by, carried forward from the previous update, which nothing records, which is why nobody can say afterwards whether the cadence held. Minutes late, the subtraction. And minutes since the previous post, because three minutes late is invisible and forty-one minutes late is a seventy-one minute silence.

The gap a status page history cannot show you

Minutes late and minutes since the previous update are different measurements. Only the second one is what the customer experienced.

Notification Log

INC-2291, Sev1, card authorisation failures on the EU acquiring route. A fictional payments API, Halvard Pay. Declared 14:18, resolved 18:02.

UpdateStatePromised byPosted atMins lateGap
U1Investigating14:3314:31013
U2Investigating15:0115:04333
U3Identified15:3415:33029
U4Identified16:0316:444171
U5Monitoring17:1417:10026
U6Resolved18:1018:02052

Six updates posted, four on time, 66.7% compliance, 44 minutes late in total, worst gap 71 minutes against a 30-minute commitment. The first update landed 13 minutes after declaration, inside the 15-minute ceiling, because that is the part everybody gets right while the incident channel is still loud.

Nothing was wrong with the wording of U3 or U4. The failover started at 15:40 and both people on the comms rota were in it. The obligation existed, the timestamp existed, and nobody was left holding the clock.

Cadence Policy by Severity

Filled in calm conditions. This is the sheet that makes a 15-minute first update possible.

SevTriggerFirst update withinThen everyWhile monitoringPublic page
Sev1Customers cannot complete a core action15 min30 min60 minYes
Sev2A core action is degraded or a secondary one is failing30 min60 min60 minYes
Sev3Customer-visible with a workaround4 hoursDailyDailyPast one working day

Three rules live on the sheet rather than in a document nobody opens. Every update states the time of the next one, and that time is the commitment. A next-update time is never allowed to be a fix time. An update with nothing new still goes out, saying what is being checked.

Status Page Component Map

Public componentInternal servicesWhat a customer cannot doDefault sev
Card authorisationauth-gateway, acquiring-router, pool-managerPayments cannot be authorisedSev1
Payoutspayout-scheduler, bank-connectorMoney does not arrive when promisedSev1
Webhooksevent-dispatcher, retry-workerMerchant systems fall out of syncSev2
Sandboxsandbox-gatewayIntegration work is blocked, no live moneySev3, not shown publicly

The third column is the one that exists nowhere and the one the first update is written from. A 40-minute first update is almost never a writing problem; it is a responder deciding what to call the thing while the clock runs.

Affected Customer Register

One row per segment, not per account, because segments are what contracts follow.

SegmentAccountsNotification clauseWindowSatisfied by the status page
EU acquiring, standard terms341Nonen/aYes, nothing owed
EU acquiring, enterprise MSA38Sev1 notice within 4 hours240 minYes, U1 at 14:31
EU acquiring, regulated PSP resellers6Named contact within 60 minutes60 minNo
EU acquiring, sandbox and test keys27Nonen/aNo live traffic in the window

412 accounts on the affected route, 385 with live traffic. 44 of those, 11.4%, carry a contractual notification clause and 38 of them are met by the public update. Six, 1.6%, need a named human inside an hour and cannot be satisfied by a public page at all.

Six accounts is a small enough number to miss and an expensive enough number to have missed. They were reached at 15:02, 44 minutes after declaration, because the register had already named the contact and the window. Nobody read a contract during the incident. That outreach happens outside this space; the register's job was to find the six and route them.

What you get

01

Notification Log

What you promised, what you posted, how late it was, and how long the actual gap ran. The promise column exists nowhere else, which is why nobody can currently say whether their cadence held during the last incident.

02

Cadence Policy by Severity

First-update ceiling, cadence, and whether it slows during monitoring, per severity. Includes who owns the updates as a role and whether that person may also hold a technical job in the same incident.

03

Status Page Component Map

Public components mapped to internal services, with what a customer cannot do when each is degraded, in the customer's own words. That third column is what makes a fifteen-minute first update possible.

04

Affected Customer Register

One row per segment with its notification clause, the shortest window in minutes, and whether a public update satisfies it. Finds the small expensive group that needs a named human inside an hour.

05

Status Update Sequence by Severity

Four states, three cadences, and the fixed shape of a first update that can be written without knowing the cause. Plus what never goes in one, which is a shorter list than people expect.

06

Internal Brief

For the people who answer customers directly, on a clock measured in hours. What a customer will have seen, pasteable sentences they may use, specific prohibitions, and the compensation question answered before somebody improvises it.

07

Resolution Note

The corrected impact window, what a customer needs to do, and your own cadence performance including the worst gap. Publishing your own longest silence is unusual and it is the most credibility-building sentence available.

08

How the Clock Works

The reference behind the log: why lateness and gap are different numbers, why the three artifacts run on three clocks, and why an update with nothing new still goes out.

How it works

  1. 1

    Send a status page history

    The timeline of your last two or three incidents with the times each update went out. That is enough on its own, because the gaps are computable from post times alone and nothing else is needed to start.

  2. 2

    Get the silence measured

    River backfills the log, derives what each update should have promised from the cadence, and reports four numbers per incident: posted, on time, total minutes late, and worst gap.

  3. 3

    Fill in the calm-conditions sheets

    The cadence policy and the component map cannot be written during an incident. River builds both, then tests the proposed cadence against your real timelines to see how many updates it would have demanded.

  4. 4

    Find the contractual clocks

    Which segments have a notification clause, which windows a public page satisfies, and which need a named human inside an hour. Nobody reads a contract mid-incident, so a clause found afterwards was breached.

Frequently asked questions

Why measure the gap instead of improving the wording?

Because wording is not what failed. Advice about jargon and fix times is universal and correct, and the incident people remember is the one where the page went quiet for over an hour. An update ending "shortly" has committed to nothing and can therefore never be late.

What goes in an update when we know nothing yet?

The public component, what a customer cannot do, the scope hedged honestly, one sentence on what is not affected, and the next update time. Every clause there is defensible fifteen minutes in, and the not-affected sentence is usually the most valuable one on the page.

Isn't posting with no news just noise?

No. Silence is not more honest than repetition, it is less informative. "Still failing over the connection pool, impact unchanged, next update by 17:14" tells a customer three useful things. This is the rule teams resist and it is the one that prevents the long gap.

Where does the three-clock structure come from?

It is the shape regulators already use. For covered communications outages, 47 CFR 4.9 requires a Notification within 120 minutes, an Initial Report within 72 hours, and a Final Report within 30 days. Minutes, hours, days, and writing one never satisfies another.

Do we really need a dedicated comms person?

The gap in the worked example happened because both people on the rota were inside the failover. Google's SRE book describes the role as the public face of the response, whose duties definitely include issuing periodic updates. It is a role, not a spare-time task.

Does this handle emailing affected customers individually?

No, deliberately. This space owns the public status page and the update rhythm. The register finds the accounts whose contractual window a public page cannot satisfy and routes them out with the contact and owner named, rather than implying an obligation was discharged by publishing.

How does this fit with the rest of our incident work?

It runs during the incident. What a responder is handed when the alert fires is the on-call handbook pack, reconstructing what happened afterwards is the timeline tool, and the monthly roll-up is the reliability review.

Measure the silence in your last incident

Send a status page history with the times each update went out. River reports what you promised, what arrived, and how long the longest gap actually ran.

Set up my updates