Customer Incident Communication Template
Four documents and two sheets that compute exactly which accounts an incident actually affects, then keep every one of them on a schedule.
Free download · No account needed
Affected Account Computation
Northweave Analytics, 640 active accounts
| Segment | Accounts | Scope |
|---|---|---|
| Shopify, EU-hosted | 47 | Affected |
| Shopify, US-hosted | 355 | Not affected |
| Other platform, EU-hosted | 140 | Not affected |
| Other platform, US-hosted | 98 | Not affected |
The 47 affected, split by symptom rather than left as one group
| Delayed — recovers automatically once the backlog drains | 34 |
| Failing — needs an active data re-pull, no automatic recovery | 13 |
The other 593 accounts, on a different platform's scheduled sync or hosted outside the EU, received no notice at all, because none of them matched the incident's actual scope.
Every incident communication template on this query follows the same fill-in-the-blank shape: state what is investigated, what customers may see, and a time for the next update. That shape is correct, and every generic guide already covers it. None of them says how to fill in the second blank, what customers may see, for one specific account rather than a hypothetical average one. The gap is not wording: most incidents expose some accounts and not others, on conditions the incident record and the account roster already hold separately, integration, region, plan tier.
This pack computes that join before drafting anything. The incident's technical scope crosses against the account roster's attributes to produce a named affected list, and just as important, a named excluded list with the specific reason each account was left off. When the incident produces more than one symptom, the affected list splits again, because a group that is merely delayed and a group that is failing outright need different information, not the same update with different names filled in.
Northweave Analytics, an order-analytics platform for e-commerce brands, had its EU order-sync consumer fall behind after a database migration held a write lock longer than planned. Of 640 active accounts, only 47 matched the incident's actual scope: hosted in the EU and connected through Shopify's webhook-based sync rather than another platform's polling. The other 593 received no notice, because none of them were exposed. Of the 47, 34 recovered automatically once the backlog drained and 13 needed an active data re-pull, because their oldest unconsumed events aged past the sync queue's retention window first.
What's in the pack
Holding Statement
The single, honest statement sent the moment an incident is declared, before the affected list is even computed. States what is being investigated and what is currently unconfirmed, and claims no scope the cross-reference has not confirmed yet.
Update Sequence
One message per symptom group rather than one message for the whole incident, each stating what changed since the last send and a specific next-update time. That time becomes a standing commitment: the next message is due whether or not anything new has happened.
Resolution Note
Sent to each group once its own fix is confirmed holding, which is not the same clock time for every group. A group that recovered automatically and a group that needed a manual data re-pull get told two different, both true, things.
Post-incident Letter
Sent days later, to the affected accounts only, never the full customer base. States the confirmed duration and, for any account whose data needed correction, the specific record of what was restored rather than a general assurance.
Affected Account List
The join between the incident's technical scope and the account roster, kept as a full register rather than a filter that discards its own near-misses. Every excluded account carries the specific condition it failed, integration or region, so the computation is checkable rather than asserted.
Notification Log
Proof that every included account actually had a usable contact channel and was actually reached, with a named fallback for the ones that did not. Complements the cadence tracking in the incident status page update template, which measures a public page's own promised-versus-posted gap rather than per-account delivery.
Why a fill-in-the-blank template is not enough on its own
Every generic incident template already covers the wording: investigating, impact, next update. None of them says how to compute the second blank, what customers may see, for one specific account instead of a hypothetical average one. That computation, not the wording, is what this pack adds, and applying it consistently across every category a business tracks is the same discipline the ticket categorization template uses for support volume.
How it works
- 1
Send the incident record and the account roster
Whatever exists right now, even a rough monitoring alert, plus the roster your system already holds with each account's integration, region and a live contact channel. The Holding Statement can go out before either is complete.
- 2
Compute the affected list and split it by symptom
Cross-reference the incident's stated scope against the roster's attributes, keeping every near-miss account on record with the reason it was excluded. If the incident describes more than one symptom, tag each affected account with which one applies before drafting anything.
- 3
Write the sequence and schedule the cadence
One Update Sequence message per symptom group, each naming a specific next-update time that becomes a standing reminder. A message with nothing new still goes out and says so, because the time was already promised.
- 4
Close it out and log who was actually reached
A Resolution Note per group once that group's own fix holds, then a Post-incident Letter to the affected accounts only. The Notification Log records whether every send actually landed, routing the accounts with no usable contact through a named fallback, while the ticket volume the incident adds to the queue gets measured separately in SLA queue performance reporting.
Frequently asked questions
Why not just post one status update for everyone?
Because it answers nobody's real question. An unaffected account reads it and cannot tell it is unaffected, so it opens a ticket asking. An affected account reads the same sentence and cannot tell what it specifically is experiencing, so it opens a different ticket asking that instead. Both land on the queue already handling the incident.
How do you decide which accounts are actually affected?
By crossing the incident record's stated technical scope, which service, which region, which integration, against the account roster's own attributes for each account. An account is included only if it matches every condition the incident actually states, not just one of them, and the ones that almost match stay on record with the specific reason they were excluded.
What if the incident affects some accounts differently than others?
Split the affected list by symptom before writing anything. On the worked example, 34 accounts recovered automatically once a backlog drained and 13 needed an active data re-pull because their events aged out of the sync queue's retention window first. Those two groups received different messages from the first send, not a shared one with a detail changed later.
Does every incident need a scheduled update cadence?
Yes, for any incident that runs long enough to have a quiet stretch. PagerDuty's own external-communication guidance states that updates should go out on a set interval and again whenever impact meaningfully changes, regardless of whether anything new has happened. A cadence that only fires when there is news goes silent exactly during the stretch nobody is watching the clock.
How is this different from a public status page?
A status page is one message for everyone watching it, and measuring whether its own updates kept their promised cadence is a different, separate mechanism. This pack's Notification Log instead checks whether each individually affected account had a usable contact channel and was actually reached, which a public page cannot do regardless of how well it keeps its own schedule.
What if an affected account has no working contact on file?
Route it through a named fallback, an account manager or a direct call, and log that routing rather than marking the account notified by default. On the worked example three of 47 affected accounts had no usable contact, and the one send that landed past its promised time happened on exactly that fallback path, not a direct one.
Even PagerDuty's own guidance mentions individual outreach. What does that leave out?
It states that accounts on certain support tiers should also receive communication delivered individually, but leaves "which accounts" as a lookup on contract tier alone. It does not compute exposure from technical scope, integration, region, plan tier, crossed against the roster, and it does not address a single incident producing more than one symptom. This pack does both before a single message goes out.
Tell the accounts an incident actually affects, not everyone
Send the incident record as it exists right now and the account roster your system already holds. River computes the affected list, splits it by symptom, and drafts the first message on a schedule.
Edit with AI