The product is down. You know it's down, half your customers know it's down, and the honest truth is you don't yet know why. Every instinct says the same thing: fix first, explain later — heads down, say something when there's something to say. And that instinct, reasonable as it feels, is how a forty-minute outage becomes a trust problem that outlives it by months. Because while you're silent, your customers aren't experiencing an outage. They're experiencing your silence during one — refreshing the page, checking their own configs, wondering if you know, wondering if anyone's there. Silence during an incident isn't neutral. It's data, and customers fill it with the worst plausible story.

The teams that come out of outages with their reputations improved — and that genuinely happens — aren't the ones who fix fastest. They're the ones who communicate on a discipline. Here it is, sized for a team where the person writing the updates may also be the person fixing the bug.

The first message: fifteen minutes, three sentences

Before the cause is known, before the scope is fully mapped, the first message goes out: “We're aware that [the thing] is failing for some customers as of [time]. We're actively investigating. Next update by [time, 30–45 minutes out] — sooner if we know more.” Every clause is doing work. “We're aware” kills the worst version of the story — that it's down AND nobody's home. The timestamp starts the honest clock. And the promised next update is the load-bearing sentence: you've converted an open-ended silence into a bounded wait, which is a different psychological event entirely. Note what the message doesn't contain: a cause you'd be guessing at, an ETA you can't stand behind, or the word “may” doing cowardly work (“some customers may be experiencing” when it's down for everyone reads as spin, and customers grade spin harshly during incidents).

The cadence is a promise — keep it even when it's empty

Whatever interval you named, hit it, including when there's nothing new: “Still investigating — we've ruled out [X]. Next update by [time].” The empty update feels pointless from inside the incident and is anything but: each kept promise, in the middle of your product failing, demonstrates that your word holds under pressure — which is the exact question the outage raised. A missed update, by contrast, reads as either things got worse or they've stopped caring; both are worse than any true status. If you're a team of three and everyone's debugging, the update is still someone's job — thirty seconds on a timer beats a heroic fix delivered into a channel full of customers who already churned emotionally an hour ago.

Saying “we don't know yet” well

The hardest sentence during an incident is the honest one about uncertainty, and there's a professional grammar for it: pair what you don't know with what you're doing to find out, and what you know it isn't. “We haven't identified the root cause yet. We've ruled out data loss — your data is intact. We're currently working through [the area].” If data is affected, or you don't know whether it is, say precisely that — the only unforgivable message in incident communication is the reassurance you later have to retract. A retracted “everything's fine” costs more trust than the outage itself, because it converts an infrastructure failure into a character question.

Resolution isn't the last message

When it's fixed: say fixed, say since when, say what a customer should do if they're still seeing trouble (and route those replies somewhere watched — stragglers are how you learn the fix was partial). Then, within a day or two, the postmortem note — short, plain, unsigned-off-on-by-a-lawyer: what happened, in one paragraph a customer can understand; what it affected and for how long; what you've changed so this class of failure is less likely; and a thank-you without groveling. Skip the vendor-blame even when it's true — customers hired you, not your stack. One good postmortem note does more durable trust-building than a quarter of flawless uptime, because uptime is invisible and character under failure isn't. If the incident cost customers something real, the note is also where the gesture lives — a credit offered before it's demanded is generosity; the same credit extracted by an angry email is a settlement.

Write the templates on a calm day

None of this composes well at 2am with your heart rate up. The whole system is four templates — first message, interim update, resolution, postmortem — written now, blanks for time and scope, living where whoever's on deck can find them. Fifteen quiet minutes buys you competence on the worst day of the quarter; it's the same logic as every checklist: the discipline is cheap, it just has to exist before it's needed.

The bottom line

You don't control when things break; you control every word that happens after. Three sentences inside fifteen minutes, a cadence kept even when empty, uncertainty stated with its edges, a resolution that invites stragglers, and a postmortem in human language — that's the entire discipline, and it's four templates and a timer. Customers forgive outages; every business has them. What they remember is who they turned out to be dealing with while it was down.

— Tom

Templates ready before the timer starts

CSByDesign holds your incident templates, fires the cadence reminders, and routes still-broken replies to a watched queue — so the worst day of the quarter runs on discipline instead of adrenaline.

See how CSByDesign works →

About the author

Tom Christian is the founder of CSByDesign, an AI-native customer support platform built for small teams — and the team of one.

He has spent twenty years inside customer service operations, training, and QA at scale — Guardian Life, ConnectiveRx, and Horizon Blue Cross Blue Shield's Service Division. He writes about running support as a team of one, de-escalation that holds under pressure, the AI-drafts/human-sends line, and the operating discipline of solo CS.