Skip to content
AlertPing

Playbooks

Incident communication examples, templates and outage communication best practices

| Playbooks | 9 min read

The core rules of incident communication: acknowledge within 10 minutes, even with nothing to report yet. Post updates on a stated cadence and keep every promise about timing. Use the four standard states. Write plainly. Publish a postmortem for anything customers noticed.

Customers forgive downtime; they do not forgive silence. During an outage your status page is being refreshed by your angriest users, your biggest accounts and, if it goes on long enough, journalists. Every one of them is deciding whether you are in control. Here is what to post, when, word for word.

The timing rules

  • Acknowledge within 10 minutes of a customer-visible incident. "We know and we are on it" posted at minute 8 buys more goodwill than a perfect explanation at minute 45. You do not need a cause to acknowledge impact.
  • Name the time of your next update, then hit it. "Next update by 14:30 UTC" turns anxious refreshing into waiting. Missing your own stated time is a second incident.
  • Update at least every 30 minutes while unresolved, even when the update is "still investigating, no change". No news reads as abandonment.
  • Timestamp everything in one timezone, UTC by convention, and keep the full history visible. Deleting updates later destroys the trust the updates bought.

None of this works if you learn about the outage after your customers do. The 10-minute acknowledgement clock starts at impact, not at discovery, which is why detection in seconds via downtime alerts is the prerequisite for every rule above.

The four status page states, with templates

Status pages converged on four states for a reason: they tell customers exactly where you are in the fight. Copy these, replace the bracketed parts, and you will never stare at an empty compose box mid-incident.

1. Investigating

You know something is wrong. You may not know why. Say what customers see, not what you suspect.

template · investigating

Investigating: Since [14:02 UTC] we are seeing [elevated error rates on checkout]. Some customers [may be unable to complete payment]. Browsing and existing orders are not affected. We are investigating and will post an update by [14:30 UTC].

2. Identified

You found the cause. Name it in one plain sentence and scope the impact precisely.

template · identified

Identified: The cause is [a failed database failover in our primary region]. [Card payments] are affected; [browsing and account access] are not. A fix is being applied now. Next update by [15:00 UTC].

3. Monitoring

The fix is in and metrics look healthy. Do not skip this state: it is what separates "we think it is fixed" from "it is fixed".

template · monitoring

Monitoring: A fix was deployed at [15:12 UTC] and [checkout error rates] have returned to normal. We are monitoring closely for the next [30 minutes] before marking this resolved.

4. Resolved

Close with the total duration and, for anything significant, a postmortem commitment with a date.

template · resolved

Resolved: This incident is resolved as of [15:45 UTC]. [Checkout] was degraded for [1 hour 43 minutes]. We are sorry for the disruption. A full postmortem will be published here within [5 business days].

What not to say

  • Do not minimize. "A small subset of users may experience intermittent issues" while checkout is down reads as spin. Customers already know how bad it is; match their reality.
  • Do not blame publicly. "Our cloud provider broke" is still your outage in your customers' eyes. Take ownership on the page; sort out vendors privately.
  • Do not promise fix times you cannot keep. Promise your next update instead. You control the cadence; you do not control the fix.
  • Do not use jargon. "Pod eviction storm in us-east" means nothing to the customer whose invoice run failed. Translate to impact: what stopped working, for whom, since when.
  • Do not go quiet after the first post. An "investigating" entry that sits untouched for two hours is worse than nothing; it proves you post and forget.

Postmortem basics

For any incident customers noticed, publish a short postmortem within 5 business days: a plain-language timeline, impact stated in customer terms (who, what, how long), the root cause, and the specific changes that make a repeat less likely. No employee names, no blame. Detection time deserves a line of its own; if it took 11 minutes to notice, say so and say what changes. Honest postmortems are how B2B buyers judge vendors, and they are the paper trail behind every SLA conversation covered in what 99.99 uptime means in an SLA. If you do not have a format yet, start from our incident postmortem template and report the recovery time using the definition in what MTTR measures.

Outage communication examples: three scenarios worth rehearsing

Templates cover the shape of an update. What trips teams up is the judgment call about which shape to use. These are the three outage communication examples that come up most often, each written the way we would actually post it.

1. Partial degradation, not a full outage

The hardest one, because the instinct is to say nothing while you work out how bad it is. Say something anyway, and be specific about who is affected. Vague sympathy generates more tickets than a narrow, concrete admission.

Investigating. API requests to /v2/reports are returning 504 errors for roughly 15% of accounts since 14:05 ET. Dashboard login, billing and all other endpoints are unaffected. We are investigating and will update by 14:40 ET.

Note the three things that do the work: the exact surface that is broken, the surfaces that are fine, and a time for the next update. A customer who reads that stops guessing and stops writing to support.

2. A total outage where you do not yet know the cause

Do not wait for the cause. The first post exists to prove a human is on it, nothing more. Teams that hold the first update until they can explain the failure routinely take 40 minutes to say anything, which is where the trust damage happens.

Investigating. The application is unavailable for all users from 09:12 ET. We have confirmed the outage from multiple regions and the engineering team is engaged. We do not yet know the cause. Next update at 09:35 ET, whether or not we have one.

"Whether or not we have one" is the most valuable clause in incident communication. It commits you to a rhythm rather than to an answer, so you can keep the promise even when the investigation stalls.

3. The false alarm you already announced

Sooner or later a single-location check fires, you post an incident, and the site was fine the whole time. Retract it plainly and quickly. Quietly deleting the incident is worse, because the people who saw it will assume you hid a real one.

Resolved. The incident posted at 03:14 ET was a monitoring false positive from a single network path, not a service outage. No customer requests failed. We have changed our checks to require confirmation from multiple regions before an incident is opened.

That last sentence is why confirmation matters before an alert ever reaches a human. A check that pages on one failed request from one location will make you publish incidents that did not happen, and each retraction spends a little of the credibility you need for the real one. Requiring several regions to agree is the cheapest fix, and it is the reason our own checks re-run a failure from three regions before anything is sent.

The infrastructure that makes this easy

Everything above gets dramatically easier when the mechanics are automated. A hosted status page that flips to "investigating" the moment a confirmed check fails starts your 10-minute clock at zero, emails subscribed customers each update so support stops answering "is it down?", and keeps the incident history that your postmortems and SLA reports are built on. You write the words; the plumbing should already be running. If you are still choosing a platform, our roundup of status page software compares what each one automates.

alertping

Your status page, updated before customers finish typing the ticket

AlertPing confirms outages from 3 regions in seconds, flips your hosted status page automatically, and emails your subscribers. Templates in, trust kept.

keep reading

More from the blog

· Comparisons

Checkly pricing 2026: how much does Checkly cost per check run, module by module

9 min read

· Comparisons

Grafana Cloud pricing 2026: how much does Grafana Cloud cost, meter by meter

10 min read

· Comparisons

Atlassian Statuspage pricing 2026: how much does Statuspage cost per subscriber, public and private

9 min read

· Comparisons

Opsgenie pricing 2026: how much does Opsgenie cost, and what you pay to replace it

8 min read

· Comparisons

PagerDuty pricing 2026: how much does PagerDuty cost per user, per plan and per year

8 min read

· Comparisons

Pingdom pricing 2026: how much does Pingdom cost per check, per plan and per year

8 min read

· Comparisons

Uptime monitoring software to pair with Datadog, New Relic or Dynatrace

8 min read

· Comparisons

How much does Splunk Observability Cloud cost? Hosts, editions and synthetic monitoring

8 min read

· Comparisons

How much does AppDynamics cost? Editions, cores and synthetic monitoring

8 min read

· Comparisons

How much does Dynatrace cost? Hosts, synthetic monitoring and log ingest

8 min read

· Comparisons

How much does New Relic cost? Users, data ingest and synthetic checks

9 min read

· Comparisons

Datadog synthetic monitoring pricing: what synthetics really cost per test run

9 min read

· Guides

Uptime guarantee vs uptime monitoring: why your host reports 99.9% when your site was down

9 min read

· Guides

Cloudflare uptime monitoring: health checks, origin monitoring, and the blind spots behind the proxy

11 min read

· Comparisons

Status page pricing: what a hosted status page actually costs in 2026

8 min read

· Guides

API monitoring best practices: what to check, how often, and how to keep alerts worth answering

10 min read

· Guides

SSL certificate 200 days: the new validity limit, and the 47-day lifetime coming next

9 min read

· Guides

SSL certificate expired: what happens and how to fix it

8 min read

· Guides

How often should you check website uptime?

7 min read

· SLAs

Error budget: the formula, burn rate alerts, and the policy that makes it work

11 min read

· Playbooks

Runbook template for incident response that gets used

8 min read

· Guides

What causes website downtime, and how to catch each cause

8 min read

· Playbooks

Incident postmortem template that teams actually use

8 min read

· Playbooks

On-call rotation best practices that keep engineers sane

8 min read

· SLAs

MTTR (mean time to recovery): what it is and how to cut it

7 min read

· Guides

Heartbeat monitoring: what it is and how it works

7 min read

· Guides

Status page examples and what the good ones get right

7 min read

· Guides

API uptime SLA: service credit tiers, downtime limits and what a good one costs

8 min read

· Guides

How to create a status page in 6 steps

7 min read

· Guides

Uptime SLA report: what to include, with a worked example

9 min read

· Guides

SLA service credits: what you get back and how to claim it

8 min read

· Guides

Synthetic monitoring vs uptime monitoring: what each one costs and when you need it

8 min read

· Guides

What is a status page?

6 min read

· Guides

Status page vs uptime monitoring: what is the difference?

6 min read

· Guides

What does 99.9% uptime mean?

6 min read

· Guides

What is five nines (99.999%) uptime?

8 min read

· Guides

How to calculate uptime percentage

7 min read

· Guides

SLA vs SLO vs SLI: what is the difference?

7 min read

· Guides

Downtime alerts: how to get notified by email, SMS or phone when your website goes down

7 min read

· Guides

How to monitor an online store for downtime

9 min read

· Guides

Why is my Shopify store unavailable?

8 min read

· Comparisons

Better Stack pricing: how much does Better Stack cost?

8 min read

· Comparisons

UptimeRobot pricing: how much does UptimeRobot cost?

7 min read

· Comparisons

Site24x7 pricing: how much does Site24x7 cost?

8 min read

· Guides

What is a dead man's switch in monitoring?

9 min read

· Guides

Why is my WordPress site down?

9 min read

· Guides

How to monitor WooCommerce uptime and checkout

8 min read

· Guides

How to monitor an API for errors, not just uptime

8 min read

· Economics

How much does website downtime cost?

8 min read

· Guides

How to monitor a cron job

9 min read

· Comparisons

Synthetic monitoring vs real user monitoring

8 min read

· Benchmarks

What is a good uptime percentage?

7 min read

· SLAs

99.99 uptime meaning: SLAs and the real cost of each nine

8 min read

· Guides

How to monitor website uptime

8 min read

Know the second your site goes down

Checks every 30 seconds, confirmed from 3 regions, alerts on every channel. Running in under a minute.

See pricing