Skip to content
AlertPing

the alertping blog

Uptime and reliability blog: SLAs, monitoring, incidents

Practical writing for the people who answer for uptime. What a good uptime percentage actually is, what each extra nine costs to promise, how to set up a website monitoring tool properly, and what to tell customers while things are broken. Concrete numbers, no filler.

Comparisons

Checkly pricing 2026: how much does Checkly cost per check run, module by module

Checkly sells executions, not monitors. The Starter allowance covers 57% of one endpoint checked every minute, running past it costs up to 62.5% more than prepaying, and one browser journey at 1-minute frequency is $175 a month.

9 min read

Comparisons

Grafana Cloud pricing 2026: how much does Grafana Cloud cost, meter by meter

The plan fee almost never decides the bill. Logs advertised at $0.050 a GB really cost $0.550 once write and retain are added, and one URL checked every minute from three probes is $64.80 a month.

10 min read

Comparisons

Atlassian Statuspage pricing 2026: how much does Statuspage cost per subscriber, public and private

Statuspage tiers are cliffs, not slopes. Subscriber number 251 costs $840 a year, subscriber 1,001 costs $3,600, and a private page for your own staff costs 13.6 times more per person than a public one.

9 min read

Comparisons

Opsgenie pricing 2026: how much does Opsgenie cost, and what you pay to replace it

Atlassian closed Opsgenie to new customers in June 2025 and shuts it off on April 5, 2027, deleting un-migrated data. Here is the rate card existing customers still pay, the worked bills, and what the replacement really costs.

8 min read

Comparisons

PagerDuty pricing 2026: how much does PagerDuty cost per user, per plan and per year

PagerDuty is $21 or $41 per user a month billed annually, free up to five users, and it does not monitor anything. Here is the full rate card, the $1,512 step the sixth user triggers, and what the add-ons really add.

8 min read

Comparisons

Pingdom pricing 2026: how much does Pingdom cost per check, per plan and per year

Pingdom charges exactly $1 per uptime check a month whether you buy ten or a thousand, so the bill scales in a straight line. The number that actually moves across the ladder is the advanced-check allowance, and it moves against you.

8 min read

Comparisons

Uptime monitoring software to pair with Datadog, New Relic or Dynatrace

All three observability platforms stop at a one-minute check interval and two of them meter every execution. Here is what uptime actually costs inside each, and how to split detection from diagnosis.

8 min read

Comparisons

How much does Splunk Observability Cloud cost? Hosts, editions and synthetic monitoring

Splunk Observability Cloud bills a flat rate per host whatever the machine, which makes it cheap on big servers and expensive on containers. Here is every meter, the second APM product hiding on the same quote, and three worked bills against Dynatrace.

8 min read

Comparisons

How much does AppDynamics cost? Editions, cores and synthetic monitoring

Splunk AppDynamics bills per virtual CPU, not per host or per user. Here is every meter, the flat browser synthetic rate that beats Dynatrace sixteen times over, the second Splunk SKU on the same quote, and three worked bills.

8 min read

Comparisons

How much does Dynatrace cost? Hosts, synthetic monitoring and log ingest

Dynatrace sells a consumption rate card, not plans. Here is every meter, the RAM-based trap in Full-Stack pricing, the synthetic request math, and three worked bills.

8 min read

Comparisons

How much does New Relic cost? Users, data ingest and synthetic checks

New Relic bills on three meters at once: platform users, data ingest and synthetic checks. Here is the rate card, the arithmetic, three worked examples, and the free ping monitor exemption most write-ups miss.

9 min read

Comparisons

Datadog synthetic monitoring pricing: what synthetics really cost per test run

Datadog bills synthetics per test run, not per monitor, and the two test types use different denominators. Here is the arithmetic per monitor, and the point where a metered bill stops making sense.

9 min read

Guides

Uptime guarantee vs uptime monitoring: why your host reports 99.9% when your site was down

Your host measures its own network at its own edge. Your monitor measures what a customer gets. Both numbers can be true in the same month, and only one of them is about your website.

9 min read

Guides

Cloudflare uptime monitoring: health checks, origin monitoring, and the blind spots behind the proxy

Cloudflare sits between your users and your server, so it can make a broken site look healthy and a healthy site look broken. The two traps that catch teams, and the setup that closes both.

11 min read

Comparisons

Status page pricing: what a hosted status page actually costs in 2026

Status page vendors meter completely different things: subscribers, pages, seats, or nothing at all. Here is what each one costs at US list prices, and the second subscription most buyers forget to budget for.

8 min read

Guides

API monitoring best practices: what to check, how often, and how to keep alerts worth answering

Most API monitoring is set up once, points at a health endpoint, and proves almost nothing. Here is what experienced teams actually check, how often, and the practices that keep an alert trustworthy enough to wake someone up for.

10 min read

Guides

SSL certificate 200 days: the new validity limit, and the 47-day lifetime coming next

Public TLS certificates now max out at 200 days, down from 398. The ceiling falls again in 2027 and 2029. Here is the schedule, what it does to a manual renewal process, and how to get ahead of it.

9 min read

Guides

SSL certificate expired: what happens and how to fix it

An expired certificate takes minutes to fix and hours to notice, and that ratio is the whole risk. What actually breaks, how to fix it in order, and how shrinking certificate lifetimes change renewals for good.

8 min read

Guides

How often should you check website uptime?

You find out about an outage roughly half a check interval after it starts, so the interval you pick decides how much downtime stays hidden. How to set it per monitor instead of guessing one for everything.

7 min read

SLAs

Error budget: the formula, burn rate alerts, and the policy that makes it work

An error budget turns a reliability target into a number you can spend: how much downtime you are allowed before you have to stop shipping features and fix reliability. Here is how to calculate and use one.

11 min read

Playbooks

Runbook template for incident response that gets used

A runbook is only useful if a tired engineer can follow it under pressure. Here is a template with the sections that matter, an example, and the habits that keep runbooks from going stale.

8 min read

Guides

What causes website downtime, and how to catch each cause

Most outages trace back to a short list of causes: traffic surges, bad deploys, expired certificates, DNS mistakes and failing dependencies. Here is what triggers each, and how to catch it fast.

8 min read

Playbooks

Incident postmortem template that teams actually use

A postmortem is only worth writing if the action items get done. Here is a blameless template with every section that matters, and the habits that turn a writeup into fixes that ship.

8 min read

Playbooks

On-call rotation best practices that keep engineers sane

A good on-call rotation is mostly about protecting people from noise: fewer false pages, clear escalation, honest handoffs, and pay for the disruption. Here is what the sustainable ones do.

8 min read

SLAs

MTTR (mean time to recovery): what it is and how to cut it

MTTR is the clock that runs from the moment something breaks to the moment it is fixed. Here is how to calculate it, how it differs from the other MTT metrics, and the levers that shorten it.

7 min read

Guides

Heartbeat monitoring: what it is and how it works

Heartbeat monitoring flips normal monitoring around: instead of pinging your service, it waits for your job to ping it, and alerts when the expected signal never arrives. Here is how it works and when to use it.

7 min read

Guides

Status page examples and what the good ones get right

The best status pages share a handful of habits: honest component health, fast incident updates, real uptime history, and plain language. Here are the patterns worth copying, with concrete examples.

7 min read

Guides

API uptime SLA: service credit tiers, downtime limits and what a good one costs

An API uptime SLA is only as good as its service credit tiers. The 10%, 25% and 50% payout thresholds, worked examples, and what separates a meaningful SLA from marketing.

8 min read

Guides

How to create a status page in 6 steps

A good status page updates itself from your monitoring instead of being edited by hand in a panic. Here are the six steps to stand one up on your own domain.

7 min read

Guides

Uptime SLA report: what to include, with a worked example

A complete uptime SLA report needs nine things, and the one teams leave out is the measurement method. Here is what belongs in one, with a worked monthly example and the arithmetic behind it.

9 min read

Guides

SLA service credits: what you get back and how to claim it

A service credit is the contractual remedy when a vendor misses its uptime promise: a capped, must-be-requested credit against a future invoice, not automatic cash. Here is how they actually work.

8 min read

Guides

Synthetic monitoring vs uptime monitoring: what each one costs and when you need it

Synthetic monitoring bills per test run, uptime monitoring runs flat. Here is what each actually costs across Datadog, Dynatrace and New Relic, and when you need one, the other, or both.

8 min read

Guides

What is a status page?

A status page tells users whether an outage is on your side, cutting support tickets and building trust. Here is what one shows and who needs it.

6 min read

Guides

Status page vs uptime monitoring: what is the difference?

Uptime monitoring is private detection that alerts your team, while a status page is public communication that tells your customers. See how they differ and why you need both.

6 min read

Guides

What does 99.9% uptime mean?

Three nines means down at most about 8h 46m a year, or 43m 50s a month. Here is the downtime table, whether it is good enough, and how to measure it.

6 min read

Guides

What is five nines (99.999%) uptime?

Five nines sounds like a badge of honor, but it means a downtime budget of 26 seconds a month, and the cost of the last nine is brutal. Here is what 99.999% really means and who genuinely needs it.

8 min read

Guides

How to calculate uptime percentage

Uptime percentage is a simple formula that hides a few traps: which window you measure, what counts as down, and how your check interval limits the precision. Here is the formula and the worked examples.

7 min read

Guides

SLA vs SLO vs SLI: what is the difference?

These three acronyms get used as if they mean the same thing. They do not. The SLI is what you measure, the SLO is what you aim for, the SLA is what you promise. Getting the order right is the whole point.

7 min read

Guides

Downtime alerts: how to get notified by email, SMS or phone when your website goes down

An alert only helps if it reaches a person. Here is how to get notified when your site goes down: external checks, second-region confirmation, and SMS or phone before email.

7 min read

Guides

How to monitor an online store for downtime

An online store can return a perfect 200 while it is functionally down for buyers, so status-code monitoring gives false comfort. Here is how to watch the pages that take money and find out before customers do.

9 min read

Guides

Why is my Shopify store unavailable?

Shopify serves the "this store is unavailable" page with a healthy 200 response, so status-code monitoring reports green while you sell nothing. Here are the 7 causes and how to fix each.

8 min read

Comparisons

Better Stack pricing: how much does Better Stack cost?

Better Stack does not sell tiers, it sells a license per person who takes alerts plus an itemized menu of everything else. That is cheap for one founder and expensive for a rotation of six. The current US rate card, with the arithmetic done.

8 min read

Comparisons

UptimeRobot pricing: how much does UptimeRobot cost?

UptimeRobot restructured its plans in 2026, so most pricing articles now describe a lineup that no longer exists. Here is the current US dollar rate card, plus the SMS credit meter that catches teams out.

7 min read

Comparisons

Site24x7 pricing: how much does Site24x7 cost?

Site24x7 has a genuinely low entry price and a generous free tier, but it is one plan in a sprawling add-on-priced Zoho suite. Here is what the plans actually cost and where the bill grows.

8 min read

Guides

What is a dead man's switch in monitoring?

A job that never runs produces no error, because there is nothing to produce one. Silence looks exactly like success. A dead man's switch is the only thing that turns absence into an alert.

9 min read

Guides

Why is my WordPress site down?

A WordPress site rarely fails with a clean error. It fails with a white screen, a database message or a browser warning while the server still answers. Here are the 9 usual causes and the fix for each.

9 min read

Guides

How to monitor WooCommerce uptime and checkout

A WooCommerce store can return a perfect 200 while the add-to-cart button is dead. Monitoring the store as one URL misses the failure that actually costs sales. Here is how to watch the paths that take money.

8 min read

Guides

How to monitor an API for errors, not just uptime

An API can return 200 and still be broken: an empty list where there should be orders, a null token, a field that quietly changed type. Status-only monitoring misses all of it. Here is how to check the payload.

8 min read

Economics

How much does website downtime cost?

Downtime cost is not one number, it is a formula: revenue lost per minute, plus payroll burn, plus recovery, plus churn. Here is the arithmetic and the costs everyone forgets.

8 min read

Guides

How to monitor a cron job

Cron jobs fail silently, and silence looks exactly like success. Heartbeat monitoring inverts the check: the job reports in when it works, and you get paged when it does not.

9 min read

Comparisons

Synthetic monitoring vs real user monitoring

Synthetic monitoring is active and runs whether or not anyone is on your site. RUM is passive and only sees what real visitors see. Here is which one you buy first, and why.

8 min read

Benchmarks

What is a good uptime percentage?

A good uptime percentage for a production web service is 99.9% or better. Here is the full nines table, what each level costs, and which target fits your product.

7 min read

SLAs

99.99 uptime meaning: SLAs and the real cost of each nine

99.99% uptime allows 52 minutes 36 seconds of downtime a year, 4.38 minutes a month. Here is the downtime-cost math CTOs use to price each extra nine.

8 min read

Guides

How to monitor website uptime

Monitoring website uptime well takes five decisions: what to check, how often, from where, who gets alerted, and what customers see. A practical setup guide.

8 min read

Playbooks

Incident communication examples, templates and outage communication best practices

Customers forgive downtime; they do not forgive silence. The timing rules, copy-paste status page templates and three worked outage communication examples that keep trust through an incident.

9 min read

Stop reading about downtime. Start catching it.

AlertPing checks your sites, APIs and ports every 30 seconds, confirms every outage from 3 regions, and pages you in under 10 seconds.

See pricing