Technical Guide · 5 min read
How to Set Up Uptime Alerts That People Don't Ignore
Starting values for confirmation, failures in a row, routing and escalation, and a 20-minute monthly review that keeps alerts useful.
Published 20 Feb 2025 · Updated 28 Sept 2026
An alert has one job: make the right person act at the right time. Every alert that does not lead to an action teaches people that alerts can wait. After enough of them, a real outage waits too.
This guide gives concrete starting values. They are starting points, not rules. Adjust them after a month of real data.
Rule 1: confirm before you alert
Most false alarms come from a single failed check from a single location. The cause is usually the network between the checker and your service, not your service.
Start with:
- Confirmation from 2 locations. A monitor is down only when at least two locations fail. If one location fails, re-check at once from the others.
- 2 failed checks in a row before an incident opens, for monitors that run every minute. That gives an alert about 2 minutes after a real outage starts, and ignores restarts shorter than a minute.
- 2 good checks in a row before an incident closes. Otherwise a service that flaps between up and down sends a new alert every few minutes.
For checks that run less often, for example every 5 minutes, use 1 failure in a row with confirmation from 2 locations. Waiting for 2 failures would add 5 minutes.
Rule 2: not every problem is an outage
Give each kind of problem its own state and its own route:
| State | Example | Where it goes |
|---|---|---|
| Down | Two locations get a 503 | Page the on-call person |
| Degraded | Response time over your threshold | Team chat channel |
| Blocked | Your firewall answers with a 403 challenge | Email once a day, with allowlisting steps |
| Maintenance | Inside a planned window | Nowhere |
Set the degraded threshold from real data. A good start is about 3 times your usual response time. If your API usually answers in 300 ms, start at 1 second.
Blocked is worth its own state. A firewall that blocks the checker says nothing about whether real users can reach you. Treat it as a configuration task, not an incident.
Rule 3: plan for deploys and maintenance
If deploys cause short outages, either make them zero-downtime or add a maintenance window. Recurring windows (for example every Tuesday 02:00 to 02:30 in your time zone) cover planned work, such as database maintenance, that you know will cause errors.
Do not use maintenance windows to hide flaky services. If a monitor fails every night at the same time and nobody knows why, that is a problem to fix, not a window to add.
Rule 4: jobs and certificates need different timing
Heartbeats (cron jobs, backups, workers): set a grace period, because jobs do not finish at the same second every day. A good start is 10 to 20% of the interval, and at least 5 minutes. For a job that runs every hour, alert if no ping arrives within about 70 minutes. For jobs on a schedule like "weekdays at 02:00", use a cron schedule, not an interval, or you will get an alert every Saturday.
Certificates: alert at 14 days and at 3 days before expiry, and at once if a certificate is invalid. As certificate lifetimes get shorter, also alert when a certificate was not renewed at its usual time. That catches a broken renewal while there are still weeks left.
Rule 5: every alert must be actionable from the message
Compare these two alerts:
Monitor "api" is DOWN.
Checkout API is down from Frankfurt and Virginia since 14:05 UTC.
GET https://api.example.com/health → 503 Service Unavailable
Acknowledge: https://… Incident: https://…
The second one answers what, where, since when, and what the error is, and it has one link to acknowledge. The person on call can decide within a few seconds whether to open a laptop.
A good alert contains:
- the monitor name and target,
- which locations failed,
- the start time with time zone,
- the status code or error,
- a link to acknowledge and a link to the incident.
Rule 6: route by ownership, then escalate
Send each monitor's alerts to the team that can fix it, not to one shared channel that everybody mutes. For anything that pages, add escalation: if nobody acknowledges within 15 minutes, alert a second person or channel. Stop escalation as soon as someone acknowledges.
One channel should get only alerts that need action. Informational messages (degraded, recovered, blocked) go somewhere else.
Rule 7: review once a month
Once a month, spend 20 minutes on these questions:
- How many alerts did we get?
- How many led to an action? Aim for most of them.
- Which monitor sent the most alerts that led to nothing? Change its settings or remove it.
- Did any real outage start without an alert? Add the missing check.
- Does the alert path still work? Pause a test monitor and make sure the alert arrives.
Alert settings drift as a product changes. The review keeps them honest.
Starting settings, in one place
- 2 locations must fail; re-check at once when one fails.
- 2 failures in a row to open, 2 successes in a row to close (1-minute checks).
- Degraded at about 3 times the usual response time.
- Blocked goes to email, not to the pager.
- Heartbeat grace of 10 to 20% of the interval, at least 5 minutes.
- Certificate alerts at 14 and 3 days, and on renewal failure.
- Escalate after 15 minutes without acknowledgement.
- Monthly 20-minute review.
StatusTick has these controls built in: confirmation from several regions, failures in a row, separate Blocked and Degraded states, maintenance windows, heartbeats with cron schedules, escalation, and alerts with an acknowledge link. See avoiding false alarms in the docs, or join the free beta before the Q1 2027 launch.