Skip to content

Business Strategy · 4 min read

How to Choose an Uptime Monitoring Tool: 12 Questions to Ask

Feature lists look the same. These twelve questions, and a one-week trial test, show how a monitoring tool behaves when something really breaks.

Published 10 Feb 2025 · Updated 28 Sept 2026

Almost every uptime monitor lists HTTP checks, email alerts and a status page. The differences show up later: at night, during a deploy, when your firewall blocks the checker, or when your team doubles and the bill does too.

These are the questions we would ask any tool, including ours. Each one comes with a way to test the answer during a trial.

Alerts you can trust

1. How many locations must fail before I get an alert?

A single check from a single location will fail now and then because of the network between the checker and your site, not your site. If the tool alerts on one failed check from one place, you will get false alarms, and your team will learn to ignore them.

Look for: confirmation from at least two locations before a monitor is marked down, and an immediate re-check when one location fails.

2. Can I require several failures in a row?

A deploy that restarts your app for 20 seconds should not page anyone. A setting like "open an incident after 2 failed checks in a row" handles this. Ask what happens on recovery too: does it close after one good check, or after a number you choose?

3. What happens when my firewall blocks the checker?

Bot protection and web application firewalls often answer monitoring requests with a 403 or a challenge page. A simple tool reports that as "down". Ask whether the tool can tell "blocked" apart from "down", and whether it publishes its IP addresses so you can allowlist them.

Test it: add a firewall rule that blocks one checker IP, and see what the tool reports.

4. Is "slow" different from "down"?

A page that takes 9 seconds is a problem, but it is a different problem from an outage, and it usually needs a different response. Look for a per-monitor response-time threshold and a separate "degraded" state.

What it can check

5. Can it check what the page says, not only the status code?

Many broken pages return status 200. You need keyword checks (must contain, must not contain), and for important flows, a real browser check that logs in or checks out. Ask whether browser checks use a standard framework such as Playwright, so you can reuse tests you already have.

6. Does it cover jobs, certificates and domains?

Cron jobs and backups need heartbeat monitors (the job calls a URL; silence means trouble), ideally with a cron schedule, not only a fixed interval. SSL certificate and domain expiry checks should be included, not an add-on.

7. Can it reach services that are not on the internet?

Internal APIs, databases and admin tools fail too. If you need to watch them, ask whether the tool offers an agent that runs inside your network and connects out, so you do not have to open inbound ports.

Working with it every day

8. Where do alerts go, and what happens when delivery fails?

Check that your channels are supported (email, Slack, Teams, a webhook, a paging tool). Then ask two less common questions: are failed deliveries retried, and is there a log of every delivery attempt? When someone says "I never got the alert", you want to be able to check.

9. What happens if nobody reacts?

Escalation (alert a second person or channel after a delay) and acknowledging from the alert itself make a big difference at night. Ask whether they are in every plan.

10. Is the status page included, and does it update itself?

Some tools charge separately for status pages, custom domains or subscribers. Ask what is included, and whether status page components follow your monitors or have to be updated by hand during an incident.

Cost and data

11. How does the price grow?

Prices are usually based on one or more of: number of monitors, check interval, number of users, add-ons (status pages, SMS, extra locations). Write down your setup today and at three times the size, and price both. Per-user pricing in particular can grow faster than you expect as more people need access during incidents.

12. How long is data kept, and can I take it with me?

Ask how long raw check results are kept, how long uptime history is kept, and whether there is an API and an export. If you leave, you want your monitors and your history.

A one-week trial test

Features on a pricing page do not tell you how a tool behaves. Run these five tests on a staging service during the trial:

  1. Kill it. Stop the service for 10 minutes. How long until the alert? Does it arrive where it should, once?
  2. Blip it. Restart the service so it is down for about 20 seconds. Did you get an alert? With good settings, you should not.
  3. Block it. Block one checker IP at your firewall. Is that shown as down, or as blocked?
  4. Slow it. Add a 5-second delay to one endpoint. Is it shown as slow, or as up?
  5. Silence a job. Stop a job that sends heartbeats. How long until you hear about it?

Write down the results. The tool that passes these tests is the one that will not wake you for nothing, and will wake you when it counts.

StatusTick is being built around the first four questions: confirmation from several regions, failures in a row, and separate Blocked and Degraded states. It launches in Q1 2027; you can join the free beta and run these tests against it yourself.

Closed beta

Know before your users do.

Find problems before your customers do, and ship with more confidence. Launching Q1 2027: join the closed beta and help decide what we build next.

We only use your email for StatusTick beta news; email us to be removed at any time.