Industry Insights · 5 min read
Why Uptime Monitoring Matters: The Failures Your Users Find First
Many outages are quiet: the server is up, and users still cannot log in or pay. What these failures look like, and which check catches each one.
Published 15 Jan 2025 · Updated 28 Sept 2026
When people picture an outage, they picture a crashed server. In practice, many outages are quieter. The server is up, the logs look normal, and users still cannot log in or pay. The first person to notice is often a customer.
Uptime monitoring checks your service from the outside, the way a user reaches it. That matters for one simple reason: a broken system cannot always report that it is broken. If DNS points to the wrong place, your application never sees the request, so it cannot log an error.
Below are the failures we see most often, what each one looks like, and the check that catches it.
The certificate expired
What users see: a full-page browser warning. Most people leave.
Why it happens: automatic renewal stopped working weeks ago, and nobody saw the error. A server was replaced and the renewal job did not move with it. A certificate on a subdomain was issued by hand and forgotten.
What catches it: an SSL check that reads the certificate's expiry date and alerts early. This is getting more important. Public TLS certificates now last at most 200 days, and the limit drops to 100 days in March 2027 and to 47 days in March 2029 (CA/Browser Forum ballot SC-081). Renewals become frequent, so a broken renewal shows up fast. The useful alert is "the certificate was not renewed when it should have been", not only "it expires tomorrow".
The domain expired
What users see: nothing loads, or a parking page from the registrar.
Why it happens: the card on file expired, or the renewal emails went to someone who left the company.
What catches it: a domain expiry check that reads the date from the registry (RDAP) and warns 30 days ahead.
DNS points to the wrong place
What users see: the site does not load, or loads an old version.
Why it happens: a record was changed during a migration, a typo in a new record, or a provider change that did not copy every record.
What catches it: an HTTP check notices that the site is unreachable. A DNS check goes further: it compares a record, such as an A or MX record, with the value you expect, so you learn that email routing changed before mail starts bouncing.
The page returns 200, but it is broken
What users see: an error message, an empty page, or a checkout button that does nothing.
Why it happens: many frameworks show an error page with status 200. A CDN serves a cached maintenance page. The HTML loads, but a script or stylesheet it needs returns 404.
What catches it:
- A keyword check: the response must contain text that only a working page has (for example "Add to cart"), or must not contain "Something went wrong".
- A page asset check that loads the scripts, styles and images the page links to.
- A browser check that runs a real journey, such as log in or check out, with Playwright. This is the only check that proves a user can finish a task.
It works from one country, not another
What users see: users in one region report errors; everyone else is fine.
Why it happens: a CDN edge problem, a routing issue at one network, or a firewall rule that blocks a range of addresses.
What catches it: checks from several regions, with each region's result kept, so you can see "down from Frankfurt only" instead of a single yes or no.
The nightly job stopped
What users see: nothing, for days. Then a backup is missing when you need it, or invoices were never sent.
Why it happens: the job crashed, the server running it was replaced, or someone commented out the cron line during a test.
What catches it: a heartbeat monitor. The job calls a unique URL when it finishes. If the call does not arrive on time, you get an alert. It takes one line:
0 2 * * * /usr/local/bin/backup.sh && curl -fsS -m 10 "$HEARTBEAT_URL" > /dev/null
The && matters: the ping is sent only when the backup succeeds.
A service you depend on is down
What users see: payments fail, emails do not arrive, or AI features time out, while your own servers are healthy.
Why it happens: your payment provider, email provider or cloud region has a problem.
What catches it: a check on the vendor's API from your side, plus the vendor's own status page. Knowing within a few minutes that the problem is theirs saves a long search through your own code.
It is slow, not down
What users see: pages that take 8 seconds. Many users leave before the page finishes loading.
Why it happens: a slow database query after a data increase, a cache that stopped working, a noisy neighbour on shared hosting.
What catches it: a response-time threshold on each monitor. A slow answer should be marked as degraded, not as up, and should not be mixed with real outages.
What uptime monitoring does not do
Uptime checks sample your service every 30 seconds to a few minutes. They tell you that something users depend on is broken, and where. They do not tell you why. For that you still need logs, error tracking and metrics from inside your application.
Checks can also miss very short failures that happen between two checks. If a 10-second failure matters to you, a check every 5 minutes is the wrong tool.
A minimal setup for a small product
If you run a small SaaS or shop, these five monitors cover most of the failures above:
- Home page: HTTP check with a keyword that only appears when the page works.
- Login or API health endpoint: HTTP check, with a response-time threshold.
- Checkout or sign-up: a browser check that runs the real flow every 10 to 15 minutes.
- SSL and domain expiry for every domain you own, including the ones only used for email.
- Heartbeats for backups and any job that sends money or email.
Then check the alert path once: pause your test endpoint and confirm the alert reaches the person who should act on it.
StatusTick is built to do exactly this: checks from several regions, keyword, asset and Playwright checks, heartbeats, SSL and domain expiry, and alerts that only fire when a failure is confirmed. It launches in Q1 2027, and you can join the free beta before then.