Website Availability Monitoring: Synthetic, RUM and Alerting
This page explains how website availability monitoring actually works: how uptime is measured and mis-measured, the difference between synthetic checks, real user monitoring, log-based detection and heartbeats, and how to design alerting and escalation so an outage reaches a human who can fix it. It is written for IT managers, digital leads and agency owners responsible for a website that generates revenue or leads.
Availability is a measurement decision before it is a technical one
Most organisations discover their monitoring is wrong during an incident, not before one. The site is down, customers are calling, and the dashboard is a wall of green. The tool did exactly what it was configured to do. It requested the homepage over HTTP, received a 200 response, and recorded the site as available. The fact that the checkout threw a 500, the search index was unreachable, or the CMS had locked every editor out was never part of the measurement.
So the first question is not which tool to buy. It is: what does "available" mean for this website? For a government information site, available might mean the homepage and a set of key service pages render with status 200 and complete content. For an ecommerce site, availability without a working add-to-cart and payment gateway is meaningless. For a lead generation site, a form that silently fails to submit is a total outage of the only function that matters, and every uptime tool on the market will call it 100% available.
Write the definition down before you configure anything. A useful definition names the specific transactions that must work, the response time above which the page is considered failed rather than slow, and the locations users actually come from. Everything else in monitoring follows from that.
How uptime percentages are calculated, and how they are gamed
Uptime is usually expressed as the percentage of measured intervals in a period where the target responded successfully. The numbers sound close together and are not.
| Target | Downtime per month | Downtime per year | Realistic for |
|---|---|---|---|
| 99.0% | 7h 18m | 3d 15h | Single server, manual recovery |
| 99.5% | 3h 39m | 1d 19h | Shared or budget hosting |
| 99.9% | 43m | 8h 45m | Managed single-region, monitored |
| 99.95% | 21m | 4h 22m | Redundant application tier |
| 99.99% | 4m 22s | 52m | Multi-AZ, automated failover |
The percentage only means something once you know what was measured and what was excluded. Common ways a reported figure flatters reality:
- Infrastructure uptime, not website uptime. The virtual machine responded to ICMP the whole time. PHP-FPM was refusing connections for forty minutes. The host reports 100%.
- Five-minute check intervals. A four-minute outage between checks did not happen as far as the record is concerned. At one-minute intervals it does.
- Excluded maintenance windows. Planned downtime is often carved out of the calculation entirely, which is defensible if the window is agreed in advance and unreasonable if it is declared retrospectively.
- Homepage-only checks. The one page most likely to be served from cache or a CDN edge, and therefore the last thing to fail.
- Status code only. A WordPress white screen of death, a Silverstripe error template or a "we'll be back soon" page can all return 200.
Insist that the SLA defines the check target, the interval, the number of consecutive failures required to declare an outage, and how the clock starts. A target without those four details is not measurable, which is usually the point.
Synthetic, real user, log-based and heartbeat monitoring
These four approaches detect different failures. Serious availability monitoring uses at least three of them, because each one is blind to something the others catch.
| Method | How it works | Detects | Blind to | Typical detection lag |
|---|---|---|---|---|
| Synthetic checks | Scheduled requests from external agents, HTTP or headless browser | Total outages, TLS failures, DNS failures, slow responses, content changes | Faults on paths not being checked, issues affecting only logged-in users | Interval plus confirmation, typically 1 to 3 minutes |
| Real user monitoring (RUM) | JavaScript beacon in the page reporting timings and errors from actual browsers | Regional degradation, browser-specific breakage, JS errors, Core Web Vitals regressions | Complete outages, because a down page loads no beacon | Minutes, and only during traffic |
| Log and metric based | Alerting on error rates, response times and queue depth from server or CDN logs | Rising 5xx rates, partial failures, one broken app node behind a load balancer | Anything upstream of your logging, including DNS and network faults | Near real time |
| Heartbeat / dead man's switch | A job pings an external endpoint on completion; absence of the ping alerts | Failed cron, stalled queue workers, backups that stopped running | Web-facing availability | Configured grace period |
RUM deserves special mention because it is frequently misunderstood as an uptime tool. It is not one. When the site is truly down, RUM goes quiet, and a quiet dashboard looks like a quiet night. RUM's value is in telling you that Perth users are seeing four second page loads while Sydney users are fine, which is a class of failure synthetic checks from a single location will never surface.
Check frequency, geography and false positive suppression
Check frequency sets the floor on how fast you can possibly know. A 60 second interval on critical paths and 5 minutes on secondary paths is a reasonable starting shape. Going below 30 seconds rarely helps, because confirmation logic eats the gain.
Run checks from at least three geographically separate locations, including at least one inside Australia and one offshore. Two reasons. First, a single vantage point cannot distinguish "the site is down" from "the route between this agent and your origin is down", and transit incidents are common. Second, if you serve Australian users through a CDN, you want to know when an Australian edge is misbehaving while the origin is perfectly healthy. Cloudflare and similar edge networks add a layer that can fail independently of your infrastructure, and monitoring the origin alone will not see it.
False positives are the real enemy, because an on-call engineer who has been woken three times by a flapping check will start ignoring the fourth. Practical suppression:
- Require two or three consecutive failures from two different locations before declaring an outage.
- Retry immediately from a second region rather than waiting for the next scheduled interval.
- Set realistic timeouts. A 5 second timeout on a page that legitimately takes 4 seconds under load generates noise, not signal.
- Exclude monitoring agents from rate limiting and bot protection rules, or your WAF will manufacture outages.
- Alert on error rate as a proportion of traffic rather than absolute counts, so a quiet Sunday and a busy Monday behave the same.
Multi-step journey checks
A single GET request tells you the web server is alive. It tells you almost nothing about whether the business function works. Scripted browser checks that walk a real journey are where availability monitoring becomes useful, because they exercise the database, session handling, third party APIs and payment gateways in the same sequence a customer does.
Journeys worth scripting, in rough order of value:
- Authenticated login. Catches session store failures, expired API credentials and SSO or IdP outages that a public page check cannot see.
- Search. Exercises Elasticsearch, Solr or the database index. Search failing silently and returning zero results is a classic invisible outage.
- Form submission. Submit a test enquiry through the real form with a tagged identifier, and assert both the confirmation page and, where possible, the record landing in the CRM.
- Add to cart and checkout to payment step. Stop short of capturing funds, but go far enough to confirm the gateway handshake succeeds.
- CMS admin login. Editors locked out of the CMS is an outage for the people who pay for the site, even if the public front end is perfect.
Scripted journeys need maintenance. A redesign changes a selector, the check breaks, and nobody notices for a month. Treat check scripts as code, version them, and review them whenever the application changes. This is one of the reasons monitoring tends to decay when it is owned by a tool vendor rather than by the team that also maintains the application.
Certificate, DNS, cron and queue monitoring nobody configures
The outages that cause the most damage are usually not server crashes. They are silent expiries and stopped background work.
- TLS certificate expiry. Alert at 30, 14 and 7 days, and monitor the full chain, not just the leaf. An expired intermediate breaks older clients while modern browsers happily rebuild the path and show green.
- Domain registration expiry. Rarer, far worse, and entirely preventable with a calendar alert and a second contact on the registrar account.
- DNS resolution and record drift. Check that the A, AAAA, CNAME and MX records resolve to expected values from multiple resolvers. Also monitor SPF, DKIM and DMARC, because transactional email failing is an availability problem for anything involving password resets or order confirmations.
- Scheduled tasks. Every cron job that matters should report completion to a heartbeat endpoint. Sitemap generation, index rebuilds, feed imports and, most importantly, backups.
- Queue depth and worker health. A queue that grows monotonically is an outage in slow motion. Alert on depth and on oldest-message age, not just on worker process count.
- Disk, inode and database connection saturation. These have a predictable trajectory, so they should be warnings hours before they become incidents.
- Backup restorability. A backup job that completes is not a backup. Periodic test restores into a scratch environment are the only check that means anything.
Alert design: thresholds, deduplication and escalation
An alert is a request for a human to do something. If nothing needs doing, it should not be an alert. Apply that test ruthlessly and most monitoring noise disappears.
Severity tiers
Three tiers are usually enough. Critical means the site or a core transaction is unavailable and someone is woken regardless of the hour. High means degraded but functional, actioned within business hours or the next morning. Informational goes to a channel, never to a phone. Certificate expiring in 30 days is informational. Certificate expiring in 48 hours is critical.
Deduplication and correlation
When an origin server fails, a naive setup fires separate alerts for the homepage check, five journey checks, the error rate threshold and three heartbeats, all within ninety seconds. Alerts should be grouped into a single incident with a single owner. Suppress downstream alerts when an upstream dependency is already known to be failing, and suppress everything inside an agreed maintenance window.
Escalation ladders and on-call
The escalation ladder is what turns a notification into a response. A workable pattern: notify the primary on-call engineer by push and SMS; if unacknowledged after 5 minutes, call them; after 10 minutes, escalate to the secondary; after 20 minutes, escalate to the technical lead and notify the client contact. Acknowledgement must be an explicit action, not an email being opened.
Ladders only work if the roster behind them is real. That means named people, a published rotation, handover between them, and cover for leave and public holidays. A three person in-house team cannot sustain genuine 24/7 on-call without burning out, which is the quiet reason so many organisations have excellent monitoring and terrible response times. If you are assessing your current arrangement, the questions in our guide to evaluating a managed web services provider are a useful place to start.
What happens between the alert firing and service being restored
This is the gap that monitoring products cannot close, and it is where the real cost of an outage sits. Detection might take ninety seconds. Restoration takes as long as it takes someone with credentials, context and the authority to change production to work out what broke and fix it.
A realistic 3am sequence looks like this. The alert fires and is acknowledged. The responder confirms the failure independently, because acting on a false positive wastes the first ten minutes. They check the obvious external causes: DNS, certificate, CDN status, upstream provider status. They check recent change: was there a deployment, a plugin update, a config push, a DNS edit? They look at application logs and the database for connection exhaustion, slow query pile-up, disk full, or a failed migration. Then they either roll back, restart the affected service, scale, or patch forward. Finally they confirm recovery through the same checks that detected the fault, and communicate.
Every step in that sequence depends on things a monitoring tool does not have: SSH access, deployment history, knowledge of the application's architecture, permission to roll back, and the ability to write code if the fix requires it. This is the split that leaves so many organisations stranded. The hosting provider will confirm the server is up and decline to look at the application. The agency that built the site is asleep, has no production access, and has no on-call obligation. The alert did its job. Nobody did anything with it. Our guide to what to do in the first hour of an outage covers the triage sequence in more detail, and if your provider missed its response commitment, what to do after an SLA breach sets out the remedies worth pursuing.
The honest way to evaluate a monitoring arrangement is to ask what happens at 3am on a Sunday of a long weekend. If the answer is that an email arrives in a shared inbox, you do not have monitoring. You have a record of when the outage began.
Frequently Asked Questions
What is website availability monitoring?
Website availability monitoring is the continuous, automated verification that a website and its critical functions respond correctly from the locations its users are in. It combines synthetic checks that request pages and run scripted journeys, real user monitoring from actual browsers, log and error rate alerting, and heartbeat checks on background jobs. Monitoring that only requests the homepage and checks for a 200 response will miss most real-world failures.
How often should a website be checked for uptime?
Critical pages and transactions should be checked every 60 seconds from at least three geographic locations, with secondary pages at 5 minute intervals. Intervals longer than 5 minutes mean short outages go unrecorded entirely, which inflates reported uptime. Pair frequent checks with a confirmation rule requiring two or three consecutive failures from different locations so transient network faults do not trigger false alarms.
What is the difference between synthetic monitoring and real user monitoring?
Synthetic monitoring sends scheduled requests from external agents and works whether or not anyone is visiting the site, which makes it the only reliable way to detect complete outages. Real user monitoring collects performance and error data from a script running in visitors' browsers, so it shows how the site actually performs by region, device and connection. RUM goes silent during a full outage, so it complements synthetic checks rather than replacing them.
What uptime percentage should an enterprise website SLA guarantee?
99.9% is a reasonable commitment for a well managed single-region deployment, allowing roughly 43 minutes of unplanned downtime per month. 99.99% requires redundancy across availability zones and automated failover, and costs considerably more to run. The percentage matters less than the definition behind it: what is measured, at what interval, how many consecutive failures declare an outage, and whether maintenance windows are excluded.
Why did my monitoring tool say the site was up when it was down?
The most common causes are checking only the homepage, which is often cached at the CDN edge and stays available after the origin fails, and checking only the HTTP status code, since error pages and blank white screens frequently return 200. Checks that ping the server rather than requesting a page will also report success while the web application is refusing connections. Fixing this requires content assertions, checks on multiple paths, and scripted journeys through authenticated and transactional flows.
Can monitoring alone prevent downtime?
No. Monitoring detects failures; it cannot diagnose or repair them. Restoring service requires someone with production access, knowledge of the application, authority to roll back or patch, and an obligation to respond outside business hours. Monitoring delivers value only when it is attached to a rostered on-call team with the technical capability to act on what it reports.
Managed uptime and availability monitoring
UnDigital runs monitoring and the response behind it as a single service. That means synthetic checks and scripted journeys tuned to your actual critical transactions, certificate, DNS, cron and queue monitoring, alert routing with deduplication and escalation, and a rostered team with production access who investigate and fix rather than forward the alert to you. It applies whether we host your infrastructure or you bring your own AWS, Azure or existing environment. You can read more about our website uptime monitoring service and the support SLAs for enterprise and government that define response and restoration commitments, or start with an infrastructure audit that tells you exactly what your current monitoring is and is not seeing.
Find out what your monitoring is missing
Reviews from our client partners.
"Thanks so much for your comprehensive strategy and execution of our digital ecosystem.
I can finally sleep at night knowing that everything is under control, secure and scalable.
Thank you!!!".
Corporate Marketing Manager, Sekisui House
"Thanks for all your help. This project was in such good hands from the beginning. We really appreciate all your hard work and expertise!!"