What the number means
Uptime is the share of time a service is available. The nines are shorthand. Two nines, 99%, allows more than three and a half days of downtime a year. Three nines allows under nine hours. Four nines, under an hour. Five nines — 99.999% — allows about five minutes and fifteen seconds across an entire year.
That is a demanding standard. One clumsy restart, one expired certificate, one botched configuration reload, and the year's allowance is gone. That is why the measurement matters as much as the figure.
How it is measured
- External probes every five minutes. Uptime Kuma checks every site and service on a five-minute cycle. It looks for real text on the page, not just an HTTP 200, because a server cheerfully returning an error page is still down as far as a visitor is concerned.
- A daily scripted health check. Dozens of checks run across the estate each day, again testing content rather than status codes: does the login page render, does the mail front end answer, does each site serve the site it should.
- A daily certificate sweep. Every TLS certificate's expiry date is checked daily. Renewal is automatic, but the sweep exists because automation can fail without saying so.
- Alerts to a human. When a probe fails, a person is told on Slack. Monitoring that only writes to a dashboard nobody is watching measures outages; it does not shorten them.
What the measurement can and cannot see
It is worth being straight about resolution. A probe every five minutes can miss a blip shorter than that. Five nines measured this way means no probe found the service unavailable beyond that allowance — not that every second of the year was observed. That is how most uptime figures are produced, whether people say so or not, and I would rather you knew.
It also measures availability, not speed or correctness in depth. A page that loads but shows yesterday's data passes an uptime check. The daily scripted checks exist partly to catch that kind of fault, and they go further than a probe can.
How it is kept that way
High uptime on one server is less about heroics and more about not causing your own outages. Configuration is tested before it is reloaded. Changes are committed to version control so they can be reversed. Certificates renew automatically and are checked anyway. Backups run nightly and are pulled off the machine, with weekly full-disk snapshots on top.
Most downtime I see in other estates is self-inflicted: an untested change on a Friday, a certificate on a host nobody remembered, a disk that filled up slowly over months. Each of those is preventable with routine, which is the unglamorous part of this work and the bit that actually earns the number.
Questions people ask
Do we need five nines?
Probably not everywhere. The useful conversation is which services genuinely need it, which can live with three or four nines, and what each level costs to deliver. Paying for five nines on an internal wiki is waste.
Can you set up the same monitoring for us?
Yes. Content-aware probes, a daily scripted check and alerts that reach a named person are a small, fast piece of work with a large payoff.
Is planned maintenance counted?
The probes do not know whether downtime was planned. If a service is unavailable when checked, it counts — which is a good reason to make maintenance not cause downtime in the first place.