On this page
"99.9% uptime" sounds airtight until you do the arithmetic. That last tenth of a percent is almost 45 minutes of downtime every month — a full outage during your busiest hour, comfortably inside the promise. Understanding what an uptime figure actually permits is the difference between a target you can hit and a number on a slide.
The short version
- Uptime % is downtime budget: 99.9% allows ~43m 50s per month; 99.99% allows ~4m 23s.
- Each additional nine cuts the allowed downtime by 10× and roughly multiplies the engineering effort.
- The percentage is meaningless without a measurement window and a definition of “down”.
- SLI is the measurement, SLO is your internal target, SLA is the customer contract (set below the SLO).
- Most teams should aim for a number they can measure and defend — not the most nines they can print.
What "uptime %" actually measures
Uptime percentage is simply the share of a time window during which your service was available: uptime = good time ÷ total time. If a month has 43,800 minutes and your service was unavailable for 44 of them, that's 99.9%. The headline number is really a downtime budget in disguise — and budgets are much easier to reason about than abstract percentages.
43m 50s
Monthly downtime budget at 99.9%
4m 23s
Monthly downtime budget at 99.99%
26s
Monthly downtime budget at 99.999%
The downtime cheat sheet
The table every SRE eventually bookmarks. Allowed downtime by availability level (using an average 30.44-day month and a 365.25-day year):
| Availability | Nickname | Per month | Per year |
|---|---|---|---|
| 99% | Two nines | 7h 18m | 3d 15h |
| 99.5% | — | 3h 39m | 1d 19h |
| 99.9% | Three nines | 43m 50s | 8h 46m |
| 99.95% | — | 21m 55s | 4h 23m |
| 99.99% | Four nines | 4m 23s | 52m 36s |
| 99.999% | Five nines | 26s | 5m 15s |
Every nine is 10× harder
The table hides a brutal truth: the cost of each nine is not linear. Going from 99% to 99.9% is mostly discipline — decent monitoring, careful deploys, a backup plan. Going from 99.9% to 99.99% usually means eliminating single points of failure: redundant instances, automated failover, multi-zone databases, zero-downtime deploys. Reaching 99.999% means the humans can't even be in the loop for recovery — 26 seconds a month isn't enough time for a pager to go off and someone to open a laptop.
Note
What the number leaves out
A bare "99.9%" is close to unfalsifiable without answers to three questions:
Over what window?
99.9% per year lets you burn all 8+ hours in a single bad afternoon. 99.9% per month, measured monthly, caps any single month at ~44 minutes. Shorter windows are stricter and much more meaningful to users.
What counts as "down"?
Is a page that loads in 9 seconds "up"? Is a 200 that renders an error page "up"? (It shouldn't be — see why a 200 can still mean broken.) The definition of the failure is where most of the honesty lives.
What's excluded?
Most real SLAs carve out scheduled maintenance and "force majeure". Fair enough — but it means the advertised number and the number your users feel can differ. Measure your own uptime from the outside so you know the difference.
SLA vs SLO vs SLI
Three terms that get used interchangeably and shouldn't be:
| Term | What it is | Example |
|---|---|---|
| SLI (Indicator) | The measurement itself | % of requests served < 500ms without error |
| SLO (Objective) | Your internal target for the SLI | 99.95% over a rolling 30 days |
| SLA (Agreement) | The external promise, with penalties | 99.9%, or customers get credits |
The healthy pattern: set your SLO stricter than your SLA. If you promise customers 99.9% but target 99.95% internally, you have headroom to catch a slipping trend before it becomes a contract breach — and a credit.
What to actually aim for
More nines is not automatically better; it's a spend. Choose the target the same way you'd choose any budget:
- Internal tools & side projects: 99% is honest and cheap. Nobody's paging at 2 a.m. for the admin dashboard.
- Standard SaaS & business sites: 99.9% is the sweet spot — achievable with good practices, and what most customers implicitly expect.
- Revenue-critical & platform services: 99.95–99.99%, and now you're paying for redundancy and automated failover.
- Five nines: reserve it for genuine infrastructure. It's a different engineering culture, not a config change.
Whatever you pick, you can't hit a target you don't measure. Uptime monitoring from outside your network is how you turn a percentage on a slide into a number you can actually defend — start with what uptime monitoring is if you're setting it up from scratch.
Frequently asked questions
- How much downtime does 99.9% uptime allow?
- 99.9% ('three nines') allows about 43 minutes and 50 seconds of downtime per month, or roughly 8 hours 46 minutes per year. 99.99% ('four nines') tightens that to about 4 minutes 23 seconds per month.
- What does 99.99% uptime mean?
- 99.99% ('four nines') means your service is unavailable no more than about 52 minutes and 36 seconds across a whole year — roughly 4 minutes 23 seconds per month. It's a serious commitment that usually requires redundancy and automated failover.
- What's the difference between an SLA, an SLO, and an SLI?
- An SLI (indicator) is the measurement, such as the percentage of successful requests. An SLO (objective) is the internal target you aim for, such as 99.95%. An SLA (agreement) is the external contract with customers, usually set below the SLO and carrying penalties or credits if you miss it.
- Is 100% uptime a realistic goal?
- No. Every layer you depend on — power, network, DNS, cloud providers — has its own failure rate, and deploys and maintenance introduce risk. Serious providers promise a number of nines and back it with credits, not perfection. Chasing 100% wastes money that's better spent detecting and recovering from the failures that will happen.