Monitoring 1018 min read

99.9% Uptime, Decoded: An SLA & Downtime Cheat Sheet

Every extra nine costs 10× the effort and buys 10× less downtime. Here's exactly what each one allows — and how to pick a realistic target.

The Ensju team
On this page

"99.9% uptime" sounds airtight until you do the arithmetic. That last tenth of a percent is almost 45 minutes of downtime every month — a full outage during your busiest hour, comfortably inside the promise. Understanding what an uptime figure actually permits is the difference between a target you can hit and a number on a slide.

The short version

  • Uptime % is downtime budget: 99.9% allows ~43m 50s per month; 99.99% allows ~4m 23s.
  • Each additional nine cuts the allowed downtime by 10× and roughly multiplies the engineering effort.
  • The percentage is meaningless without a measurement window and a definition of “down”.
  • SLI is the measurement, SLO is your internal target, SLA is the customer contract (set below the SLO).
  • Most teams should aim for a number they can measure and defend — not the most nines they can print.

What "uptime %" actually measures

Uptime percentage is simply the share of a time window during which your service was available: uptime = good time ÷ total time. If a month has 43,800 minutes and your service was unavailable for 44 of them, that's 99.9%. The headline number is really a downtime budget in disguise — and budgets are much easier to reason about than abstract percentages.

43m 50s

Monthly downtime budget at 99.9%

4m 23s

Monthly downtime budget at 99.99%

26s

Monthly downtime budget at 99.999%

The downtime cheat sheet

The table every SRE eventually bookmarks. Allowed downtime by availability level (using an average 30.44-day month and a 365.25-day year):

AvailabilityNicknamePer monthPer year
99%Two nines7h 18m3d 15h
99.5%3h 39m1d 19h
99.9%Three nines43m 50s8h 46m
99.95%21m 55s4h 23m
99.99%Four nines4m 23s52m 36s
99.999%Five nines26s5m 15s
Allowed downtime shrinks by a factor of ~10 with each added nine.

Every nine is 10× harder

The table hides a brutal truth: the cost of each nine is not linear. Going from 99% to 99.9% is mostly discipline — decent monitoring, careful deploys, a backup plan. Going from 99.9% to 99.99% usually means eliminating single points of failure: redundant instances, automated failover, multi-zone databases, zero-downtime deploys. Reaching 99.999% means the humans can't even be in the loop for recovery — 26 seconds a month isn't enough time for a pager to go off and someone to open a laptop.

Note

A useful rule of thumb: each additional nine multiplies the allowed downtime reduction by 10 and the engineering investment by roughly the same. Buy nines only where the revenue or risk justifies them.

What the number leaves out

A bare "99.9%" is close to unfalsifiable without answers to three questions:

Over what window?

99.9% per year lets you burn all 8+ hours in a single bad afternoon. 99.9% per month, measured monthly, caps any single month at ~44 minutes. Shorter windows are stricter and much more meaningful to users.

What counts as "down"?

Is a page that loads in 9 seconds "up"? Is a 200 that renders an error page "up"? (It shouldn't be — see why a 200 can still mean broken.) The definition of the failure is where most of the honesty lives.

What's excluded?

Most real SLAs carve out scheduled maintenance and "force majeure". Fair enough — but it means the advertised number and the number your users feel can differ. Measure your own uptime from the outside so you know the difference.

SLA vs SLO vs SLI

Three terms that get used interchangeably and shouldn't be:

TermWhat it isExample
SLI (Indicator)The measurement itself% of requests served < 500ms without error
SLO (Objective)Your internal target for the SLI99.95% over a rolling 30 days
SLA (Agreement)The external promise, with penalties99.9%, or customers get credits

The healthy pattern: set your SLO stricter than your SLA. If you promise customers 99.9% but target 99.95% internally, you have headroom to catch a slipping trend before it becomes a contract breach — and a credit.

What to actually aim for

More nines is not automatically better; it's a spend. Choose the target the same way you'd choose any budget:

  • Internal tools & side projects: 99% is honest and cheap. Nobody's paging at 2 a.m. for the admin dashboard.
  • Standard SaaS & business sites: 99.9% is the sweet spot — achievable with good practices, and what most customers implicitly expect.
  • Revenue-critical & platform services: 99.95–99.99%, and now you're paying for redundancy and automated failover.
  • Five nines: reserve it for genuine infrastructure. It's a different engineering culture, not a config change.

Whatever you pick, you can't hit a target you don't measure. Uptime monitoring from outside your network is how you turn a percentage on a slide into a number you can actually defend — start with what uptime monitoring is if you're setting it up from scratch.

Frequently asked questions

How much downtime does 99.9% uptime allow?
99.9% ('three nines') allows about 43 minutes and 50 seconds of downtime per month, or roughly 8 hours 46 minutes per year. 99.99% ('four nines') tightens that to about 4 minutes 23 seconds per month.
What does 99.99% uptime mean?
99.99% ('four nines') means your service is unavailable no more than about 52 minutes and 36 seconds across a whole year — roughly 4 minutes 23 seconds per month. It's a serious commitment that usually requires redundancy and automated failover.
What's the difference between an SLA, an SLO, and an SLI?
An SLI (indicator) is the measurement, such as the percentage of successful requests. An SLO (objective) is the internal target you aim for, such as 99.95%. An SLA (agreement) is the external contract with customers, usually set below the SLO and carrying penalties or credits if you miss it.
Is 100% uptime a realistic goal?
No. Every layer you depend on — power, network, DNS, cloud providers — has its own failure rate, and deploys and maintenance introduce risk. Serious providers promise a number of nines and back it with credits, not perfection. Chasing 100% wastes money that's better spent detecting and recovering from the failures that will happen.

Get the next one in your inbox

Practical, vendor-neutral monitoring tips — what to watch and how to alert on it — plus the occasional Ensju update. No spam, unsubscribe anytime.

Monitoring tips + Ensju updates. No spam, unsubscribe anytime.