System Design (High-Level Design) Module 3 – Performance and reliability fundamentals
Availability maths: nines, series and parallel
Turn nines into allowed downtime, compute the availability of parts in series and in parallel, and see why shared causes break the redundancy you counted.
What you will learn
- Convert availability targets into allowed downtime over a day, a month and a year
- Compute the availability of components in series and in parallel
- Explain why independence assumptions often fail, and what removes shared causes
Before you start
On this page
Availability is the share of time, or of requests, for which a service works. It is written in “nines”: 99.9% is three nines and allows 43.2 minutes of downtime in a 30-day month; 99.99% is four nines and allows 4.32 minutes. To estimate a design’s availability, combine its parts. Parts that are all needed, one after another, are in series, and their availabilities multiply, so the whole is weaker than its weakest part. Redundant parts of which any one is enough are in parallel, and the whole is down only when all of them are, so two 99% replicas give 99.99%. Both formulas assume that parts fail independently, and real failures often share a cause. This lesson works through each step with numbers.
Nines as downtime
A target means little until it is turned into minutes, and the period it is measured over matters as much as the number. This calculator turns targets into downtime for five periods, and a month’s outages back into a percentage; change the lists at the top and run it again:
"""Availability targets as allowed downtime, and downtime as availability.
Edit TARGETS or OUTAGES and run it again. The year is 365.25 days and the quarter a quarter of it; the month is
30 days, the month most availability tables use. "Nines" is -log10(1 - availability): 99.9% is 3 nines.
"""
import math
TARGETS = [99.0, 99.5, 99.9, 99.95, 99.99, 99.999] # percent
OUTAGES = [(1, 43.2), (2, 12.0), (3, 4.5)] # (outages in a 30-day month, minutes each)
PERIODS = [("day", 1), ("week", 7), ("30-day month", 30), ("quarter", 365.25 / 4), ("year", 365.25)]
def readable(minutes):
"""Minutes as the largest unit that keeps the number at 1 or more."""
if minutes >= 60:
return f"{minutes / 60:.2f} h"
if minutes >= 1:
return f"{minutes:.2f} min"
return f"{minutes * 60:.1f} s"
print(f"{'target':>8}{'nines':>7}" + "".join(f"{name:>14}" for name, _ in PERIODS))
for target in TARGETS:
down = 1 - target / 100
nines = -math.log10(down)
print(f"{target:>7}%{nines:>7.1f}" + "".join(f"{readable(days * 1440 * down):>14}" for _, days in PERIODS))
print("\nWhat a month of outages leaves")
for count, minutes in OUTAGES:
down = count * minutes
availability = 100 * (1 - down / (30 * 1440))
outages = "1 outage" if count == 1 else f"{count} outages"
print(f"{outages} of {minutes:g} min = {down:g} min down: {availability:.3f}% for the month") Output
target nines day week 30-day month quarter year 99.0% 2.0 14.40 min 1.68 h 7.20 h 21.92 h 87.66 h 99.5% 2.3 7.20 min 50.40 min 3.60 h 10.96 h 43.83 h 99.9% 3.0 1.44 min 10.08 min 43.20 min 2.19 h 8.77 h 99.95% 3.3 43.2 s 5.04 min 21.60 min 1.10 h 4.38 h 99.99% 4.0 8.6 s 1.01 min 4.32 min 13.15 min 52.60 min 99.999% 5.0 0.9 s 6.0 s 25.9 s 1.31 min 5.26 min What a month of outages leaves 1 outage of 43.2 min = 43.2 min down: 99.900% for the month 2 outages of 12 min = 24 min down: 99.944% for the month 3 outages of 4.5 min = 13.5 min down: 99.969% for the month
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 nines.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
Each extra nine divides the allowed downtime by ten. At 99.9%, a day allows 1.44 minutes and a year 8.77 hours (with a year of 365.25 days); at 99.99%, a year allows 52.6 minutes and a day only 8.6 seconds. The second block runs the other way: one 43.2-minute outage spends the whole of a month’s 99.9%, while three outages of 4.5 minutes leave 99.969%. The quality-attributes lesson showed what each budget means for how a team works; this lesson is about how the parts of a design add up to one number.
Availability can be counted in two ways. AWS’s availability whitepaper defines it by time, as uptime divided by uptime plus downtime. Google’s SRE book counts requests instead: the share of requests that succeed, because a service that runs in many places is almost never down everywhere at once, so minutes of total outage say little about what users met. A service that answers 2.5 million requests a day with a daily target of 99.99% may fail 250 of them. Counting requests also captures the partial outages that a time-based count misses, which the SLO lesson of this module builds on.
Parts in series multiply
If a request needs every part of a chain, it succeeds only when all of them work at once. With independent parts, the availability of the chain is the product of theirs:
Two 99.9% services in a row give 99.8001%: their unavailabilities, 0.1% each, roughly add up. That rule of thumb is the quickest check in a design round: a load balancer at 99.99%, an app tier at 99.95% and a database at 99.9% leave about 0.01 + 0.05 + 0.1 = 0.16% of failure, so the path is about 99.84%. Every extra hard dependency lowers the ceiling, which is why AWS’s whitepaper says that reducing dependencies improves availability, and that a workload’s dependencies should have goals at least as high as its own. The same paper adds that the product is only a rough estimate: published figures are targets, and dependencies often do better than their stated SLAs.
Redundant parts in parallel
If any one of several parts is enough, the group fails only when all of them have failed at the same time. With independent parts, multiply the chances of failure instead:
Two 99% replicas give 1 − 0.01 × 0.01 = 99.99%, and three give 99.9999%. Two things change this in practice. First, redundancy often has to cover capacity as well: if the peak needs two of three instances, the tier is up only while at least two are, which is a weaker condition than “any one”. Second, failing over takes time: a standby that needs a minute to take over adds a minute of downtime to every failure of the primary. This program puts both into one checkout path:
"""Availability of parts in series and in parallel, for one checkout path, and what recovery time buys.
Every availability below is an assumption for the sketch, written as a fraction (0.999 is 99.9%). The formulas
assume that the parts fail independently; the next program shows what happens when they do not.
"""
from math import comb
def series(*parts):
"""All parts are needed: multiply the availabilities."""
result = 1.0
for a in parts:
result *= a
return result
def parallel(*parts):
"""Any one part is enough: the group is down only when every part is down."""
down = 1.0
for a in parts:
down *= 1 - a
return 1 - down
def at_least(k, n, a):
"""At least k of n identical parts (each available a) must be up."""
return sum(comb(n, up) * a**up * (1 - a) ** (n - up) for up in range(k, n + 1))
def pct(a):
return f"{a * 100:.4f}%"
print("Basics")
print(f" two 99.9% services in series: {pct(series(0.999, 0.999))}")
print(f" two 99% replicas in parallel: {pct(parallel(0.99, 0.99))}")
print(f" three 99% replicas in parallel: {pct(parallel(0.99, 0.99, 0.99))}")
print(f" any 2 of 3 replicas at 99%: {pct(at_least(2, 3, 0.99))}")
# The checkout path: load balancer -> app tier -> database -> payment provider.
lb = 0.9999
app = at_least(2, 3, 0.995) # 3 instances at 99.5%; the peak needs any 2
failover_down = 2 * 1 / (30 * 24 * 60) # about 2 failovers a month, 1 minute of downtime each
db = parallel(0.999, 0.999) - failover_down # primary and standby, minus the failover gaps
payments = 0.9995 # an outside provider, a hard dependency
path = series(lb, app, db, payments)
print("\nCheckout path")
for name, a in [("load balancer", lb), ("app tier (2 of 3)", app), ("database pair", db), ("payment provider", payments), ("whole path", path)]:
print(f" {name:20}{pct(a):>12} down {(1 - a) * 30 * 24 * 60:6.1f} min a month")
print("\nRarer failures or faster recovery (A = MTBF / (MTBF + MTTR))")
for label, mtbf_h, mttr_h in [
("fails every 30 days, 60 min to recover", 720, 1),
("fails every 60 days, 60 min to recover", 1440, 1),
("fails every 30 days, 5 min to recover", 720, 5 / 60),
]:
print(f" {label:40}{pct(mtbf_h / (mtbf_h + mttr_h)):>10}") Output
Basics two 99.9% services in series: 99.8001% two 99% replicas in parallel: 99.9900% three 99% replicas in parallel: 99.9999% any 2 of 3 replicas at 99%: 99.9702% Checkout path load balancer 99.9900% down 4.3 min a month app tier (2 of 3) 99.9925% down 3.2 min a month database pair 99.9953% down 2.0 min a month payment provider 99.9500% down 21.6 min a month whole path 99.9278% down 31.2 min a month Rarer failures or faster recovery (A = MTBF / (MTBF + MTTR)) fails every 30 days, 60 min to recover 99.8613% fails every 60 days, 60 min to recover 99.9306% fails every 30 days, 5 min to recover 99.9884%
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 composite.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
One checkout path: four parts in series, two of them redundant inside
Text description of the diagram
The diagram shows the parts a checkout request passes through, from top to bottom. The request needs every part, so the parts are in series.
- A load balancer, 99.99% available.
- An app tier of three instances, each 99.5% available. The peak needs any two of them, which makes the tier 99.9925% available.
- A database pair, a primary and a standby, each 99.9% available. Either one can serve, so the pair is down only when both are, plus about two minutes a month of failover gaps: 99.9953% in all.
- An outside payment provider, 99.95% available.
The whole path is about 99.928% available, and the payment provider alone accounts for about two thirds of its downtime. All the numbers are the assumptions of the program composite.py in this lesson.
Two of three 99% instances give 99.9702%, against 99.9999% when any one would do. On the checkout path, the redundant app tier and database pair are down only 3.2 and 2.0 minutes a month, while the payment provider, a single outside dependency at 99.95%, accounts for 21.6 of the path’s 31.2 minutes. The path is 99.928% available, and more app instances would not change that: the next improvement has to come from the payment provider, for example a second provider to fail over to, or from not making the payment a hard dependency of every checkout.
AWS’s whitepaper makes the same point about spares from the other side: each spare costs as much as the original, and beyond about three spares for a part that is at least 99% available the extra availability is not worth it. It also warns about the unit of failure: ten instances in one Availability Zone are one failure if the zone goes, so the spare should be a second zone, or three zones of five instances each, which still leaves ten when one zone is lost.
Faster recovery buys more nines
Availability can also be written in mean times: a part that runs for its mean time between failures (MTBF) and then takes its mean time to recover (MTTR) is available
The last block of the program compares two ways to improve a service that fails once every 30 days and needs an hour to recover (99.8613%). Halving how often it fails gives 99.9306%. Recovering in 5 minutes instead of 60 gives 99.9884%, more than either, and often at a lower cost: automatic detection, automatic failover and a fast rollback are usually easier to build than a system that fails half as often. AWS’s whitepaper names the same three levers: fewer failures, faster detection, which is part of recovery, and faster repair.
Correlated failures break the parallel formula
The parallel formula’s promise, 99.99% from two 99% parts, holds only if one part’s failure tells you nothing about the other’s. Replicas that share a power feed, a network switch, a configuration push, a deploy or a software bug fail together. This simulation gives two replicas the same individual availability twice, once with independent failures and once with a shared cause that takes both down about four times a year for about an hour:
"""Two replicas, each up 99% of the time: independent failures against a shared cause.
Each replica fails on its own about once a month and takes about 7.3 hours to repair, so it is down about 1% of
the time. In the second case the pair also shares a cause (one power feed, one configuration push, one region)
that takes both down at once about four times a year, for about an hour. Failures are simulated over 2,000 years
with a fixed seed; the numbers are assumptions for the sketch.
"""
import random
HOURS = 2_000 * 8_766 # 2,000 years of 365.25 days
rng = random.Random(3)
def outages(mtbf_h, mttr_h):
"""Random down intervals (start, end) in hours: up for about mtbf_h, then down for about mttr_h."""
t, spans = 0.0, []
while True:
t += rng.expovariate(1 / mtbf_h)
if t >= HOURS:
return spans
down = rng.expovariate(1 / mttr_h)
spans.append((t, min(t + down, HOURS)))
t += down
def merged(spans):
"""The same down time as non-overlapping intervals, in order."""
out = []
for start, end in sorted(spans):
if out and start <= out[-1][1]:
out[-1] = (out[-1][0], max(out[-1][1], end))
else:
out.append((start, end))
return out
def both_down(a, b):
"""Hours during which both replicas are down."""
a, b, total, i, j = merged(a), merged(b), 0.0, 0, 0
while i < len(a) and j < len(b):
total += max(0.0, min(a[i][1], b[j][1]) - max(a[i][0], b[j][0]))
if a[i][1] < b[j][1]:
i += 1
else:
j += 1
return total
replica_a, replica_b = outages(720, 7.27), outages(720, 7.27)
shared = outages(8_766 / 4, 1) # about four times a year, about an hour each
alone = sum(end - start for start, end in merged(replica_a)) / HOURS
print(f"one replica on its own: {100 * (1 - alone):.3f}% available")
for name, extra in [("two replicas, independent", []), ("two replicas, shared cause", shared)]:
down = both_down(replica_a + extra, replica_b + extra) / HOURS
print(f"{name + ':':33}{100 * (1 - down):.3f}% available, down {down * 8_766:.1f} h a year")
print(f"formula for independent replicas: {100 * (1 - alone**2):.3f}%") Output
one replica on its own: 98.992% available two replicas, independent: 99.989% available, down 0.9 h a year two replicas, shared cause: 99.944% available, down 4.9 h a year formula for independent replicas: 99.990%
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 correlated.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
Each replica on its own is about 99% available, and independently the pair reaches 99.989%, about 0.9 hours of downtime a year, close to the formula’s 99.990%. Four hour-long shared events a year cost the pair five times as much downtime, 4.9 hours, and the shared cause now dominates: once failures are correlated, adding replicas that share the cause adds almost nothing.
Measurements of real fleets agree. A year-long study of Google’s storage clusters by Ford and colleagues found that a large share of node failures came in bursts, that most large bursts were tied to a rack or a group of racks, and that modelling the same failures as independent typically overestimated how long data stays available by two orders of magnitude or more. Correlation also shrank the gain from adding replicas.
Amazon’s Builders’ Library describes correlated failures as several components failing from the same underlying cause, and lists causes beyond the obvious ones: data-centre power, networking and cooling; shared dependencies such as DNS; deployments and tools that act on many servers at once; every server running the same software with the same limits; and many clients changing their behaviour at the same moment. Its remedies are the design moves that the rest of this track uses:
- Independent zones. Run separate copies of the service in each zone, sharing as little as possible with other zones, so a facility problem stays in one zone.
- Staggered changes. Deploy and change configuration in one zone, or one cell, at a time, so a bad change fails small.
- Cells and shuffle sharding. Split the fleet into independent cells, and give customers different combinations of servers, so that no single failure or bad workload reaches everyone.
- Jitter. Add a little randomness to timers, retry delays, cache expiry and limits, so that servers do not all hit the same edge at the same moment.
Practice
Exercise · Easy · Python
Downtime budgets and composite availability
Write four helpers in availability.py. Availabilities go in and come out as percentages, so 99.9 means 99.9%.
downtime_minutes(target_pct, days) returns the minutes of downtime that a target allows over days days. For example, downtime_minutes(99.9, 30) is 43.2.
series(*parts_pct) returns the availability of parts that are all needed for a request to succeed, such as a load balancer, an app tier and a database called one after the other. For example, series(99.9, 99.9) is 99.8001.
parallel(*parts_pct) returns the availability of redundant parts of which any one is enough, assuming they fail independently. For example, parallel(99, 99) is 99.99.
from_mtbf(mtbf_hours, mttr_hours) returns the availability of a part that runs for mtbf_hours on average between failures and takes mttr_hours on average to recover. For example, from_mtbf(720, 1) is about 99.8613.
Starter code · availability.py
def downtime_minutes(target_pct, days):
"""Minutes of downtime that target_pct allows over `days` days."""
# Replace this line with your code.
return 0
def series(*parts_pct):
"""Availability, in percent, of parts that are all needed."""
# Replace this line with your code.
return 0
def parallel(*parts_pct):
"""Availability, in percent, of independent redundant parts of which any one is enough."""
# Replace this line with your code.
return 0
def from_mtbf(mtbf_hours, mttr_hours):
"""Availability, in percent, from the mean time between failures and the mean time to recover."""
# Replace this line with your code.
return 0 The sample tests · test_availability.py
import math
from availability import downtime_minutes, from_mtbf, parallel, series
def close(a, b):
return math.isclose(a, b, abs_tol=1e-6)
def test_downtime_minutes():
"""turns a target into minutes of downtime over a period"""
assert close(downtime_minutes(99.9, 30), 43.2)
assert close(downtime_minutes(99.95, 30), 21.6)
assert close(downtime_minutes(99.99, 365.25), 52.596)
assert close(downtime_minutes(100, 30), 0)
def test_series():
"""multiplies the availabilities of parts that are all needed"""
assert close(series(99.9, 99.9), 99.8001)
assert close(series(99.99, 99.95, 99.9), 99.840064995)
assert close(series(99.5), 99.5)
def test_parallel():
"""is down only when every redundant part is down"""
assert close(parallel(99, 99), 99.99)
assert close(parallel(99.5, 99.5), 99.9975)
assert close(parallel(90, 90, 90), 99.9)
assert close(parallel(80), 80)
def test_from_mtbf():
"""is the share of time up between failures"""
assert close(from_mtbf(720, 1), 100 * 720 / 721)
assert close(from_mtbf(720, 5 / 60), 100 * 720 / (720 + 5 / 60))
assert close(from_mtbf(99, 1), 99) A hint
Turn each percentage into a fraction first (p / 100) and back at the end. In series, multiply the fractions. In parallel, multiply the chances that each part is down (1 - fraction) and subtract the product from 1. For the mean times, the part is up for MTBF out of every MTBF + MTTR hours.
Results of the sample tests
| Test | Result | Details |
|---|
What your code printed
The sample tests run on this device, in your browser (Pyodide): nothing is sent to mysmartcopilot.com. The first run downloads Python (about 13.5 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.
Check yourself
6 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.
Interview questions
Warm-up (fresher to mid level): two services, each 99.9% available, must both work for a request to succeed. What availability can the pair offer, and why? About 99.8%. The request needs both, so it succeeds only when both are up at once: 0.999 × 0.999 = 0.998001, assuming they fail independently. A quick way to see it is that the unavailabilities roughly add up: 0.1% + 0.1% = 0.2% of failures, about 86 minutes a month instead of 43. That is why every hard dependency lowers the ceiling of a design, and why designs remove dependencies from the critical path, add redundancy inside them, or make them soft dependencies that the service can work without for a while.
Key takeaways
- Turn targets into downtime over a stated period: 99.9% is 43.2 minutes per 30-day month and 8.77 hours per year; each nine divides it by ten.
- In series, availabilities multiply and unavailabilities roughly add: every hard dependency lowers the ceiling.
- In parallel, the group fails only when every part has failed: two independent 99% replicas give 99.99%, but “any 2 of 3” and failover time give less.
- A = MTBF ÷ (MTBF + MTTR): faster detection and recovery often buy more nines than rarer failures.
- Shared causes (power, deploys, configuration, software, zones) make replicas fail together; separate what they share before adding more of them.
References
- Availability and Beyond: Understanding and Improving the Resilience of Distributed Systems on AWS (Amazon Web Services)
- Availability and Beyond: Understanding availability (MTBF and MTTR) (Amazon Web Services)
- Availability and Beyond: Availability with dependencies (Amazon Web Services)
- Availability and Beyond: Availability with redundancy (Amazon Web Services)
- Minimizing correlated failures in distributed systems (Amazon Builders' Library) (Amazon Web Services)
- Site Reliability Engineering, chapter 3: Embracing Risk (Google (O'Reilly Media))
- Availability in Globally Distributed Storage Systems (Ford et al.) (USENIX OSDI 2010)
Related tools
Report a problem with this lesson
Kept only in this browser. Your Learn progress