Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

System Design (High-Level Design)  Module 3 – Performance and reliability fundamentals

Availability maths: nines, series and parallel

Turn nines into allowed downtime, compute the availability of parts in series and in parallel, and see why shared causes break the redundancy you counted.

  • Beginner
  • 25 minutes
  • Examples run with Python 3.14.8, Pyodide 314.0.7, Node.js 24.21.0 and quickjs 0.32.0
  • By MySmartCoPilot

What you will learn

  • Convert availability targets into allowed downtime over a day, a month and a year
  • Compute the availability of components in series and in parallel
  • Explain why independence assumptions often fail, and what removes shared causes

Before you start

On this page

Availability is the share of time, or of requests, for which a service works. It is written in “nines”: 99.9% is three nines and allows 43.2 minutes of downtime in a 30-day month; 99.99% is four nines and allows 4.32 minutes. To estimate a design’s availability, combine its parts. Parts that are all needed, one after another, are in series, and their availabilities multiply, so the whole is weaker than its weakest part. Redundant parts of which any one is enough are in parallel, and the whole is down only when all of them are, so two 99% replicas give 99.99%. Both formulas assume that parts fail independently, and real failures often share a cause. This lesson works through each step with numbers.

Nines as downtime

A target means little until it is turned into minutes, and the period it is measured over matters as much as the number. This calculator turns targets into downtime for five periods, and a month’s outages back into a percentage; change the lists at the top and run it again:

Allowed downtime per period, and what a month of outages leaves Python · nines.py
"""Availability targets as allowed downtime, and downtime as availability.

Edit TARGETS or OUTAGES and run it again. The year is 365.25 days and the quarter a quarter of it; the month is
30 days, the month most availability tables use. "Nines" is -log10(1 - availability): 99.9% is 3 nines.
"""

import math

TARGETS = [99.0, 99.5, 99.9, 99.95, 99.99, 99.999]  # percent
OUTAGES = [(1, 43.2), (2, 12.0), (3, 4.5)]  # (outages in a 30-day month, minutes each)

PERIODS = [("day", 1), ("week", 7), ("30-day month", 30), ("quarter", 365.25 / 4), ("year", 365.25)]


def readable(minutes):
    """Minutes as the largest unit that keeps the number at 1 or more."""
    if minutes >= 60:
        return f"{minutes / 60:.2f} h"
    if minutes >= 1:
        return f"{minutes:.2f} min"
    return f"{minutes * 60:.1f} s"


print(f"{'target':>8}{'nines':>7}" + "".join(f"{name:>14}" for name, _ in PERIODS))
for target in TARGETS:
    down = 1 - target / 100
    nines = -math.log10(down)
    print(f"{target:>7}%{nines:>7.1f}" + "".join(f"{readable(days * 1440 * down):>14}" for _, days in PERIODS))

print("\nWhat a month of outages leaves")
for count, minutes in OUTAGES:
    down = count * minutes
    availability = 100 * (1 - down / (30 * 1440))
    outages = "1 outage" if count == 1 else f"{count} outages"
    print(f"{outages} of {minutes:g} min = {down:g} min down: {availability:.3f}% for the month")

Output

  target  nines           day          week  30-day month       quarter          year
   99.0%    2.0     14.40 min        1.68 h        7.20 h       21.92 h       87.66 h
   99.5%    2.3      7.20 min     50.40 min        3.60 h       10.96 h       43.83 h
   99.9%    3.0      1.44 min     10.08 min     43.20 min        2.19 h        8.77 h
  99.95%    3.3        43.2 s      5.04 min     21.60 min        1.10 h        4.38 h
  99.99%    4.0         8.6 s      1.01 min      4.32 min     13.15 min     52.60 min
 99.999%    5.0         0.9 s         6.0 s        25.9 s      1.31 min      5.26 min

What a month of outages leaves
1 outage of 43.2 min = 43.2 min down: 99.900% for the month
2 outages of 12 min = 24 min down: 99.944% for the month
3 outages of 4.5 min = 13.5 min down: 99.969% for the month

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 nines.py

Each extra nine divides the allowed downtime by ten. At 99.9%, a day allows 1.44 minutes and a year 8.77 hours (with a year of 365.25 days); at 99.99%, a year allows 52.6 minutes and a day only 8.6 seconds. The second block runs the other way: one 43.2-minute outage spends the whole of a month’s 99.9%, while three outages of 4.5 minutes leave 99.969%. The quality-attributes lesson showed what each budget means for how a team works; this lesson is about how the parts of a design add up to one number.

Availability can be counted in two ways. AWS’s availability whitepaper defines it by time, as uptime divided by uptime plus downtime. Google’s SRE book counts requests instead: the share of requests that succeed, because a service that runs in many places is almost never down everywhere at once, so minutes of total outage say little about what users met. A service that answers 2.5 million requests a day with a daily target of 99.99% may fail 250 of them. Counting requests also captures the partial outages that a time-based count misses, which the SLO lesson of this module builds on.

Parts in series multiply

If a request needs every part of a chain, it succeeds only when all of them work at once. With independent parts, the availability of the chain is the product of theirs:

Aseries=A1×A2×⋯×AnA_{\text{series}} = A_1 \times A_2 \times \cdots \times A_n

Two 99.9% services in a row give 99.8001%: their unavailabilities, 0.1% each, roughly add up. That rule of thumb is the quickest check in a design round: a load balancer at 99.99%, an app tier at 99.95% and a database at 99.9% leave about 0.01 + 0.05 + 0.1 = 0.16% of failure, so the path is about 99.84%. Every extra hard dependency lowers the ceiling, which is why AWS’s whitepaper says that reducing dependencies improves availability, and that a workload’s dependencies should have goals at least as high as its own. The same paper adds that the product is only a rough estimate: published figures are targets, and dependencies often do better than their stated SLAs.

Redundant parts in parallel

If any one of several parts is enough, the group fails only when all of them have failed at the same time. With independent parts, multiply the chances of failure instead:

Aparallel=1−(1−A1)(1−A2)⋯(1−An)A_{\text{parallel}} = 1 - (1 - A_1)(1 - A_2)\cdots(1 - A_n)

Two 99% replicas give 1 − 0.01 × 0.01 = 99.99%, and three give 99.9999%. Two things change this in practice. First, redundancy often has to cover capacity as well: if the peak needs two of three instances, the tier is up only while at least two are, which is a weaker condition than “any one”. Second, failing over takes time: a standby that needs a minute to take over adds a minute of downtime to every failure of the primary. This program puts both into one checkout path:

Series, parallel and two of three, for one checkout path Python · composite.py
"""Availability of parts in series and in parallel, for one checkout path, and what recovery time buys.

Every availability below is an assumption for the sketch, written as a fraction (0.999 is 99.9%). The formulas
assume that the parts fail independently; the next program shows what happens when they do not.
"""

from math import comb


def series(*parts):
    """All parts are needed: multiply the availabilities."""
    result = 1.0
    for a in parts:
        result *= a
    return result


def parallel(*parts):
    """Any one part is enough: the group is down only when every part is down."""
    down = 1.0
    for a in parts:
        down *= 1 - a
    return 1 - down


def at_least(k, n, a):
    """At least k of n identical parts (each available a) must be up."""
    return sum(comb(n, up) * a**up * (1 - a) ** (n - up) for up in range(k, n + 1))


def pct(a):
    return f"{a * 100:.4f}%"


print("Basics")
print(f"  two 99.9% services in series:     {pct(series(0.999, 0.999))}")
print(f"  two 99% replicas in parallel:     {pct(parallel(0.99, 0.99))}")
print(f"  three 99% replicas in parallel:   {pct(parallel(0.99, 0.99, 0.99))}")
print(f"  any 2 of 3 replicas at 99%:       {pct(at_least(2, 3, 0.99))}")

# The checkout path: load balancer -> app tier -> database -> payment provider.
lb = 0.9999
app = at_least(2, 3, 0.995)  # 3 instances at 99.5%; the peak needs any 2
failover_down = 2 * 1 / (30 * 24 * 60)  # about 2 failovers a month, 1 minute of downtime each
db = parallel(0.999, 0.999) - failover_down  # primary and standby, minus the failover gaps
payments = 0.9995  # an outside provider, a hard dependency
path = series(lb, app, db, payments)
print("\nCheckout path")
for name, a in [("load balancer", lb), ("app tier (2 of 3)", app), ("database pair", db), ("payment provider", payments), ("whole path", path)]:
    print(f"  {name:20}{pct(a):>12}  down {(1 - a) * 30 * 24 * 60:6.1f} min a month")

print("\nRarer failures or faster recovery (A = MTBF / (MTBF + MTTR))")
for label, mtbf_h, mttr_h in [
    ("fails every 30 days, 60 min to recover", 720, 1),
    ("fails every 60 days, 60 min to recover", 1440, 1),
    ("fails every 30 days, 5 min to recover", 720, 5 / 60),
]:
    print(f"  {label:40}{pct(mtbf_h / (mtbf_h + mttr_h)):>10}")

Output

Basics
  two 99.9% services in series:     99.8001%
  two 99% replicas in parallel:     99.9900%
  three 99% replicas in parallel:   99.9999%
  any 2 of 3 replicas at 99%:       99.9702%

Checkout path
  load balancer           99.9900%  down    4.3 min a month
  app tier (2 of 3)       99.9925%  down    3.2 min a month
  database pair           99.9953%  down    2.0 min a month
  payment provider        99.9500%  down   21.6 min a month
  whole path              99.9278%  down   31.2 min a month

Rarer failures or faster recovery (A = MTBF / (MTBF + MTTR))
  fails every 30 days, 60 min to recover    99.8613%
  fails every 60 days, 60 min to recover    99.9306%
  fails every 30 days, 5 min to recover     99.9884%

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 composite.py

A checkout path in series: load balancer, an app tier of three instances, a database pair, and a payment provider.Load balancer99.99%App tier: any 2 of 3 instances99.9925%Database pair: either one99.9953% after failoversPayment provider99.95%App 199.5%App 299.5%App 399.5%Primary99.9%Standby99.9%

One checkout path: four parts in series, two of them redundant inside

Text description of the diagram

The diagram shows the parts a checkout request passes through, from top to bottom. The request needs every part, so the parts are in series.

  1. A load balancer, 99.99% available.
  2. An app tier of three instances, each 99.5% available. The peak needs any two of them, which makes the tier 99.9925% available.
  3. A database pair, a primary and a standby, each 99.9% available. Either one can serve, so the pair is down only when both are, plus about two minutes a month of failover gaps: 99.9953% in all.
  4. An outside payment provider, 99.95% available.

The whole path is about 99.928% available, and the payment provider alone accounts for about two thirds of its downtime. All the numbers are the assumptions of the program composite.py in this lesson.

Two of three 99% instances give 99.9702%, against 99.9999% when any one would do. On the checkout path, the redundant app tier and database pair are down only 3.2 and 2.0 minutes a month, while the payment provider, a single outside dependency at 99.95%, accounts for 21.6 of the path’s 31.2 minutes. The path is 99.928% available, and more app instances would not change that: the next improvement has to come from the payment provider, for example a second provider to fail over to, or from not making the payment a hard dependency of every checkout.

AWS’s whitepaper makes the same point about spares from the other side: each spare costs as much as the original, and beyond about three spares for a part that is at least 99% available the extra availability is not worth it. It also warns about the unit of failure: ten instances in one Availability Zone are one failure if the zone goes, so the spare should be a second zone, or three zones of five instances each, which still leaves ten when one zone is lost.

Faster recovery buys more nines

Availability can also be written in mean times: a part that runs for its mean time between failures (MTBF) and then takes its mean time to recover (MTTR) is available

A=MTBFMTBF+MTTRA = \frac{\text{MTBF}}{\text{MTBF} + \text{MTTR}}

The last block of the program compares two ways to improve a service that fails once every 30 days and needs an hour to recover (99.8613%). Halving how often it fails gives 99.9306%. Recovering in 5 minutes instead of 60 gives 99.9884%, more than either, and often at a lower cost: automatic detection, automatic failover and a fast rollback are usually easier to build than a system that fails half as often. AWS’s whitepaper names the same three levers: fewer failures, faster detection, which is part of recovery, and faster repair.

Correlated failures break the parallel formula

The parallel formula’s promise, 99.99% from two 99% parts, holds only if one part’s failure tells you nothing about the other’s. Replicas that share a power feed, a network switch, a configuration push, a deploy or a software bug fail together. This simulation gives two replicas the same individual availability twice, once with independent failures and once with a shared cause that takes both down about four times a year for about an hour:

Two 99% replicas, independent and with a shared cause Python · correlated.py
"""Two replicas, each up 99% of the time: independent failures against a shared cause.

Each replica fails on its own about once a month and takes about 7.3 hours to repair, so it is down about 1% of
the time. In the second case the pair also shares a cause (one power feed, one configuration push, one region)
that takes both down at once about four times a year, for about an hour. Failures are simulated over 2,000 years
with a fixed seed; the numbers are assumptions for the sketch.
"""

import random

HOURS = 2_000 * 8_766  # 2,000 years of 365.25 days
rng = random.Random(3)


def outages(mtbf_h, mttr_h):
    """Random down intervals (start, end) in hours: up for about mtbf_h, then down for about mttr_h."""
    t, spans = 0.0, []
    while True:
        t += rng.expovariate(1 / mtbf_h)
        if t >= HOURS:
            return spans
        down = rng.expovariate(1 / mttr_h)
        spans.append((t, min(t + down, HOURS)))
        t += down


def merged(spans):
    """The same down time as non-overlapping intervals, in order."""
    out = []
    for start, end in sorted(spans):
        if out and start <= out[-1][1]:
            out[-1] = (out[-1][0], max(out[-1][1], end))
        else:
            out.append((start, end))
    return out


def both_down(a, b):
    """Hours during which both replicas are down."""
    a, b, total, i, j = merged(a), merged(b), 0.0, 0, 0
    while i < len(a) and j < len(b):
        total += max(0.0, min(a[i][1], b[j][1]) - max(a[i][0], b[j][0]))
        if a[i][1] < b[j][1]:
            i += 1
        else:
            j += 1
    return total


replica_a, replica_b = outages(720, 7.27), outages(720, 7.27)
shared = outages(8_766 / 4, 1)  # about four times a year, about an hour each
alone = sum(end - start for start, end in merged(replica_a)) / HOURS

print(f"one replica on its own:          {100 * (1 - alone):.3f}% available")
for name, extra in [("two replicas, independent", []), ("two replicas, shared cause", shared)]:
    down = both_down(replica_a + extra, replica_b + extra) / HOURS
    print(f"{name + ':':33}{100 * (1 - down):.3f}% available, down {down * 8_766:.1f} h a year")
print(f"formula for independent replicas: {100 * (1 - alone**2):.3f}%")

Output

one replica on its own:          98.992% available
two replicas, independent:       99.989% available, down 0.9 h a year
two replicas, shared cause:      99.944% available, down 4.9 h a year
formula for independent replicas: 99.990%

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 correlated.py

Each replica on its own is about 99% available, and independently the pair reaches 99.989%, about 0.9 hours of downtime a year, close to the formula’s 99.990%. Four hour-long shared events a year cost the pair five times as much downtime, 4.9 hours, and the shared cause now dominates: once failures are correlated, adding replicas that share the cause adds almost nothing.

Measurements of real fleets agree. A year-long study of Google’s storage clusters by Ford and colleagues found that a large share of node failures came in bursts, that most large bursts were tied to a rack or a group of racks, and that modelling the same failures as independent typically overestimated how long data stays available by two orders of magnitude or more. Correlation also shrank the gain from adding replicas.

Amazon’s Builders’ Library describes correlated failures as several components failing from the same underlying cause, and lists causes beyond the obvious ones: data-centre power, networking and cooling; shared dependencies such as DNS; deployments and tools that act on many servers at once; every server running the same software with the same limits; and many clients changing their behaviour at the same moment. Its remedies are the design moves that the rest of this track uses:

  • Independent zones. Run separate copies of the service in each zone, sharing as little as possible with other zones, so a facility problem stays in one zone.
  • Staggered changes. Deploy and change configuration in one zone, or one cell, at a time, so a bad change fails small.
  • Cells and shuffle sharding. Split the fleet into independent cells, and give customers different combinations of servers, so that no single failure or bad workload reaches everyone.
  • Jitter. Add a little randomness to timers, retry delays, cache expiry and limits, so that servers do not all hit the same edge at the same moment.
Time Unit Converter Turn a downtime budget such as 52.6 minutes a year into seconds, hours or days.

Practice

Exercise · Easy · Python

Downtime budgets and composite availability

Write four helpers in availability.py. Availabilities go in and come out as percentages, so 99.9 means 99.9%.

downtime_minutes(target_pct, days) returns the minutes of downtime that a target allows over days days. For example, downtime_minutes(99.9, 30) is 43.2.

series(*parts_pct) returns the availability of parts that are all needed for a request to succeed, such as a load balancer, an app tier and a database called one after the other. For example, series(99.9, 99.9) is 99.8001.

parallel(*parts_pct) returns the availability of redundant parts of which any one is enough, assuming they fail independently. For example, parallel(99, 99) is 99.99.

from_mtbf(mtbf_hours, mttr_hours) returns the availability of a part that runs for mtbf_hours on average between failures and takes mttr_hours on average to recover. For example, from_mtbf(720, 1) is about 99.8613.

Starter code · availability.py

def downtime_minutes(target_pct, days):
    """Minutes of downtime that target_pct allows over `days` days."""
    # Replace this line with your code.
    return 0


def series(*parts_pct):
    """Availability, in percent, of parts that are all needed."""
    # Replace this line with your code.
    return 0


def parallel(*parts_pct):
    """Availability, in percent, of independent redundant parts of which any one is enough."""
    # Replace this line with your code.
    return 0


def from_mtbf(mtbf_hours, mttr_hours):
    """Availability, in percent, from the mean time between failures and the mean time to recover."""
    # Replace this line with your code.
    return 0
The sample tests · test_availability.py
import math

from availability import downtime_minutes, from_mtbf, parallel, series


def close(a, b):
    return math.isclose(a, b, abs_tol=1e-6)


def test_downtime_minutes():
    """turns a target into minutes of downtime over a period"""
    assert close(downtime_minutes(99.9, 30), 43.2)
    assert close(downtime_minutes(99.95, 30), 21.6)
    assert close(downtime_minutes(99.99, 365.25), 52.596)
    assert close(downtime_minutes(100, 30), 0)


def test_series():
    """multiplies the availabilities of parts that are all needed"""
    assert close(series(99.9, 99.9), 99.8001)
    assert close(series(99.99, 99.95, 99.9), 99.840064995)
    assert close(series(99.5), 99.5)


def test_parallel():
    """is down only when every redundant part is down"""
    assert close(parallel(99, 99), 99.99)
    assert close(parallel(99.5, 99.5), 99.9975)
    assert close(parallel(90, 90, 90), 99.9)
    assert close(parallel(80), 80)


def test_from_mtbf():
    """is the share of time up between failures"""
    assert close(from_mtbf(720, 1), 100 * 720 / 721)
    assert close(from_mtbf(720, 5 / 60), 100 * 720 / (720 + 5 / 60))
    assert close(from_mtbf(99, 1), 99)
A hint

Turn each percentage into a fraction first (p / 100) and back at the end. In series, multiply the fractions. In parallel, multiply the chances that each part is down (1 - fraction) and subtract the product from 1. For the mean times, the part is up for MTBF out of every MTBF + MTTR hours.

The sample tests run on this device, in your browser (Pyodide): nothing is sent to mysmartcopilot.com. The first run downloads Python (about 13.5 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.

Check yourself

6 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.

  1. Question 1 of 6 How many minutes of downtime does a 99.95% target allow in a 30-day month?

    Type a number.

    Show the answer to question 1

    Answer: 21.6 minutes (anything from 21.55 to 21.65 counts)

    A 30-day month has 30 × 24 × 60 = 43,200 minutes, and 0.05% of them is 21.6 minutes. Two incidents of a quarter of an hour already break it.

  2. Question 2 of 6 A request passes a load balancer (99.99%), an app tier (99.95%) and a database (99.9%), and needs all three. What is the path's availability, in percent to two decimal places?

    Type a number.

    Show the answer to question 2

    Answer: 99.84 % (anything from 99.83 to 99.85 counts)

    In series the availabilities multiply: 0.9999 × 0.9995 × 0.999 ≈ 0.9984, or 99.84%. A quick check: the unavailabilities add up for small values, 0.01 + 0.05 + 0.1 = 0.16%, so 99.84%.

  3. Question 3 of 6 Two replicas, each 99.5% available and failing independently, serve the same data; either one is enough. What is the pair's availability, in percent to four decimal places?

    Type a number.

    Show the answer to question 3

    Answer: 99.9975 % (anything from 99.9974 to 99.9976 counts)

    Both are down at once with probability 0.005 × 0.005 = 0.000025, so the pair is up 99.9975% of the time, if their failures really are independent.

  4. Question 4 of 6 Three app instances are each 99% available, and the peak needs any two of them. What is the tier's availability, in percent to two decimal places?

    Type a number.

    Show the answer to question 4

    Answer: 99.97 % (anything from 99.96 to 99.98 counts)

    The tier is up when all three are up (0.99³ = 0.970299) or exactly two are (3 × 0.99² × 0.01 = 0.029403), which adds up to 0.999702, or 99.97%. Needing two of three is weaker than needing one of three (99.9999%).

  5. Question 5 of 6 A service fails on average every 10 days and takes 1 hour to recover. What is its availability, in percent to two decimal places?

    Type a number.

    Show the answer to question 5

    Answer: 99.59 % (anything from 99.58 to 99.6 counts)

    Availability is MTBF ÷ (MTBF + MTTR) = 240 ÷ 241 ≈ 0.99585, or 99.59%. Cutting recovery to 6 minutes would give 240 ÷ 240.1 ≈ 99.96% without making failures any rarer.

  6. Question 6 of 6 Two database replicas sit in the same rack, on one power feed, and receive configuration from the same push. Which change adds the most availability?

    Choose one answer.

    Show the answer to question 6

    Answer: Moving one replica to another zone, with its own power and network, and pushing configuration to one replica at a time

    The parallel formula assumes the replicas fail independently, and these share a rack, a power feed and a configuration push, so one event takes both down. Separating what they share removes the common causes; a third replica in the same rack fails with the other two.

Interview questions

Warm-up (fresher to mid level): two services, each 99.9% available, must both work for a request to succeed. What availability can the pair offer, and why? About 99.8%. The request needs both, so it succeeds only when both are up at once: 0.999 × 0.999 = 0.998001, assuming they fail independently. A quick way to see it is that the unavailabilities roughly add up: 0.1% + 0.1% = 0.2% of failures, about 86 minutes a month instead of 43. That is why every hard dependency lowers the ceiling of a design, and why designs remove dependencies from the critical path, add redundancy inside them, or make them soft dependencies that the service can work without for a while.

Key takeaways

  • Turn targets into downtime over a stated period: 99.9% is 43.2 minutes per 30-day month and 8.77 hours per year; each nine divides it by ten.
  • In series, availabilities multiply and unavailabilities roughly add: every hard dependency lowers the ceiling.
  • In parallel, the group fails only when every part has failed: two independent 99% replicas give 99.99%, but “any 2 of 3” and failover time give less.
  • A = MTBF ÷ (MTBF + MTTR): faster detection and recovery often buy more nines than rarer failures.
  • Shared causes (power, deploys, configuration, software, zones) make replicas fail together; separate what they share before adding more of them.

References

Related tools

Report a problem with this lesson

Quick answers and tool search

Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.