Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

System Design (High-Level Design) Module 1 – Start here: how system design works

Quality attributes and how to trade them off

Name the quality attributes of a distributed service, see why improving one usually costs another, and record each trade-off so a reviewer can check it.

  • Beginner
  • 25 minutes
  • Examples run with Python 3.14.8, Pyodide 314.0.7, Node.js 24.21.0 and quickjs 0.32.0
  • By MySmartCoPilot

What you will learn

  • Name the main quality attributes of a distributed service
  • Explain why improving one attribute usually costs another
  • Record a trade-off with its reason and the condition that would reverse it

Before you start

On this page

The previous lesson ended with a list of non-functional requirements, each a quality with a number attached. Those qualities are called quality attributes, and they are what high-level design is really about: almost any set of boxes can do what the functional requirements ask, but only some can do it fast enough, often enough and cheaply enough. This lesson names the attributes you meet in nearly every design, shows why they pull against each other, and gives you a way to write down which one wins.

The attributes you will meet in every design

Each attribute answers one question about the system, and each is stated as something you can measure:

  • Latency: how long does one request take? A percentile in milliseconds, measured at a named point.
  • Throughput: how much work can it do per second? Requests, messages or bytes per second at the peak.
  • Availability: how often does a request succeed? The share of successful requests over a month.
  • Durability: is confirmed data still there later? The failures a confirmed write survives.
  • Consistency: does every reader see the latest write? The promise readers get, and how old a read may be.
  • Scalability: can it grow by adding machines, and at what cost? The cost of each extra unit of load, and the largest step it can take before a redesign.
  • Cost: what does it cost to run? Money a month, or per million requests.
  • Security: who can do what, and what happens if one part is compromised? The threats handled and the data protected.
  • Operability: can a small team deploy, watch and repair it? The time to deploy, to notice a failure and to recover.
  • Evolvability: how hard is the next change? How many parts a typical change touches.

The large cloud providers group the same concerns into the “pillars” of their architecture frameworks. AWS has six (operational excellence, security, reliability, performance efficiency, cost optimization and sustainability), Azure five (reliability, security, cost optimization, operational excellence and performance efficiency) and Google Cloud six (operational excellence; security, privacy and compliance; reliability; cost optimization; performance optimization; and sustainability). The words differ, the concerns do not: latency and throughput sit under performance, availability and durability under reliability, operability under operational excellence.

Turn adjectives into numbers

“Highly available” becomes useful the moment it becomes a number, because the number says how much failure you are allowed to spend. This program works out that budget for common targets, and how many incidents of 20 minutes, from the first failed request to the fix, fit into one month:

Downtime budgets for common availability targets Python · nines.py
"""How much unavailability an availability target allows, per 30-day month and per year, and how many
incidents of one length fit into the monthly budget."""

from decimal import ROUND_HALF_UP, Decimal

TARGETS = ["99", "99.5", "99.9", "99.95", "99.99", "99.999"]
MONTH = Decimal(30 * 24 * 60)  # minutes in a 30-day month, the month availability tables use
YEAR = Decimal(365 * 24 * 60)  # minutes in a 365-day year
INCIDENT = 20  # minutes from the first failed request to the fix: an assumption, change it


def budget(target, period_minutes):
    """Minutes of unavailability that `target` percent allows in a period (Decimal keeps 99.9 exact)."""
    return period_minutes * (100 - Decimal(target)) / 100


def readable(minutes):
    """The duration in the largest unit that keeps it at 1 or more, to two decimals at most."""
    for unit, size in (("days", 1440), ("hours", 60), ("minutes", 1)):
        if minutes >= size:
            value = minutes / size
            break
    else:
        unit, value = "seconds", minutes * 60
    value = value.quantize(Decimal("0.01"), rounding=ROUND_HALF_UP).normalize()
    return f"{value:f} {unit}"


print(f"{'Target':<9}{'Per 30-day month':<18}{'Per year':<15}{INCIDENT}-minute incidents a month")
for target in TARGETS:
    month = budget(target, MONTH)
    fits = int(month // INCIDENT)
    print(f"{target + '%':<9}{readable(month):<18}{readable(budget(target, YEAR)):<15}{fits}")

Output

Target   Per 30-day month  Per year       20-minute incidents a month
99%      7.2 hours         3.65 days      21
99.5%    3.6 hours         1.83 days      10
99.9%    43.2 minutes      8.76 hours     2
99.95%   21.6 minutes      4.38 hours     1
99.99%   4.32 minutes      52.56 minutes  0
99.999%  25.92 seconds     5.26 minutes   0

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 nines.py

The budgets match the availability table in Google’s Site Reliability Engineering book, whose monthly figures also assume a 30-day month; the program uses Python’s Decimal so that 99.9 is exactly 99.9 and not the nearest binary fraction. The last column is the one to read:

  • At 99.9% a month, two 20-minute incidents fit. A person can be woken, think and fix the problem twice.
  • At 99.95%, one fits.
  • At 99.99%, none does: 4.32 minutes is gone before anyone has opened a laptop. That target means automatic detection and automatic failover, and every deploy must be invisible to users. Each extra nine changes how the system is operated, not only how much hardware it has.

Two edge cases. A target of 100% is not a target: it allows no failed deploy, failover or network fault, ever, so nothing can be built to meet it. And a budget only means something for the period it is stated over: 99.9% a month allows one 43-minute outage, while 99.9% measured over each day allows only 1.44 minutes in any one day. For a service that is often partly up, such as one with a few unhealthy servers, counting failed requests says more than counting minutes of outage; the reliability module of this track measures availability that way.

Why improving one attribute usually costs another

Quality attributes are not independent sliders. The mechanism that buys one of them usually spends another, and naming the mechanism is how you explain a trade-off instead of just asserting it.

  • Durability costs write latency. A database that confirms a write only once it is on disk is slower per transaction than one that confirms first. PostgreSQL makes this a choice per transaction: with asynchronous commit, transactions finish sooner, but the most recent ones can be lost if the server crashes (lost, not corrupted). Copying each write to another machine before confirming it moves the same trade-off onto the network.
  • Consistency costs latency. Once data is copied, a read can go to a nearby copy that may be slightly behind, or wait until the system is sure it has the latest value. Daniel Abadi’s PACELC formulation points out that this latency-or-consistency choice applies to every replicated system all the time, not only during a network failure.
  • Availability costs money. Replicas, spare capacity that waits for a failover and deployments that run old and new versions side by side all have to be paid for, as Azure’s guidance on reliability trade-offs spells out.
  • Reliability costs operability. Every component added to survive failures is one more thing to deploy, monitor and understand during an incident, and running in several regions adds the work of keeping them in step.
  • Security costs latency. Inspecting traffic, verifying the identity behind every call and encrypting data add work, and often network hops, to each request.
  • Flexibility costs simplicity. A plugin system or a general workflow engine makes some future changes cheap and every present change dearer. Build the flexibility the requirements ask for, not the flexibility you can imagine.

The first two are easiest to see in numbers. A write is safe from the loss of a machine only once a copy exists on another machine, and the further away that copy is, the more failures it survives and the longer the write waits:

A write goes from the client to an app server and the database leader, which copies it to a replica nearby and to one far away.ClientApp serverDatabase leaderzone A, region 1Replicazone B, region 1Replicaregion 21. save this2. commit: 1 ms3a. copy, 1 ms3b. copy, 80 ms

Where a write can wait before it is confirmed

Text description of the diagram

The diagram follows one write from top to bottom.

  1. The client asks the app server to save something.
  2. The app server sends the write to the database leader in zone A of region 1, which commits it to its own disk in about 1 ms.
  3. The leader copies the write to two replicas: one in zone B of the same region, about 1 ms away for a round trip (step 3a), and one in region 2, about 80 ms away (step 3b).

If a copy is synchronous, the client is told the write is saved only after that copy has arrived, so it waits for the round trip as well. If a copy is asynchronous, the client is told at once and the copy follows later. The times are the assumptions of the program replication_tradeoff.py in this lesson, not measurements.

The program below puts assumed numbers on that picture for three ways of making the copy. Every value at its top is an assumption for the sketch; change them and run it again.

Three ways to copy a write, and what each costs Python · replication_tradeoff.py
"""Write latency against the writes at risk, for three ways of copying every write to a second machine.
Each number below is an assumption for a sketch, not a measurement: change them and run it again."""

LOCAL_COMMIT_MS = 1  # the leader writes the change to its own disk
ZONE_RTT_MS = 1  # round trip to a replica in another data centre (zone) of the same region
REGION_RTT_MS = 80  # round trip to a replica in another region, thousands of kilometres away
ASYNC_LAG_MS = 200  # how far behind the leader an asynchronous replica usually is
WRITES_PER_SECOND = 500
WRITES_PER_REQUEST = 3  # writes that one request makes, one after another

MODES = [
    # how the copy is made, extra wait before a write is confirmed, where the replica is
    ("asynchronous", 0, "another zone"),
    ("synchronous", ZONE_RTT_MS, "another zone"),
    ("synchronous", REGION_RTT_MS, "another region"),
]

for how, wait_ms, where in MODES:
    write_ms = LOCAL_COMMIT_MS + wait_ms
    # An asynchronous replica lacks the writes of the last ASYNC_LAG_MS; a synchronous one lacks none.
    at_risk = WRITES_PER_SECOND * ASYNC_LAG_MS // 1000 if how == "asynchronous" else 0
    print(f"{how} copy to {where}")
    print(f"  a write is confirmed after {write_ms} ms; a request with {WRITES_PER_REQUEST} writes waits {write_ms * WRITES_PER_REQUEST} ms")
    print(f"  if the leader's machine dies: {f'up to {at_risk}' if at_risk else 'no'} confirmed writes lost")
    print(f"  if its whole region goes down: {'the replica carries on' if where == 'another region' else 'no copy is outside the region'}")
    print()

Output

asynchronous copy to another zone
  a write is confirmed after 1 ms; a request with 3 writes waits 3 ms
  if the leader's machine dies: up to 100 confirmed writes lost
  if its whole region goes down: no copy is outside the region

synchronous copy to another zone
  a write is confirmed after 2 ms; a request with 3 writes waits 6 ms
  if the leader's machine dies: no confirmed writes lost
  if its whole region goes down: no copy is outside the region

synchronous copy to another region
  a write is confirmed after 81 ms; a request with 3 writes waits 243 ms
  if the leader's machine dies: no confirmed writes lost
  if its whole region goes down: the replica carries on

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 replication_tradeoff.py

The asynchronous copy is the fastest and loses up to 100 confirmed writes when the leader’s machine dies, because the replica trails by 200 ms at 500 writes a second. A synchronous copy in another zone costs 1 ms more per write and loses nothing when the leader dies, but a whole-region outage still takes everything down. A synchronous copy in another region survives even that, at 81 ms per write, and a request that makes three writes one after another now waits 243 ms: the cost multiplies. One compromise is a synchronous copy inside the region and an asynchronous one to a second region: a small wait on every write, nothing lost when a machine dies, and only the last moments of writes at risk if a whole region is lost.

CAP is not “pick any two”

The CAP theorem is often summed up as “consistency, availability, partition tolerance: pick two”. Eric Brewer, who stated it, wrote that this “2 of 3” reading was always misleading. Partitions are not something a distributed system can opt out of; the real choice is what to give up while one lasts, it can differ from one operation to the next, and while there is no partition nothing forces you to give up either consistency or availability. The consistency module of this track treats CAP and PACELC properly.

A worked trade-off: where to keep sessions

A logged-in user’s session has to live somewhere between requests, and the three usual places trade the attributes above against each other:

  • In the app server’s memory. Nothing extra happens per request, and logging a user out is instant. But the sessions on a server vanish when it restarts or is redeployed, and the load balancer must send each user back to the same server (“sticky sessions”), so servers cannot be added or removed freely. Choose it for one server or a prototype.
  • In a shared store with expiry. Sessions survive restarts and deploys, any server can serve any request, and logging out means deleting one entry. The price is one round trip to the store on every request, and the store is one more component that must stay up. Choose it for most services with more than one server.
  • In a signed token in a cookie. Any server can check the token without a lookup, and nothing is kept on the servers. But the token travels with every request, browsers need only keep about 4 KB per cookie, and a token cannot be taken back before it expires unless the servers check a list of revoked tokens, which brings the lookup back. Choose it when many servers or regions make a lookup per request too costly and short-lived tokens are acceptable.

The Twelve-Factor App, a short methodology for building web services, calls sticky sessions a violation of its rules and suggests a store with time-based expiry instead; that is the second option, and it is why the app servers in later lessons are stateless. The cookie limit comes from the HTTP cookie standard, RFC 6265, which asks browsers to support cookies of at least 4,096 bytes; a signed token such as a JSON Web Token (RFC 7519) carries its claims with it, which is what lets any server check it without a lookup, and also why it cannot simply be taken back.

Write the trade-off down

A trade-off that lives only in someone’s head is lost when they leave, and one written as “we chose X because it is good” cannot be checked. Write each one as a sentence with four parts:

We chose X over Y because Z; we will revisit this if W.

  • Weak: “We chose PostgreSQL because it is reliable.” There is no alternative, the reason is an adjective, and nothing says when the choice would stop being right.
  • Strong: “We chose a synchronous replica in a second zone over asynchronous replication, because a confirmed link must survive the loss of a data centre and 1 ms more per write fits the latency budget; we will revisit this if we add a second region or the write rate grows tenfold.”

The revisit condition, W, is the part people leave out, and the most useful one: it names the assumption that would flip the decision, so a reader years later knows whether the decision still holds. The architecture decision records community uses a similar one-sentence form, the Y-statement: “In the context of a use case, facing a concern, we decided for an option to achieve a quality, accepting a downside.” The lesson on communicating designs shows how to keep such decisions as records next to the code.

Time Unit Converter Turn a downtime budget such as 43.2 minutes into seconds, hours or days.

Key takeaways

  • The attributes that recur are latency, throughput, availability, durability, consistency, scalability, cost, security, operability and evolvability; the cloud frameworks group them into pillars with different names.
  • An attribute is useful only as a number: 99.9% a month is 43.2 minutes, and each extra nine changes how the system must be operated.
  • Improving one attribute usually spends another through a mechanism you can name: waiting for a copy, reading a nearby replica, paying for spare capacity, adding parts to run.
  • CAP is not “pick any two”; the everyday trade-off of replicated data is latency against consistency.
  • Write every trade-off as “we chose X over Y because Z; we will revisit this if W”.

Exercise

Exercise · Easy · Python

Turn an availability target into a downtime budget

An availability target only becomes useful when you know how much failure it allows. Write two functions in budget.py.

downtime_budget(target_percent, days=30) returns the minutes of unavailability that the target allows in a period of days days, rounded to two decimals. A 30-day month has 43,200 minutes, so 99.9% allows 0.1% of them:

downtime_budget(99.9)        # 43.2
downtime_budget(99.99, 365)  # 52.56

A target must be greater than 0 and less than 100, and the period at least one day; anything else raises ValueError (100% would allow no failure at all, ever, which no real system can promise).

incidents_that_fit(target_percent, incident_minutes, days=30) returns how many whole incidents of incident_minutes minutes fit into that budget, as an int: at 99.9% a month, two 20-minute incidents fit and a third would break the target. An incident of 0 minutes or less raises ValueError.

The sample tests import both functions from budget.py and run in your browser.

Starter code · budget.py

def downtime_budget(target_percent, days=30):
    """Minutes of unavailability that the target allows in `days` days, rounded to 2 decimals."""
    # Replace this line with your code.
    return 0.0


def incidents_that_fit(target_percent, incident_minutes, days=30):
    """How many whole incidents of `incident_minutes` minutes fit into that budget."""
    # Replace this line with your code.
    return 0
The sample tests · test_budget.py
from budget import downtime_budget, incidents_that_fit


def raises_value_error(call):
    try:
        call()
    except ValueError:
        return True
    return False


def test_monthly_budget():
    """99.9% allows 43.2 minutes in a 30-day month"""
    assert downtime_budget(99.9) == 43.2
    assert downtime_budget(99.95, 30) == 21.6


def test_yearly_budget():
    """99.99% allows 52.56 minutes in a 365-day year"""
    assert downtime_budget(99.99, 365) == 52.56


def test_rejects_impossible_targets():
    """100%, 0% and a period of no days raise ValueError"""
    assert raises_value_error(lambda: downtime_budget(100))
    assert raises_value_error(lambda: downtime_budget(0))
    assert raises_value_error(lambda: downtime_budget(99.9, 0))


def test_whole_incidents():
    """counts whole incidents only"""
    assert incidents_that_fit(99.9, 20) == 2
    assert incidents_that_fit(99.5, 30) == 7
    assert incidents_that_fit(99.99, 20) == 0


def test_incident_length():
    """an incident of no minutes raises ValueError"""
    assert raises_value_error(lambda: incidents_that_fit(99.9, 0))
A hint

The budget is the period in minutes (days * 24 * 60) times the share of it that may fail, (100 - target_percent) / 100. Check the arguments first and raise ValueError(...) before computing anything. For whole incidents, divide the budget by the incident length with // and turn the result into an int.

The sample tests run on this device, in your browser (Pyodide): nothing is sent to mysmartcopilot.com. The first run downloads Python (about 13.5 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.

Check yourself

7 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.

  1. Question 1 of 7 A payments team writes: "a payment the API has confirmed is still there after any one data centre is lost". Which quality attribute is that requirement about?

    Choose one answer.

    Show the answer to question 1

    Answer: Durability

    Durability asks whether confirmed data is still there later, and the requirement names the failure a confirmed write must survive. Availability would be about how often a request succeeds, consistency about how fresh a read is, and throughput about how much work per second the system can do.

  2. Question 2 of 7 How many minutes of unavailability does a target of 99.95% allow in a 30-day month?

    Type a number.

    Show the answer to question 2

    Answer: 21.6 minutes (anything from 21.55 to 21.65 counts)

    A 30-day month has 30 × 24 × 60 = 43,200 minutes, and 99.95% leaves 0.05% of them: 43,200 × 0.0005 = 21.6 minutes. That is one 20-minute incident and very little else.

  3. Question 3 of 7 Which change makes writes faster but can lose the most recently confirmed writes if the database server crashes?

    Choose one answer.

    Show the answer to question 3

    Answer: Confirming a transaction before its log records are flushed to disk (asynchronous commit)

    Asynchronous commit tells the client "done" before the change is on disk, so a crash in that short window loses the last few transactions. Waiting for a replica does the opposite: it makes writes slower and safer.

  4. Question 4 of 7 A team makes every write wait for a replica in another region before it is confirmed. What do they gain, and what do they pay?

    Choose every answer that is right.

    Show the answer to question 4

    Answer:

    • Every write waits at least one round trip between the regions
    • A confirmed write survives the loss of a whole region

    The far replica is what lets confirmed data survive a regional outage, and the wait for it is paid on every write, multiplied when a request makes several writes in a row. Reads, running costs and the second region's capacity are separate decisions with their own costs.

  5. Question 5 of 7 Which trade-off statement could a reviewer actually check?

    Choose one answer.

    Show the answer to question 5

    Answer: We chose a cache in front of the database over more read replicas, because 95% of reads hit 1% of the links and staleness of 60 seconds is allowed; we will revisit this if links become editable.

    The statement about read replicas names the alternative, ties the reason to numbers in the requirements and says what would change the decision. The others give no alternative, no evidence, or no decision at all.

  6. Question 6 of 7 Why is "consistency, availability, partition tolerance, pick any two" a misleading summary of CAP?

    Choose one answer.

    Show the answer to question 6

    Answer: Partitions cannot be ruled out, so the real choice is what to give up while one lasts; without a partition nothing forces you to give up either

    A distributed system cannot opt out of network partitions. CAP describes the choice between consistency and availability during one, and that choice can differ by operation. The everyday trade-off of replicated data, when the network is fine, is latency against consistency.

  7. Question 7 of 7 A service runs on many identical app servers behind a load balancer, and users must be logged out at once when they change their password. Where should sessions live?

    Choose one answer.

    Show the answer to question 7

    Answer: In a shared store with expiry, which every server can read and from which a session can be deleted

    A shared store lets any server serve any request and lets the service delete a session immediately. Server memory breaks when that server restarts and ties users to one server; a long-lived token cannot be taken back before it expires; unsigned data in the browser can be changed by anyone.

References

Related tools

Report a problem with this lesson

Quick answers and tool search

Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.