System Design (High-Level Design) Module 1 – Start here: how system design works
Quality attributes and how to trade them off
Name the quality attributes of a distributed service, see why improving one usually costs another, and record each trade-off so a reviewer can check it.
What you will learn
- Name the main quality attributes of a distributed service
- Explain why improving one attribute usually costs another
- Record a trade-off with its reason and the condition that would reverse it
Before you start
On this page
The previous lesson ended with a list of non-functional requirements, each a quality with a number attached. Those qualities are called quality attributes, and they are what high-level design is really about: almost any set of boxes can do what the functional requirements ask, but only some can do it fast enough, often enough and cheaply enough. This lesson names the attributes you meet in nearly every design, shows why they pull against each other, and gives you a way to write down which one wins.
The attributes you will meet in every design
Each attribute answers one question about the system, and each is stated as something you can measure:
- Latency: how long does one request take? A percentile in milliseconds, measured at a named point.
- Throughput: how much work can it do per second? Requests, messages or bytes per second at the peak.
- Availability: how often does a request succeed? The share of successful requests over a month.
- Durability: is confirmed data still there later? The failures a confirmed write survives.
- Consistency: does every reader see the latest write? The promise readers get, and how old a read may be.
- Scalability: can it grow by adding machines, and at what cost? The cost of each extra unit of load, and the largest step it can take before a redesign.
- Cost: what does it cost to run? Money a month, or per million requests.
- Security: who can do what, and what happens if one part is compromised? The threats handled and the data protected.
- Operability: can a small team deploy, watch and repair it? The time to deploy, to notice a failure and to recover.
- Evolvability: how hard is the next change? How many parts a typical change touches.
The large cloud providers group the same concerns into the “pillars” of their architecture frameworks. AWS has six (operational excellence, security, reliability, performance efficiency, cost optimization and sustainability), Azure five (reliability, security, cost optimization, operational excellence and performance efficiency) and Google Cloud six (operational excellence; security, privacy and compliance; reliability; cost optimization; performance optimization; and sustainability). The words differ, the concerns do not: latency and throughput sit under performance, availability and durability under reliability, operability under operational excellence.
Turn adjectives into numbers
“Highly available” becomes useful the moment it becomes a number, because the number says how much failure you are allowed to spend. This program works out that budget for common targets, and how many incidents of 20 minutes, from the first failed request to the fix, fit into one month:
"""How much unavailability an availability target allows, per 30-day month and per year, and how many
incidents of one length fit into the monthly budget."""
from decimal import ROUND_HALF_UP, Decimal
TARGETS = ["99", "99.5", "99.9", "99.95", "99.99", "99.999"]
MONTH = Decimal(30 * 24 * 60) # minutes in a 30-day month, the month availability tables use
YEAR = Decimal(365 * 24 * 60) # minutes in a 365-day year
INCIDENT = 20 # minutes from the first failed request to the fix: an assumption, change it
def budget(target, period_minutes):
"""Minutes of unavailability that `target` percent allows in a period (Decimal keeps 99.9 exact)."""
return period_minutes * (100 - Decimal(target)) / 100
def readable(minutes):
"""The duration in the largest unit that keeps it at 1 or more, to two decimals at most."""
for unit, size in (("days", 1440), ("hours", 60), ("minutes", 1)):
if minutes >= size:
value = minutes / size
break
else:
unit, value = "seconds", minutes * 60
value = value.quantize(Decimal("0.01"), rounding=ROUND_HALF_UP).normalize()
return f"{value:f} {unit}"
print(f"{'Target':<9}{'Per 30-day month':<18}{'Per year':<15}{INCIDENT}-minute incidents a month")
for target in TARGETS:
month = budget(target, MONTH)
fits = int(month // INCIDENT)
print(f"{target + '%':<9}{readable(month):<18}{readable(budget(target, YEAR)):<15}{fits}") Output
Target Per 30-day month Per year 20-minute incidents a month 99% 7.2 hours 3.65 days 21 99.5% 3.6 hours 1.83 days 10 99.9% 43.2 minutes 8.76 hours 2 99.95% 21.6 minutes 4.38 hours 1 99.99% 4.32 minutes 52.56 minutes 0 99.999% 25.92 seconds 5.26 minutes 0
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 nines.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
The budgets match the availability table in Google’s Site Reliability Engineering book, whose monthly figures also
assume a 30-day month; the program uses Python’s Decimal so that 99.9 is exactly 99.9 and not the nearest binary fraction. The
last column is the one to read:
- At 99.9% a month, two 20-minute incidents fit. A person can be woken, think and fix the problem twice.
- At 99.95%, one fits.
- At 99.99%, none does: 4.32 minutes is gone before anyone has opened a laptop. That target means automatic detection and automatic failover, and every deploy must be invisible to users. Each extra nine changes how the system is operated, not only how much hardware it has.
Two edge cases. A target of 100% is not a target: it allows no failed deploy, failover or network fault, ever, so nothing can be built to meet it. And a budget only means something for the period it is stated over: 99.9% a month allows one 43-minute outage, while 99.9% measured over each day allows only 1.44 minutes in any one day. For a service that is often partly up, such as one with a few unhealthy servers, counting failed requests says more than counting minutes of outage; the reliability module of this track measures availability that way.
Why improving one attribute usually costs another
Quality attributes are not independent sliders. The mechanism that buys one of them usually spends another, and naming the mechanism is how you explain a trade-off instead of just asserting it.
- Durability costs write latency. A database that confirms a write only once it is on disk is slower per transaction than one that confirms first. PostgreSQL makes this a choice per transaction: with asynchronous commit, transactions finish sooner, but the most recent ones can be lost if the server crashes (lost, not corrupted). Copying each write to another machine before confirming it moves the same trade-off onto the network.
- Consistency costs latency. Once data is copied, a read can go to a nearby copy that may be slightly behind, or wait until the system is sure it has the latest value. Daniel Abadi’s PACELC formulation points out that this latency-or-consistency choice applies to every replicated system all the time, not only during a network failure.
- Availability costs money. Replicas, spare capacity that waits for a failover and deployments that run old and new versions side by side all have to be paid for, as Azure’s guidance on reliability trade-offs spells out.
- Reliability costs operability. Every component added to survive failures is one more thing to deploy, monitor and understand during an incident, and running in several regions adds the work of keeping them in step.
- Security costs latency. Inspecting traffic, verifying the identity behind every call and encrypting data add work, and often network hops, to each request.
- Flexibility costs simplicity. A plugin system or a general workflow engine makes some future changes cheap and every present change dearer. Build the flexibility the requirements ask for, not the flexibility you can imagine.
The first two are easiest to see in numbers. A write is safe from the loss of a machine only once a copy exists on another machine, and the further away that copy is, the more failures it survives and the longer the write waits:
Where a write can wait before it is confirmed
Text description of the diagram
The diagram follows one write from top to bottom.
- The client asks the app server to save something.
- The app server sends the write to the database leader in zone A of region 1, which commits it to its own disk in about 1 ms.
- The leader copies the write to two replicas: one in zone B of the same region, about 1 ms away for a round trip (step 3a), and one in region 2, about 80 ms away (step 3b).
If a copy is synchronous, the client is told the write is saved only after that copy has arrived, so it waits for the round trip as well. If a copy is asynchronous, the client is told at once and the copy follows later. The times are the assumptions of the program replication_tradeoff.py in this lesson, not measurements.
The program below puts assumed numbers on that picture for three ways of making the copy. Every value at its top is an assumption for the sketch; change them and run it again.
"""Write latency against the writes at risk, for three ways of copying every write to a second machine.
Each number below is an assumption for a sketch, not a measurement: change them and run it again."""
LOCAL_COMMIT_MS = 1 # the leader writes the change to its own disk
ZONE_RTT_MS = 1 # round trip to a replica in another data centre (zone) of the same region
REGION_RTT_MS = 80 # round trip to a replica in another region, thousands of kilometres away
ASYNC_LAG_MS = 200 # how far behind the leader an asynchronous replica usually is
WRITES_PER_SECOND = 500
WRITES_PER_REQUEST = 3 # writes that one request makes, one after another
MODES = [
# how the copy is made, extra wait before a write is confirmed, where the replica is
("asynchronous", 0, "another zone"),
("synchronous", ZONE_RTT_MS, "another zone"),
("synchronous", REGION_RTT_MS, "another region"),
]
for how, wait_ms, where in MODES:
write_ms = LOCAL_COMMIT_MS + wait_ms
# An asynchronous replica lacks the writes of the last ASYNC_LAG_MS; a synchronous one lacks none.
at_risk = WRITES_PER_SECOND * ASYNC_LAG_MS // 1000 if how == "asynchronous" else 0
print(f"{how} copy to {where}")
print(f" a write is confirmed after {write_ms} ms; a request with {WRITES_PER_REQUEST} writes waits {write_ms * WRITES_PER_REQUEST} ms")
print(f" if the leader's machine dies: {f'up to {at_risk}' if at_risk else 'no'} confirmed writes lost")
print(f" if its whole region goes down: {'the replica carries on' if where == 'another region' else 'no copy is outside the region'}")
print() Output
asynchronous copy to another zone a write is confirmed after 1 ms; a request with 3 writes waits 3 ms if the leader's machine dies: up to 100 confirmed writes lost if its whole region goes down: no copy is outside the region synchronous copy to another zone a write is confirmed after 2 ms; a request with 3 writes waits 6 ms if the leader's machine dies: no confirmed writes lost if its whole region goes down: no copy is outside the region synchronous copy to another region a write is confirmed after 81 ms; a request with 3 writes waits 243 ms if the leader's machine dies: no confirmed writes lost if its whole region goes down: the replica carries on
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 replication_tradeoff.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
The asynchronous copy is the fastest and loses up to 100 confirmed writes when the leader’s machine dies, because the replica trails by 200 ms at 500 writes a second. A synchronous copy in another zone costs 1 ms more per write and loses nothing when the leader dies, but a whole-region outage still takes everything down. A synchronous copy in another region survives even that, at 81 ms per write, and a request that makes three writes one after another now waits 243 ms: the cost multiplies. One compromise is a synchronous copy inside the region and an asynchronous one to a second region: a small wait on every write, nothing lost when a machine dies, and only the last moments of writes at risk if a whole region is lost.
CAP is not “pick any two”
The CAP theorem is often summed up as “consistency, availability, partition tolerance: pick two”. Eric Brewer, who stated it, wrote that this “2 of 3” reading was always misleading. Partitions are not something a distributed system can opt out of; the real choice is what to give up while one lasts, it can differ from one operation to the next, and while there is no partition nothing forces you to give up either consistency or availability. The consistency module of this track treats CAP and PACELC properly.
A worked trade-off: where to keep sessions
A logged-in user’s session has to live somewhere between requests, and the three usual places trade the attributes above against each other:
- In the app server’s memory. Nothing extra happens per request, and logging a user out is instant. But the sessions on a server vanish when it restarts or is redeployed, and the load balancer must send each user back to the same server (“sticky sessions”), so servers cannot be added or removed freely. Choose it for one server or a prototype.
- In a shared store with expiry. Sessions survive restarts and deploys, any server can serve any request, and logging out means deleting one entry. The price is one round trip to the store on every request, and the store is one more component that must stay up. Choose it for most services with more than one server.
- In a signed token in a cookie. Any server can check the token without a lookup, and nothing is kept on the servers. But the token travels with every request, browsers need only keep about 4 KB per cookie, and a token cannot be taken back before it expires unless the servers check a list of revoked tokens, which brings the lookup back. Choose it when many servers or regions make a lookup per request too costly and short-lived tokens are acceptable.
The Twelve-Factor App, a short methodology for building web services, calls sticky sessions a violation of its rules and suggests a store with time-based expiry instead; that is the second option, and it is why the app servers in later lessons are stateless. The cookie limit comes from the HTTP cookie standard, RFC 6265, which asks browsers to support cookies of at least 4,096 bytes; a signed token such as a JSON Web Token (RFC 7519) carries its claims with it, which is what lets any server check it without a lookup, and also why it cannot simply be taken back.
Write the trade-off down
A trade-off that lives only in someone’s head is lost when they leave, and one written as “we chose X because it is good” cannot be checked. Write each one as a sentence with four parts:
We chose X over Y because Z; we will revisit this if W.
- Weak: “We chose PostgreSQL because it is reliable.” There is no alternative, the reason is an adjective, and nothing says when the choice would stop being right.
- Strong: “We chose a synchronous replica in a second zone over asynchronous replication, because a confirmed link must survive the loss of a data centre and 1 ms more per write fits the latency budget; we will revisit this if we add a second region or the write rate grows tenfold.”
The revisit condition, W, is the part people leave out, and the most useful one: it names the assumption that would flip the decision, so a reader years later knows whether the decision still holds. The architecture decision records community uses a similar one-sentence form, the Y-statement: “In the context of a use case, facing a concern, we decided for an option to achieve a quality, accepting a downside.” The lesson on communicating designs shows how to keep such decisions as records next to the code.
Time Unit Converter Turn a downtime budget such as 43.2 minutes into seconds, hours or days.Key takeaways
- The attributes that recur are latency, throughput, availability, durability, consistency, scalability, cost, security, operability and evolvability; the cloud frameworks group them into pillars with different names.
- An attribute is useful only as a number: 99.9% a month is 43.2 minutes, and each extra nine changes how the system must be operated.
- Improving one attribute usually spends another through a mechanism you can name: waiting for a copy, reading a nearby replica, paying for spare capacity, adding parts to run.
- CAP is not “pick any two”; the everyday trade-off of replicated data is latency against consistency.
- Write every trade-off as “we chose X over Y because Z; we will revisit this if W”.
Exercise
Exercise · Easy · Python
Turn an availability target into a downtime budget
An availability target only becomes useful when you know how much failure it allows. Write two functions in budget.py.
downtime_budget(target_percent, days=30) returns the minutes of unavailability that the target allows in a period of days days, rounded to two decimals. A 30-day month has 43,200 minutes, so 99.9% allows 0.1% of them:
downtime_budget(99.9) # 43.2
downtime_budget(99.99, 365) # 52.56
A target must be greater than 0 and less than 100, and the period at least one day; anything else raises ValueError (100% would allow no failure at all, ever, which no real system can promise).
incidents_that_fit(target_percent, incident_minutes, days=30) returns how many whole incidents of incident_minutes minutes fit into that budget, as an int: at 99.9% a month, two 20-minute incidents fit and a third would break the target. An incident of 0 minutes or less raises ValueError.
The sample tests import both functions from budget.py and run in your browser.
Starter code · budget.py
def downtime_budget(target_percent, days=30):
"""Minutes of unavailability that the target allows in `days` days, rounded to 2 decimals."""
# Replace this line with your code.
return 0.0
def incidents_that_fit(target_percent, incident_minutes, days=30):
"""How many whole incidents of `incident_minutes` minutes fit into that budget."""
# Replace this line with your code.
return 0 The sample tests · test_budget.py
from budget import downtime_budget, incidents_that_fit
def raises_value_error(call):
try:
call()
except ValueError:
return True
return False
def test_monthly_budget():
"""99.9% allows 43.2 minutes in a 30-day month"""
assert downtime_budget(99.9) == 43.2
assert downtime_budget(99.95, 30) == 21.6
def test_yearly_budget():
"""99.99% allows 52.56 minutes in a 365-day year"""
assert downtime_budget(99.99, 365) == 52.56
def test_rejects_impossible_targets():
"""100%, 0% and a period of no days raise ValueError"""
assert raises_value_error(lambda: downtime_budget(100))
assert raises_value_error(lambda: downtime_budget(0))
assert raises_value_error(lambda: downtime_budget(99.9, 0))
def test_whole_incidents():
"""counts whole incidents only"""
assert incidents_that_fit(99.9, 20) == 2
assert incidents_that_fit(99.5, 30) == 7
assert incidents_that_fit(99.99, 20) == 0
def test_incident_length():
"""an incident of no minutes raises ValueError"""
assert raises_value_error(lambda: incidents_that_fit(99.9, 0)) A hint
The budget is the period in minutes (days * 24 * 60) times the share of it that may fail, (100 - target_percent) / 100. Check the arguments first and raise ValueError(...) before computing anything. For whole incidents, divide the budget by the incident length with // and turn the result into an int.
Results of the sample tests
| Test | Result | Details |
|---|
What your code printed
The sample tests run on this device, in your browser (Pyodide): nothing is sent to mysmartcopilot.com. The first run downloads Python (about 13.5 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.
Check yourself
7 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.
References
- AWS Well-Architected Framework: the pillars of the framework (Amazon Web Services)
- Azure Well-Architected Framework pillars (Microsoft)
- Reliability tradeoffs (Azure Well-Architected Framework) (Microsoft)
- Security tradeoffs (Azure Well-Architected Framework) (Microsoft)
- Google Cloud Well-Architected Framework (Google Cloud)
- Site Reliability Engineering, appendix A: Availability Table (Google (O'Reilly Media))
- PostgreSQL 18 documentation: 28.4 Asynchronous Commit (The PostgreSQL Global Development Group)
- Consistency Tradeoffs in Modern Distributed Database System Design (Daniel J. Abadi) (IEEE Computer (author's copy))
- CAP Twelve Years Later: How the "Rules" Have Changed (Eric Brewer) (InfoQ, first published in IEEE Computer)
- The Twelve-Factor App: VI. Processes (The Twelve-Factor App (Adam Wiggins))
- RFC 6265: HTTP State Management Mechanism (section 6.1, limits) (IETF)
- RFC 7519: JSON Web Token (JWT) (IETF)
- ADR templates (Nygard ADRs, MADR and Y-statements) (ADR GitHub organization)
Related tools
Report a problem with this lesson
Kept only in this browser. Your Learn progress