Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

System Design (High-Level Design) Module 1 – Start here: how system design works

Diagrams, design docs and decision records

Draw context and container diagrams a reviewer can read in a minute, show one request as a sequence diagram, and record each decision as an ADR.

  • Beginner
  • 25 minutes
  • Examples run with Python 3.14.8, Pyodide 314.0.7, Node.js 24.21.0 and quickjs 0.32.0
  • By MySmartCoPilot

What you will learn

  • Draw a system context and container diagram that a reviewer can read in one minute
  • Use a sequence diagram for one critical request flow
  • Write an architecture decision record with context, decision and consequences

Before you start

On this page

A design nobody else can read cannot be reviewed. In a design round, a diagram the interviewer cannot follow is a design they cannot assess; at work, a decision nobody wrote down gets argued again every time someone new joins. Three artefacts carry most of the weight: a diagram of the structure, a sequence diagram of the one request that matters most, and a record of each decision and why it was taken.

Boxes, arrows and labels

The C4 model prescribes no notation, but the advice on its notation page makes a good checklist for any boxes-and-arrows diagram, whatever tool draws it:

  • Every box says what it is. A name, the kind of thing it is (a person, an outside system, an application, a data store), a short line about its responsibility and, for applications and stores, the technology.
  • Data stores look different from services. Shapes are your choice, and C4 asks only that the key explain the ones that carry meaning. This track draws a cylinder for a database and a queue shape for a queue, so a reader sees at a glance where data rests.
  • Every arrow points one way and says what it does. “Reads and writes orders” says something; “uses” says nothing. Between applications and stores, add the protocol: “JSON over HTTPS”, “SQL over TLS”.
  • A title and a key. Say what the diagram shows and at which level of detail, and explain any shape, line style or colour that carries meaning. Spell out abbreviations or put them in the key.
  • Colour is optional. If it means something, use it the same way everywhere and never as the only signal: some readers do not see colours apart, and printouts are often black and white.

Levels of detail: context, container, component

The C4 model describes a system at four levels of zoom, and each answers a different question:

  1. System context: your system as one box, with the people who use it and the outside systems it talks to. It is the picture you can show anyone, technical or not.
  2. Container: inside the system boundary, the separately running applications and data stores, their technology and how they talk to each other. “Container” here has nothing to do with Docker: it means a web app, a service, a database, a queue, a bucket.
  3. Component: the parts inside one container.
  4. Code: classes and functions, which are rarely worth drawing by hand.

A design round mostly needs the container level plus one sequence diagram. One difference from C4 is worth knowing: C4 keeps load balancers, replicas and failover out of container diagrams and puts them in a separate deployment diagram, because they differ from one environment to the next. In a round you usually draw them on the same diagram, which is fine as long as each box is labelled for what it is. Here is a container diagram of an online shop’s checkout:

Container diagram of a shop checkout: a shopper, five parts of the shop with labelled arrows, and two outside providers.Shopper[person]Shop (ours)Not oursWeb app[single-page app in the browser]basket and checkout pagesCheckout API[stateless service]places and charges ordersOrders database[relational database]orders and their statusOrder events[message queue]Receipt worker[background service]emails receiptsPayment providercharges cardsEmail providerdelivers emailplaces an orderJSON over HTTPSreads and writes ordersSQL over TLSpublishes order-paidAMQP over TLSdelivers order-paidAMQP over TLSpays for a basketcharges the cardHTTPSsends the receiptHTTPS

Container diagram: the checkout of an online shop

Text description of the diagram

At the top is the shopper, a person, who pays for a basket in the web app.

The large box in the middle, labelled "Shop (ours)", holds the five parts of the system being designed, each with its kind in square brackets and one line about what it does: - the web app, a single-page app in the browser with the basket and checkout pages, which places an order with the Checkout API, sending JSON over HTTPS; - the Checkout API, a stateless service that places and charges orders; it reads and writes orders in the orders database using SQL over TLS, and publishes an order-paid event to the order events queue over AMQP over TLS; - the orders database, a relational database drawn as a cylinder, holding orders and their status; - order events, a message queue drawn as a horizontal cylinder, which delivers order-paid events to the receipt worker over AMQP over TLS; - the receipt worker, a background service that emails receipts.

At the bottom, a box labelled "Not ours" holds two outside systems: a payment provider, which the Checkout API calls over HTTPS to charge the card, and an email provider, which the receipt worker calls over HTTPS to send the receipt.

Try the one-minute test on it. Who uses it? A shopper. What runs? A web app, an API, a background worker. Where does data rest? An orders database and a queue. What is outside the team’s control? The payment and email providers. What crosses the network, and how? Every arrow says. A reviewer who can answer those five questions in a minute can spend the rest of the review on whether the design is right.

One request as a sequence diagram

A container diagram shows what exists. A sequence diagram shows what happens, in order, for one request: who calls whom, what is stored when, and what the caller hears back. C4 calls diagrams of behaviour “dynamic diagrams” and allows them in this sequence style. Draw one for the request that matters most, usually the hot path or the one that moves money:

Sequence diagram of one checkout: seven numbered messages between the web app, Checkout API, orders database, payments and an events queue.Web appCheckout APIOrders databasePayment providerOrder events queue1. place order k72. save k7 as pending3. charge the card, key k74. approved5. mark k7 as paid6. publish order-paid7. 201 Created

Sequence diagram: one checkout request, in order

Text description of the diagram

Five participants stand in columns, from left to right: the web app, the Checkout API, the orders database, the payment provider and the order events queue. Time runs from top to bottom, and the messages are numbered in order.

  1. The web app asks the Checkout API to place order k7; k7 is the order's idempotency key.
  2. The Checkout API saves order k7 in the orders database with the status pending.
  3. The Checkout API asks the payment provider to charge the card, passing the same key, k7.
  4. The payment provider answers that the charge is approved.
  5. The Checkout API marks order k7 as paid in the orders database.
  6. The Checkout API publishes an order-paid event to the order events queue, which the receipt worker reads later.
  7. The Checkout API answers the web app with 201 Created.

The order of the steps is the design. The order is saved as pending before the card is charged, so if the API crashes between steps 3 and 5 there is a record to reconcile with the payment provider instead of a charge nobody knows about. The same idempotency key, k7, travels with the order and the charge, so a retry after a timeout finds the existing order rather than charging twice; the APIs module of this track shows how idempotency keys work. A container diagram cannot say any of this, which is why the two diagrams go together.

Keep diagrams as text, next to the code

A diagram can be a short text file that a tool turns into an image. In the D2 language, for example, the sequence diagram above takes seven lines of messages, and the order in which messages are written is the order in which they are drawn. Diagrams kept as text sit in the repository next to the code, change in the same pull request as the design they describe, and show a reviewable difference when they change.

Spot the problems

Before reading on, look at this diagram and list what you would change. The first question of the quiz at the end of this lesson asks for the problems.

A small diagram for the quiz: six plain boxes from Mobile app to Orders, joined by five arrows, some of them labelled.Mobile appLBAppCacheOrdersUserService.validate()usesdatacalls

For the quiz: what would you change in this diagram?

Text description of the diagram

Six boxes, all plain rectangles of the same style, are labelled Mobile app, LB, App, Cache, Orders and UserService.validate().

The arrows are: - from Mobile app to LB, with no label; - from LB to App, labelled "uses"; - between App and Cache, a line with an arrowhead at both ends, labelled "data"; - from App to UserService.validate(), labelled "calls"; - from App to Orders, with no label.

Architecture decision records

An architecture decision record (ADR) is a short text file that records one decision. Michael Nygard’s original format, from the blog post that made the idea popular, has five parts:

  • Title: a number and a short phrase that names the decision.
  • Status: proposed, accepted, deprecated, or superseded by a later record.
  • Context: the forces at play, stated as facts: requirements, constraints, what the team knows.
  • Decision: what the team will do, in full sentences (“We will …”).
  • Consequences: everything that follows from the decision, the bad and the neutral as well as the good.

Records are numbered in order and numbers are never reused. A decision that is reversed is not edited away: its record stays, marked superseded, and a new record explains the new decision, so the history of the design survives. Each record is short, one or two pages at most. The MADR template, from the architecture decision records community, adds a list of the options that were considered with their pros and cons, which is often the most useful part for a later reader.

Here are two records for the link shortener of the earlier lessons, and a checker that looks for the parts above:

Two decision records and a checker

adr/adr_check.py

"""Check architecture decision records: a numbered title, the four sections of Michael Nygard's format,
a known status, and more than one consequence."""

import re
from pathlib import Path

REQUIRED = ["Status", "Context", "Decision", "Consequences"]
STATUSES = ("Proposed", "Accepted", "Deprecated", "Superseded")


def sections(text):
    """The lines under each '## ' heading, keyed by the heading."""
    found, current = {}, None
    for line in text.splitlines():
        if line.startswith("## "):
            current = line[3:].strip()
            found[current] = []
        elif current is not None and line.strip():
            found[current].append(line.strip())
    return found


def check(path):
    text = path.read_text(encoding="utf-8")
    title = re.match(r"# (\d+)\. (.+)", text)
    found = sections(text)
    status = (found.get("Status") or [""])[0]
    consequences = [line for line in found.get("Consequences", []) if line.startswith("- ")]

    problems = []
    if not title:
        problems.append('the title is not "# <number>. <decision>"')
    missing = [name for name in REQUIRED if name not in found]
    if missing:
        problems.append(f"missing {', '.join(missing)}")
    if status and not status.startswith(STATUSES):
        problems.append(f'status "{status}" is not one of {", ".join(STATUSES)}')
    if "Consequences" in found and len(consequences) < 2:
        problems.append("only one consequence: list the bad ones too")

    print(path.name)
    if title:
        print(f"  ADR {int(title.group(1))}: {title.group(2)}")
    print(f"  sections: {', '.join(found)}")
    print(f"  consequences: {len(consequences)}")
    print("  " + ("complete" if not problems else "needs work: " + "; ".join(problems)))


for name in sorted(p.name for p in Path(".").glob("[0-9][0-9][0-9][0-9]-*.md")):
    check(Path(name))

Output

0001-record-clicks-in-a-log.md
  ADR 1: Record clicks in a durable log, not in the links database
  sections: Status, Context, Options considered, Decision, Consequences
  consequences: 5
  complete
0002-cache-redirects.md
  ADR 2: Cache redirects for one hour
  sections: Status, Context, Decision
  consequences: 0
  needs work: missing Consequences; status "Draft" is not one of Proposed, Accepted, Deprecated, Superseded

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 adr_check.py

adr/0001-record-clicks-in-a-log.md

# 1. Record clicks in a durable log, not in the links database

## Status

Accepted

## Context

Every redirect of the link shortener is a click, and the owner of a link sees how many clicks it has had. The
requirements allow a count to be up to 60 seconds behind, and losing a few clicks when a server crashes is
acceptable; losing links is not. Redirects outnumber new links about 100 to 1, and 99% of them must be answered
within 100 ms. Writing a row to the links database for every click would put the busiest path of the system behind
the database's write capacity.

## Options considered

- Write each click to the links database during the redirect.
- Count clicks in the cache and copy the counts to the database every minute.
- Append each click to a durable, partitioned log, and count clicks in a separate consumer.

## Decision

We will append every click to a durable log from the redirect service, without making the redirect wait for it,
and run a consumer that adds clicks up and writes one count per link per minute to the links database.

## Consequences

- Redirects no longer wait for a database write, and the links database gets one write per active link per minute
  instead of one per click.
- Counts are behind by up to a minute in normal running; we alert when the lag passes 60 seconds.
- A click is lost if a redirect server dies before handing it to the log. We accept this.
- The log is a new component to run, monitor and pay for.
- If clicks ever become billable, this decision must be revisited, because billing needs every click counted.

adr/0002-cache-redirects.md

# 2. Cache redirects for one hour

## Status

Draft

## Context

Most redirects are for a small share of the links.

## Decision

We will cache every redirect for one hour.

The first record passes, and its consequences are what make it honest: it says that clicks can be lost and that a new component has to be run, not only that redirects get faster. The second fails twice: “Draft” is not a status the team agreed on, and it has no consequences at all. The missing section hides a real problem. A redirect cached for an hour keeps working for up to an hour after its link expires or is deleted, while the requirements say an expired link answers 410 Gone. Writing the consequences down is how that conflict would have surfaced before the code did.

Design documents

At work, a design document gathers all of this for reviewers who were not in the room. This track uses an outline like the following; adapt it to your team’s habits:

  1. Context and goals: the problem and the requirements, with numbers.
  2. Non-goals: what the design leaves out on purpose.
  3. The design: the container diagram, the sequence diagram of the critical request, the API and the data model.
  4. Alternatives considered: the options that lost, and why.
  5. Risks and open questions: what could go wrong, and how you will know it is working (the metrics and alerts).
  6. Rollout: how the change reaches production, and how it is undone if it fails.

Decisions taken while writing it become decision records, and the document links to them, so that the document can age while the records stay true.

Markdown Editor Draft a decision record or a design document in Markdown, with a live preview.

Key takeaways

  • Label every box with what it is and every arrow with what it does, one way, with the protocol; draw data stores as stores, add a title and a key, and never let colour carry meaning alone.
  • Use the C4 zoom levels: context for everyone, container for most design discussions, component and code rarely.
  • Pair the container diagram with a sequence diagram of the request that matters most; the order of its steps is part of the design.
  • Keep diagrams as text next to the code, so they change and are reviewed with it.
  • Record each decision as a numbered ADR with status, context, decision and consequences, including the bad ones; supersede decisions instead of editing them away.

Exercise

Exercise · Easy · Python

Number and name a new decision record

Decision records live in a folder as NNNN-short-title.md: a four-digit number, then the title in lower case with hyphens. Numbers go up one at a time and are never reused, even when a record is deleted. Write two functions in adr_names.py that a small command-line tool could use to add the next record.

next_number(file_names) takes the names of the files in the folder, in any order, and returns the number for a new record: one more than the highest number in use. Only names that start with four digits and a hyphen and end in .md count; other files, such as README.md, are ignored. An empty folder starts at 1. A gap does not get filled:

next_number(["0001-use-postgres.md", "0003-cache-redirects.md", "README.md"])  # 4, not 2

file_name(number, title) returns the file name for a record: the number with four digits, a hyphen, then the words of the title (letters and digits only) in lower case, joined by single hyphens, and .md:

file_name(4, "Record clicks in a durable log")  # '0004-record-clicks-in-a-durable-log.md'
file_name(12, "Cache redirects: for 1 hour?")   # '0012-cache-redirects-for-1-hour.md'

The sample tests import both functions from adr_names.py and run in your browser.

Starter code · adr_names.py

def next_number(file_names):
    """The number for a new decision record: one more than the highest number in use, or 1."""
    # Replace this line with your code.
    return 0


def file_name(number, title):
    """'NNNN-words-of-the-title.md' for a record."""
    # Replace this line with your code.
    return ""
The sample tests · test_adr_names.py
from adr_names import file_name, next_number


def test_next_after_the_highest():
    """the next number is one more than the highest, in any order"""
    assert next_number(["0002-b.md", "0001-a.md"]) == 3
    assert next_number(["0010-j.md", "0009-i.md"]) == 11


def test_gaps_are_not_filled():
    """a number that was used once is never used again"""
    assert next_number(["0001-use-postgres.md", "0003-cache-redirects.md"]) == 4


def test_other_files_are_ignored():
    """only NNNN-....md names count, and an empty folder starts at 1"""
    assert next_number(["README.md", "template.md", "0001-a.md", "0005-draft.txt"]) == 2
    assert next_number([]) == 1
    assert next_number(["README.md"]) == 1


def test_file_name():
    """four digits, then the title in lower case with hyphens"""
    assert file_name(4, "Record clicks in a durable log") == "0004-record-clicks-in-a-durable-log.md"


def test_file_name_punctuation():
    """punctuation and extra spaces disappear"""
    assert file_name(12, "Cache redirects: for 1 hour?") == "0012-cache-redirects-for-1-hour.md"
    assert file_name(120, "  Use   UTC  everywhere ") == "0120-use-utc-everywhere.md"
A hint

For next_number, re.match(r"(\d{4})-.*\.md$", name) matches the names that count, and int(m.group(1)) is the number; the answer is max(...) + 1, or 1 when nothing matched. For file_name, re.findall(r"[a-z0-9]+", title.lower()) gives the words, "-".join(words) joins them, and f"{number:04d}" writes the number with four digits.

The sample tests run on this device, in your browser (Pyodide): nothing is sent to mysmartcopilot.com. The first run downloads Python (about 13.5 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.

Check yourself

6 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.

  1. Question 1 of 6 Look at the diagram in "Spot the problems". Which of these are problems that make it hard to read?

    Choose every answer that is right.

    Show the answer to question 1

    Answer:

    • The arrow between App and Cache points both ways, so you cannot tell who calls whom
    • Two arrows have no label at all
    • "LB" is an abbreviation that no key explains
    • The labels "uses" and "data" say nothing about what is sent or done
    • "Orders" is drawn like the other boxes, so you cannot tell whether it is a database, a service or a queue
    • It mixes levels of zoom, putting a method name next to whole applications and a load balancer

    Unlabelled and vague arrows, a two-way arrow, mixed zoom levels, a store that does not look like one and an unexplained abbreviation all make the reader guess. Colour and direction are free choices: one colour is fine, and top to bottom fits narrow screens better.

  2. Question 2 of 6 You must explain a new system to the finance team, who are not engineers. Which C4 diagram do you draw?

    Choose one answer.

    Show the answer to question 2

    Answer: A system context diagram, with the system as one box among its users and the systems it talks to

    The context diagram leaves out technology and inner structure and shows the system in its surroundings, which is exactly what people outside the team need. The other levels are for technical readers.

  3. Question 3 of 6 Which label is best for an arrow from the Checkout API to the orders database?

    Choose one answer.

    Show the answer to question 3

    Answer: Reads and writes orders (SQL over TLS)

    A good label says what the arrow does, in the direction it points, and between containers which protocol it uses. "Uses", "data" and "connection" would fit any arrow on any diagram, so they tell the reader nothing.

  4. Question 4 of 6 Your team reverses a decision recorded in ADR 3. What should happen to ADR 3?

    Choose one answer.

    Show the answer to question 4

    Answer: Keep it, mark it as superseded, and write a new ADR that explains the new decision and points back to it

    Records are history. The old one stays, marked superseded by the new one, so a reader can see what was decided, when it changed and why; numbers are never reused.

  5. Question 5 of 6 What belongs in the Consequences section of an ADR?

    Choose one answer.

    Show the answer to question 5

    Answer: Everything that follows from the decision, including the costs and risks the team accepts

    A record that lists only benefits cannot be checked. The costs and risks are what a future reader needs most, because they are the reasons the decision may need to change. Rejected options belong in their own section.

  6. Question 6 of 6 In the checkout sequence diagram, why is the order saved as "pending" before the card is charged?

    Choose one answer.

    Show the answer to question 6

    Answer: If the API crashes after the charge, there is still a record of the order to reconcile with the payment provider

    Writing the order first means no charge can happen without a record of what it was for. A crash between the charge and "mark paid" leaves a pending order that a reconciliation job can match with the provider's records.

References

Related tools

Report a problem with this lesson

Quick answers and tool search

Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.