System Design (High-Level Design) Module 1 – Start here: how system design works
What system design is and what a design round tests
What high-level design decides, what it leaves to low-level design, what a system design round checks, and how to pick your route through this track.
What you will learn
- Explain what high-level design decides and what it leaves to low-level design
- Describe the signals a design round looks for in generic terms
- Choose a learning route through this track for your goal and level
On this page
System design is the work of deciding how a piece of software is put together from parts so that it does what its users need, as fast, as reliably and as cheaply as they need it. A single program on a laptop needs little of it. A service that millions of people use at once, that must keep their data safe for years and that a team has to run every day needs a great deal, and the decisions are expensive to change once data and users depend on them.
This track is about high-level design: the decisions about the system as a whole. This first lesson says what those decisions are, how they differ from low-level design, what a design interview is really checking, and how to plan your own route through the track.
What high-level design decides
High-level design (HLD) answers questions about the whole system rather than about any one piece of code:
- The parts. Which services, databases, caches, queues and outside systems exist, and what each one is responsible for.
- The flows. How a request or a piece of data travels between those parts for each main operation, in which order and over which protocol.
- The data. Where each kind of data lives, how it is stored, how it is split across machines and how many copies are kept.
- The failures. What happens when a part is slow or down, what the user sees, and how the system recovers.
- The growth. What has to change when traffic or data grows ten times, and which part runs out first.
The result is usually a diagram, a handful of numbers (requests per second, data per year) and a short list of decisions, each with its reason. The diagram below shows the high-level design of a small link shortener, and then zooms into one of its parts.
High-level and low-level design of the same link shortener
Text description of the diagram
The diagram has two boxes, one above the other, joined by an arrow labelled "zoom into the Links API".
The upper box, high-level design, shows the parts of a link shortener and how they talk: - an app or browser calls the Links API over HTTPS; - the Links API runs on stateless servers; to redirect, it first looks the short code up in a cache that maps codes to long URLs (step 1); - on a cache miss it reads the links database (step 2); - for every redirect it sends a click to a click log.
The lower box, low-level design, shows the classes inside the Links API alone: - a RedirectHandler calls a LinkService, which has a resolve(code) method; - the LinkService calls a LinkStore with a find(code) method, and uses a ShortCode class that encodes and decodes short codes.
What it leaves to low-level design
Low-level design (LLD) works inside one part: the classes and interfaces, the data structures, the method signatures and error types, and the patterns that keep the code easy to test and change. In the diagram, the upper box is high-level design and the lower box is low-level design of the Links API alone. This site teaches it in a separate Low-Level Design track.
The same question often has an answer at each level:
- Where do links live? High level: in a key-value table keyed by short code, with a cache in front of it. Low
level: behind a
LinkStoreinterface, with an in-memory version for tests. - How are short codes made? High level: each server takes blocks of numeric IDs from a counter and writes them
in base 62. Low level: a
ShortCodeclass withencode(number)anddecode(code). - What if the database is down? High level: redirects are served from the cache while it lasts, and creating
links fails fast. Low level:
resolve(code)raises aStoreUnavailableerror that the handler turns into a 503 response.
The line between the two is not sharp, and nobody checks it with a ruler. The C4 model, Simon Brown’s approach to drawing software architecture, gives a useful vocabulary: it describes a system at four levels of zoom, the system in its context, its containers (applications and data stores), the components inside one container, and the code. High-level design lives mostly at the first two levels, low-level design at the last two.
Why there is rarely one right answer
Give two teams the prompt “design a link shortener” and they can both be right with very different designs. For an internal tool used by 200 colleagues, one small server and an ordinary relational database are plenty: simple to back up, simple to run. For a public service creating 100 million links a month, the same design would fall over, and caching, partitioning and abuse controls become the heart of the work.
So the design is never judged on its own. It is judged against the requirements it was meant to meet, and on whether each choice is the reasoned result of those requirements. A choice that is right for one set of numbers is a mistake for another, which is why the next lesson starts with requirements, not with boxes.
What a design round tests
A system design interview, or design round, is a conversation in which you design a system out loud, with a diagram, while the interviewer asks questions and changes the problem. This track plans for rounds of 45 to 60 minutes. Formats differ between employers and change over time, so the list below describes the signals in generic terms rather than any company’s process. A good round shows:
- Scoping. You ask the questions that change the design, say what you assume, and say what you will leave out.
- Quantifying. You put rough numbers on load and data, and the numbers decide something (one database or ten, a cache or not).
- A working whole first. You get an end-to-end design that serves the main operations before you polish any one part.
- Trade-offs. You name the alternatives, say why one wins here, and know what would change your mind.
- Failure thinking. You say what breaks, how you would notice, what the user sees, and how the system recovers.
- Depth where it matters. You can go deep on the riskiest part, such as the data model, a consistency boundary or a hot spot, when asked or when it is the crux.
- Communication. You structure the conversation, keep the diagram in step with what you say, check in, and adapt when the interviewer steers.
A round does not test product trivia, the exact settings of a particular database, or whether you reach a “correct” architecture. How much of each signal is expected grows with seniority, and the interview-practice module of this track sets out the rubric it uses for self-assessment, level by level: earlier in a career, reaching a sound design with some prompting; later, driving the discussion and going deep without being led.
Two audiences, one method
Interview candidates and working engineers use the same method at different speeds. In an interview the design is spoken and time-boxed. At work it is written down as a design document, reviewed by colleagues and recorded as decisions that a future reader can follow. The steps are the same: requirements, rough numbers, interfaces, data, a diagram, the risky parts, and what could go wrong.
The large cloud providers each publish a framework for reviewing real designs in this spirit. The AWS Well-Architected Framework describes its review as a constructive conversation about architectural decisions, not an audit, and the Azure and Google Cloud frameworks organise the same concerns (reliability, security, cost, performance and operations) into “pillars”. The third lesson of this module uses them to name the qualities a design trades against each other.
Choose your route through this track
The track is long, and you probably do not need all of it at once. The program below suggests a route from your goal, your level and the time you have. Its estimates are the track’s own reading-and-doing times for each module; practice on your own comes on top. It prints four sample learners:
"""Suggest a route through this track from a goal, a level and the time available.
The minutes are the track's own estimates for reading and doing each module's lessons;
practice on your own comes on top of them.
"""
MODULES = {
1: ("Start here: how system design works", 135),
2: ("Back-of-the-envelope estimation", 145),
3: ("Performance and reliability fundamentals", 255),
5: ("APIs and communication patterns", 200),
6: ("Data modelling and storage engines", 245),
7: ("Caching", 160),
8: ("Replication", 145),
9: ("Partitioning and sharding", 160),
10: ("Consistency, time and coordination", 175),
12: ("Messaging and stream processing", 175),
16: ("Observability, incidents and lessons from outages", 190),
18: ("Delivery, multi-region, disaster recovery and cost", 175),
19: ("Building blocks: traffic, storage and messaging", 385),
21: ("Recurring problems and the moves that solve them", 215),
22: ("Case studies: links, files and media", 340),
23: ("Case studies: social, feeds and search", 330),
30: ("Interview practice and study plans", 85),
}
# Module numbers in the order to study them.
ROUTES = {
"interview": [1, 2, 3, 7, 8, 9, 12, 21, 22, 30, 23, 19, 10],
"interview, experienced": [1, 30, 21, 22, 23, 19, 10, 12],
"build": [1, 3, 5, 6, 7, 8, 16, 18, 2, 9, 10, 12],
}
CASE_STUDIES = {22, 23}
# Someone new to the subject reads more slowly: every estimate is stretched by this factor.
PACE = {"new": 1.25, "some experience": 1.0, "experienced": 0.8}
def hm(minutes):
"""120 -> '2 h', 135 -> '2 h 15 min', 45 -> '45 min'."""
h, m = divmod(round(minutes), 60)
return " ".join(part for part in (f"{h} h" if h else "", f"{m} min" if m else "") if part) or "0 min"
def plan(goal, level, weeks, hours_per_week):
"""Fill the time available with modules in route order; stop at the first one that does not fit."""
route = ROUTES["interview, experienced" if goal == "interview" and level == "experienced" else goal]
week_minutes = hours_per_week * 60
used = 0
chosen, later = [], []
for number in route:
title, minutes = MODULES[number]
need = round(minutes * PACE[level])
if not later and used + need <= weeks * week_minutes:
used += need
week = -(-used // week_minutes) # ceiling division: the week in which this module ends
chosen.append((week, number, title, need))
else:
later.append(number)
return route, used, chosen, later
def show(goal, level, weeks, hours_per_week):
route, used, chosen, later = plan(goal, level, weeks, hours_per_week)
print(f"Goal: {goal}. Level: {level}. Time: {weeks} week{'s' * (weeks != 1)}, {hours_per_week} hours a week.")
for week, number, title, need in chosen:
print(f" week {week}: module {number}, {title} ({hm(need)})")
print(f" Planned: {hm(used)} of {weeks * hours_per_week} h.")
if later:
print(f" Next, when you have time: modules {', '.join(str(n) for n in later)}.")
if goal == "interview" and not CASE_STUDIES & {n for _, n, _, _ in chosen}:
first = next(i for i, n in enumerate(route) if n in CASE_STUDIES)
needed = sum(round(MODULES[n][1] * PACE[level]) for n in route[: first + 1])
print(f" Warning: no case study fits. The first one ends {hm(needed)} into this route.")
print()
show("interview", "some experience", 4, 8)
show("interview", "experienced", 1, 12)
show("build", "new", 2, 5)
show("interview", "new", 2, 5) Output
Goal: interview. Level: some experience. Time: 4 weeks, 8 hours a week. week 1: module 1, Start here: how system design works (2 h 15 min) week 1: module 2, Back-of-the-envelope estimation (2 h 25 min) week 2: module 3, Performance and reliability fundamentals (4 h 15 min) week 2: module 7, Caching (2 h 40 min) week 2: module 8, Replication (2 h 25 min) week 3: module 9, Partitioning and sharding (2 h 40 min) week 3: module 12, Messaging and stream processing (2 h 55 min) week 3: module 21, Recurring problems and the moves that solve them (3 h 35 min) week 4: module 22, Case studies: links, files and media (5 h 40 min) week 4: module 30, Interview practice and study plans (1 h 25 min) Planned: 30 h 15 min of 32 h. Next, when you have time: modules 23, 19, 10. Goal: interview. Level: experienced. Time: 1 week, 12 hours a week. week 1: module 1, Start here: how system design works (1 h 48 min) week 1: module 30, Interview practice and study plans (1 h 8 min) week 1: module 21, Recurring problems and the moves that solve them (2 h 52 min) week 1: module 22, Case studies: links, files and media (4 h 32 min) Planned: 10 h 20 min of 12 h. Next, when you have time: modules 23, 19, 10, 12. Goal: build. Level: new. Time: 2 weeks, 5 hours a week. week 1: module 1, Start here: how system design works (2 h 49 min) week 2: module 3, Performance and reliability fundamentals (5 h 19 min) Planned: 8 h 8 min of 10 h. Next, when you have time: modules 5, 6, 7, 8, 16, 18, 2, 9, 10, 12. Goal: interview. Level: new. Time: 2 weeks, 5 hours a week. week 1: module 1, Start here: how system design works (2 h 49 min) week 2: module 2, Back-of-the-envelope estimation (3 h 1 min) Planned: 5 h 50 min of 10 h. Next, when you have time: modules 3, 7, 8, 9, 12, 21, 22, 30, 23, 19, 10. Warning: no case study fits. The first one ends 36 h 3 min into this route.
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 route_planner.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
Read the output as a schedule: “week 2: module 3” means module 3 should be finished by the end of week 2.
- Interview, some experience, 4 weeks of 8 hours. The interview route covers the method, estimation, reliability, caching, replication, sharding and messaging, then the recurring problems, a first group of case studies and interview practice, with 1 h 45 min to spare.
- Interview, experienced, 1 week of 12 hours. Experienced engineers start with the method and interview practice and go straight to the recurring problems and case studies; the fundamentals move to “when you have time”.
- Build, new to the subject, 2 weeks of 5 hours. Building services at work starts with reliability, APIs and data instead of case studies. Ten hours at a beginner’s pace covers only two modules, so the rest waits.
- Interview, new to the subject, 2 weeks of 5 hours. This is the edge case. The planner never skips ahead to a module that happens to fit, because the order matters, so it stops after two modules and warns that the first case study is about 36 hours into the route. Ten hours is not enough for a route that ends in case studies; plan for more weeks rather than jumping straight to the case studies without the fundamentals they rely on.
Press Edit to change the profiles at the bottom of the program to your own goal, level and hours, and run it again.
Study Planner & Revision Timetable Turn your route into a day-by-day plan with spaced reviews.Key takeaways
- High-level design decides the parts of a system, how data flows between them, where data lives, how failures are handled and how the system grows; low-level design works inside one part.
- There is rarely one right design: a design is judged against its requirements and by the reasons behind each choice.
- A design round looks for scoping, rough numbers, a working whole before details, named trade-offs, failure thinking, depth on the risky part and clear communication.
- Interviews and design documents at work use the same steps; only the speed and the medium differ.
- Pick a route for your goal and time, and follow it in order: case studies build on the fundamentals.
Check yourself
6 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.
References
- The C4 model for visualising software architecture (Simon Brown (c4model.com))
- AWS Well-Architected Framework (Amazon Web Services)
- Azure Well-Architected Framework (Microsoft)
- Google Cloud Well-Architected Framework (Google Cloud)
Related tools
Report a problem with this lesson
Kept only in this browser. Your Learn progress