System Design (High-Level Design) Module 1 – Start here: how system design works
A step-by-step framework for a design round
Seven time-boxed steps for a 45 to 60 minute design round, when to go broad and when to go deep, and how to close with bottlenecks and failure modes.
What you will learn
- Run a 45-60 minute design discussion in seven time-boxed steps
- Decide when to go broad and when to go deep
- Summarise bottlenecks, failure modes and what you would do next
Before you start
On this page
A design round gives you 45 to 60 minutes, a vague prompt and a listener who will interrupt. Without a plan, the time goes to whatever you started with: twenty minutes of requirements, or a beautiful database schema and no time left to say how the system survives a failure. This lesson gives the seven steps that every case study in this track follows, the time each one deserves, and a way to choose where to go deep.
The seven steps
The seven steps and their time boxes for 45 minutes
Text description of the diagram
The diagram is a column of seven boxes joined by arrows, from top to bottom:
- Requirements, 5 minutes: what the system does, how well, and the non-goals.
- Estimates, 5 minutes: only the numbers that decide something.
- API, 5 minutes: the main operations.
- Data model, 5 minutes: entities, keys and access patterns.
- High-level diagram, 10 minutes: the whole path, with every arrow labelled.
- Deep dives, 12 minutes: the two or three riskiest parts.
- Wrap-up, 3 minutes: bottlenecks, failures and next steps.
An extra arrow leads from the deep dives back to the high-level diagram, labelled "update it as the design changes".
For each step, here is what to produce and when to move on, with the minutes it gets in a 45-minute round:
- Requirements, 5 minutes. The functional requirements, the non-functional ones with numbers, and the non-goals. Move on when the list is agreed, or when your assumptions are said out loud.
- Estimates, 5 minutes. Traffic, storage and bandwidth at the peak. Move on when each number has decided something, such as one database or many.
- API, 5 minutes. The main operations, with their key inputs and outputs. Move on when every functional requirement maps to an operation.
- Data model, 5 minutes. The entities, their keys, how each is read and written, and where each lives. Move on when you know the hot path’s main read and what serves it.
- High-level diagram, 10 minutes. The parts and the path of each main operation, at the level the C4 model calls containers: the separately running applications and the data stores, and what each arrow carries. Move on when you can walk one request end to end along labelled arrows.
- Deep dives, 12 minutes. Two or three risky parts, worked out. Move on when each has a decision and the trade-off behind it.
- Wrap-up, 3 minutes. The bottlenecks, the failure modes and what you would do next, until time is up.
For a 60-minute round, give the extra quarter of an hour to the parts that show judgement: 5, 5, 5, 5, 15, 20 and 5 minutes. Keep the first four steps short whatever the length; they set up the design, they are not the design.
Say which step you are on. “I’ll take five minutes on requirements, then estimate.” Signposting costs a sentence and lets the interviewer redirect you early: “skip the estimates, assume it’s large” or “I’d like to spend the time on the data model”. When they steer, follow. The steps are a habit that stops you from forgetting a part, not a script to recite, and an interviewer who moves you on is giving you information about what they want to see.
Practise against the clock
Time boxes only help if you notice when you blow one. The program below compares one recorded practice round with
the 45-minute plan; the minute at which each step ended is in ENDED_AT.
// Compare one recorded practice round with this track's seven-step plan for 45 minutes.
const PLAN = [
['Requirements', 5],
['Estimates', 5],
['API', 5],
['Data model', 5],
['High-level diagram', 10],
['Deep dives', 12],
['Wrap-up', 3],
];
// The minute at which each step ended in one recorded 45-minute practice round.
const ENDED_AT = [11, 16, 20, 25, 37, 45, 45];
const rows = [];
let plannedEnd = 0;
let start = 0;
PLAN.forEach(([step, minutes], i) => {
plannedEnd += minutes;
rows.push({ step, minutes, took: ENDED_AT[i] - start, behind: ENDED_AT[i] - plannedEnd });
start = ENDED_AT[i];
});
const cell = (text, width) => String(text).padEnd(width);
console.log(cell('Step', 20) + cell('Planned', 9) + cell('Took', 8) + 'Behind plan after it');
for (const r of rows) {
console.log(cell(r.step, 20) + cell(`${r.minutes} min`, 9) + cell(`${r.took} min`, 8) + `${r.behind} min`);
}
const signed = (n) => (n > 0 ? `+${n}` : `${n}`);
const over = rows.filter((r) => r.took > r.minutes);
const squeezed = rows.filter((r) => r.took < r.minutes);
console.log('');
console.log(`Over plan: ${over.map((r) => `${r.step} ${signed(r.took - r.minutes)}`).join(', ')}`);
console.log(`Paid for by: ${squeezed.map((r) => `${r.step} ${signed(r.took - r.minutes)}`).join(', ')}`);
console.log(`Skipped: ${rows.filter((r) => r.took === 0).map((r) => r.step).join(', ') || 'none'}`); Output
Step Planned Took Behind plan after it Requirements 5 min 11 min 6 min Estimates 5 min 5 min 6 min API 5 min 4 min 5 min Data model 5 min 5 min 5 min High-level diagram 10 min 12 min 7 min Deep dives 12 min 8 min 3 min Wrap-up 3 min 0 min 0 min Over plan: Requirements +6, High-level diagram +2 Paid for by: API -1, Deep dives -4, Wrap-up -3 Skipped: Wrap-up
Recorded with Node.js 24.21.0 on macOS 26 arm64. To run it yourself: mise exec node@24.21.0 -- node round_timer.mjs
Runs on this device, in your browser. The first run downloads JavaScript (about 0.6 MB), which is kept for the next runs.
Your run, in this browser
Requirements took 11 minutes instead of 5, and the last column shows what that cost: the round ran five to seven minutes behind for the next four steps, and only caught up by cutting the deep dives by four minutes and skipping the wrap-up entirely. That is the usual pattern. Time lost early is never won back; it comes out of the steps at the end, which are the ones that show judgement. The fix is to treat the requirements box as hard: when it ends, say the assumptions you have not checked (“I’ll assume reads far outnumber writes”) and move on.
Press Edit, put your own end minutes into ENDED_AT, and run it after each practice round.
Go broad first, then deep
Broad first. Get a design that works end to end, however plainly, before you improve any part of it. An interviewer cannot judge the depth of a design that does not yet work, and the riskiest part is often visible only once the whole path is drawn.
Then deep, chosen by risk. Pick the deep dives by asking where the design is most likely to break or be wrong, not which part you know best:
- The hot path: the operation with the most traffic, such as the redirect of a link shortener. A small cost per request is multiplied by every request.
- The contended resource: many users wanting the same thing at once, such as the last seat of a show. This is where double booking and lock storms live.
- The consistency boundary: where money or state must be exact, such as a wallet balance, and where retries can do damage.
- The biggest unknown: a number you had to guess that would change the design if it were ten times larger.
Go deep earlier than the plan says when the interviewer asks, or when one requirement is unusual enough to be the real problem (“no seat may ever be sold twice”). The common mistake is the opposite: spending the deep dives on the login service or the logging library because they are familiar, while the part that decides the design gets one sentence.
Close with what could go wrong. The wrap-up is three minutes, so prepare it as you go: the first bottleneck as traffic grows, the failure you would test first, and the next thing you would build. Saying them shows you know where your own design is weak.
A worked outline: a paste service in 45 minutes
Here is the outline of one round for “design a paste service”, a site where people paste text and share a link to it. The numbers come from this program, so every estimate in the outline can be checked:
// The numbers behind the worked outline of a paste service. Every input is an assumption said out loud in the round.
const NEW_PASTES_PER_DAY = 1000000;
const READS_PER_PASTE = 10; // reads for every new paste
const PEAK_FACTOR = 5; // the busiest second of the day compared with the average second
const PASTE_BYTES = 10000; // 10 KB of text in an average paste
const COPIES = 3; // copies kept of every paste
const ID_LENGTH = 8; // random characters from a-z, A-Z and 0-9 in each paste ID
const YEARS = 5;
const SECONDS_PER_DAY = 86400;
const groups = (n) => String(Math.round(n)).replace(/\B(?=(\d{3})+(?!\d))/g, ',');
const writes = NEW_PASTES_PER_DAY / SECONDS_PER_DAY;
const reads = writes * READS_PER_PASTE;
const bytesPerYear = NEW_PASTES_PER_DAY * PASTE_BYTES * 365;
console.log(`Writes: ${writes.toFixed(1)} a second on average, ${(writes * PEAK_FACTOR).toFixed(0)} at the peak`);
console.log(`Reads: ${reads.toFixed(0)} a second on average, ${(reads * PEAK_FACTOR).toFixed(0)} at the peak`);
console.log(`New text: ${(NEW_PASTES_PER_DAY * PASTE_BYTES / 1e9).toFixed(0)} GB a day, ${(bytesPerYear / 1e12).toFixed(2)} TB a year`);
console.log(`Stored with ${COPIES} copies: ${(bytesPerYear * COPIES / 1e12).toFixed(2)} TB a year`);
// How likely is a new random ID to be one that already exists, after YEARS of pastes?
const possibleIds = 62 ** ID_LENGTH;
const existing = NEW_PASTES_PER_DAY * 365 * YEARS;
console.log(`IDs of ${ID_LENGTH} characters: ${groups(possibleIds)} possible`);
console.log(`After ${YEARS} years, a new random ID matches an existing one about 1 time in ${groups(possibleIds / existing)}`); Output
Writes: 11.6 a second on average, 58 at the peak Reads: 116 a second on average, 579 at the peak New text: 10 GB a day, 3.65 TB a year Stored with 3 copies: 10.95 TB a year IDs of 8 characters: 218,340,105,584,896 possible After 5 years, a new random ID matches an existing one about 1 time in 119,638
Recorded with Node.js 24.21.0 on macOS 26 arm64. To run it yourself: mise exec node@24.21.0 -- node paste_numbers.mjs
Runs on this device, in your browser. The first run downloads JavaScript (about 0.6 MB), which is kept for the next runs.
Your run, in this browser
- Requirements (minutes 0 to 5). Users paste text of up to 1 MB and get a short link; anyone with the link can read it; a paste can expire after an hour, a day, a week or never; pastes cannot be edited. Accounts, search and syntax highlighting are out of scope. Assumptions said out loud: 1 million new pastes a day, 10 reads for each, 99% of reads within 200 ms, 99.9% monthly availability for reads.
- Estimates (5 to 10). About 12 writes and 116 reads a second on average, 58 and 579 at a peak of five times the average. 10 GB of new text a day, 3.65 TB a year, about 11 TB a year with three copies. So: one database handles the writes easily, the text itself belongs in object storage, and reads are where caching pays.
- API (10 to 15).
POST /pasteswith the text and an expiry answers201 Createdwith the paste’s ID and URL.GET /p/{id}answers with the text,404 Not Foundfor an ID that never existed, and410 Goneonce the paste has expired, the HTTP status for something removed that is not coming back. - Data model (15 to 20). A
pastestable keyed by ID, holding the creation time, the expiry time, the size and where the text is stored. The main read is one row by its ID. The text goes in object storage under the same ID. - High-level diagram (20 to 30). Client, load balancer, stateless API servers, the
pastestable and object storage, with a cache in front of reads and a cleanup job that deletes expired pastes. Every arrow is labelled with what it carries. - Deep dives (30 to 42). First, IDs: 8 random characters from 62 give 218 trillion IDs, and after five years a new ID still matches an existing one about once in 120,000 tries, so insert only if the ID is new and draw again if it is not. Second, expiry: reads check the expiry time, so an expired paste answers 410 at once even if the cleanup job is behind. Third, a paste that goes viral: the cache absorbs it, and since HTTP lets a 410 response be cached by default, even a viral paste that has expired does not reach the database every time.
- Wrap-up (42 to 45). The first bottleneck is read traffic on popular pastes, which the cache takes; writes are far from any limit. If object storage is unavailable, reads of uncached pastes fail and new pastes cannot be saved, so creation should fail fast with a clear error. If the cleanup job falls behind, storage grows but no expired paste is served. Next steps: limits on how fast one client may create pastes, and alerts on the cache hit rate and the cleanup backlog.
A real round will not run this smoothly: the interviewer will interrupt, and you will go back to the diagram during the deep dives, as the loop in the first figure shows. The outline shows where each part of an answer belongs.
Key takeaways
- Use the same seven steps every time: requirements, estimates, API, data model, diagram, deep dives, wrap-up.
- For 45 minutes, spend about 5, 5, 5, 5, 10, 12 and 3 minutes; give the extra time of a 60-minute round to the diagram and the deep dives.
- Say which step you are on, and follow the interviewer when they steer.
- Time lost early comes out of the deep dives and the wrap-up, so keep the requirements box hard.
- Go broad first, then deep on the riskiest parts: the hot path, the contended resource, the consistency boundary and the biggest unknown.
Exercise
Exercise · Easy · JavaScript
Tell a practice timer which step you should be on
A practice timer is more useful when it tells you which step you should be on, not only how long you have been talking. Write stepAt(plan, minute) in timer.mjs and export it.
plan is a list of steps in order, each a pair of a name and its time box in minutes, such as [['Requirements', 5], ['Estimates', 5], ['API', 5]]. minute is the time since the round started, in minutes, and may have a fraction. Return the name of the step whose time box contains that minute:
- each step starts where the previous one ended, and its box includes its start but not its end, so in the plan above minute 0 and minute 4.5 are 'Requirements' and minute 5 is already 'Estimates'; - a step of 0 minutes contains no minute, so it is never returned; - at or after the end of the last step, return 'Over time'; - a negative minute is a mistake: throw a RangeError.
The sample tests use this track's 45-minute plan, import stepAt from timer.mjs and run in your browser.
Starter code · timer.mjs
/** The name of the step of `plan` whose time box contains `minute`, or 'Over time' after the last one. */
export function stepAt(plan, minute) {
// Replace this line with your code.
return '';
} The sample tests · timer.test.mjs
import { test, assert } from 'toolverse:test';
import { stepAt } from './timer.mjs';
const PLAN = [
['Requirements', 5],
['Estimates', 5],
['API', 5],
['Data model', 5],
['High-level diagram', 10],
['Deep dives', 12],
['Wrap-up', 3],
];
test('starts with the first step', () => {
assert.equal(stepAt(PLAN, 0), 'Requirements');
assert.equal(stepAt(PLAN, 4.5), 'Requirements');
});
test('a boundary belongs to the next step', () => {
assert.equal(stepAt(PLAN, 5), 'Estimates');
assert.equal(stepAt(PLAN, 30), 'Deep dives');
});
test('the last minutes, then over time', () => {
assert.equal(stepAt(PLAN, 44.9), 'Wrap-up');
assert.equal(stepAt(PLAN, 45), 'Over time');
assert.equal(stepAt(PLAN, 60), 'Over time');
});
test('a step of 0 minutes is never the answer', () => {
assert.equal(stepAt([['Estimates', 0], ['API', 5]], 0), 'API');
});
test('a negative minute throws a RangeError', () => {
assert.throws(() => stepAt(PLAN, -1), RangeError);
}); A hint
Walk through the plan once, keeping the minute at which the current step ends: end += minutes. The first step whose end is greater than minute is the answer; a 0-minute step never passes that test, so it is skipped for you. If the loop finishes, the round is over time. Check for a negative minute before the loop and throw new RangeError(...).
Results of the sample tests
| Test | Result | Details |
|---|
What your code printed
The sample tests run on this device, in your browser (QuickJS): nothing is sent to mysmartcopilot.com. The first run downloads JavaScript (about 0.6 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.
Check yourself
5 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.
References
- RFC 9110: HTTP Semantics (status codes 201, 404 and 410) (IETF)
- C4 model: container diagram (Simon Brown (c4model.com))
Related tools
Report a problem with this lesson
Kept only in this browser. Your Learn progress