Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

System Design (High-Level Design) Module 1 – Start here: how system design works

A step-by-step framework for a design round

Seven time-boxed steps for a 45 to 60 minute design round, when to go broad and when to go deep, and how to close with bottlenecks and failure modes.

  • Beginner
  • 25 minutes
  • Examples run with Python 3.14.8, Pyodide 314.0.7, Node.js 24.21.0 and quickjs 0.32.0
  • By MySmartCoPilot

What you will learn

  • Run a 45-60 minute design discussion in seven time-boxed steps
  • Decide when to go broad and when to go deep
  • Summarise bottlenecks, failure modes and what you would do next

Before you start

On this page

A design round gives you 45 to 60 minutes, a vague prompt and a listener who will interrupt. Without a plan, the time goes to whatever you started with: twenty minutes of requirements, or a beautiful database schema and no time left to say how the system survives a failure. This lesson gives the seven steps that every case study in this track follows, the time each one deserves, and a way to choose where to go deep.

The seven steps

Seven steps of a 45-minute design round in order, from requirements to wrap-up, with a loop from the deep dives back to the diagram.1. Requirements, 5 minwhat it does, how well, non-goals2. Estimates, 5 minonly numbers that decide something3. API, 5 minthe main operations4. Data model, 5 minentities, keys, access patterns5. High-level diagram, 10 minthe whole path, every arrow labelled6. Deep dives, 12 minthe two or three riskiest parts7. Wrap-up, 3 minbottlenecks, failures, next stepsupdate it as the design changes

The seven steps and their time boxes for 45 minutes

Text description of the diagram

The diagram is a column of seven boxes joined by arrows, from top to bottom:

  1. Requirements, 5 minutes: what the system does, how well, and the non-goals.
  2. Estimates, 5 minutes: only the numbers that decide something.
  3. API, 5 minutes: the main operations.
  4. Data model, 5 minutes: entities, keys and access patterns.
  5. High-level diagram, 10 minutes: the whole path, with every arrow labelled.
  6. Deep dives, 12 minutes: the two or three riskiest parts.
  7. Wrap-up, 3 minutes: bottlenecks, failures and next steps.

An extra arrow leads from the deep dives back to the high-level diagram, labelled "update it as the design changes".

For each step, here is what to produce and when to move on, with the minutes it gets in a 45-minute round:

  1. Requirements, 5 minutes. The functional requirements, the non-functional ones with numbers, and the non-goals. Move on when the list is agreed, or when your assumptions are said out loud.
  2. Estimates, 5 minutes. Traffic, storage and bandwidth at the peak. Move on when each number has decided something, such as one database or many.
  3. API, 5 minutes. The main operations, with their key inputs and outputs. Move on when every functional requirement maps to an operation.
  4. Data model, 5 minutes. The entities, their keys, how each is read and written, and where each lives. Move on when you know the hot path’s main read and what serves it.
  5. High-level diagram, 10 minutes. The parts and the path of each main operation, at the level the C4 model calls containers: the separately running applications and the data stores, and what each arrow carries. Move on when you can walk one request end to end along labelled arrows.
  6. Deep dives, 12 minutes. Two or three risky parts, worked out. Move on when each has a decision and the trade-off behind it.
  7. Wrap-up, 3 minutes. The bottlenecks, the failure modes and what you would do next, until time is up.

For a 60-minute round, give the extra quarter of an hour to the parts that show judgement: 5, 5, 5, 5, 15, 20 and 5 minutes. Keep the first four steps short whatever the length; they set up the design, they are not the design.

Say which step you are on. “I’ll take five minutes on requirements, then estimate.” Signposting costs a sentence and lets the interviewer redirect you early: “skip the estimates, assume it’s large” or “I’d like to spend the time on the data model”. When they steer, follow. The steps are a habit that stops you from forgetting a part, not a script to recite, and an interviewer who moves you on is giving you information about what they want to see.

Practise against the clock

Time boxes only help if you notice when you blow one. The program below compares one recorded practice round with the 45-minute plan; the minute at which each step ended is in ENDED_AT.

A practice round against the 45-minute plan JavaScript · round_timer.mjs
// Compare one recorded practice round with this track's seven-step plan for 45 minutes.
const PLAN = [
  ['Requirements', 5],
  ['Estimates', 5],
  ['API', 5],
  ['Data model', 5],
  ['High-level diagram', 10],
  ['Deep dives', 12],
  ['Wrap-up', 3],
];

// The minute at which each step ended in one recorded 45-minute practice round.
const ENDED_AT = [11, 16, 20, 25, 37, 45, 45];

const rows = [];
let plannedEnd = 0;
let start = 0;
PLAN.forEach(([step, minutes], i) => {
  plannedEnd += minutes;
  rows.push({ step, minutes, took: ENDED_AT[i] - start, behind: ENDED_AT[i] - plannedEnd });
  start = ENDED_AT[i];
});

const cell = (text, width) => String(text).padEnd(width);
console.log(cell('Step', 20) + cell('Planned', 9) + cell('Took', 8) + 'Behind plan after it');
for (const r of rows) {
  console.log(cell(r.step, 20) + cell(`${r.minutes} min`, 9) + cell(`${r.took} min`, 8) + `${r.behind} min`);
}

const signed = (n) => (n > 0 ? `+${n}` : `${n}`);
const over = rows.filter((r) => r.took > r.minutes);
const squeezed = rows.filter((r) => r.took < r.minutes);
console.log('');
console.log(`Over plan: ${over.map((r) => `${r.step} ${signed(r.took - r.minutes)}`).join(', ')}`);
console.log(`Paid for by: ${squeezed.map((r) => `${r.step} ${signed(r.took - r.minutes)}`).join(', ')}`);
console.log(`Skipped: ${rows.filter((r) => r.took === 0).map((r) => r.step).join(', ') || 'none'}`);

Output

Step                Planned  Took    Behind plan after it
Requirements        5 min    11 min  6 min
Estimates           5 min    5 min   6 min
API                 5 min    4 min   5 min
Data model          5 min    5 min   5 min
High-level diagram  10 min   12 min  7 min
Deep dives          12 min   8 min   3 min
Wrap-up             3 min    0 min   0 min

Over plan: Requirements +6, High-level diagram +2
Paid for by: API -1, Deep dives -4, Wrap-up -3
Skipped: Wrap-up

Recorded with Node.js 24.21.0 on macOS 26 arm64. To run it yourself: mise exec node@24.21.0 -- node round_timer.mjs

Requirements took 11 minutes instead of 5, and the last column shows what that cost: the round ran five to seven minutes behind for the next four steps, and only caught up by cutting the deep dives by four minutes and skipping the wrap-up entirely. That is the usual pattern. Time lost early is never won back; it comes out of the steps at the end, which are the ones that show judgement. The fix is to treat the requirements box as hard: when it ends, say the assumptions you have not checked (“I’ll assume reads far outnumber writes”) and move on.

Press Edit, put your own end minutes into ENDED_AT, and run it after each practice round.

Interval Timer Practise with seven intervals of 5, 5, 5, 5, 10, 12 and 3 minutes.

Go broad first, then deep

Broad first. Get a design that works end to end, however plainly, before you improve any part of it. An interviewer cannot judge the depth of a design that does not yet work, and the riskiest part is often visible only once the whole path is drawn.

Then deep, chosen by risk. Pick the deep dives by asking where the design is most likely to break or be wrong, not which part you know best:

  • The hot path: the operation with the most traffic, such as the redirect of a link shortener. A small cost per request is multiplied by every request.
  • The contended resource: many users wanting the same thing at once, such as the last seat of a show. This is where double booking and lock storms live.
  • The consistency boundary: where money or state must be exact, such as a wallet balance, and where retries can do damage.
  • The biggest unknown: a number you had to guess that would change the design if it were ten times larger.

Go deep earlier than the plan says when the interviewer asks, or when one requirement is unusual enough to be the real problem (“no seat may ever be sold twice”). The common mistake is the opposite: spending the deep dives on the login service or the logging library because they are familiar, while the part that decides the design gets one sentence.

Close with what could go wrong. The wrap-up is three minutes, so prepare it as you go: the first bottleneck as traffic grows, the failure you would test first, and the next thing you would build. Saying them shows you know where your own design is weak.

A worked outline: a paste service in 45 minutes

Here is the outline of one round for “design a paste service”, a site where people paste text and share a link to it. The numbers come from this program, so every estimate in the outline can be checked:

The estimates of the paste service JavaScript · paste_numbers.mjs
// The numbers behind the worked outline of a paste service. Every input is an assumption said out loud in the round.
const NEW_PASTES_PER_DAY = 1000000;
const READS_PER_PASTE = 10; // reads for every new paste
const PEAK_FACTOR = 5; // the busiest second of the day compared with the average second
const PASTE_BYTES = 10000; // 10 KB of text in an average paste
const COPIES = 3; // copies kept of every paste
const ID_LENGTH = 8; // random characters from a-z, A-Z and 0-9 in each paste ID
const YEARS = 5;
const SECONDS_PER_DAY = 86400;

const groups = (n) => String(Math.round(n)).replace(/\B(?=(\d{3})+(?!\d))/g, ',');

const writes = NEW_PASTES_PER_DAY / SECONDS_PER_DAY;
const reads = writes * READS_PER_PASTE;
const bytesPerYear = NEW_PASTES_PER_DAY * PASTE_BYTES * 365;
console.log(`Writes: ${writes.toFixed(1)} a second on average, ${(writes * PEAK_FACTOR).toFixed(0)} at the peak`);
console.log(`Reads: ${reads.toFixed(0)} a second on average, ${(reads * PEAK_FACTOR).toFixed(0)} at the peak`);
console.log(`New text: ${(NEW_PASTES_PER_DAY * PASTE_BYTES / 1e9).toFixed(0)} GB a day, ${(bytesPerYear / 1e12).toFixed(2)} TB a year`);
console.log(`Stored with ${COPIES} copies: ${(bytesPerYear * COPIES / 1e12).toFixed(2)} TB a year`);

// How likely is a new random ID to be one that already exists, after YEARS of pastes?
const possibleIds = 62 ** ID_LENGTH;
const existing = NEW_PASTES_PER_DAY * 365 * YEARS;
console.log(`IDs of ${ID_LENGTH} characters: ${groups(possibleIds)} possible`);
console.log(`After ${YEARS} years, a new random ID matches an existing one about 1 time in ${groups(possibleIds / existing)}`);

Output

Writes: 11.6 a second on average, 58 at the peak
Reads: 116 a second on average, 579 at the peak
New text: 10 GB a day, 3.65 TB a year
Stored with 3 copies: 10.95 TB a year
IDs of 8 characters: 218,340,105,584,896 possible
After 5 years, a new random ID matches an existing one about 1 time in 119,638

Recorded with Node.js 24.21.0 on macOS 26 arm64. To run it yourself: mise exec node@24.21.0 -- node paste_numbers.mjs

  1. Requirements (minutes 0 to 5). Users paste text of up to 1 MB and get a short link; anyone with the link can read it; a paste can expire after an hour, a day, a week or never; pastes cannot be edited. Accounts, search and syntax highlighting are out of scope. Assumptions said out loud: 1 million new pastes a day, 10 reads for each, 99% of reads within 200 ms, 99.9% monthly availability for reads.
  2. Estimates (5 to 10). About 12 writes and 116 reads a second on average, 58 and 579 at a peak of five times the average. 10 GB of new text a day, 3.65 TB a year, about 11 TB a year with three copies. So: one database handles the writes easily, the text itself belongs in object storage, and reads are where caching pays.
  3. API (10 to 15). POST /pastes with the text and an expiry answers 201 Created with the paste’s ID and URL. GET /p/{id} answers with the text, 404 Not Found for an ID that never existed, and 410 Gone once the paste has expired, the HTTP status for something removed that is not coming back.
  4. Data model (15 to 20). A pastes table keyed by ID, holding the creation time, the expiry time, the size and where the text is stored. The main read is one row by its ID. The text goes in object storage under the same ID.
  5. High-level diagram (20 to 30). Client, load balancer, stateless API servers, the pastes table and object storage, with a cache in front of reads and a cleanup job that deletes expired pastes. Every arrow is labelled with what it carries.
  6. Deep dives (30 to 42). First, IDs: 8 random characters from 62 give 218 trillion IDs, and after five years a new ID still matches an existing one about once in 120,000 tries, so insert only if the ID is new and draw again if it is not. Second, expiry: reads check the expiry time, so an expired paste answers 410 at once even if the cleanup job is behind. Third, a paste that goes viral: the cache absorbs it, and since HTTP lets a 410 response be cached by default, even a viral paste that has expired does not reach the database every time.
  7. Wrap-up (42 to 45). The first bottleneck is read traffic on popular pastes, which the cache takes; writes are far from any limit. If object storage is unavailable, reads of uncached pastes fail and new pastes cannot be saved, so creation should fail fast with a clear error. If the cleanup job falls behind, storage grows but no expired paste is served. Next steps: limits on how fast one client may create pastes, and alerts on the cache hit rate and the cleanup backlog.

A real round will not run this smoothly: the interviewer will interrupt, and you will go back to the diagram during the deep dives, as the loop in the first figure shows. The outline shows where each part of an answer belongs.

Key takeaways

  • Use the same seven steps every time: requirements, estimates, API, data model, diagram, deep dives, wrap-up.
  • For 45 minutes, spend about 5, 5, 5, 5, 10, 12 and 3 minutes; give the extra time of a 60-minute round to the diagram and the deep dives.
  • Say which step you are on, and follow the interviewer when they steer.
  • Time lost early comes out of the deep dives and the wrap-up, so keep the requirements box hard.
  • Go broad first, then deep on the riskiest parts: the hot path, the contended resource, the consistency boundary and the biggest unknown.

Exercise

Exercise · Easy · JavaScript

Tell a practice timer which step you should be on

A practice timer is more useful when it tells you which step you should be on, not only how long you have been talking. Write stepAt(plan, minute) in timer.mjs and export it.

plan is a list of steps in order, each a pair of a name and its time box in minutes, such as [['Requirements', 5], ['Estimates', 5], ['API', 5]]. minute is the time since the round started, in minutes, and may have a fraction. Return the name of the step whose time box contains that minute:

- each step starts where the previous one ended, and its box includes its start but not its end, so in the plan above minute 0 and minute 4.5 are 'Requirements' and minute 5 is already 'Estimates'; - a step of 0 minutes contains no minute, so it is never returned; - at or after the end of the last step, return 'Over time'; - a negative minute is a mistake: throw a RangeError.

The sample tests use this track's 45-minute plan, import stepAt from timer.mjs and run in your browser.

Starter code · timer.mjs

/** The name of the step of `plan` whose time box contains `minute`, or 'Over time' after the last one. */
export function stepAt(plan, minute) {
  // Replace this line with your code.
  return '';
}
The sample tests · timer.test.mjs
import { test, assert } from 'toolverse:test';
import { stepAt } from './timer.mjs';

const PLAN = [
  ['Requirements', 5],
  ['Estimates', 5],
  ['API', 5],
  ['Data model', 5],
  ['High-level diagram', 10],
  ['Deep dives', 12],
  ['Wrap-up', 3],
];

test('starts with the first step', () => {
  assert.equal(stepAt(PLAN, 0), 'Requirements');
  assert.equal(stepAt(PLAN, 4.5), 'Requirements');
});

test('a boundary belongs to the next step', () => {
  assert.equal(stepAt(PLAN, 5), 'Estimates');
  assert.equal(stepAt(PLAN, 30), 'Deep dives');
});

test('the last minutes, then over time', () => {
  assert.equal(stepAt(PLAN, 44.9), 'Wrap-up');
  assert.equal(stepAt(PLAN, 45), 'Over time');
  assert.equal(stepAt(PLAN, 60), 'Over time');
});

test('a step of 0 minutes is never the answer', () => {
  assert.equal(stepAt([['Estimates', 0], ['API', 5]], 0), 'API');
});

test('a negative minute throws a RangeError', () => {
  assert.throws(() => stepAt(PLAN, -1), RangeError);
});
A hint

Walk through the plan once, keeping the minute at which the current step ends: end += minutes. The first step whose end is greater than minute is the answer; a 0-minute step never passes that test, so it is skipped for you. If the loop finishes, the round is over time. Check for a negative minute before the loop and throw new RangeError(...).

The sample tests run on this device, in your browser (QuickJS): nothing is sent to mysmartcopilot.com. The first run downloads JavaScript (about 0.6 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.

Check yourself

5 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.

  1. Question 1 of 5 These moments come from one round about a paste service. Put them in the order of the seven steps.

    Give each item its position, from 1 (first).

    Show the answer to question 1

    Answer:

    1. Ask whether pastes can be edited and how long they live
    2. Work out that 1 million pastes a day is about 12 writes a second
    3. Write POST /pastes and GET /p/{id} with their status codes
    4. Decide that the text goes in object storage and its details in a table keyed by ID
    5. Draw the client, load balancer, API servers and both stores with labelled arrows
    6. Walk through what happens when one paste goes viral
    7. Name the first bottleneck and the failure you would test first

    Requirements, estimates, API, data model, high-level diagram, deep dives, wrap-up. Each step uses what the earlier ones produced: the estimates need the requirements, the data model needs the API, and a deep dive needs a diagram to point at.

  2. Question 2 of 5 About how many writes a second, on average, do 1 million new pastes a day make?

    Type a number.

    Show the answer to question 2

    Answer: 11.57 writes a second (anything from 10.97 to 12.17 counts)

    A day has 86,400 seconds, so 1,000,000 ÷ 86,400 ≈ 11.6 writes a second. The busiest second is higher; the worked outline assumes five times the average, about 58.

  3. Question 3 of 5 In a design for selling concert tickets, which part most deserves a deep dive?

    Choose one answer.

    Show the answer to question 3

    Answer: How two buyers are stopped from getting the same seat when a popular show opens

    The seats of a popular show are a contended resource with a strict consistency requirement: that is where the design is most likely to be wrong. The other parts matter, but none of them decides the design.

  4. Question 4 of 5 Fifteen minutes in, you are still clarifying requirements in a 45-minute round. What is the best move?

    Choose one answer.

    Show the answer to question 4

    Answer: Say the assumptions you have not confirmed, and move on to the estimates

    Time lost early comes out of the deep dives and the wrap-up. Stating your remaining assumptions keeps the design honest and gets you moving; skipping the diagram would leave the deep dives with nothing to stand on.

  5. Question 5 of 5 Why say out loud which step you are on?

    Choose one answer.

    Show the answer to question 5

    Answer: It lets the interviewer redirect you early, before you spend time on something they do not need

    Signposting costs a sentence and gives the interviewer a chance to steer. The steps are a habit, not a script: when the interviewer moves you on, follow them.

References

Related tools

Report a problem with this lesson

Quick answers and tool search

Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.