System Design (High-Level Design) Module 1 – Start here: how system design works
Turning a vague prompt into requirements
Separate functional from non-functional requirements, ask the clarifying questions that change a design, and write a scope with explicit non-goals.
What you will learn
- Separate functional requirements from non-functional requirements
- Ask clarifying questions that change the design, not ones that only fill time
- Write a one-paragraph scope statement with explicit non-goals
Before you start
On this page
“Design a link shortener.” Four words, and every number that matters is missing: how many links, how many visitors, how fast, for how long. Interview prompts are vague on purpose, and so are most feature requests at work. Before the first box goes on the board, turn the prompt into requirements: a short list of what the system must do and how well, written so that someone could check the finished system against it.
Two kinds of requirements
Functional requirements describe behaviour that a user or another system can see. Write each one as a sentence with an actor and a verb: “A signed-in user can create a short link for a long URL”, “Anyone who opens a short link is redirected to its long URL”. Each one is something a test could exercise.
Non-functional requirements describe how well the system must behave: how fast, how often available, how much data it keeps and for how long, how quickly a change becomes visible, what it must survive. They are also called quality attributes, and the next lesson is about them. NASA’s systems engineering handbook draws the same line: it separates the functions a system must perform from how well it must perform them, expects performance to be stated as a quantity, and asks that every requirement be written so that it can be verified once the product exists.
That last point is the test of a good non-functional requirement: could it fail? “Redirects must be fast” cannot fail, because nobody agreed what fast means. “99% of redirects are answered within 100 ms, measured at the server” can. The Site Reliability Engineering book calls the measured quantity a service level indicator (latency, error rate, throughput, availability) and the target for it a service level objective, and it prefers percentiles to averages, because an average can look healthy while a slow tail of requests is not.
| Vague | Measurable |
|---|---|
| Redirects must be fast | 99% of redirects are answered within 100 ms, measured at the server |
| Highly available | 99.95% of redirect requests succeed in each 30-day month |
| It must scale | Up to 100 million new links a month, with about 100 redirects for every new link |
| Stored reliably | A link the API has confirmed survives the loss of any one data centre |
| Real-time counts | A redirect shows in its link’s click count within 60 seconds |
Questions that change the design
A clarifying question is worth asking when its answer changes a number or a box on the diagram. These are the ones that most often do:
| Ask | Why the answer matters |
|---|---|
| How many new items and reads a day, and at the peak? | It sizes every part and decides whether one database is enough |
| How many reads for each write? | Read-heavy systems lean on caches and replicas; write-heavy ones on partitioning and the write path |
| How soon must a write be visible to readers? | It decides whether a cache or a replica may answer with slightly old data |
| Who are the users, and where are they? | It decides regions, edge caching and where data may live |
| How long is data kept, and can users delete it? | It sets storage growth and needs a deletion path through every copy |
| What happens when someone abuses it? | It adds rate limits, sign-in rules and content scanning |
| Is any of the data regulated, such as payments, health or personal data? | It constrains where data is stored, how it is encrypted and how long it is kept |
Some questions only fill time. “Which language should I use?” and “Should this be microservices?” are choices you make later, from the requirements; asking them first hands the design back to the interviewer. Ask about a constraint only if it is one: “Must this run on the company’s existing database?” is a fair question, “Which database do you like?” is not.
Regulated data in India
Two Indian rules are worth asking about early, because they decide regions, deletion pipelines and audit trails. For personal data of people in India, section 8 of the Digital Personal Data Protection Act, 2023 makes the “Data Fiduciary” take reasonable security safeguards against a breach, inform the Data Protection Board and each affected person when a breach happens, and erase personal data when consent is withdrawn or its purpose is no longer served, unless another law requires keeping it. For payment systems, the Reserve Bank of India’s directive on the storage of payment system data requires the full end-to-end transaction data to be stored in systems in India only; for the foreign leg of a cross-border payment, a copy may also be kept abroad.
A worked example: the link shortener
Here is a first draft of the link shortener’s requirements, the version the same author wrote after asking those questions, and a small checker that reads both. The checker counts the requirements in each section, flags every non-functional requirement that has no number with a unit (or a percentage or a ratio) and names the vague words in it, and looks for a list of what is out of scope.
requirements/scope_check.py
"""Check a requirements file before designing against it: is every non-functional requirement
measurable, and does the scope say what is left out?"""
import re
from pathlib import Path
# A number with a unit, a percentage or a ratio: "100 ms", "99.95%", "5 years", "100 to 1".
MEASURABLE = re.compile(
r"\d[\d,.]*\s*(?:%|ms\b|s\b|seconds?\b|minutes?\b|hours?\b|days?\b|weeks?\b|months?\b|years?\b"
r"|[KMGT]B\b|million\b|billion\b|requests?\b|links?\b|users?\b|data centres?\b|to \d)",
re.IGNORECASE,
)
# Words that sound like requirements but cannot be checked without a number.
VAGUE = ("fast", "quick", "highly available", "scalable", "scale", "lots of",
"reliable", "reliably", "real-time", "secure", "robust")
def sections(text):
"""The list items under each '## ' heading of a Markdown file."""
found, current = {}, None
for line in text.splitlines():
if line.startswith("## "):
current = line[3:].strip().lower()
found[current] = []
elif line.startswith("- ") and current is not None:
found[current].append(line[2:].strip())
return found
def vague_words(line):
return [w for w in VAGUE if re.search(rf"(?<![\w-]){re.escape(w)}(?![\w-])", line, re.IGNORECASE)]
def check(path):
found = sections(path.read_text(encoding="utf-8"))
functional = found.get("functional requirements", [])
quality = found.get("non-functional requirements", [])
non_goals = found.get("out of scope", [])
unmeasurable = [line for line in quality if not MEASURABLE.search(line)]
print(path.name)
print(f" functional requirements: {len(functional)}")
if unmeasurable:
print(f" non-functional requirements: {len(quality)}, not measurable: {len(unmeasurable)}")
for line in unmeasurable:
words = vague_words(line)
print(f' "{line}"' + (f" (vague: {', '.join(words)})" if words else ""))
else:
print(f" non-functional requirements: {len(quality)}, all measurable")
print(f" non-goals: {len(non_goals)}" if non_goals else ' non-goals: none (add an "Out of scope" section)')
ready = functional and quality and not unmeasurable and non_goals
print(f" verdict: {'ready for estimates' if ready else 'send back to the author'}")
print()
for name in ("link_shortener_draft.md", "link_shortener.md"):
check(Path(name)) Output
link_shortener_draft.md
functional requirements: 3
non-functional requirements: 5, not measurable: 5
"Redirects must be fast." (vague: fast)
"The service should be highly available." (vague: highly available)
"It must scale to lots of users." (vague: scale, lots of)
"Links are stored reliably." (vague: reliably)
"Click counts are real-time." (vague: real-time)
non-goals: none (add an "Out of scope" section)
verdict: send back to the author
link_shortener.md
functional requirements: 6
non-functional requirements: 7, all measurable
non-goals: 4
verdict: ready for estimates
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 scope_check.py
requirements/link_shortener_draft.md
# Link shortener: requirements (first draft)
## Functional requirements
- Users can shorten links.
- Short links redirect.
- Links can expire.
## Non-functional requirements
- Redirects must be fast.
- The service should be highly available.
- It must scale to lots of users.
- Links are stored reliably.
- Click counts are real-time. requirements/link_shortener.md
# Link shortener: requirements
## Functional requirements
- A signed-in user can create a short link for a long URL.
- A user can choose a custom alias if no one else has it.
- Anyone who opens a short link is redirected to its long URL.
- A link can have an expiry time; after it, the short link answers 410 Gone.
- A user can delete their own links.
- Every redirect is counted, and the owner sees the count for each link.
## Non-functional requirements
- Scale: up to 100 million new links a month; redirects outnumber new links about 100 to 1.
- Latency: 99% of redirects are answered within 100 ms, measured at the server.
- Availability: 99.95% of redirect requests succeed in each 30-day month.
- Durability: a link the API has confirmed survives the loss of any 1 data centre.
- Freshness: a new link redirects everywhere within 5 seconds of being created.
- Click counts: a redirect shows in its link's count within 60 seconds.
- Retention: links are kept for 5 years unless their owner deletes them.
## Out of scope
- Analytics dashboards beyond the click count for each link.
- Custom domains.
- Changing where an existing short link points.
- Link previews. Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
The draft goes back to its author. Every one of its non-functional requirements is an adjective, and nothing says what is left out. The second version passes, and comparing the two shows what changed:
- “Users can shorten links” became six functional requirements, including the ones the draft forgot: custom
aliases, deleting a link, and what an expired link does. Answering
410 Gonerather than404 Not Foundtells clients that the link has been removed on purpose and is likely to stay gone, which is what HTTP defines status code 410 to mean. - Each adjective became a number with a unit and, where it matters, where it is measured.
- “Click counts are real-time” turned into two decisions: counting every redirect is a functional requirement, “within 60 seconds” is its quality, and a dashboard is explicitly out of scope.
A checker like this finds missing numbers, and nothing more. It matches words, so it flags a requirement whose figure is spelled out (“any one data centre” rather than “any 1 data centre”) even though a failure test could check it. And it cannot tell whether 100 ms is the right number; only the people who need the system can. Treat the numbers as agreements: if the interviewer gives none, propose them out loud (“I’ll assume 100 million new links a month”) and carry on.
Write the scope down, with non-goals
Finish the requirements step with a short scope statement: one paragraph that says what this design covers and, just as important, what it does not.
Scope statement for the link shortener
This design covers the public link shortener: creating, resolving, expiring and deleting short links and counting their clicks, for up to 100 million new links a month, with redirects as the path that must stay fast. It does not cover analytics dashboards, custom domains, changing where a link points, or link previews; each can be added later without changing how redirects work.
Non-goals do three jobs. They stop the design from growing while you talk; they show the reviewer that you left something out on purpose rather than forgot it; and they turn a later addition into a decision someone takes, rather than something that creeps in. The requirements that shape the architecture most, the ones the architecture decision records community calls architecturally significant, are often on the non-functional list: scale, latency, durability and where data may live.
A common mistake: choosing the database first
“We’ll use a NoSQL database,” two minutes into a round, is a guess. A store is chosen by the way the system reads and writes its data, and those access patterns come from the requirements. For the link shortener they are now clear: the hot path is reading one row by its short code, about 100 times as often as a link is created; deletes are rare; nothing on the hot path needs a join; click counts are written constantly and may lag by a minute. Those facts, not a product’s reputation, decide where links and clicks are stored, and the data-modelling module of this track shows how.
Markdown Editor Write your own requirements and scope in Markdown, with a live preview.Key takeaways
- Functional requirements say what the system does, as sentences with an actor and a verb; non-functional requirements say how well, as numbers with units that a test or a measurement could fail.
- Ask the questions whose answers change a number or a box: volume, read-to-write ratio, freshness, geography, retention, abuse and regulated data. Leave implementation choices for later.
- If the interviewer gives no numbers, propose them out loud and carry on.
- End with a scope statement that names the non-goals.
- Do not pick a database before the access patterns are known.
Exercise
Exercise · Easy · Python
Send back the requirements that cannot be checked
A reviewer reads a list of requirement lines and sends back every line that uses a vague quality word and gives no number to check it against. Write that reviewer as a function.
Write flag_vague(lines) in vague.py. It takes a list of strings and returns a new list with the lines to send back, in their original order and unchanged. A line is sent back when both are true:
- it contains one of the words or phrases in VAGUE (already in the starter file), compared without regard to case and as whole words: "fast" counts in "Redirects must be fast." but not in "Show the breakfast menu."; - it contains no digit at all.
For example:
flag_vague(["Redirects must be fast.", "p99 redirect latency is under 100 ms.", "Users can delete links."])
# ['Redirects must be fast.']
The second line has a number, so it can be checked; the third has no vague word. An empty list gives an empty list. The sample tests import flag_vague from vague.py and run in your browser.
Starter code · vague.py
VAGUE = ("fast", "quick", "scalable", "highly available", "reliable", "real-time", "secure", "robust",
"efficient", "user-friendly")
def flag_vague(lines):
"""Return the lines that use a word from VAGUE and contain no digit, in their original order."""
# Replace this line with your code.
return [] The sample tests · test_vague.py
from vague import flag_vague
def test_flags_a_vague_line():
"""sends back a line that only says the service is fast"""
assert flag_vague(["Redirects must be fast."]) == ["Redirects must be fast."]
def test_a_number_makes_it_checkable():
"""keeps a line that gives a number to check"""
assert flag_vague(["Redirects are fast: p99 under 100 ms."]) == []
def test_whole_words_only():
"""does not find a vague word inside a longer word"""
assert flag_vague(["Show the breakfast menu in the app."]) == []
def test_case_and_phrases():
"""ignores case and finds two-word phrases"""
lines = ["The API is Highly Available.", "Users can delete their links.", "Search must be REAL-TIME."]
assert flag_vague(lines) == ["The API is Highly Available.", "Search must be REAL-TIME."]
def test_order_and_duplicates():
"""keeps the input order and adds a line once even when it has two vague words"""
lines = ["Secure and robust storage.", "Uploads are quick."]
assert flag_vague(lines) == ["Secure and robust storage.", "Uploads are quick."]
assert flag_vague([]) == [] A hint
any(ch.isdigit() for ch in line) tells you whether a line has a digit. For whole words, the re module helps: re.search(r"\b" + re.escape(term) + r"\b", line, re.IGNORECASE) finds term only where it is not part of a longer word. Stop looking at a line as soon as one term matches, so it is not added twice.
Results of the sample tests
| Test | Result | Details |
|---|
What your code printed
The sample tests run on this device, in your browser (Pyodide): nothing is sent to mysmartcopilot.com. The first run downloads Python (about 13.5 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.
Check yourself
6 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.
References
- NASA Systems Engineering Handbook, 4.2 Technical Requirements Definition (NASA)
- Site Reliability Engineering, chapter 4: Service Level Objectives (Google (O'Reilly Media))
- Architectural Decision Records (motivation and definitions) (ADR GitHub organization)
- RFC 9110: HTTP Semantics, section 15.5.11 (410 Gone) (IETF)
- The Digital Personal Data Protection Act, 2023 (Ministry of Electronics and Information Technology, Government of India)
- Storage of Payment System Data (RBI/2017-18/153) (Reserve Bank of India)
Related tools
Report a problem with this lesson
Kept only in this browser. Your Learn progress