System Design (High-Level Design) Module 2 – Back-of-the-envelope estimation
Latency numbers, measured on real hardware
Order operations from nanoseconds to hundreds of milliseconds, measure memory, SQLite and disk latency with scripts, and work out the speed-of-light floor.
What you will learn
- Order common operations by latency, from a memory lookup to a round trip between continents
- Measure lookups, database reads and durable disk writes with documented scripts
- Explain why a latency table copied without its hardware, software and method misleads
- Calculate the speed-of-light floor of a network round trip between two places
Before you start
On this page
A design is mostly decided by orders of magnitude. If a page must load in 200 milliseconds, it cannot make ten calls one after another to a server on another continent, and no amount of tuning will change that. A cache in the same process answers in nanoseconds to microseconds, while a cache in another data centre costs a network round trip of milliseconds before it answers anything. You do not need exact values to see this, but you do need the right order of magnitude for each operation, and the safest way to get it is to measure.
This lesson measures a few operations on a real machine, shows the method next to every number, and works out the one latency that no hardware can improve: the time light takes to cover the distance.
A ladder of orders of magnitude
Each of the first three rungs is about a thousand times slower than the one above it, and the last is ten to a hundred times slower again. The examples on each rung are the ones this lesson measures or computes.
The latency ladder, with operations this lesson measures
Text description of the diagram
The diagram is a ladder of four rungs, read from top to bottom. Each rung is a unit of time with examples from this lesson.
- Nanoseconds: a Python dictionary lookup and a function call.
- Microseconds, about a thousand times slower: hashing 1 KiB with SHA-256, reading one SQLite row by its primary key, and writing 4 KiB into the operating system's cache.
- Milliseconds, about a thousand times slower again: scanning 20,000 SQLite rows without an index, and writing 4 KiB that is flushed all the way to the drive.
- Tens to hundreds of milliseconds, ten to a hundred times slower still: network round trips between cities and continents, whose floor is set by the speed of light.
Read the ladder by its rungs, not by its details. Once you know that a key lookup in memory is nanoseconds, a durable write to a drive is up to milliseconds and a cross-continent round trip is tens to hundreds of milliseconds, you can tell at a glance which part of a design dominates its latency.
Measuring work inside a process
Python’s timeit module runs a statement many times in a loop and reports the total. Its documentation recommends
looking at the fastest of several repeats, because slower repeats are usually caused by other programs competing
for the machine rather than by the code itself, and it switches off garbage collection while it times
(Python documentation). This script times five everyday
operations that way and prints the fastest repeat per operation:
# Time a few everyday operations with timeit and report the fastest of five repeats, per operation.
import hashlib
import json
import timeit
def show(seconds):
"""A duration in the unit that suits it."""
if seconds < 1e-6:
return f"{seconds * 1e9:.0f} ns"
if seconds < 1e-3:
return f"{seconds * 1e6:.1f} µs"
return f"{seconds * 1e3:.2f} ms"
def noop():
return None
names = {
"table": {n: n * 2 for n in range(1_000)},
"noop": noop,
"sha256": hashlib.sha256,
"kilobyte": bytes(range(256)) * 4,
"loads": json.loads,
"document": json.dumps({"id": 42, "name": "Asha", "tags": ["a", "b", "c"], "scores": list(range(100))}),
"numbers": list(range(1_000, 0, -1)),
}
CASES = [
("dict lookup, key present", "table[500]", 200_000),
("call a function that does nothing", "noop()", 200_000),
("SHA-256 of 1 KiB", "sha256(kilobyte).digest()", 5_000),
(f"json.loads of a {len(names['document'])}-byte document", "loads(document)", 2_000),
("sort 1,000 integers", "sorted(numbers)", 1_000),
]
for label, statement, number in CASES:
best = min(timeit.repeat(statement, repeat=5, number=number, globals=names)) / number
print(f"{label:<44}{show(best):>10}") Output
dict lookup, key present 25 ns call a function that does nothing 21 ns SHA-256 of 1 KiB 1.0 µs json.loads of a 453-byte document 7.7 µs sort 1,000 integers 4.6 µs
This output changes from run to run: Timings change on every run and on every computer; the values shown come from one run on the computer named below.
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 bench_python.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
Version note
The Run button uses Pyodide 314.0.7, a WebAssembly build of CPython 3.14.2, so in your browser the same operations usually take longer than in the native CPython above, and they depend on your device. That is the lesson of this page in one example: a latency number belongs to the machine and the runtime that produced it.
In the recorded run, a dictionary lookup and a function call took tens of nanoseconds, hashing a kilobyte about a microsecond, and parsing a small JSON document or sorting a thousand numbers several microseconds. Run it twice and the values move, sometimes by a factor of two or more, because the operating system decides which core runs the program and other programs compete for it. The order of magnitude stays put, and that is the part an estimate uses.
A database read: by key or by scan
The same small table can answer a query in a microsecond or in a millisecond, depending on whether the database can
go straight to the row. SQLite finds a row by its primary key with a binary search, at a cost that grows with the
logarithm of the number of rows; a query on a column without an index makes it read every row
(SQLite query planner). EXPLAIN QUERY PLAN shows which one it chose:
# Point reads and scans in an in-memory SQLite table of 20,000 rows, with the plan SQLite chose for each.
import sqlite3
import timeit
db = sqlite3.connect(":memory:")
db.execute("CREATE TABLE users (id INTEGER PRIMARY KEY, email TEXT, name TEXT)")
db.executemany(
"INSERT INTO users VALUES (?, ?, ?)",
((i, f"user{i}@example.com", f"User {i}") for i in range(1, 20_001)),
)
QUERIES = [
("by primary key", "SELECT name FROM users WHERE id = ?", (12_345,)),
("by email, no index", "SELECT name FROM users WHERE email = ?", ("user12345@example.com",)),
]
def report(label, sql, args):
plan = db.execute("EXPLAIN QUERY PLAN " + sql, args).fetchall()[0][3]
run = lambda: db.execute(sql, args).fetchone()
number = 2_000 if "SCAN" not in plan else 20
best = min(timeit.repeat(run, repeat=5, number=number)) / number
shown = f"{best * 1e6:.1f} µs" if best < 1e-3 else f"{best * 1e3:.2f} ms"
print(f"{label:<22}{shown:>10} plan: {plan}")
return best
key = report(*QUERIES[0])
scan = report(*QUERIES[1])
db.execute("CREATE INDEX users_email ON users (email)")
indexed = report("by email, with index", QUERIES[1][1], QUERIES[1][2])
print(f"The scan took {scan / key:,.0f} times as long as the primary-key read; the index made it {scan / indexed:,.0f} times faster.") Output
by primary key 1.3 µs plan: SEARCH users USING INTEGER PRIMARY KEY (rowid=?) by email, no index 790.6 µs plan: SCAN users by email, with index 1.7 µs plan: SEARCH users USING INDEX users_email (email=?) The scan took 597 times as long as the primary-key read; the index made it 476 times faster.
This output changes from run to run: Timings change on every run and on every computer.
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 bench_sqlite.py
The table is in memory, so no disk is involved: the difference of several hundred times comes only from how much
work each query does. A plan that says SCAN on a large table is a warning sign in any design review, and it grows
with the table, while the indexed read barely changes. The database is the same; the access path decides its
latency.
Durable writes: where milliseconds come from
Writing to a file usually only copies the bytes into the operating system’s cache, which is fast. Making the write
durable means waiting until it is safe on the drive, and what “safe” means depends on the system. On macOS,
fsync() moves the data to the drive, but the drive may keep it in its own cache and write it later; the
F_FULLFSYNC control asks the drive to write everything it has buffered to permanent storage
(Apple fsync manual).
This script times 4 KiB writes all three ways:
# Write 4 KiB blocks to a file 200 times each way and time every write: into the operating system's
# cache only, then also waiting for fsync(), then (on macOS) for F_FULLFSYNC, which also empties the drive's cache.
import os
import time
BLOCK = b"x" * 4096
WRITES = 200
def show(ns):
return f"{ns / 1000:.1f} µs" if ns < 1_000_000 else f"{ns / 1_000_000:.1f} ms"
def measure(flush):
times = []
with open("blocks.bin", "wb") as f:
for _ in range(WRITES):
start = time.perf_counter_ns()
f.write(BLOCK)
f.flush()
flush(f.fileno())
times.append(time.perf_counter_ns() - start)
os.remove("blocks.bin")
times.sort()
return times[len(times) // 2], times[int(len(times) * 0.99)]
ways = [("write, OS cache only", lambda fd: None), ("write + fsync()", os.fsync)]
try:
import fcntl
if hasattr(fcntl, "F_FULLFSYNC"):
ways.append(("write + F_FULLFSYNC", lambda fd: fcntl.fcntl(fd, fcntl.F_FULLFSYNC)))
except ImportError:
pass
print(f"{WRITES} writes of 4 KiB each {'median':>9} {'p99':>9} writes/s, one at a time")
for name, flush in ways:
median, p99 = measure(flush)
print(f"{name:<26}{show(median):>9} {show(p99):>9} {1e9 / median:>10,.0f}") Output
200 writes of 4 KiB each median p99 writes/s, one at a time write, OS cache only 10.1 µs 103.6 µs 99,167 write + fsync() 122.7 µs 455.5 µs 8,152 write + F_FULLFSYNC 4.2 ms 22.1 ms 240
This output changes from run to run: Timings change on every run and on every computer, and the F_FULLFSYNC line appears only on macOS.
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 bench_disk.py
On the machine that recorded this output, a Mac with Apple silicon writing to its internal SSD, a full flush took
tens of times longer than fsync() and hundreds of times longer than a cached write. The last column is the
consequence for a design: a system that waits for one full flush per request, one request at a time, can only make
a few hundred durable writes a second on this drive. Databases care about exactly this. PostgreSQL’s documentation
warns that many SSDs have volatile write-back caches, and that on macOS write caching can be prevented by setting
wal_sync_method to fsync_writethrough
(PostgreSQL WAL reliability). Its group commit
makes several transactions durable with a single flush to raise throughput
(commit_delay).
Test on the system you will run
The same fsync() call promises different things on different operating systems and file systems, so a durable
write measured on a laptop says little about a server. Measure on the machine type, operating system and storage
you will deploy to, and state all three next to the number.
The network: distance sets the floor
The speed of light in a vacuum is exactly 299,792,458 metres a second (NIST), and in optical fibre light travels at roughly two thirds of that speed (Singla and others), about 200 km per millisecond. A round trip covers the distance twice, so every 1,000 km between client and server adds at least 10 ms. This script computes the floor for a few routes from their great-circle distance, using the Earth’s mean radius from the WGS 84 standard (NGA) and city coordinates from GeoNames:
# The fastest possible round trip between two cities: twice the great-circle distance at the speed of light,
# in a vacuum and in optical fibre (about two thirds as fast). Real routes are longer and add queues and processing.
from math import asin, cos, radians, sin, sqrt
C_KM_PER_S = 299_792.458 # speed of light in a vacuum, exact
FIBRE_KM_PER_S = C_KM_PER_S * 2 / 3
EARTH_RADIUS_KM = 6_371.0087714 # WGS 84 mean radius
CITIES = { # latitude, longitude in degrees (GeoNames)
"Mumbai": (19.0728, 72.8826),
"Delhi": (28.6519, 77.2315),
"Singapore": (1.2897, 103.8501),
"Frankfurt": (50.1155, 8.6842),
"London": (51.5085, -0.1257),
"New York": (40.7143, -74.0060),
"San Francisco": (37.7749, -122.4194),
"Sydney": (-33.8678, 151.2073),
}
def great_circle_km(a, b):
"""Distance along the Earth's surface (haversine formula, spherical Earth)."""
(lat1, lon1), (lat2, lon2) = [(radians(lat), radians(lon)) for lat, lon in (CITIES[a], CITIES[b])]
h = sin((lat2 - lat1) / 2) ** 2 + cos(lat1) * cos(lat2) * sin((lon2 - lon1) / 2) ** 2
return 2 * EARTH_RADIUS_KM * asin(sqrt(h))
print(f"{'Route':<26}{'distance':>10}{'vacuum':>9}{'fibre':>9} (round trip)")
for a, b in [("Mumbai", "Delhi"), ("Mumbai", "Singapore"), ("Mumbai", "Frankfurt"), ("Mumbai", "London"), ("Mumbai", "New York"), ("London", "New York"), ("San Francisco", "Sydney")]:
km = great_circle_km(a, b)
vacuum_ms = 2 * km / C_KM_PER_S * 1000
fibre_ms = 2 * km / FIBRE_KM_PER_S * 1000
print(f"{a + ' - ' + b:<26}{km:>7,.0f} km{vacuum_ms:>6.0f} ms{fibre_ms:>6.0f} ms")
print(f"Fibre adds at least {2 * 1000 / FIBRE_KM_PER_S * 1000:.0f} ms of round trip for every 1,000 km of distance.") Output
Route distance vacuum fibre (round trip) Mumbai - Delhi 1,153 km 8 ms 12 ms Mumbai - Singapore 3,910 km 26 ms 39 ms Mumbai - Frankfurt 6,564 km 44 ms 66 ms Mumbai - London 7,192 km 48 ms 72 ms Mumbai - New York 12,538 km 84 ms 125 ms London - New York 5,570 km 37 ms 56 ms San Francisco - Sydney 11,948 km 80 ms 120 ms Fibre adds at least 10 ms of round trip for every 1,000 km of distance.
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 light_floor.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
Real round trips are slower than these floors. Cables do not follow great circles, and every router, queue and handshake adds time: in the measurements of Singla and colleagues, the median ping was about 3.2 times the vacuum floor and fetching just the HTML of a popular page took a median of 34 times the floor. Even so, these floors are the most useful numbers on this page: nothing can beat the vacuum figure, and no route over fibre can beat the fibre figure. A user in Mumbai talking to a server in New York over fibre waits at least 125 ms for every round trip (84 ms even at the speed of light in a vacuum), so a design for that user must reduce the number of round trips or move the server closer.
To measure real network timings yourself, curl’s --write-out option prints when each phase of a request finished
(curl manual), for example
curl -s -o /dev/null -w '%{time_namelookup} %{time_connect} %{time_appconnect} %{time_starttransfer}\n' https://example.com/
for the DNS lookup, the TCP connection, the TLS handshake and the first byte, in seconds from the start. Run it a
few times, and from more than one place: one measurement from one network is an anecdote.
Why copied tables mislead
Tables of “latency numbers” travel from slide to slide without the context that made them true. A well-known one, in Peter Norvig’s essay Teach Yourself Programming in Ten Years, gives approximate timings for “a typical PC” of its time, including 8 ms for a disk seek. That figure describes a spinning disk moving its head to a new place; an SSD has no head to move, and the drive measured above finished a cached write in microseconds and even a fully flushed one in a few milliseconds. Its 150 ms for a packet from the US to Europe and back is a different kind of number: distance sets a floor under it that no new hardware moves (37 ms in a vacuum and 56 ms over fibre between London and New York, from the script above), so it remains the right order of magnitude.
So treat any latency number as a measurement that needs its context:
- Hardware and software. The machine, the drive, the operating system and the runtime. The same Python code ran at different speeds natively and in the browser above.
- The exact operation. “A write” can mean into a cache, to the drive, or to permanent storage, and those differed by hundreds of times above.
- The statistic. The fastest of several runs, the median and the 99th percentile answer different questions. The disk example prints both a median and a p99, and the p99 is much higher.
- When and how it was measured. Hardware changes; a script that anyone can run again keeps a number honest.
Key takeaways
- Think in rungs: nanoseconds in memory, microseconds for small work and indexed reads, milliseconds for durable writes and scans, tens to hundreds of milliseconds across cities and continents.
- Measure with a documented method (
timeit, the fastest of several repeats) and publish the hardware and runtime with the number. - The access path, not the database, decides a read’s latency: a scan grows with the table, an indexed read barely does.
- Durability has a price, and its meaning differs between systems: one full flush per request caps a drive at a few hundred durable writes a second unless writes are grouped.
- Distance sets a floor of about 10 ms of round trip per 1,000 km over fibre, and real round trips are slower still: in one large study the median ping took 3.2 times the vacuum floor.
Exercise
Exercise · Medium · Python
Compute the speed-of-light floor of a round trip
No network can beat the speed of light, so the shortest possible round trip between two places is a useful floor for any latency estimate. Write two functions in floor.py.
great_circle_km(lat1, lon1, lat2, lon2) returns the distance in kilometres along the Earth's surface between two points given in degrees, with the haversine formula on a sphere of radius EARTH_RADIUS_KM:
- Convert all four angles to radians with
math.radians. - Compute
h = sin(Δlat / 2)² + cos(lat1) · cos(lat2) · sin(Δlon / 2)². - The distance is
2 · R · asin(√h).
fibre_rtt_ms(distance_km) returns the shortest round trip in milliseconds over optical fibre, where light travels at FIBRE_KM_PER_S, about two thirds of its speed in a vacuum: there and back is twice the distance.
The sample tests check a few points whose distances you can work out by hand, such as a quarter of the way round the equator, and a real pair of cities.
Starter code · floor.py
import math
EARTH_RADIUS_KM = 6_371.0
FIBRE_KM_PER_S = 299_792.458 * 2 / 3
def great_circle_km(lat1, lon1, lat2, lon2):
"""Distance in km along the Earth's surface between two points given in degrees (haversine formula)."""
# Replace this line with your code.
return 0.0
def fibre_rtt_ms(distance_km):
"""The fastest possible round trip in milliseconds over fibre for a one-way distance in km."""
# Replace this line with your code.
return 0.0 The sample tests · test_floor.py
import math
from floor import EARTH_RADIUS_KM, fibre_rtt_ms, great_circle_km
def test_same_point():
"""is zero from a point to itself"""
assert great_circle_km(19.07, 72.88, 19.07, 72.88) == 0
def test_quarter_of_the_equator():
"""measures a quarter of the way round the equator"""
expected = math.pi / 2 * EARTH_RADIUS_KM
assert math.isclose(great_circle_km(0, 0, 0, 90), expected, rel_tol=1e-9)
def test_pole_to_pole():
"""measures half the way round, from pole to pole"""
assert math.isclose(great_circle_km(90, 0, -90, 0), math.pi * EARTH_RADIUS_KM, rel_tol=1e-9)
def test_real_cities():
"""measures London to New York at about 5,570 km"""
assert math.isclose(great_circle_km(51.5085, -0.1257, 40.7143, -74.0060), 5_570, rel_tol=0.005)
def test_round_trip():
"""turns a one-way distance into a round trip over fibre"""
assert fibre_rtt_ms(0) == 0
assert math.isclose(fibre_rtt_ms(1_000), 10.0069, rel_tol=1e-4)
assert math.isclose(fibre_rtt_ms(5_570), 55.74, rel_tol=1e-3) A hint
Work in radians from the start: lat1, lon1, lat2, lon2 = map(math.radians, (lat1, lon1, lat2, lon2)). Then Δlat is lat2 - lat1 and Δlon is lon2 - lon1. For the round trip, time is distance divided by speed, so 2 * distance_km / FIBRE_KM_PER_S is in seconds: multiply by 1,000 for milliseconds.
Results of the sample tests
| Test | Result | Details |
|---|
What your code printed
The sample tests run on this device, in your browser (Pyodide): nothing is sent to mysmartcopilot.com. The first run downloads Python (about 13.5 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.
Check yourself
5 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.
References
- timeit, measure execution time of small code snippets (Python Software Foundation)
- Query planning (SQLite documentation) (SQLite)
- fsync(2), the macOS manual page (Apple)
- PostgreSQL documentation: Reliability and the Write-Ahead Log (The PostgreSQL Global Development Group)
- PostgreSQL documentation: Write Ahead Log settings (commit_delay) (The PostgreSQL Global Development Group)
- CODATA value: speed of light in vacuum (National Institute of Standards and Technology (NIST))
- The Internet at the Speed of Light (ACM HotNets (Singla, Chandrasekaran, Godfrey and Maggs))
- World Geodetic System 1984: Its Definition and Relationships with Local Geodetic Systems (National Geospatial-Intelligence Agency (NGA))
- GeoNames geographical database (GeoNames)
- curl manual page (--write-out variables) (curl project)
- Teach Yourself Programming in Ten Years (approximate timings) (Peter Norvig)
Related tools
Report a problem with this lesson
Kept only in this browser. Your Learn progress