System Design (High-Level Design) Module 2 – Back-of-the-envelope estimation
Little's law, utilisation and server counts
Use Little's law to size concurrency and connection pools, see why queueing delay explodes near full utilisation, and count servers with headroom for failures.
What you will learn
- Apply Little's law to find requests in flight, thread pool sizes and connection pool sizes
- Explain why queueing delay rises sharply as utilisation approaches 100%
- Estimate server and connection counts with headroom for bursts, failures and deploys
Before you start
On this page
Requests per second tell you how fast work arrives. To size threads, connections and servers you also need to know how much work is in progress at once, and how busy you can let each server get before waiting times run away. Two results from queueing theory answer both questions, and neither needs more than school arithmetic: Little’s law for how much is in flight, and the shape of queueing delay for how much headroom to keep.
Little’s law: rate times time is what is inside
Take any system that requests enter and leave: a web server, a database, a queue with its workers. If λ requests arrive per second and each spends W seconds inside on average, then on average L requests are inside:
The law asks for remarkably little: it does not matter in which order requests are served or how their arrival and service times are spread, only that the system is stable, with requests leaving as fast as they arrive (Gallager, section 4.5.4).
Little's law: L = λ × W
Text description of the diagram
The diagram shows a system from top to bottom.
- Requests arrive at a rate of λ (lambda) a second.
- They enter the system, a box labelled "Inside: L requests, W seconds each". It holds a queue, where requests wait for a server, and the servers, which work on them. Requests move from the queue to the servers in turn.
- Requests leave the system, at λ a second too once the system is stable.
On average there are L requests inside the system, waiting or being served, and each spends W seconds inside. Little's law says that L = λ × W.
That makes it the quickest way to turn a rate and a latency into a count of things you have to provide:
# Little's law: items in a system = arrival rate x average time each item spends in it (L = lambda x W).
CASES = [
# what arrives, arrivals per second, average seconds inside, what the result counts
("web requests", 2_000, 0.050, "requests in flight"),
("database queries", 3_000, 0.004, "connections busy"),
("calls to a slow partner API", 200, 0.800, "threads waiting on it"),
("jobs through a queue and its workers", 500, 30.0, "jobs in the system"),
]
for name, rate, seconds, unit in CASES:
print(f"{name:<38}{rate:>6,}/s x {seconds * 1000:>6,.0f} ms = {rate * seconds:>7,.0f} {unit}") Output
web requests 2,000/s x 50 ms = 100 requests in flight database queries 3,000/s x 4 ms = 12 connections busy calls to a slow partner API 200/s x 800 ms = 160 threads waiting on it jobs through a queue and its workers 500/s x 30,000 ms = 15,000 jobs in the system
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 in_flight.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
Read each line as a sizing decision. A web tier at 2,000 requests a second and 50 ms per request has 100 requests in progress, so it needs at least 100 concurrent workers across the fleet. The slow partner API is the dangerous line: at 800 ms per call, a modest 200 calls a second hold 160 threads, and if the partner slows to 4 seconds the same traffic holds 800. A latency increase anywhere turns into a concurrency increase everywhere upstream of it.
Connection pools are Little’s law too
A database connection is busy for as long as a query holds it, so the connections in use on average are the query rate times the time each query holds a connection. A pool needs that much plus headroom for bursts. If the pool is smaller than Little’s law demands, no timeout or queue setting can save it: queries arrive faster than connections free up, and the line grows for as long as the load lasts.
// Connection pools by Little's law, and what happens when a pool is too small for the load.
const workloads = [
// name, queries a second at the peak, milliseconds each query holds a connection
["checkout service", 1200, 5],
["reporting API", 40, 250],
["search suggestions", 6000, 2],
];
for (const [name, qps, ms] of workloads) {
const busy = (qps * ms) / 1000; // Little's law: connections in use on average
const pool = Math.ceil(busy * 1.5); // 50% headroom for bursts
console.log(`${name.padEnd(20)} ${String(qps).padStart(5)}/s x ${String(ms).padStart(3)} ms = ${busy.toFixed(1).padStart(4)} busy, pool of ${pool}`);
}
// A pool of 4 connections for 1,200 queries a second of 5 ms each: 6 connections' worth of work.
// Queries arrive evenly; each takes the first free connection, or waits in line for one.
const POOL = 4;
const RATE = 1200; // queries a second
const HOLD = 0.005; // seconds each query holds a connection
const freeAt = new Array(POOL).fill(0); // when each connection is next free
console.log(`\nPool of ${POOL} for ${RATE} queries a second of ${HOLD * 1000} ms each:`);
let n = 0;
for (const stop of [1, 2, 5, 10]) {
let lastWait = 0;
for (; n < RATE * stop; n++) {
const arrives = n / RATE;
const first = freeAt.indexOf(Math.min(...freeAt));
const start = Math.max(arrives, freeAt[first]);
freeAt[first] = start + HOLD;
lastWait = start - arrives;
}
const waiting = Math.round(lastWait / HOLD * POOL);
console.log(`after ${String(stop).padStart(2)} s a new query waits ${(lastWait * 1000).toFixed(0).padStart(5)} ms, about ${waiting} queries ahead of it`);
} Output
checkout service 1200/s x 5 ms = 6.0 busy, pool of 9 reporting API 40/s x 250 ms = 10.0 busy, pool of 15 search suggestions 6000/s x 2 ms = 12.0 busy, pool of 18 Pool of 4 for 1200 queries a second of 5 ms each: after 1 s a new query waits 498 ms, about 399 queries ahead of it after 2 s a new query waits 998 ms, about 799 queries ahead of it after 5 s a new query waits 2498 ms, about 1999 queries ahead of it after 10 s a new query waits 4998 ms, about 3999 queries ahead of it
Recorded with Node.js 24.21.0 on macOS 26 arm64. To run it yourself: mise exec node@24.21.0 -- node pool_size.mjs
Runs on this device, in your browser. The first run downloads JavaScript (about 0.6 MB), which is kept for the next runs.
Your run, in this browser
The second half of the output is the failure in slow motion: a pool of 4 offered 6 connections’ worth of work falls further behind every second, by about 400 queries. Real pools then time out and fail requests, which is a better outcome than an unbounded wait, but still an outage.
The same law explains a classic failure of fleets that scale out quickly. If every worker, container or function
copy opens its own database connection, the connections grow with the number of copies, not with the queries in
progress. PostgreSQL’s max_connections is typically 100 by default
(PostgreSQL documentation), so a burst to a
few hundred copies can exhaust it even when only a dozen queries are running. A pooler in front of the database fixes
this by sharing a few server connections among many clients. PgBouncer, for example, can return a server connection
to the pool at the end of each transaction instead of when the client disconnects
(PgBouncer usage), and its default_pool_size caps server connections per
user and database pair at 20 unless configured otherwise (PgBouncer configuration).
Why queues explode near full utilisation
Utilisation is the share of time a server is busy: arrival rate times service time, for one server. It feels efficient to run servers at 90% or more, but waiting time does not grow in proportion to it. For the simplest queue, one server with random arrivals and random service times (the M/M/1 queue), the mean wait before service is
where ρ is the utilisation. This follows from the Pollaczek-Khinchin formula for queues with random arrivals (Gallager, section 4.5.5). At 50% the wait equals one service time, at 80% four, at 90% nine and at 95% nineteen. The script simulates the queue to check the formula and to show the tail, which the formula’s mean hides:
# One server, requests arriving at random (a Poisson process) and service times that vary at random (exponential):
# the M/M/1 queue. Simulate it at four utilisations and compare the mean wait with the formula.
import random
SERVICE_MS = 10.0 # average service time: the server can do 100 requests a second
REQUESTS = 200_000
def simulate(utilisation, seed):
rng = random.Random(seed)
arrival_rate = utilisation / SERVICE_MS # requests per millisecond
wait = 0.0 # how long the current request waits before its service starts
waits, totals = [], []
for _ in range(REQUESTS):
service = rng.expovariate(1 / SERVICE_MS)
waits.append(wait)
totals.append(wait + service)
# The next request waits for whatever work is still ahead of it when it arrives (zero if the server is idle).
wait = max(0.0, wait + service - rng.expovariate(arrival_rate))
totals.sort()
return sum(waits) / REQUESTS, sum(totals) / REQUESTS, totals[int(REQUESTS * 0.99)]
print(f"{'utilisation':>11}{'mean wait':>11}{'formula':>9}{'mean total':>12}{'p99 total':>11}")
for rho in (0.5, 0.8, 0.9, 0.95):
mean_wait, mean_total, p99 = simulate(rho, seed=1)
formula = rho / (1 - rho) * SERVICE_MS # mean wait in queue for M/M/1
print(f"{rho:>10.0%}{mean_wait:>8.0f} ms{formula:>6.0f} ms{mean_total:>9.0f} ms{p99:>8.0f} ms") Output
utilisation mean wait formula mean total p99 total
50% 10 ms 10 ms 20 ms 90 ms
80% 39 ms 40 ms 49 ms 228 ms
90% 86 ms 90 ms 96 ms 418 ms
95% 186 ms 190 ms 196 ms 712 ms
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 queue_sim.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
The simulated means land within a few percent of the formula, and the 99th percentile is several times the mean at every utilisation. Going from 50% to 95% busy multiplies the mean wait by nineteen and pushes the p99 from under a tenth of a second to most of a second, on a server whose work takes 10 ms. Real servers also slow themselves down as they get busy, which the model leaves out. Amazon’s Builders’ Library reports this from its load tests: threads compete for locks and for input and output, the processor spends more of its time switching between them and on garbage collection, and beyond some load latency climbs faster still (load shedding).
So choose a target utilisation from the latency you must meet, not from the cost you would like. A latency goal of a few service times means staying well below 80%; a background batch system with no latency goal can run hotter.
Each run of the simulation is one sample
The simulation uses a fixed seed, so it prints the same every time on the same version of Python. Change the seed and the rows at high utilisation move by several percent, because a busy queue swings far from its average for long stretches. That instability is part of the lesson: near saturation, even the average is hard to pin down.
Counting servers
The server count follows from the peak rate, what one server sustains, the utilisation you allow, and spares:
Measure the per-server rate with a load test on the real code instead of guessing it. Google’s SRE book counts regular load tests among the steps that capacity planning cannot skip, because they tell you how much traffic a given set of servers really serves (SRE book, introduction). Then add spares for the servers that will be down at the worst moment. Two spares leave room for one server to fail while another is out for a deploy; the SRE book uses “N + 2 redundancy” as its example of a capacity requirement of this kind (SRE book, chapter 18). For services that must survive the loss of a whole zone, the spares are a zone’s worth of servers: the Builders’ Library article above describes services scaled so that losing their servers in one Availability Zone still leaves enough capacity to meet the latency goals.
# Servers for the peak: divide by what one server sustains at a safe utilisation, then add spares.
import math
SERVICES = [
# name, peak requests a second (rounded up), requests a second one server sustains (from a load test), spares
("chat API", 210_000, 4_000, 2),
("news feed reads", 29_000, 1_500, 2),
("payments", 8_400, 500, 2),
]
print(f"{'Service':<17}{'peak/s':>9}{'at 50%':>8}{'at 70%':>8}{'at 90%':>8} (servers, with 2 spares)")
for name, peak, per_server, spares in SERVICES:
counts = [math.ceil(peak / (per_server * target)) + spares for target in (0.5, 0.7, 0.9)]
print(f"{name:<17}{peak:>9,}" + "".join(f"{n:>8}" for n in counts)) Output
Service peak/s at 50% at 70% at 90% (servers, with 2 spares) chat API 210,000 107 77 61 news feed reads 29,000 41 30 24 payments 8,400 36 26 21
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 servers.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
The peaks are the traffic lesson’s estimates, rounded up to two significant figures (up, because this is load), and the per-server rates are assumptions that a load test would replace. The table shows the trade: running at 90% instead of 50% saves about two fifths of the fleet, and costs the latency that the queue simulation measured.
Key takeaways
- Little’s law, L = λ × W, turns a rate and a latency into how much is in progress: requests in flight, threads, connections, jobs in a queue.
- A pool or fleet smaller than Little’s law demands falls further behind every second; latency increases upstream become concurrency increases.
- Queueing delay grows like ρ / (1 − ρ): four service times at 80% busy, nine at 90%, nineteen at 95%, with a tail several times the mean.
- Servers = ⌈peak ÷ (per-server rate × target utilisation)⌉ + spares, with the per-server rate from a load test and spares for failures and deploys.
Exercise
Exercise · Medium · Python
Size concurrency, a connection pool and a fleet
Write three capacity helpers in capacity.py.
in_flight(rate_per_s, latency_ms) returns how many requests are inside a system on average, by Little's law: the arrival rate times the average time each request spends inside. Mind the units: the latency is in milliseconds and the rate is per second.
pool_size(qps, query_ms, headroom=1.0) returns the connections a pool needs: the connections busy on average, multiplied by headroom (1.5 means 50% spare for bursts), rounded up to a whole connection, because a pool cannot hold half a connection.
servers_needed(peak_rps, per_server_rps, target_utilisation, spares) returns how many servers carry peak_rps when no server may run busier than target_utilisation (0.6 means 60% of what one server sustains), rounded up, plus spares servers for failures and deploys.
For example, in_flight(2000, 50) is 100, pool_size(1200, 5) is 6, and servers_needed(12_000, 800, 0.6, 2) is 27.
Starter code · capacity.py
import math
def in_flight(rate_per_s, latency_ms):
"""Average requests inside a system (Little's law)."""
# Replace this line with your code.
return 0
def pool_size(qps, query_ms, headroom=1.0):
"""Connections a pool needs, with headroom, rounded up."""
# Replace this line with your code.
return 0
def servers_needed(peak_rps, per_server_rps, target_utilisation, spares):
"""Servers for the peak at the target utilisation, rounded up, plus spares."""
# Replace this line with your code.
return 0 The sample tests · test_capacity.py
import math
from capacity import in_flight, pool_size, servers_needed
def test_in_flight():
"""multiplies the rate by the time inside, in seconds"""
assert math.isclose(in_flight(2_000, 50), 100, rel_tol=1e-9)
assert math.isclose(in_flight(3_000, 4), 12, rel_tol=1e-9)
assert math.isclose(in_flight(500, 30_000), 15_000, rel_tol=1e-9)
def test_pool_size():
"""sizes a pool by Little's law and rounds up"""
assert pool_size(1_200, 5) == 6
assert pool_size(1_000, 3) == 3
assert pool_size(10, 15) == 1
def test_pool_headroom():
"""adds headroom before rounding up"""
assert pool_size(1_200, 5, headroom=1.5) == 9
assert pool_size(40, 250, headroom=1.5) == 15
assert pool_size(100, 25, headroom=1.5) == 4 # 2.5 busy x 1.5 = 3.75, so 4 (rounding first would give 4.5)
def test_servers_needed():
"""divides the peak by what one server may take, rounds up and adds spares"""
assert servers_needed(12_000, 800, 0.6, 2) == 27
assert servers_needed(10_000, 1_000, 0.5, 0) == 20
assert servers_needed(10_001, 1_000, 0.5, 1) == 22 A hint
Convert milliseconds to seconds first (latency_ms / 1000) and multiply by the rate. math.ceil rounds up. For servers, one server may take per_server_rps * target_utilisation requests a second, so divide the peak by that, round up, then add the spares.
Results of the sample tests
| Test | Result | Details |
|---|
What your code printed
The sample tests run on this device, in your browser (Pyodide): nothing is sent to mysmartcopilot.com. The first run downloads Python (about 13.5 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.
Check yourself
5 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.
References
- Discrete Stochastic Processes, chapter 4: Renewal processes (Little's theorem, M/G/1 queueing delay) (MIT OpenCourseWare (Robert Gallager))
- Using load shedding to avoid overload (Amazon Builders' Library) (Amazon Web Services)
- PgBouncer usage (pooling modes) (PgBouncer project)
- PgBouncer configuration (default_pool_size) (PgBouncer project)
- PostgreSQL documentation: Connections and Authentication (max_connections) (The PostgreSQL Global Development Group)
- Site Reliability Engineering, chapter 18: Software Engineering in SRE (intent-based capacity planning) (Google (O'Reilly Media))
- Site Reliability Engineering, chapter 1: Introduction (demand forecasting and capacity planning) (Google (O'Reilly Media))
Related tools
Report a problem with this lesson
Kept only in this browser. Your Learn progress