Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

System Design (High-Level Design) Module 2 – Back-of-the-envelope estimation

Little's law, utilisation and server counts

Use Little's law to size concurrency and connection pools, see why queueing delay explodes near full utilisation, and count servers with headroom for failures.

  • Intermediate
  • 25 minutes
  • Examples run with Python 3.14.8, Pyodide 314.0.7, Node.js 24.21.0 and quickjs 0.32.0
  • By MySmartCoPilot

What you will learn

  • Apply Little's law to find requests in flight, thread pool sizes and connection pool sizes
  • Explain why queueing delay rises sharply as utilisation approaches 100%
  • Estimate server and connection counts with headroom for bursts, failures and deploys

Before you start

On this page

Requests per second tell you how fast work arrives. To size threads, connections and servers you also need to know how much work is in progress at once, and how busy you can let each server get before waiting times run away. Two results from queueing theory answer both questions, and neither needs more than school arithmetic: Little’s law for how much is in flight, and the shape of queueing delay for how much headroom to keep.

Little’s law: rate times time is what is inside

Take any system that requests enter and leave: a web server, a database, a queue with its workers. If λ requests arrive per second and each spends W seconds inside on average, then on average L requests are inside:

L=λ×WL = \lambda \times W

The law asks for remarkably little: it does not matter in which order requests are served or how their arrival and service times are spread, only that the system is stable, with requests leaving as fast as they arrive (Gallager, section 4.5.4).

Requests arrive at λ a second, wait in a queue, are served and leave; L requests are inside the system, each for W seconds.Requests arriveλ a secondInside: L requests,W seconds eachRequests leaveλ a second, once stableQueuewaiting for a serverServersworking on requestsnext in line

Little's law: L = λ × W

Text description of the diagram

The diagram shows a system from top to bottom.

  1. Requests arrive at a rate of λ (lambda) a second.
  2. They enter the system, a box labelled "Inside: L requests, W seconds each". It holds a queue, where requests wait for a server, and the servers, which work on them. Requests move from the queue to the servers in turn.
  3. Requests leave the system, at λ a second too once the system is stable.

On average there are L requests inside the system, waiting or being served, and each spends W seconds inside. Little's law says that L = λ × W.

That makes it the quickest way to turn a rate and a latency into a count of things you have to provide:

Little's law for four kinds of system Python · in_flight.py
# Little's law: items in a system = arrival rate x average time each item spends in it (L = lambda x W).
CASES = [
    # what arrives, arrivals per second, average seconds inside, what the result counts
    ("web requests", 2_000, 0.050, "requests in flight"),
    ("database queries", 3_000, 0.004, "connections busy"),
    ("calls to a slow partner API", 200, 0.800, "threads waiting on it"),
    ("jobs through a queue and its workers", 500, 30.0, "jobs in the system"),
]

for name, rate, seconds, unit in CASES:
    print(f"{name:<38}{rate:>6,}/s x {seconds * 1000:>6,.0f} ms = {rate * seconds:>7,.0f} {unit}")

Output

web requests                           2,000/s x     50 ms =     100 requests in flight
database queries                       3,000/s x      4 ms =      12 connections busy
calls to a slow partner API              200/s x    800 ms =     160 threads waiting on it
jobs through a queue and its workers     500/s x 30,000 ms =  15,000 jobs in the system

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 in_flight.py

Read each line as a sizing decision. A web tier at 2,000 requests a second and 50 ms per request has 100 requests in progress, so it needs at least 100 concurrent workers across the fleet. The slow partner API is the dangerous line: at 800 ms per call, a modest 200 calls a second hold 160 threads, and if the partner slows to 4 seconds the same traffic holds 800. A latency increase anywhere turns into a concurrency increase everywhere upstream of it.

Connection pools are Little’s law too

A database connection is busy for as long as a query holds it, so the connections in use on average are the query rate times the time each query holds a connection. A pool needs that much plus headroom for bursts. If the pool is smaller than Little’s law demands, no timeout or queue setting can save it: queries arrive faster than connections free up, and the line grows for as long as the load lasts.

Pool sizes, and a pool that is too small JavaScript · pool_size.mjs
// Connection pools by Little's law, and what happens when a pool is too small for the load.
const workloads = [
  // name, queries a second at the peak, milliseconds each query holds a connection
  ["checkout service", 1200, 5],
  ["reporting API", 40, 250],
  ["search suggestions", 6000, 2],
];

for (const [name, qps, ms] of workloads) {
  const busy = (qps * ms) / 1000; // Little's law: connections in use on average
  const pool = Math.ceil(busy * 1.5); // 50% headroom for bursts
  console.log(`${name.padEnd(20)} ${String(qps).padStart(5)}/s x ${String(ms).padStart(3)} ms = ${busy.toFixed(1).padStart(4)} busy, pool of ${pool}`);
}

// A pool of 4 connections for 1,200 queries a second of 5 ms each: 6 connections' worth of work.
// Queries arrive evenly; each takes the first free connection, or waits in line for one.
const POOL = 4;
const RATE = 1200; // queries a second
const HOLD = 0.005; // seconds each query holds a connection
const freeAt = new Array(POOL).fill(0); // when each connection is next free
console.log(`\nPool of ${POOL} for ${RATE} queries a second of ${HOLD * 1000} ms each:`);
let n = 0;
for (const stop of [1, 2, 5, 10]) {
  let lastWait = 0;
  for (; n < RATE * stop; n++) {
    const arrives = n / RATE;
    const first = freeAt.indexOf(Math.min(...freeAt));
    const start = Math.max(arrives, freeAt[first]);
    freeAt[first] = start + HOLD;
    lastWait = start - arrives;
  }
  const waiting = Math.round(lastWait / HOLD * POOL);
  console.log(`after ${String(stop).padStart(2)} s a new query waits ${(lastWait * 1000).toFixed(0).padStart(5)} ms, about ${waiting} queries ahead of it`);
}

Output

checkout service      1200/s x   5 ms =  6.0 busy, pool of 9
reporting API           40/s x 250 ms = 10.0 busy, pool of 15
search suggestions    6000/s x   2 ms = 12.0 busy, pool of 18

Pool of 4 for 1200 queries a second of 5 ms each:
after  1 s a new query waits   498 ms, about 399 queries ahead of it
after  2 s a new query waits   998 ms, about 799 queries ahead of it
after  5 s a new query waits  2498 ms, about 1999 queries ahead of it
after 10 s a new query waits  4998 ms, about 3999 queries ahead of it

Recorded with Node.js 24.21.0 on macOS 26 arm64. To run it yourself: mise exec node@24.21.0 -- node pool_size.mjs

The second half of the output is the failure in slow motion: a pool of 4 offered 6 connections’ worth of work falls further behind every second, by about 400 queries. Real pools then time out and fail requests, which is a better outcome than an unbounded wait, but still an outage.

The same law explains a classic failure of fleets that scale out quickly. If every worker, container or function copy opens its own database connection, the connections grow with the number of copies, not with the queries in progress. PostgreSQL’s max_connections is typically 100 by default (PostgreSQL documentation), so a burst to a few hundred copies can exhaust it even when only a dozen queries are running. A pooler in front of the database fixes this by sharing a few server connections among many clients. PgBouncer, for example, can return a server connection to the pool at the end of each transaction instead of when the client disconnects (PgBouncer usage), and its default_pool_size caps server connections per user and database pair at 20 unless configured otherwise (PgBouncer configuration).

JavaScript & TypeScript Online Compiler Change the pool size or the query time in the simulation above and watch the queue.

Why queues explode near full utilisation

Utilisation is the share of time a server is busy: arrival rate times service time, for one server. It feels efficient to run servers at 90% or more, but waiting time does not grow in proportion to it. For the simplest queue, one server with random arrivals and random service times (the M/M/1 queue), the mean wait before service is

Wq=ρ1−ρ×average service timeW_q = \frac{\rho}{1 - \rho} \times \text{average service time}

where ρ is the utilisation. This follows from the Pollaczek-Khinchin formula for queues with random arrivals (Gallager, section 4.5.5). At 50% the wait equals one service time, at 80% four, at 90% nine and at 95% nineteen. The script simulates the queue to check the formula and to show the tail, which the formula’s mean hides:

One server at four utilisations, simulated Python · queue_sim.py
# One server, requests arriving at random (a Poisson process) and service times that vary at random (exponential):
# the M/M/1 queue. Simulate it at four utilisations and compare the mean wait with the formula.
import random

SERVICE_MS = 10.0  # average service time: the server can do 100 requests a second
REQUESTS = 200_000


def simulate(utilisation, seed):
    rng = random.Random(seed)
    arrival_rate = utilisation / SERVICE_MS  # requests per millisecond
    wait = 0.0  # how long the current request waits before its service starts
    waits, totals = [], []
    for _ in range(REQUESTS):
        service = rng.expovariate(1 / SERVICE_MS)
        waits.append(wait)
        totals.append(wait + service)
        # The next request waits for whatever work is still ahead of it when it arrives (zero if the server is idle).
        wait = max(0.0, wait + service - rng.expovariate(arrival_rate))
    totals.sort()
    return sum(waits) / REQUESTS, sum(totals) / REQUESTS, totals[int(REQUESTS * 0.99)]


print(f"{'utilisation':>11}{'mean wait':>11}{'formula':>9}{'mean total':>12}{'p99 total':>11}")
for rho in (0.5, 0.8, 0.9, 0.95):
    mean_wait, mean_total, p99 = simulate(rho, seed=1)
    formula = rho / (1 - rho) * SERVICE_MS  # mean wait in queue for M/M/1
    print(f"{rho:>10.0%}{mean_wait:>8.0f} ms{formula:>6.0f} ms{mean_total:>9.0f} ms{p99:>8.0f} ms")

Output

utilisation  mean wait  formula  mean total  p99 total
       50%      10 ms    10 ms       20 ms      90 ms
       80%      39 ms    40 ms       49 ms     228 ms
       90%      86 ms    90 ms       96 ms     418 ms
       95%     186 ms   190 ms      196 ms     712 ms

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 queue_sim.py

The simulated means land within a few percent of the formula, and the 99th percentile is several times the mean at every utilisation. Going from 50% to 95% busy multiplies the mean wait by nineteen and pushes the p99 from under a tenth of a second to most of a second, on a server whose work takes 10 ms. Real servers also slow themselves down as they get busy, which the model leaves out. Amazon’s Builders’ Library reports this from its load tests: threads compete for locks and for input and output, the processor spends more of its time switching between them and on garbage collection, and beyond some load latency climbs faster still (load shedding).

So choose a target utilisation from the latency you must meet, not from the cost you would like. A latency goal of a few service times means staying well below 80%; a background batch system with no latency goal can run hotter.

Each run of the simulation is one sample

The simulation uses a fixed seed, so it prints the same every time on the same version of Python. Change the seed and the rows at high utilisation move by several percent, because a busy queue swings far from its average for long stretches. That instability is part of the lesson: near saturation, even the average is hard to pin down.

Counting servers

The server count follows from the peak rate, what one server sustains, the utilisation you allow, and spares:

servers=⌈peak requests a secondper-server requests a second×target utilisation⌉+spares\text{servers} = \left\lceil \frac{\text{peak requests a second}}{\text{per-server requests a second} \times \text{target utilisation}} \right\rceil + \text{spares}

Measure the per-server rate with a load test on the real code instead of guessing it. Google’s SRE book counts regular load tests among the steps that capacity planning cannot skip, because they tell you how much traffic a given set of servers really serves (SRE book, introduction). Then add spares for the servers that will be down at the worst moment. Two spares leave room for one server to fail while another is out for a deploy; the SRE book uses “N + 2 redundancy” as its example of a capacity requirement of this kind (SRE book, chapter 18). For services that must survive the loss of a whole zone, the spares are a zone’s worth of servers: the Builders’ Library article above describes services scaled so that losing their servers in one Availability Zone still leaves enough capacity to meet the latency goals.

Servers for three services at three target utilisations Python · servers.py
# Servers for the peak: divide by what one server sustains at a safe utilisation, then add spares.
import math

SERVICES = [
    # name, peak requests a second (rounded up), requests a second one server sustains (from a load test), spares
    ("chat API", 210_000, 4_000, 2),
    ("news feed reads", 29_000, 1_500, 2),
    ("payments", 8_400, 500, 2),
]

print(f"{'Service':<17}{'peak/s':>9}{'at 50%':>8}{'at 70%':>8}{'at 90%':>8}   (servers, with 2 spares)")
for name, peak, per_server, spares in SERVICES:
    counts = [math.ceil(peak / (per_server * target)) + spares for target in (0.5, 0.7, 0.9)]
    print(f"{name:<17}{peak:>9,}" + "".join(f"{n:>8}" for n in counts))

Output

Service             peak/s  at 50%  at 70%  at 90%   (servers, with 2 spares)
chat API           210,000     107      77      61
news feed reads     29,000      41      30      24
payments             8,400      36      26      21

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 servers.py

The peaks are the traffic lesson’s estimates, rounded up to two significant figures (up, because this is load), and the per-server rates are assumptions that a load test would replace. The table shows the trade: running at 90% instead of 50% saves about two fifths of the fleet, and costs the latency that the queue simulation measured.

Key takeaways

  • Little’s law, L = λ × W, turns a rate and a latency into how much is in progress: requests in flight, threads, connections, jobs in a queue.
  • A pool or fleet smaller than Little’s law demands falls further behind every second; latency increases upstream become concurrency increases.
  • Queueing delay grows like ρ / (1 − ρ): four service times at 80% busy, nine at 90%, nineteen at 95%, with a tail several times the mean.
  • Servers = ⌈peak ÷ (per-server rate × target utilisation)⌉ + spares, with the per-server rate from a load test and spares for failures and deploys.

Exercise

Exercise · Medium · Python

Size concurrency, a connection pool and a fleet

Write three capacity helpers in capacity.py.

in_flight(rate_per_s, latency_ms) returns how many requests are inside a system on average, by Little's law: the arrival rate times the average time each request spends inside. Mind the units: the latency is in milliseconds and the rate is per second.

pool_size(qps, query_ms, headroom=1.0) returns the connections a pool needs: the connections busy on average, multiplied by headroom (1.5 means 50% spare for bursts), rounded up to a whole connection, because a pool cannot hold half a connection.

servers_needed(peak_rps, per_server_rps, target_utilisation, spares) returns how many servers carry peak_rps when no server may run busier than target_utilisation (0.6 means 60% of what one server sustains), rounded up, plus spares servers for failures and deploys.

For example, in_flight(2000, 50) is 100, pool_size(1200, 5) is 6, and servers_needed(12_000, 800, 0.6, 2) is 27.

Starter code · capacity.py

import math


def in_flight(rate_per_s, latency_ms):
    """Average requests inside a system (Little's law)."""
    # Replace this line with your code.
    return 0


def pool_size(qps, query_ms, headroom=1.0):
    """Connections a pool needs, with headroom, rounded up."""
    # Replace this line with your code.
    return 0


def servers_needed(peak_rps, per_server_rps, target_utilisation, spares):
    """Servers for the peak at the target utilisation, rounded up, plus spares."""
    # Replace this line with your code.
    return 0
The sample tests · test_capacity.py
import math

from capacity import in_flight, pool_size, servers_needed


def test_in_flight():
    """multiplies the rate by the time inside, in seconds"""
    assert math.isclose(in_flight(2_000, 50), 100, rel_tol=1e-9)
    assert math.isclose(in_flight(3_000, 4), 12, rel_tol=1e-9)
    assert math.isclose(in_flight(500, 30_000), 15_000, rel_tol=1e-9)


def test_pool_size():
    """sizes a pool by Little's law and rounds up"""
    assert pool_size(1_200, 5) == 6
    assert pool_size(1_000, 3) == 3
    assert pool_size(10, 15) == 1


def test_pool_headroom():
    """adds headroom before rounding up"""
    assert pool_size(1_200, 5, headroom=1.5) == 9
    assert pool_size(40, 250, headroom=1.5) == 15
    assert pool_size(100, 25, headroom=1.5) == 4  # 2.5 busy x 1.5 = 3.75, so 4 (rounding first would give 4.5)


def test_servers_needed():
    """divides the peak by what one server may take, rounds up and adds spares"""
    assert servers_needed(12_000, 800, 0.6, 2) == 27
    assert servers_needed(10_000, 1_000, 0.5, 0) == 20
    assert servers_needed(10_001, 1_000, 0.5, 1) == 22
A hint

Convert milliseconds to seconds first (latency_ms / 1000) and multiply by the rate. math.ceil rounds up. For servers, one server may take per_server_rps * target_utilisation requests a second, so divide the peak by that, round up, then add the spares.

The sample tests run on this device, in your browser (Pyodide): nothing is sent to mysmartcopilot.com. The first run downloads Python (about 13.5 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.

Check yourself

5 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.

  1. Question 1 of 5 A service receives 2,000 requests a second and each spends 50 ms inside it. How many requests are inside at any moment, on average?

    Type a number.

    Show the answer to question 1

    Answer: 100 requests (anything from 95 to 105 counts)

    Little's law: L = λ × W = 2,000 a second × 0.05 s = 100. Convert milliseconds to seconds first; 2,000 × 50 gives a number a thousand times too big.

  2. Question 2 of 5 A service runs 1,200 queries a second against its database, and each holds a connection for 5 ms. How many connections are busy on average?

    Type a number.

    Show the answer to question 2

    Answer: 6 connections

    1,200 × 0.005 = 6 connections busy on average. A pool needs more than that for bursts, but a pool of 4 can never keep up: the queue for connections grows for as long as the load lasts.

  3. Question 3 of 5 For an M/M/1 queue, how does the mean wait change when utilisation goes from 80% to 90%?

    Choose one answer.

    Show the answer to question 3

    Answer: It more than doubles, from 4 to 9 times the average service time

    The mean wait in queue is ρ / (1 − ρ) service times: 0.8 / 0.2 = 4 and 0.9 / 0.1 = 9. Every step towards 100% costs more than the one before, which is why fleets run well below saturation.

  4. Question 4 of 5 The peak is 12,000 requests a second, one server sustains 800 a second in a load test, no server should run above 60% of that, and you want 2 spares. How many servers?

    Type a number.

    Show the answer to question 4

    Answer: 27 servers

    One server may take 800 × 0.6 = 480 a second, so the peak needs 12,000 / 480 = 25 servers, plus 2 spares for failures and deploys: 27.

  5. Question 5 of 5 A function platform runs 500 copies of a function at the peak, and each copy opens its own database connection. The database allows about 100 connections by default. What helps most?

    Choose one answer.

    Show the answer to question 5

    Answer: A connection pooler in front of the database, so that copies share a small pool of server connections

    Concurrency is what fills connections: 500 copies holding one each need 500. A pooler such as PgBouncer hands a few server connections to many clients, for example one per transaction, so the database sees only the connections Little's law says are busy.

References

Related tools

Report a problem with this lesson

Quick answers and tool search

Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.