System Design (High-Level Design) Module 2 – Back-of-the-envelope estimation
Estimating traffic: users to requests per second
Turn daily active users and actions per user into average and peak requests per second, split reads from writes and pick a peak factor from the traffic shape.
What you will learn
- Calculate average and peak requests per second from daily active users and actions per user
- Separate read traffic from write traffic and justify the ratio you assume
- Choose a peak factor from the shape of the traffic and the window it applies to
- Check an estimate against a published total by the share it implies
Before you start
On this page
The first sizing question of almost every design is how much traffic it must carry: how many requests a second, of which kind, at the worst moment. Everything else follows from that answer. A thousand requests a second fits comfortably on a few servers; a few hundred thousand writes a second needs a partitioned database and a careful write path. This lesson turns the numbers a product team gives you (users and what they do) into the numbers an engineer sizes things with (requests per second), and shows where the peak comes from instead of guessing it.
From users to requests per second
Start from daily active users (DAU): the people who use the product on a typical day. If you are given monthly active users instead, assume what share of them come on a given day and say so, because daily users can never be more than monthly ones. Then count what each of them makes in a day. Count requests to your servers, not taps on a screen: one screen may load with several API calls, and an app may sync in the background.
From users to requests per second
Text description of the diagram
The diagram is a chain of boxes from top to bottom.
- Daily active users, multiplied by the reads and writes each user makes in a day, give the requests a day.
- The requests a day, divided by 86,400 seconds, give the average requests a second.
- The average, multiplied by a peak factor for a stated window such as the busiest hour or second, gives the peak requests a second.
- The peak splits into two boxes: peak reads a second and peak writes a second, each the peak multiplied by its share of the traffic.
In symbols, with the 86,400 seconds of a day from the previous lesson:
This script applies the formula to three services with assumptions you can change:
# From users to requests per second, for three services. Every number in ASSUMPTIONS is an assumption to state
# out loud; change one and run the script again to see what it moves.
DAY = 86_400
ASSUMPTIONS = {
# daily active users, reads and writes per user per day, peak factor
"chat": {"dau": 50_000_000, "reads": 80, "writes": 40, "peak": 3},
"news feed": {"dau": 20_000_000, "reads": 50, "writes": 0.5, "peak": 2.5},
"payments": {"dau": 30_000_000, "reads": 4, "writes": 2, "peak": 4},
}
def estimate(dau, reads, writes, peak):
"""Average and peak requests per second, with the peak split into reads and writes."""
daily = dau * (reads + writes)
average = daily / DAY
return {
"daily": daily,
"average": average,
"peak": average * peak,
"peak_reads": dau * reads / DAY * peak,
"peak_writes": dau * writes / DAY * peak,
"ratio": reads / writes,
}
print("requests a day = daily active users x (reads + writes) per user")
print("average per second = requests a day / 86,400; peak = average x peak factor")
print()
print(f"{'Service':<11}{'per day':>15}{'average/s':>11}{'peak/s':>10}{'reads/s':>10}{'writes/s':>10}{'reads:writes':>14}")
for name, a in ASSUMPTIONS.items():
e = estimate(**a)
print(f"{name:<11}{e['daily']:>15,.0f}{e['average']:>11,.0f}{e['peak']:>10,.0f}{e['peak_reads']:>10,.0f}{e['peak_writes']:>10,.0f}{e['ratio']:>12,.0f}:1") Output
requests a day = daily active users x (reads + writes) per user average per second = requests a day / 86,400; peak = average x peak factor Service per day average/s peak/s reads/s writes/s reads:writes chat 6,000,000,000 69,444 208,333 138,889 69,444 2:1 news feed 1,010,000,000 11,690 29,225 28,935 289 100:1 payments 180,000,000 2,083 8,333 5,556 2,778 2:1
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 traffic.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
The three answers differ by more than an order of magnitude, and they point to different designs. That is the purpose of the exercise: not the exact numbers, but which part of the system the numbers say will be under pressure.
Reads and writes shape different parts of a system
Split every estimate into reads and writes, because they load different components. Reads can be answered from copies: a cache holds data close to the application to serve repeated reads faster (Cache-Aside pattern), and read replicas spread reads over several databases. Writes must reach the copy that owns the data, so heavy write traffic pushes towards partitioning, because a single server has finite storage, computing power and network bandwidth (Sharding pattern).
The ratio should come from how the product is used, not from habit:
- A news feed is read far more than it is written: a post is written once and read by many followers, and people scroll more than they post. The script assumes 100 reads per write, so the design is about the read path.
- Chat writes every message once and reads it once per recipient, so reads and writes stay within a small factor of each other, and both paths need attention.
- Payments have few reads per write compared with a feed, and every write must be durable and recorded exactly once. Their peak write rate, not their read rate, sets the hardest requirement.
A peak factor comes from the traffic’s shape
Traffic that people generate follows their days: low at night and higher while they are awake, with the busiest hours depending on the product. If you have real request counts, the peak factor is simply the busiest window divided by the average window. Without them, assume a daily shape, say why, and derive the factor from it. The window matters as much as the shape: the busiest second is busier than the busiest minute, which is busier than the busiest hour.
The script below assumes a daily curve with an evening peak, then simulates the busiest hour with requests that
arrive at random, independently of each other and at a steady average rate (a Poisson process,
Gallager).
Its random numbers come from Python’s random module, seeded with a fixed number, so any one Python version
prints identical output on every run (Python documentation):
# How big is the peak? It depends on the window you measure it over, and on how big the service is.
# An assumed day: the share of daily requests that arrives in each hour (a quiet night, an evening peak).
import random
HOURLY_SHARE = [ # percent of the day's requests, midnight to midnight
1.5, 1.0, 0.7, 0.6, 0.6, 0.9, 1.8, 3.2, 4.5, 5.0, 5.2, 5.3,
5.3, 5.2, 5.1, 5.0, 5.2, 5.4, 6.4, 7.4, 8.2, 7.9, 5.6, 3.0,
]
assert round(sum(HOURLY_SHARE), 6) == 100
busiest = max(HOURLY_SHARE)
print(f"Busiest hour: {busiest}% of the day, {busiest / (100 / 24):.2f} times an average hour")
def busiest_windows(daily_requests, seed):
"""Simulate the busiest hour with requests arriving at random (a Poisson process) and return the peak
factors over that hour, its busiest minute and its busiest second, each relative to the day's average."""
rng = random.Random(seed)
day_average = daily_requests / 86_400
rate = daily_requests * busiest / 100 / 3_600 # requests per second during the busiest hour
per_second = [0] * 3_600
t = rng.expovariate(rate)
while t < 3_600:
per_second[int(t)] += 1
t += rng.expovariate(rate)
per_minute = [sum(per_second[m * 60:(m + 1) * 60]) / 60 for m in range(60)]
return day_average, rate / day_average, max(per_minute) / day_average, max(per_second) / day_average
print()
print(f"{'Requests a day':>15}{'average/s':>11}{'peak hour':>11}{'peak minute':>13}{'peak second':>13}")
for daily in [20_000_000, 2_000_000, 200_000]:
average, hour, minute, second = busiest_windows(daily, seed=7)
print(f"{daily:>15,}{average:>11,.1f}{hour:>10.2f}x{minute:>12.2f}x{second:>12.2f}x") Output
Busiest hour: 8.2% of the day, 1.97 times an average hour
Requests a day average/s peak hour peak minute peak second
20,000,000 231.5 1.97x 2.00x 2.31x
2,000,000 23.1 1.97x 2.05x 3.11x
200,000 2.3 1.97x 2.23x 6.91x
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 peak_factor.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
Three results are worth carrying into every estimate:
- The daily shape sets the hourly factor. With this curve, the busiest hour carries about twice the traffic of an average hour, whatever the size of the service.
- Shorter windows have higher peaks. Within that hour, the busiest minute and the busiest second are higher again, because random arrivals bunch together.
- Small services are spikier. For the service with 200,000 requests a day, the busiest second was about seven times the daily average; for the one with 20 million, about 2.3 times. Bunching that adds a few requests is a large share of a small average.
Real traffic can be much burstier than this model, because people do not always act independently: a push notification, a goal in a match or the opening minute of a sale makes thousands act at once, and retries after an outage add more. Google’s SRE book asks capacity plans to cover both kinds of demand: the steady growth that comes from people adopting a product, and the jumps that launches, marketing pushes and other business events cause (SRE book, introduction). For a known event, estimate its peak on its own from the people expected and the time over which they arrive.
SQLite Online (SQL Playground) Import an exported request log as a CSV file and count requests per hour or per minute with SQL, to measure a real peak factor. Line Graph Maker Plot hourly request counts to see the shape of a day before you choose a factor.Check the share an estimate implies
When a public total exists for the whole market, compare your estimate with it. The question is not whether the numbers match, because they should not, but whether the share your estimate implies is believable for the product you are designing.
Checking a payments estimate against published volumes
The Reserve Bank of India’s Payment System Indicators give the monthly volume of each payment system, UPI included, in lakh. One edition’s UPI row reads 2,45,089.58 lakh for a 31-day month. Spread over that month, it is the average rate of the whole network, and the payments estimate above can be set against it:
# Compare an estimate with a published total: is the share it implies believable?
LAKH = 10**5
DAY = 86_400
published = 2_45_089.58 * LAKH # a month's volume as the payment tables print it, in lakh
network = published / (31 * DAY) # spread over a 31-day month
ours = 30_000_000 * 2 / DAY # the payments estimate above: 30 million users, 2 payments each a day
print(f"Whole network: {network:,.0f} payments a second on average")
print(f"Our estimate: {ours:,.0f} payments a second on average, {ours / network:.1%} of the network") Output
Whole network: 9,151 payments a second on average Our estimate: 694 payments a second on average, 7.6% of the network
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 sanity_check.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
For a new product, handling about one payment in every thirteen made across the whole country is far too high a share, so the number of daily users is the assumption to revisit. Festival sales and big cricket matches are the kind of known events that deserve their own peak estimate, made from the expected audience and the minutes in which it arrives.
Mistakes to check
- Monthly users counted as daily users. That inflates every rate by the ratio between them.
- Taps counted instead of requests. One action can be several requests, and background sync adds more.
- A peak factor without a window. “3 times” means little until you say whether it is the busiest hour or the busiest second.
- The average used for capacity. The average is the lowest load your busiest hour will bring, not a target to size for.
Key takeaways
- Average requests per second = DAU × requests per user per day ÷ 86,400; the peak is the average times a peak factor for a stated window.
- Split reads from writes: read-heavy traffic shapes caches and replicas, write-heavy traffic shapes partitioning and the write path.
- Derive the peak factor from the traffic’s shape: about 2 for the busiest hour of the daily curve assumed here, more for shorter windows and for small services, and a separate estimate for known events.
- Sanity-check an estimate by the share of a published total it implies.
Exercise
Exercise · Easy · Python
Turn users into average and peak requests per second
Write the two functions of a traffic estimate in estimate.py.
estimate(dau, reads_per_user, writes_per_user, peak_factor) returns a dictionary with four rates, all in requests per second:
"average": the day's requests,dau × (reads_per_user + writes_per_user), spread evenly over 86,400 seconds;"peak": the average multiplied bypeak_factor;"peak_reads"and"peak_writes": the peak split in the same proportion as reads and writes, so that the two add up to the peak.
For example, 1 million users who each make 8 reads and 2 writes a day, with a peak factor of 3, give an average of about 115.7 requests a second and a peak of about 347.2, of which 277.8 are reads.
peak_factor(counts) measures a peak factor from data: given a list of request counts, one per window (per hour, say), it returns the busiest window's count divided by the average count. An empty list raises ValueError.
Starter code · estimate.py
DAY = 86_400
def estimate(dau, reads_per_user, writes_per_user, peak_factor):
"""Return {"average", "peak", "peak_reads", "peak_writes"} in requests per second."""
# Replace this line with your code.
return {}
def peak_factor(counts):
"""Return the busiest window's count divided by the average count; raise ValueError for an empty list."""
# Replace this line with your code.
return 0 The sample tests · test_estimate.py
import math
from estimate import estimate, peak_factor
def raises_value_error(function, *args):
"""True when function(*args) raises ValueError."""
try:
function(*args)
except ValueError:
return True
return False
def test_average():
"""spreads the day's requests over 86,400 seconds"""
result = estimate(1_000_000, 8, 2, 3)
assert math.isclose(result["average"], 115.740741, rel_tol=1e-6)
def test_peak():
"""multiplies the average by the peak factor"""
result = estimate(1_000_000, 8, 2, 3)
assert math.isclose(result["peak"], 347.222222, rel_tol=1e-6)
assert math.isclose(estimate(86_400, 1, 0.5, 2)["peak"], 3, rel_tol=1e-9)
def test_split():
"""splits the peak into reads and writes that add up to it"""
result = estimate(1_000_000, 8, 2, 3)
assert math.isclose(result["peak_reads"], 277.777778, rel_tol=1e-6)
assert math.isclose(result["peak_writes"], 69.444444, rel_tol=1e-6)
assert math.isclose(result["peak_reads"] + result["peak_writes"], result["peak"], rel_tol=1e-9)
def test_peak_factor():
"""divides the busiest window by the average window"""
assert peak_factor([10, 10, 10, 10]) == 1
assert peak_factor([5, 15]) == 1.5
assert peak_factor([0, 0, 6]) == 3
def test_empty_counts():
"""refuses an empty list of counts"""
assert raises_value_error(peak_factor, []) A hint
Compute the daily total once, dau * (reads_per_user + writes_per_user), and divide by DAY for the average. For the split, the reads' share of the traffic is reads_per_user / (reads_per_user + writes_per_user). For the peak factor, max(counts) divided by sum(counts) / len(counts) is the whole calculation.
Results of the sample tests
| Test | Result | Details |
|---|
What your code printed
The sample tests run on this device, in your browser (Pyodide): nothing is sent to mysmartcopilot.com. The first run downloads Python (about 13.5 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.
Check yourself
5 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.
References
- Site Reliability Engineering, chapter 1: Introduction (demand forecasting and capacity planning) (Google (O'Reilly Media))
- Cache-Aside pattern (Azure Architecture Center) (Microsoft)
- Sharding pattern (Azure Architecture Center) (Microsoft)
- random, generate pseudo-random numbers (Python Software Foundation)
- Discrete Stochastic Processes, chapter 2: Poisson processes (MIT OpenCourseWare (Robert Gallager))
- Payment System Indicators (Reserve Bank of India)
Related tools
Report a problem with this lesson
Kept only in this browser. Your Learn progress