System Design (High-Level Design) Module 4 – Networking for system designers
The life of a web request, hop by hop
Follow one HTTPS request through DNS, TCP, TLS, a load balancer, an app server and a database, with recorded curl timings that show where the time goes.
What you will learn
- Trace a request through DNS, TCP, TLS, HTTP, a load balancer, a service and a database
- Measure where the time goes with curl's timing variables
- Identify which hops a cache or a CDN can remove
Before you start
On this page
When you open https://example.com/, your browser first asks DNS for the server’s address, then opens a TCP
connection, then runs a TLS handshake to agree on keys and check the certificate, and only then sends the HTTP
request. In the data centre the request usually passes a load balancer, reaches an app server, and the app asks a
database or a cache before the answer travels back. Every hop adds time and is one more thing that can fail. On a
long route most of that time is spent waiting for round trips before the server does any work at all, which is why
connection reuse, CDNs and caches matter so much.
This lesson follows one request hop by hop, measures each phase on real servers near and far, and shows which hops a cache or a CDN can take away.
The hops, in order
One HTTPS request on a new connection, hop by hop
Text description of the diagram
Five participants stand in columns, from left to right: the browser, a DNS resolver, the load balancer (where TLS ends), the app server and the database. Time runs from top to bottom.
- The browser asks the DNS resolver for the address of the name in the URL.
- The resolver answers with the address; the answer may be cached for the record's TTL.
- The browser sends a TCP SYN to the load balancer.
- The load balancer answers with SYN-ACK. One round trip has passed.
- The browser completes the TCP handshake with an ACK and starts TLS with a ClientHello.
- The load balancer answers with its ServerHello, certificate and Finished message. A second round trip has passed.
- The browser sends its own Finished message and, right behind it, the HTTP request GET /.
- The load balancer passes the request to an app server over a connection it keeps open.
- The app server queries the database over a pooled connection.
- The database returns rows.
- The app server answers the load balancer with 200 OK.
- The load balancer sends the response to the browser. The first byte arrives one more round trip after the request left, plus the time the servers spent on it.
- The name lookup. The browser needs an IP address for the name in the URL. It asks a recursive resolver, which answers from its cache or asks the servers that are authoritative for the name. Answers carry a TTL, the time limit on how long they may be kept in a cache (RFC 1034), so the browser, the operating system and the resolver all keep copies. A cached answer costs almost nothing; an uncached one can cost several round trips.
- The TCP handshake. TCP opens a connection with a three-way handshake: SYN, SYN-ACK, ACK (RFC 9293). The client’s side of the connection is open as soon as the SYN-ACK arrives and it can send right behind its ACK, so the handshake costs one round trip before anything useful can be sent.
- The TLS handshake. TLS 1.3 agrees on keys and proves the server’s identity in one round trip: the client’s first message already carries its key share, and the server answers with its own share, its certificate and a Finished message (RFC 8446). If the server cannot use the client’s key share, it asks for another one with a HelloRetryRequest, and the handshake costs a second round trip.
- The request and the response. The request goes out with the client’s Finished message, the servers do their work, and the response comes back: one more round trip plus the work. HTTP’s meaning, its methods and status codes, is defined once in RFC 9110, whatever version carries it.
- The hops inside the data centre. The address usually belongs to a load balancer or a reverse proxy, often the place where TLS ends. It passes the request to an app server, which reads from a cache or a database and may call other services. Round trips inside one data centre take around a millisecond or less, and these hops’ connections are kept open in pools, so they add little time; but each is a component with its own way to fail.
So a request on a new connection pays about three round trips before its first byte arrives: TCP, TLS and the request itself. The servers’ own work comes on top.
Measuring each phase with curl
curl’s --write-out option prints timing variables after a transfer. Each is the time from the start of the
transfer until a phase finished, so a phase’s length is the difference between two of them
(curl manual):
time_namelookup: until the name was resolved;time_connect: until the TCP connection to the host was complete;time_appconnect: until the TLS handshake was complete;time_starttransfer: until the first byte of the response arrived, including the server’s work;time_total: until the whole transfer ended.num_connectscounts the new connections the transfer made.
curl reuses a connection only for the addresses of one command, so measure.sh names each address twice: the
second transfer runs on the connection the first one opened. The script was run once, from one computer on one
network, against three public servers: example.com (a name reserved for documentation by
RFC 6761), a software mirror in Japan and one in Australia, five
times each. curl_runs.txt is what it printed, and timeline.py turns those cumulative times into phases:
timeline/timeline.py
"""Turns curl's timings (curl_runs.txt, recorded by measure.sh) into the time each phase of a request took."""
from collections import defaultdict
from pathlib import Path
from statistics import median
from urllib.parse import urlsplit
RUNS = Path(__file__).resolve().with_name("curl_runs.txt")
def phases(new_connection, dns, connect, tls_done, first_byte):
"""Milliseconds per phase. curl's times are cumulative seconds from the start of the transfer."""
ms = lambda seconds: seconds * 1000
if not new_connection:
# A reused connection has no TCP or TLS handshake: curl reports 0 for both.
return {"dns": ms(dns), "tcp": 0.0, "tls": 0.0, "wait": ms(first_byte - dns), "first byte": ms(first_byte)}
return {
"dns": ms(dns),
"tcp": ms(connect - dns),
"tls": ms(tls_done - connect),
"wait": ms(first_byte - tls_done),
"first byte": ms(first_byte),
}
runs = defaultdict(list)
versions = {}
for line in RUNS.read_text().splitlines():
url, version, connects, dns, connect, tls_done, first_byte, _total = line.split()
host = urlsplit(url).hostname
new = connects != "0"
versions[host] = version
runs[host, new].append(phases(new, float(dns), float(connect), float(tls_done), float(first_byte)))
hosts = list(dict.fromkeys(host for host, _ in runs))
mid = lambda host, new, key: median(p[key] for p in runs[host, new])
print(f"Medians of {len(runs[hosts[0], True])} runs, in milliseconds. A new connection:")
print(f"{'address':<22}{'HTTP':<6}{'DNS':>4}{'TCP':>6}{'TLS':>6}{'wait':>6}{'first byte':>12}{'TLS / TCP':>11}{'handshakes':>12}")
for host in hosts:
dns, tcp, tls, wait, first = (mid(host, True, k) for k in ("dns", "tcp", "tls", "wait", "first byte"))
share = (dns + tcp + tls) / first
print(f"{host:<22}{versions[host]:<6}{dns:>4.0f}{tcp:>6.0f}{tls:>6.0f}{wait:>6.0f}{first:>12.0f}{tls / tcp:>11.1f}{share:>12.0%}")
print()
print("The same address again, on the connection the first request opened:")
print(f"{'address':<22}{'first byte':>12}{'saved':>8}{'first byte / TCP':>18}")
for host in hosts:
first, again, tcp = mid(host, True, "first byte"), mid(host, False, "first byte"), mid(host, True, "tcp")
print(f"{host:<22}{again:>12.0f}{first - again:>8.0f}{again / tcp:>18.1f}") Output
Medians of 5 runs, in milliseconds. A new connection: address HTTP DNS TCP TLS wait first byte TLS / TCP handshakes example.com 2 3 9 17 18 46 1.8 62% ftp.jaist.ac.jp 1.1 2 140 293 245 680 2.1 64% mirror.aarnet.edu.au 2 2 293 295 324 913 1.0 65% The same address again, on the connection the first request opened: address first byte saved first byte / TCP example.com 17 29 1.9 ftp.jaist.ac.jp 176 505 1.3 mirror.aarnet.edu.au 344 568 1.2
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 timeline.py
timeline/curl_runs.txt
https://example.com/ 2 1 0.014236 0.023428 0.042566 0.060607 0.060759
https://example.com/ 2 0 0.000009 0.000000 0.000000 0.017412 0.019786
https://ftp.jaist.ac.jp/ 1.1 1 0.002938 0.139269 0.422442 0.562435 0.562561
https://ftp.jaist.ac.jp/ 1.1 0 0.000011 0.000000 0.000000 0.138319 0.138409
https://mirror.aarnet.edu.au/ 2 1 0.002485 0.303998 0.611080 0.947506 0.947605
https://mirror.aarnet.edu.au/ 2 0 0.000007 0.000000 0.000000 0.352795 0.352880
https://example.com/ 2 1 0.002777 0.013391 0.040009 0.065077 0.065276
https://example.com/ 2 0 0.000008 0.000000 0.000000 0.015161 0.017584
https://ftp.jaist.ac.jp/ 1.1 1 0.002539 0.164622 0.503889 0.675654 0.675797
https://ftp.jaist.ac.jp/ 1.1 0 0.000008 0.000000 0.000000 0.233633 0.233720
https://mirror.aarnet.edu.au/ 2 1 0.001834 0.294950 0.589363 0.912957 0.913080
https://mirror.aarnet.edu.au/ 2 0 0.000007 0.000000 0.000000 0.347612 0.347747
https://example.com/ 2 1 0.002952 0.011532 0.025747 0.041621 0.041737
https://example.com/ 2 0 0.000008 0.000000 0.000000 0.017630 0.017741
https://ftp.jaist.ac.jp/ 1.1 1 0.002437 0.142863 0.435800 0.680456 0.680544
https://ftp.jaist.ac.jp/ 1.1 0 0.000007 0.000000 0.000000 0.263183 0.263278
https://mirror.aarnet.edu.au/ 2 1 0.002357 0.292631 0.605226 0.932900 0.932993
https://mirror.aarnet.edu.au/ 2 0 0.000008 0.000000 0.000000 0.344341 0.344477
https://example.com/ 2 1 0.002633 0.011386 0.028187 0.046270 0.048678
https://example.com/ 2 0 0.000010 0.000000 0.000000 0.018484 0.020657
https://ftp.jaist.ac.jp/ 1.1 1 0.002184 0.141438 0.427447 0.835079 0.835186
https://ftp.jaist.ac.jp/ 1.1 0 0.000008 0.000000 0.000000 0.156694 0.156790
https://mirror.aarnet.edu.au/ 2 1 0.002092 0.296457 0.591279 0.912852 0.913007
https://mirror.aarnet.edu.au/ 2 0 0.000008 0.000000 0.000000 0.344485 0.344668
https://example.com/ 2 1 0.002098 0.011815 0.025337 0.038308 0.038427
https://example.com/ 2 0 0.000010 0.000000 0.000000 0.014357 0.014493
https://ftp.jaist.ac.jp/ 1.1 1 0.001781 0.152723 0.465717 0.731328 0.731471
https://ftp.jaist.ac.jp/ 1.1 0 0.000008 0.000000 0.000000 0.175738 0.175816
https://mirror.aarnet.edu.au/ 2 1 0.001886 0.294106 0.589307 0.910253 0.910445
https://mirror.aarnet.edu.au/ 2 0 0.000010 0.000000 0.000000 0.342327 0.342458 timeline/measure.sh
#!/bin/sh
# Records when each phase of an HTTPS request finished, in seconds from the start of the transfer, with curl's
# --write-out variables. One curl command fetches each address twice, so the second transfer reuses the connection
# that the first one opened (curl reuses connections only within one command).
FORMAT='%{url} %{http_version} %{num_connects} %{time_namelookup} %{time_connect} %{time_appconnect} %{time_starttransfer} %{time_total}\n'
for run in 1 2 3 4 5; do
for url in https://example.com/ https://ftp.jaist.ac.jp/ https://mirror.aarnet.edu.au/; do
curl --silent --output /dev/null --output /dev/null --write-out "$FORMAT" "$url" "$url"
sleep 2
done
done Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
The program reads recorded measurements, so its output is the same every time it runs. Each column is the median of its own five values, which is why the columns do not add up exactly.
Reading the numbers
TCP is one round trip. The connect times, 9, 140 and 293 ms, are the round-trip times to the three servers: nothing else happens in a TCP handshake. They are the most useful numbers in the table, because every other phase can be read in round trips.
TLS 1.3 is one round trip, sometimes two. To the server in Australia, TLS took as long as TCP, 295 ms against
293. To the
server in Japan it took twice as long, and curl’s verbose output (curl -v) for that server shows why: the client’s
first hello is followed by a second one, because the server answered with a HelloRetryRequest and asked for another
key share. To example.com it took 1.8 round trips for a different reason: when the round trip is only 9 ms, the
work at both ends (the key exchange, the server’s signature and the client’s checks of the certificate chain) is no
longer small next to it.
About two thirds of the wait is setup. The last column is the share of the time to the first byte spent on DNS, TCP and TLS: 62 to 65 % for all three servers, before any of them saw the request. To Australia, a small page took 913 ms to start arriving, and 588 ms of that was the two handshakes. No faster server could have changed that; only fewer round trips or a shorter distance can.
A reused connection pays one round trip. The second table shows the same requests on an open connection: 344 ms
instead of 913 to Australia, and 176 ms instead of 680 to Japan. What is left is one round trip plus the server’s
work, about 1.2 to 1.3 round trips to the far servers. To example.com the server’s few milliseconds of work are
as long as the round trip itself, so the ratio is closer to 2.
Reuse is the cheapest optimisation
Those numbers explain why every layer of the web works hard to keep connections open. HTTP/1.1 uses persistent connections by default, so many requests and responses travel over one connection (RFC 9112). HTTP/2 and HTTP/3 go further and carry many requests at once over a single connection, which the next lesson covers. Browsers keep connections to a site open between requests, and load balancers, app servers and database drivers keep pools of open connections to the hops behind them, so the hops inside the data centre never pay a handshake on the hot path.
The first request of a visit is the expensive one. A page that needs resources from five different hosts pays five sets of handshakes, which is one reason designs keep the critical resources of a page on as few hosts as they can.
Which hops a cache or a CDN removes
A CDN puts servers, called edges, close to users. The browser’s DNS lookup returns the address of a nearby edge, and the TCP and TLS handshakes run against that edge instead of the distant origin. This model takes the measured round trips, 9 ms to a nearby server and 293 ms to the far one, and computes the time to the first byte of a small page on six paths. The work times at the top are assumptions, and each is labelled as one:
"""Time to the first byte for one small page, on six paths, from the round trips measured by measure.sh."""
# Measured: the median TCP connect time (one round trip) to a nearby edge and to the server in Australia.
RTT_EDGE_MS = 9
RTT_ORIGIN_MS = 293
# Assumptions, labelled as such: the origin's own work per request (the measured reused-connection wait minus one
# round trip, rounded), the edge's work on a cache hit, and a DNS answer from a nearby resolver's cache.
ORIGIN_WORK_MS = 50
EDGE_WORK_MS = 2
DNS_MS = 3
HANDSHAKE_ROUND_TRIPS = 2 # TCP (1) and TLS 1.3 (1) before the request can be sent
def new_connection(rtt):
return HANDSHAKE_ROUND_TRIPS * rtt
paths = [
("origin, new connection", DNS_MS + new_connection(RTT_ORIGIN_MS) + RTT_ORIGIN_MS + ORIGIN_WORK_MS, 3),
("origin, reused connection", RTT_ORIGIN_MS + ORIGIN_WORK_MS, 1),
("edge hit, new connection", DNS_MS + new_connection(RTT_EDGE_MS) + RTT_EDGE_MS + EDGE_WORK_MS, 0),
("edge hit, reused connection", RTT_EDGE_MS + EDGE_WORK_MS, 0),
# On a miss the edge asks the origin. If it keeps a connection to the origin open, that costs one long round trip.
("edge miss, warm edge-origin link", DNS_MS + new_connection(RTT_EDGE_MS) + RTT_EDGE_MS + RTT_ORIGIN_MS + ORIGIN_WORK_MS, 1),
("edge miss, cold edge-origin link", DNS_MS + new_connection(RTT_EDGE_MS) + RTT_EDGE_MS + new_connection(RTT_ORIGIN_MS) + RTT_ORIGIN_MS + ORIGIN_WORK_MS, 3),
]
direct = paths[0][1]
print(f"{'path':<34}{'first byte':>11}{'long round trips':>18}{'vs direct':>11}")
for name, ms, long_trips in paths:
print(f"{name:<34}{ms:>8} ms{long_trips:>18}{ms / direct:>11.0%}") Output
path first byte long round trips vs direct origin, new connection 932 ms 3 100% origin, reused connection 343 ms 1 37% edge hit, new connection 32 ms 0 3% edge hit, reused connection 11 ms 0 1% edge miss, warm edge-origin link 373 ms 1 40% edge miss, cold edge-origin link 959 ms 3 103%
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 cdn_paths.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
The model’s first line, 932 ms, is within 2 % of the 913 ms measured to Australia, so its assumptions are close enough to reason with. Read the rest by what each path removes:
- An edge hit removes everything long. All three round trips are short ones, and the origin does no work: 32 ms instead of 932.
- An edge miss still saves the handshakes, as long as the edge keeps its connection to the origin open: one long round trip remains, 373 ms. This is why CDNs hold warm connections to origins.
- A cold edge-to-origin link makes a miss slower than going direct, 959 ms against 932, because the request now pays both sets of handshakes. An edge in front of a far origin helps only if it hits often or stays connected.
Caches remove hops at every layer, not only at the edge:
| Cache | What it removes |
|---|---|
| The browser’s HTTP cache, with a fresh response | Every hop: the response is reused without contacting any server (RFC 9111) |
| DNS caches in the browser, the operating system and the resolver | The name lookup and its round trips |
| A CDN edge with the response cached | The long round trip and all of the origin’s work |
| A cache in front of the database, inside the app’s data centre | The database query, often the slowest hop inside |
Where a request can fail
Every hop can fail in its own way, and what the user sees depends on which one did. The status codes of the 5xx class tell the hops apart: 502 Bad Gateway means a proxy or gateway received an invalid response from the server behind it, 503 Service Unavailable means a server is temporarily overloaded or down for maintenance, and 504 Gateway Timeout means a gateway did not get a timely answer from the server behind it (RFC 9110).
- DNS. The resolver times out, or a removed record is still cached. The user sees “server not found”, or reaches the old address. The defence is more than one resolver, and TTLs chosen for how fast changes must spread.
- TCP. Nothing listens on the address, or packets are dropped. The user sees a connection error or a long hang. Health checks take dead addresses out of service.
- TLS. The certificate has expired or does not match the name. The user sees a certificate error that no retry can fix. Certificates are renewed automatically, with alerts well before they expire.
- Load balancer. No healthy app server is left behind it, and the user gets 502 or 503. Health checks and spare capacity are the defence.
- App server. A bug, or more load than it can take, shows as 500, 503 or slow answers. Timeouts, load shedding and more servers help.
- Database. A slow query, or no free connection in the pool, ends as a 504 from the gateway once its timeout passes. Indexes, pool sizing and timeouts shorter than the caller’s are the defence.
The last item hides a common trap: each hop has its own timeout, and when an inner hop waits longer than the hop in front of it, the outer one gives up first and the user gets a 504 while the work goes on behind it, holding a database connection for nothing. Timeouts should shrink as a request travels inward.
HTTP Status Codes Reference Look up what 502, 503 and 504 mean and which hop usually sends each of them.Interview questions
Warm-up (fresher to mid level): what happens, step by step, when you type a URL into a browser and press Enter? The browser resolves the name to an IP address through DNS, unless it has the answer cached. It opens a TCP connection to that address, one round trip, and runs a TLS handshake to agree on keys and check the server’s certificate, one more round trip in TLS 1.3. It then sends the HTTP request. A load balancer usually receives it and passes it to an app server, which reads from a cache or a database and builds the response. The response travels back, the browser parses the HTML and requests the scripts, styles and images it names, mostly over the same connection, and renders the page. A strong answer adds where the time goes: on a long route, about three round trips before the first byte, which caching, CDNs and connection reuse reduce.
Key takeaways
- A request on a new connection pays a DNS lookup (often cached), one round trip for TCP, one for TLS 1.3 (two after a HelloRetryRequest) and one for the request, before the server’s work.
- Measure it with curl’s
--write-outvariables: they are cumulative, so subtract to get each phase, and read the TCP connect time as one round trip. - In the recorded runs, 62 to 65 % of the time to the first byte went on setup; a reused connection cut 913 ms to 344 ms on a long route.
- A CDN edge hit removes every long round trip; a miss still saves the handshakes only if the edge keeps a warm connection to the origin.
- Caches remove hops at every layer: the browser, DNS, the edge and the database tier.
- Each hop fails in its own way: 502, 503 and 504 point at different hops, and timeouts should shrink inward.
Exercise
Exercise · Easy · Python
Predict the time to the first byte from round trips
Write first_byte_ms(rtt_ms, server_ms, dns_ms=0, tls_round_trips=1, reused=False) in first_byte.py. It predicts how many milliseconds pass before the first byte of a response arrives:
- On a new connection, the client first waits for the DNS answer (dns_ms), then one round trip for the TCP handshake, then tls_round_trips round trips for TLS (1 for TLS 1.3, 2 when the server asks for another key share, 0 for plain HTTP), then one round trip for the request and response, plus the server's own work (server_ms). - On a reused connection (reused=True), the name is already resolved and both handshakes are done, so only the request's round trip and the server's work remain.
Return the result rounded to one decimal. Raise ValueError when a time or tls_round_trips is negative.
For the server in Australia measured in this lesson, with a round trip of 293 ms, 50 ms of server work and a 2 ms DNS answer:
first_byte_ms(293, 50, dns_ms=2) # 931.0
first_byte_ms(293, 50, dns_ms=2, reused=True) # 343.0
The sample tests import first_byte_ms from first_byte.py and run in your browser.
Starter code · first_byte.py
def first_byte_ms(rtt_ms, server_ms, dns_ms=0, tls_round_trips=1, reused=False):
"""Milliseconds until the first byte of the response arrives, rounded to 1 decimal."""
# Replace this line with your code.
return 0.0 The sample tests · test_first_byte.py
from first_byte import first_byte_ms
def raises_value_error(call):
try:
call()
except ValueError:
return True
return False
def test_new_connection():
"""a new connection pays DNS, TCP, TLS and the request's round trip"""
assert first_byte_ms(293, 50, dns_ms=2) == 931.0
def test_reused_connection():
"""a reused connection pays only the request's round trip"""
assert first_byte_ms(293, 50, dns_ms=2, reused=True) == 343.0
def test_tls_round_trips():
"""counts two TLS round trips after a retry, and none for plain HTTP"""
assert first_byte_ms(140, 25, tls_round_trips=2) == 585.0
assert first_byte_ms(140, 25, tls_round_trips=0) == 305.0
def test_rounding():
"""rounds to one decimal"""
assert first_byte_ms(9.4, 1.25, dns_ms=3.33) == 32.8
def test_negative_inputs():
"""raises ValueError for a negative time or round-trip count"""
assert raises_value_error(lambda: first_byte_ms(-1, 50))
assert raises_value_error(lambda: first_byte_ms(100, 50, tls_round_trips=-1))
assert raises_value_error(lambda: first_byte_ms(100, 50, dns_ms=-5)) A hint
Count the round trips first. A new connection pays 1 + tls_round_trips + 1 of them (TCP, TLS, the request) and a reused one pays only the last. Multiply by rtt_ms, add server_ms, and add dns_ms only for a new connection. Check the inputs before you compute anything.
Results of the sample tests
| Test | Result | Details |
|---|
What your code printed
The sample tests run on this device, in your browser (Pyodide): nothing is sent to mysmartcopilot.com. The first run downloads Python (about 13.5 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.
Check yourself
7 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.
References
- curl manual page (--write-out variables and connection reuse) (curl project)
- RFC 1034: Domain Names, Concepts and Facilities (IETF)
- RFC 9293: Transmission Control Protocol (TCP) (IETF)
- RFC 8446: The Transport Layer Security (TLS) Protocol Version 1.3 (IETF)
- RFC 9110: HTTP Semantics (IETF)
- RFC 9112: HTTP/1.1 (IETF)
- RFC 9111: HTTP Caching (IETF)
- RFC 6761: Special-Use Domain Names (IETF)
Related tools
Report a problem with this lesson
Kept only in this browser. Your Learn progress