System Design (High-Level Design) Module 4 – Networking for system designers
CDNs and edge caching
How CDNs serve from edges near users, pull against push, the cache keys and TTLs that set the hit ratio, purging, origin shields and edge compute.
What you will learn
- Explain pull and push CDN models
- Design cache keys, TTLs and purge strategies
- Calculate how much a shield tier reduces origin load
- Use tiered caching and edge compute where they help
Before you start
On this page
A CDN is a network of caching reverse proxies, called edges, placed close to users. DNS or anycast sends each user to a nearby edge; the edge answers from its cache when it holds a fresh copy and asks the origin when it does not, which is the pull model. How much load the CDN takes off the origin depends on three choices: the cache key, which decides which requests count as the same; the TTL, which decides how long a copy may be served; and how changes reach the caches, by new versioned addresses or by purges. A shield tier between the edges and the origin turns the misses of many edges into one origin request, and edge compute runs small programs inside the edges, for redirects or token checks, without a trip to the origin.
This lesson measures each of those choices: a cache-key experiment, a shield simulation, and header recipes for the responses a typical site sends.
How a CDN answers a request
Edges near users, a shield in front of the origin
Text description of the diagram
Users in three cities, A, B and C, each reach the CDN edge nearest to them. Each edge keeps its own cache and answers a request itself when it holds a fresh copy: a hit.
On a miss, an edge does not go to the origin directly. All three edges ask the same shield tier, a larger cache placed near the origin. Only when the shield misses too does a request reach the origin's app servers and storage. A file that is requested in all three cities therefore costs the origin one request instead of three.
An edge is a shared cache in the sense of the HTTP caching standard: it stores responses and reuses them for many users, unlike the browser’s private cache, which serves one (RFC 9111). The previous lessons placed the pieces: DNS or anycast brings the user to a nearby edge, the edge is a reverse proxy that ends TLS close to the user, and on a miss it fetches from the origin over a connection it keeps open.
For every request the edge does the same three things: it builds the request’s cache key, looks for a stored response under that key that is still fresh, and either serves it (a hit) or forwards the request to the next tier and, if the response may be stored, stores it on the way back (a miss).
Pull or push. A pull CDN works like that: it fetches on the first miss and keeps the copy for its TTL, so there is nothing to prepare, and most websites use it. A push CDN is loaded ahead: the publisher uploads files to the CDN’s storage before anyone asks, which suits large media libraries, software releases and launch days, when a burst of misses from every edge at once could overwhelm the origin. Push costs work and storage: the publisher must manage which files are where, and pays to keep files nobody requests.
The cache key decides the hit ratio
The HTTP caching standard defines the cache key as at least the request method and the target URI, query string
included, and many caches that store only GET responses key on the URI alone. A response that varies by request
headers names them in Vary, and the cache then adds those headers’ values to the key
(RFC 9111). Both rules cut both ways: every irrelevant difference in
a URL or a varied header is a new key and a new miss, and every relevant difference left out of the key serves one
user’s response to another.
This experiment makes 1,000 requests for 50 product pages that come in two colours and two languages, so there are 200 different responses. Half the links carry a tracking parameter, some an advertising click id, some have their parameters in another order, and a few spell the host in capitals. It builds the cache key in five ways, each step on top of the last:
// How tracking parameters and parameter order split one page into many cache keys, and how normalising the key
// wins the hits back. A seeded generator makes the same 1,000 requests on every run.
function seeded(seed) {
// mulberry32: a small, well-known pseudo-random generator, so the trace is identical everywhere
return () => {
seed = (seed + 0x6d2b79f5) | 0;
let t = Math.imul(seed ^ (seed >>> 15), 1 | seed);
t = (t + Math.imul(t ^ (t >>> 7), 61 | t)) ^ t;
return ((t ^ (t >>> 14)) >>> 0) / 4294967296;
};
}
const rnd = seeded(2026);
const pick = (list) => list[Math.floor(rnd() * list.length)];
// 1,000 requests for 50 product pages. A page's content depends only on its id, its colour and its language.
const requests = [];
for (let i = 0; i < 1000; i++) {
const params = [`color=${pick(['red', 'blue'])}`, `lang=${pick(['en', 'hi'])}`];
if (rnd() < 0.5) params.push(`utm_source=${pick(['mail', 'ads', 'social', 'partner'])}`);
if (rnd() < 0.2) params.push(`gclid=${Math.floor(rnd() * 1e6)}`);
if (rnd() < 0.5) params.reverse();
const host = rnd() < 0.1 ? 'Shop.Example' : 'shop.example';
requests.push({ url: `https://${host}/products/${1 + Math.floor(rnd() * 50)}?${params.join('&')}`, color: params.find((p) => p.startsWith('color=')) });
}
const split = (url) => {
const [base, query = ''] = url.split('?');
const [scheme, , host, ...path] = base.split('/');
return { scheme, host, path: '/' + path.join('/'), params: query ? query.split('&') : [] };
};
const join = (u) => `${u.scheme}//${u.host}${u.path}${u.params.length ? '?' + u.params.join('&') : ''}`;
const TRACKING = (p) => p.startsWith('utm_') || p.startsWith('gclid=') || p.startsWith('fbclid=');
const steps = [
['the URL as it arrived', (u) => u],
['host in lower case', (u) => ({ ...u, host: u.host.toLowerCase() })],
['tracking parameters dropped', (u) => ({ ...u, params: u.params.filter((p) => !TRACKING(p)) })],
['parameters sorted', (u) => ({ ...u, params: [...u.params].sort() })],
['colour dropped too (wrong!)', (u) => ({ ...u, params: u.params.filter((p) => !p.startsWith('color=')) })],
];
console.log(`${requests.length} requests for 50 product pages in 2 colours and 2 languages (200 different responses)`);
console.log(`${'cache key'.padEnd(30)}${'keys'.padStart(6)}${'hit ratio'.padStart(11)}${'wrong colour served'.padStart(21)}`);
let normalise = (u) => u;
for (const [name, step] of steps) {
const previous = normalise;
normalise = (u) => step(previous(u));
const cache = new Map(); // key -> the colour of the response stored under it
let hits = 0;
let wrong = 0;
for (const r of requests) {
const key = join(normalise(split(r.url)));
if (cache.has(key)) {
hits += 1;
if (cache.get(key) !== r.color) wrong += 1;
} else cache.set(key, r.color);
}
const ratio = `${((100 * hits) / requests.length).toFixed(1)}%`;
console.log(`${name.padEnd(30)}${String(cache.size).padStart(6)}${ratio.padStart(11)}${String(wrong).padStart(21)}`);
} Output
1000 requests for 50 product pages in 2 colours and 2 languages (200 different responses) cache key keys hit ratio wrong colour served the URL as it arrived 826 17.4% 0 host in lower case 802 19.8% 0 tracking parameters dropped 365 63.5% 0 parameters sorted 200 80.0% 0 colour dropped too (wrong!) 100 90.0% 445
Recorded with Node.js 24.21.0 on macOS 26 arm64. To run it yourself: mise exec node@24.21.0 -- node cache_key_demo.mjs
Runs on this device, in your browser. The first run downloads JavaScript (about 0.6 MB), which is kept for the next runs.
Your run, in this browser
Keyed on the raw URL, the cache hit 17.4 % of the time. Lower-casing the host helped a little, dropping tracking parameters lifted the hit ratio to 63.5 %, and sorting what remained brought it to 80.0 %, the best possible here, since each of the 200 responses must miss once. The last row is the opposite mistake: dropping the colour too lifts the “hit ratio” to 90 %, while 445 users get a page in a colour they did not ask for. A cache key should contain everything that changes the response and nothing else, which is why an allowlist of the parameters that matter is safer than a list of the ones to drop.
Vary works the same way. Vary: Accept-Encoding is cheap, because there are only a few encodings. Vary: User-Agent would store a separate copy for each of hundreds of browser versions, which brings the hit ratio close to
zero; normalise such a header into a few classes at the edge first, or leave it out of the key.
TTLs: how long a copy may be served
A stored response stays fresh for the lifetime the origin gives it, and a cache may serve a fresh copy without asking anyone (RFC 9111):
max-age=Nsets the lifetime for every cache;s-maxage=Noverrides it for shared caches only, so a CDN can keep a copy longer than browsers do.no-cachedoes not mean “do not store”: it means a cache must check with the origin before every reuse. The check is cheap when the response has a validator, such as anETag, because an unchanged resource comes back as a short 304 Not Modified (RFC 9110).no-storemeans do not store at all, andprivatemeans only a single user’s cache may store the response.- Without any lifetime, a cache may guess one from
Last-Modified; the standard mentions 10 % of the time since the last change as a typical setting. Leaving the lifetime to guesswork is how stale pages happen. stale-while-revalidate=Nlets a cache serve a stale copy for up to N more seconds while it refreshes in the background, andstale-if-error=Nlets it serve a stale copy when the origin fails (RFC 5861).CDN-Cache-Controlcarries directives for CDN caches only, separately from what browsers are told (RFC 9213).
Together they give a recipe for each kind of response a site sends:
- A script or style at a versioned address (
app.3f9c2b.js):public, max-age=31536000, immutable. The address changes whenever the content does, so the file can stay fresh for a year. - An image at a fixed address:
public, max-age=86400, stale-while-revalidate=600. Changes are rare, and a few minutes of staleness are harmless. - An HTML page:
no-cachefor browsers, withCDN-Cache-Control: max-age=60. Browsers revalidate every time, and the CDN absorbs the load for a minute. - A public API list, such as today’s offers:
public, s-maxage=30, stale-while-revalidate=30. Shared caches serve it and refresh it every 30 seconds. - A signed-in user’s account data:
private, no-store. No shared cache may ever store it.
The last item deserves its warning. Responses to requests with an Authorization header are off limits to shared
caches by default, and public, s-maxage and must-revalidate are exactly the directives that lift that
protection (RFC 9111). Adding public to a personal API response to
“speed it up” can serve one user’s data to the next.
Versioned addresses beat purges
CDNs let you purge: remove a URL from their caches before its TTL runs out, and many can also purge every response labelled with a tag. Purges are useful for mistakes and legal takedowns, but they are a poor release mechanism:
- A purge cannot reach browsers. A file served with a day’s
max-agestays in each browser’s own cache for that day; clearing the CDN changes nothing for a browser that does not ask again. - A purge sends the next requests to the origin. Purging a popular file at every edge at once turns into a burst of misses, the stampede a shield or request collapsing has to absorb.
- Purges and deploys race. If the purge lands before the new file is on every origin server, an edge may fetch and cache the old one again.
A versioned address avoids all three. When the content changes, the file is published under a new name that embeds its version or hash, and the HTML that links it changes at the same time; the old copies age out unused. The immutable directive was standardised for exactly this pattern: a response that will not change while it is fresh need not be revalidated, even on an ordinary reload of the page (RFC 8246). Only the HTML, with its short TTL, still needs to change in place.
Shields: one origin request instead of one per edge
Each edge fills its cache on its own. A file that becomes popular everywhere misses once at every edge, so a CDN with many edges sends the origin many requests for the same file. A shield is an extra cache tier that every edge asks on a miss; one CDN’s documentation describes it as an additional caching layer that all requests to the origin pass through, which can combine simultaneous requests for the same object into as few as one origin request (Amazon CloudFront Developer Guide).
This simulation sends 100,000 requests for 10,000 files, popular in the way real catalogues are (a few files get most requests), spread over 5, 20 or 50 edges, each of which can keep a tenth of the catalogue:
"""Origin load behind a CDN, with and without a shield tier: a seeded trace of requests over a catalogue whose
popularity follows a Zipf curve, spread over many edges. Every number at the top is an assumption."""
import random
from collections import OrderedDict
OBJECTS = 10_000 # assumption: distinct files on the site
REQUESTS = 100_000 # requests in the trace
ZIPF_S = 0.9 # assumption: how steeply popularity falls from the most popular file
EDGE_SLOTS = 1_000 # assumption: files each edge keeps (10 % of the catalogue)
SHIELD_SLOTS = 5_000 # assumption: files the shield keeps (50 %)
class LRU:
"""A cache that keeps the most recently used keys and forgets the least recently used one when full."""
def __init__(self, slots):
self.slots, self.items = slots, OrderedDict()
def get_or_fill(self, key):
"""True on a hit; on a miss, store the key and return False."""
if key in self.items:
self.items.move_to_end(key)
return True
self.items[key] = True
if len(self.items) > self.slots:
self.items.popitem(last=False)
return False
rng = random.Random(42)
weights = [1 / (rank ** ZIPF_S) for rank in range(1, OBJECTS + 1)]
trace = rng.choices(range(OBJECTS), weights=weights, k=REQUESTS)
print(f"{REQUESTS:,} requests for {OBJECTS:,} files; each edge keeps {EDGE_SLOTS:,}, the shield {SHIELD_SLOTS:,}")
print(f"{'edges':>5} {'edge hit ratio':>14} {'origin requests':>15} {'with a shield':>13} {'shield cuts origin load by':>26}")
for edge_count in (5, 20, 50):
edges = [LRU(EDGE_SLOTS) for _ in range(edge_count)]
shield = LRU(SHIELD_SLOTS)
pick = random.Random(7) # which edge each request lands on: users spread evenly over the edges
edge_misses = shield_misses = 0
for key in trace:
if not edges[pick.randrange(edge_count)].get_or_fill(key):
edge_misses += 1
if not shield.get_or_fill(key):
shield_misses += 1
hit_ratio = 1 - edge_misses / REQUESTS
print(f"{edge_count:>5} {hit_ratio:>14.1%} {edge_misses:>15,} {shield_misses:>13,} {edge_misses / shield_misses:>25.1f}x") Output
100,000 requests for 10,000 files; each edge keeps 1,000, the shield 5,000
edges edge hit ratio origin requests with a shield shield cuts origin load by
5 54.6% 45,377 17,422 2.6x
20 51.6% 48,432 17,174 2.8x
50 45.4% 54,593 17,083 3.2x
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 cdn_hit_ratio.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
Two numbers move in opposite directions. As the edges multiply, each sees a smaller share of the traffic and its hit ratio falls, from 54.6 % with 5 edges to 45.4 % with 50, so the origin’s load without a shield grows from 45,377 to 54,593 requests. With the shield, the origin’s load stays almost flat at about 17,000, because the shield sees every edge’s misses and keeps the popular files for all of them: 2.6 times fewer origin requests with 5 edges, 3.2 times with 50. Without any CDN, the origin would have answered all 100,000.
A shield is not free. It is one more hop on every miss, it adds cost, and it helps little for content that is rarely requested or cannot be cached. The same guide names the cases where it pays most: users spread across many regions, origins that do expensive work per request such as packaging video or resizing images, and origins with little spare capacity.
Edge compute
CDNs also run code at the edge. One CDN’s documentation lists what such functions do: change requests and responses as they pass, perform basic authentication and authorisation, and generate whole responses at the edge, all close to the user and without servers to manage (Amazon CloudFront Developer Guide). Typical uses in a design:
- normalising cache keys before the lookup, as in the experiment above;
- redirects and A/B assignment that would otherwise cost a trip to the origin;
- checking a signed token on every request, so that the origin sees only valid callers;
- assembling a cached page with a small personal fragment fetched separately.
State is where edge compute gets hard. An edge is one of many copies spread across the world, so a value written at one edge is not instantly visible at the others; a counter, a stock level or a session that must be correct everywhere at once still belongs in the origin’s database. Read-mostly data, such as configuration or the keys that verify tokens, suits the edge well.
To see what each layer did, the Cache-Status response header records how every cache on the path handled a
response, such as a hit, or a miss and why it was forwarded (RFC 9211),
and Age says how long a served copy has been in a cache.
Interview questions
Warm-up (fresher to mid level): how do a pull CDN and a push CDN differ, and when would you use each? A pull CDN fetches a file from the origin the first time an edge is asked for it and keeps it for its TTL, so the publisher only sets cache headers and the CDN fills itself. Most websites use pull, because it needs no preparation and adapts to whatever users request; its cost is that the first request for each file at each edge misses. A push CDN is loaded ahead: the publisher uploads files to the CDN’s storage before users ask, so even the first request hits. Push suits large media libraries, software downloads and launches, when a wave of first misses at every edge could overload the origin, at the price of managing what is stored where and paying for storage. A strong answer adds that a shield tier gives a pull CDN much of push’s protection for the origin.
Key takeaways
- A CDN edge is a shared cache near users: it answers hits itself and sends misses to the origin, often through a shield.
- The cache key is the hit ratio: in the experiment, normalising the key raised hits from 17.4 % to 80 %, and dropping a parameter that mattered served 445 wrong pages.
- Set lifetimes explicitly:
max-ageands-maxagefor freshness,no-cacheto force revalidation,privateorno-storefor personal data, and neverpublicon authenticated responses. - Release by versioned addresses with
immutable, not purges: a purge cannot reach browser caches. - A shield kept origin load flat at about 17,000 requests while edges grew from 5 to 50; without it, origin load grew with every edge.
- Edge compute suits stateless work such as redirects, token checks and key normalisation; shared mutable state stays at the origin.
Exercise
Exercise · Medium · Python
Normalise a URL into an edge cache key
Requests that should get the same response must get the same cache key, and requests that should not must not. Write cache_key(url, keep) in cache_key.py. keep is the set of query parameter names that change the response; every other parameter is dropped. Build the key from the URL like this:
- the scheme and the host in lower case; - no port when it is the scheme's default (:443 for https, :80 for http), and the port kept otherwise; - the path exactly as it is, upper case included (/Products and /products can be different pages), or / when it is empty; - only the parameters named in keep, sorted by name and then by value, with repeated parameters kept; - no ? when no parameter is left, and never the fragment (the part after #, which browsers do not send).
cache_key("https://Shop.Example:443/p/42?utm_source=mail&size=m&color=red", {"color", "size"})
# "https://shop.example/p/42?color=red&size=m"
Keep the parameters' values exactly as they appear in the URL. The sample tests import cache_key from cache_key.py and run in your browser.
Starter code · cache_key.py
def cache_key(url, keep):
"""The edge cache key of `url`, keeping only the query parameters named in `keep`."""
# Replace this line with your code.
return url The sample tests · test_cache_key.py
from cache_key import cache_key
def test_tracking_parameters_dropped():
"""drops every parameter that is not in keep"""
assert cache_key("https://shop.example/p/42?utm_source=mail&color=red&gclid=8812", {"color"}) == "https://shop.example/p/42?color=red"
def test_order_does_not_matter():
"""sorts the parameters, so their order makes no new key"""
a = cache_key("https://shop.example/p/42?size=m&color=red", {"color", "size"})
b = cache_key("https://shop.example/p/42?color=red&size=m", {"color", "size"})
assert a == b == "https://shop.example/p/42?color=red&size=m"
def test_host_scheme_and_ports():
"""lower-cases the scheme and host, drops default ports and keeps others"""
assert cache_key("HTTPS://Shop.Example:443/p/42", set()) == "https://shop.example/p/42"
assert cache_key("http://shop.example:80/p/42", set()) == "http://shop.example/p/42"
assert cache_key("https://shop.example:8443/p/42", set()) == "https://shop.example:8443/p/42"
def test_path_and_fragment():
"""keeps the path's case, uses / for an empty path and drops the fragment"""
assert cache_key("https://shop.example/Products/42#reviews", set()) == "https://shop.example/Products/42"
assert cache_key("https://shop.example?utm_source=ads", set()) == "https://shop.example/"
def test_repeated_parameters():
"""keeps repeated parameters, sorted by value"""
assert cache_key("https://shop.example/s?tag=b&tag=a&page=2", {"tag"}) == "https://shop.example/s?tag=a&tag=b" A hint
urllib.parse.urlsplit(url) gives you the parts you need: scheme, hostname (already in lower case), port, path and query. Split the query on & and each part on the first =, keep the pairs whose name is in keep, sort the pairs, and join them back with = and &.
Results of the sample tests
| Test | Result | Details |
|---|
What your code printed
The sample tests run on this device, in your browser (Pyodide): nothing is sent to mysmartcopilot.com. The first run downloads Python (about 13.5 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.
Check yourself
7 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.
References
- RFC 9111: HTTP Caching (IETF)
- RFC 9110: HTTP Semantics (conditional requests and validators) (IETF)
- RFC 8246: HTTP Immutable Responses (IETF)
- RFC 5861: HTTP Cache-Control Extensions for Stale Content (IETF)
- RFC 9213: Targeted HTTP Cache Control (IETF)
- RFC 9211: The Cache-Status HTTP Response Header Field (IETF)
- Use Amazon CloudFront Origin Shield (Amazon CloudFront Developer Guide) (Amazon Web Services)
- Customize at the edge with functions (Amazon CloudFront Developer Guide) (Amazon Web Services)
Related tools
Report a problem with this lesson
Kept only in this browser. Your Learn progress