Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

System Design (High-Level Design)  Module 4 – Networking for system designers

CDNs and edge caching

How CDNs serve from edges near users, pull against push, the cache keys and TTLs that set the hit ratio, purging, origin shields and edge compute.

  • Intermediate
  • 30 minutes
  • Examples run with Python 3.14.8, Pyodide 314.0.7, Node.js 24.21.0 and quickjs 0.32.0
  • By MySmartCoPilot

What you will learn

  • Explain pull and push CDN models
  • Design cache keys, TTLs and purge strategies
  • Calculate how much a shield tier reduces origin load
  • Use tiered caching and edge compute where they help

Before you start

On this page

A CDN is a network of caching reverse proxies, called edges, placed close to users. DNS or anycast sends each user to a nearby edge; the edge answers from its cache when it holds a fresh copy and asks the origin when it does not, which is the pull model. How much load the CDN takes off the origin depends on three choices: the cache key, which decides which requests count as the same; the TTL, which decides how long a copy may be served; and how changes reach the caches, by new versioned addresses or by purges. A shield tier between the edges and the origin turns the misses of many edges into one origin request, and edge compute runs small programs inside the edges, for redirects or token checks, without a trip to the origin.

This lesson measures each of those choices: a cache-key experiment, a shield simulation, and header recipes for the responses a typical site sends.

How a CDN answers a request

Users in three cities reach their nearest edge cache; edges send misses to one shield cache, and only its misses reach the origin.Users in city AUsers in city BUsers in city CEdge Aanswers its hitsEdge Banswers its hitsEdge Canswers its hitsShield tierone larger cache near the originOriginapp servers and storagemissmissmissonly the shield's misses

Edges near users, a shield in front of the origin

Text description of the diagram

Users in three cities, A, B and C, each reach the CDN edge nearest to them. Each edge keeps its own cache and answers a request itself when it holds a fresh copy: a hit.

On a miss, an edge does not go to the origin directly. All three edges ask the same shield tier, a larger cache placed near the origin. Only when the shield misses too does a request reach the origin's app servers and storage. A file that is requested in all three cities therefore costs the origin one request instead of three.

An edge is a shared cache in the sense of the HTTP caching standard: it stores responses and reuses them for many users, unlike the browser’s private cache, which serves one (RFC 9111). The previous lessons placed the pieces: DNS or anycast brings the user to a nearby edge, the edge is a reverse proxy that ends TLS close to the user, and on a miss it fetches from the origin over a connection it keeps open.

For every request the edge does the same three things: it builds the request’s cache key, looks for a stored response under that key that is still fresh, and either serves it (a hit) or forwards the request to the next tier and, if the response may be stored, stores it on the way back (a miss).

Pull or push. A pull CDN works like that: it fetches on the first miss and keeps the copy for its TTL, so there is nothing to prepare, and most websites use it. A push CDN is loaded ahead: the publisher uploads files to the CDN’s storage before anyone asks, which suits large media libraries, software releases and launch days, when a burst of misses from every edge at once could overwhelm the origin. Push costs work and storage: the publisher must manage which files are where, and pays to keep files nobody requests.

The cache key decides the hit ratio

The HTTP caching standard defines the cache key as at least the request method and the target URI, query string included, and many caches that store only GET responses key on the URI alone. A response that varies by request headers names them in Vary, and the cache then adds those headers’ values to the key (RFC 9111). Both rules cut both ways: every irrelevant difference in a URL or a varied header is a new key and a new miss, and every relevant difference left out of the key serves one user’s response to another.

This experiment makes 1,000 requests for 50 product pages that come in two colours and two languages, so there are 200 different responses. Half the links carry a tracking parameter, some an advertising click id, some have their parameters in another order, and a few spell the host in capitals. It builds the cache key in five ways, each step on top of the last:

One page, many cache keys JavaScript · cache_key_demo.mjs
// How tracking parameters and parameter order split one page into many cache keys, and how normalising the key
// wins the hits back. A seeded generator makes the same 1,000 requests on every run.

function seeded(seed) {
  // mulberry32: a small, well-known pseudo-random generator, so the trace is identical everywhere
  return () => {
    seed = (seed + 0x6d2b79f5) | 0;
    let t = Math.imul(seed ^ (seed >>> 15), 1 | seed);
    t = (t + Math.imul(t ^ (t >>> 7), 61 | t)) ^ t;
    return ((t ^ (t >>> 14)) >>> 0) / 4294967296;
  };
}
const rnd = seeded(2026);
const pick = (list) => list[Math.floor(rnd() * list.length)];

// 1,000 requests for 50 product pages. A page's content depends only on its id, its colour and its language.
const requests = [];
for (let i = 0; i < 1000; i++) {
  const params = [`color=${pick(['red', 'blue'])}`, `lang=${pick(['en', 'hi'])}`];
  if (rnd() < 0.5) params.push(`utm_source=${pick(['mail', 'ads', 'social', 'partner'])}`);
  if (rnd() < 0.2) params.push(`gclid=${Math.floor(rnd() * 1e6)}`);
  if (rnd() < 0.5) params.reverse();
  const host = rnd() < 0.1 ? 'Shop.Example' : 'shop.example';
  requests.push({ url: `https://${host}/products/${1 + Math.floor(rnd() * 50)}?${params.join('&')}`, color: params.find((p) => p.startsWith('color=')) });
}

const split = (url) => {
  const [base, query = ''] = url.split('?');
  const [scheme, , host, ...path] = base.split('/');
  return { scheme, host, path: '/' + path.join('/'), params: query ? query.split('&') : [] };
};
const join = (u) => `${u.scheme}//${u.host}${u.path}${u.params.length ? '?' + u.params.join('&') : ''}`;
const TRACKING = (p) => p.startsWith('utm_') || p.startsWith('gclid=') || p.startsWith('fbclid=');

const steps = [
  ['the URL as it arrived', (u) => u],
  ['host in lower case', (u) => ({ ...u, host: u.host.toLowerCase() })],
  ['tracking parameters dropped', (u) => ({ ...u, params: u.params.filter((p) => !TRACKING(p)) })],
  ['parameters sorted', (u) => ({ ...u, params: [...u.params].sort() })],
  ['colour dropped too (wrong!)', (u) => ({ ...u, params: u.params.filter((p) => !p.startsWith('color=')) })],
];

console.log(`${requests.length} requests for 50 product pages in 2 colours and 2 languages (200 different responses)`);
console.log(`${'cache key'.padEnd(30)}${'keys'.padStart(6)}${'hit ratio'.padStart(11)}${'wrong colour served'.padStart(21)}`);
let normalise = (u) => u;
for (const [name, step] of steps) {
  const previous = normalise;
  normalise = (u) => step(previous(u));
  const cache = new Map(); // key -> the colour of the response stored under it
  let hits = 0;
  let wrong = 0;
  for (const r of requests) {
    const key = join(normalise(split(r.url)));
    if (cache.has(key)) {
      hits += 1;
      if (cache.get(key) !== r.color) wrong += 1;
    } else cache.set(key, r.color);
  }
  const ratio = `${((100 * hits) / requests.length).toFixed(1)}%`;
  console.log(`${name.padEnd(30)}${String(cache.size).padStart(6)}${ratio.padStart(11)}${String(wrong).padStart(21)}`);
}

Output

1000 requests for 50 product pages in 2 colours and 2 languages (200 different responses)
cache key                       keys  hit ratio  wrong colour served
the URL as it arrived            826      17.4%                    0
host in lower case               802      19.8%                    0
tracking parameters dropped      365      63.5%                    0
parameters sorted                200      80.0%                    0
colour dropped too (wrong!)      100      90.0%                  445

Recorded with Node.js 24.21.0 on macOS 26 arm64. To run it yourself: mise exec node@24.21.0 -- node cache_key_demo.mjs

Keyed on the raw URL, the cache hit 17.4 % of the time. Lower-casing the host helped a little, dropping tracking parameters lifted the hit ratio to 63.5 %, and sorting what remained brought it to 80.0 %, the best possible here, since each of the 200 responses must miss once. The last row is the opposite mistake: dropping the colour too lifts the “hit ratio” to 90 %, while 445 users get a page in a colour they did not ask for. A cache key should contain everything that changes the response and nothing else, which is why an allowlist of the parameters that matter is safer than a list of the ones to drop.

Vary works the same way. Vary: Accept-Encoding is cheap, because there are only a few encodings. Vary: User-Agent would store a separate copy for each of hundreds of browser versions, which brings the hit ratio close to zero; normalise such a header into a few classes at the edge first, or leave it out of the key.

TTLs: how long a copy may be served

A stored response stays fresh for the lifetime the origin gives it, and a cache may serve a fresh copy without asking anyone (RFC 9111):

  • max-age=N sets the lifetime for every cache; s-maxage=N overrides it for shared caches only, so a CDN can keep a copy longer than browsers do.
  • no-cache does not mean “do not store”: it means a cache must check with the origin before every reuse. The check is cheap when the response has a validator, such as an ETag, because an unchanged resource comes back as a short 304 Not Modified (RFC 9110).
  • no-store means do not store at all, and private means only a single user’s cache may store the response.
  • Without any lifetime, a cache may guess one from Last-Modified; the standard mentions 10 % of the time since the last change as a typical setting. Leaving the lifetime to guesswork is how stale pages happen.
  • stale-while-revalidate=N lets a cache serve a stale copy for up to N more seconds while it refreshes in the background, and stale-if-error=N lets it serve a stale copy when the origin fails (RFC 5861).
  • CDN-Cache-Control carries directives for CDN caches only, separately from what browsers are told (RFC 9213).

Together they give a recipe for each kind of response a site sends:

  • A script or style at a versioned address (app.3f9c2b.js): public, max-age=31536000, immutable. The address changes whenever the content does, so the file can stay fresh for a year.
  • An image at a fixed address: public, max-age=86400, stale-while-revalidate=600. Changes are rare, and a few minutes of staleness are harmless.
  • An HTML page: no-cache for browsers, with CDN-Cache-Control: max-age=60. Browsers revalidate every time, and the CDN absorbs the load for a minute.
  • A public API list, such as today’s offers: public, s-maxage=30, stale-while-revalidate=30. Shared caches serve it and refresh it every 30 seconds.
  • A signed-in user’s account data: private, no-store. No shared cache may ever store it.

The last item deserves its warning. Responses to requests with an Authorization header are off limits to shared caches by default, and public, s-maxage and must-revalidate are exactly the directives that lift that protection (RFC 9111). Adding public to a personal API response to “speed it up” can serve one user’s data to the next.

HTTP Header Checker Check what Cache-Control, Age and Vary a URL really sends, before and after it passes your CDN.

Versioned addresses beat purges

CDNs let you purge: remove a URL from their caches before its TTL runs out, and many can also purge every response labelled with a tag. Purges are useful for mistakes and legal takedowns, but they are a poor release mechanism:

  • A purge cannot reach browsers. A file served with a day’s max-age stays in each browser’s own cache for that day; clearing the CDN changes nothing for a browser that does not ask again.
  • A purge sends the next requests to the origin. Purging a popular file at every edge at once turns into a burst of misses, the stampede a shield or request collapsing has to absorb.
  • Purges and deploys race. If the purge lands before the new file is on every origin server, an edge may fetch and cache the old one again.

A versioned address avoids all three. When the content changes, the file is published under a new name that embeds its version or hash, and the HTML that links it changes at the same time; the old copies age out unused. The immutable directive was standardised for exactly this pattern: a response that will not change while it is fresh need not be revalidated, even on an ordinary reload of the page (RFC 8246). Only the HTML, with its short TTL, still needs to change in place.

Shields: one origin request instead of one per edge

Each edge fills its cache on its own. A file that becomes popular everywhere misses once at every edge, so a CDN with many edges sends the origin many requests for the same file. A shield is an extra cache tier that every edge asks on a miss; one CDN’s documentation describes it as an additional caching layer that all requests to the origin pass through, which can combine simultaneous requests for the same object into as few as one origin request (Amazon CloudFront Developer Guide).

This simulation sends 100,000 requests for 10,000 files, popular in the way real catalogues are (a few files get most requests), spread over 5, 20 or 50 edges, each of which can keep a tenth of the catalogue:

Origin requests with and without a shield Python · cdn_hit_ratio.py
"""Origin load behind a CDN, with and without a shield tier: a seeded trace of requests over a catalogue whose
popularity follows a Zipf curve, spread over many edges. Every number at the top is an assumption."""
import random
from collections import OrderedDict

OBJECTS = 10_000        # assumption: distinct files on the site
REQUESTS = 100_000      # requests in the trace
ZIPF_S = 0.9            # assumption: how steeply popularity falls from the most popular file
EDGE_SLOTS = 1_000      # assumption: files each edge keeps (10 % of the catalogue)
SHIELD_SLOTS = 5_000    # assumption: files the shield keeps (50 %)


class LRU:
    """A cache that keeps the most recently used keys and forgets the least recently used one when full."""

    def __init__(self, slots):
        self.slots, self.items = slots, OrderedDict()

    def get_or_fill(self, key):
        """True on a hit; on a miss, store the key and return False."""
        if key in self.items:
            self.items.move_to_end(key)
            return True
        self.items[key] = True
        if len(self.items) > self.slots:
            self.items.popitem(last=False)
        return False


rng = random.Random(42)
weights = [1 / (rank ** ZIPF_S) for rank in range(1, OBJECTS + 1)]
trace = rng.choices(range(OBJECTS), weights=weights, k=REQUESTS)

print(f"{REQUESTS:,} requests for {OBJECTS:,} files; each edge keeps {EDGE_SLOTS:,}, the shield {SHIELD_SLOTS:,}")
print(f"{'edges':>5}  {'edge hit ratio':>14}  {'origin requests':>15}  {'with a shield':>13}  {'shield cuts origin load by':>26}")
for edge_count in (5, 20, 50):
    edges = [LRU(EDGE_SLOTS) for _ in range(edge_count)]
    shield = LRU(SHIELD_SLOTS)
    pick = random.Random(7)  # which edge each request lands on: users spread evenly over the edges
    edge_misses = shield_misses = 0
    for key in trace:
        if not edges[pick.randrange(edge_count)].get_or_fill(key):
            edge_misses += 1
            if not shield.get_or_fill(key):
                shield_misses += 1
    hit_ratio = 1 - edge_misses / REQUESTS
    print(f"{edge_count:>5}  {hit_ratio:>14.1%}  {edge_misses:>15,}  {shield_misses:>13,}  {edge_misses / shield_misses:>25.1f}x")

Output

100,000 requests for 10,000 files; each edge keeps 1,000, the shield 5,000
edges  edge hit ratio  origin requests  with a shield  shield cuts origin load by
    5           54.6%           45,377         17,422                        2.6x
   20           51.6%           48,432         17,174                        2.8x
   50           45.4%           54,593         17,083                        3.2x

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 cdn_hit_ratio.py

Two numbers move in opposite directions. As the edges multiply, each sees a smaller share of the traffic and its hit ratio falls, from 54.6 % with 5 edges to 45.4 % with 50, so the origin’s load without a shield grows from 45,377 to 54,593 requests. With the shield, the origin’s load stays almost flat at about 17,000, because the shield sees every edge’s misses and keeps the popular files for all of them: 2.6 times fewer origin requests with 5 edges, 3.2 times with 50. Without any CDN, the origin would have answered all 100,000.

A shield is not free. It is one more hop on every miss, it adds cost, and it helps little for content that is rarely requested or cannot be cached. The same guide names the cases where it pays most: users spread across many regions, origins that do expensive work per request such as packaging video or resizing images, and origins with little spare capacity.

Edge compute

CDNs also run code at the edge. One CDN’s documentation lists what such functions do: change requests and responses as they pass, perform basic authentication and authorisation, and generate whole responses at the edge, all close to the user and without servers to manage (Amazon CloudFront Developer Guide). Typical uses in a design:

  • normalising cache keys before the lookup, as in the experiment above;
  • redirects and A/B assignment that would otherwise cost a trip to the origin;
  • checking a signed token on every request, so that the origin sees only valid callers;
  • assembling a cached page with a small personal fragment fetched separately.

State is where edge compute gets hard. An edge is one of many copies spread across the world, so a value written at one edge is not instantly visible at the others; a counter, a stock level or a session that must be correct everywhere at once still belongs in the origin’s database. Read-mostly data, such as configuration or the keys that verify tokens, suits the edge well.

To see what each layer did, the Cache-Status response header records how every cache on the path handled a response, such as a hit, or a miss and why it was forwarded (RFC 9211), and Age says how long a served copy has been in a cache.

Page Speed Test (Lighthouse via PageSpeed Insights) See how fast a page loads from a lab browser, and which of its files are slow or uncached.

Interview questions

Warm-up (fresher to mid level): how do a pull CDN and a push CDN differ, and when would you use each? A pull CDN fetches a file from the origin the first time an edge is asked for it and keeps it for its TTL, so the publisher only sets cache headers and the CDN fills itself. Most websites use pull, because it needs no preparation and adapts to whatever users request; its cost is that the first request for each file at each edge misses. A push CDN is loaded ahead: the publisher uploads files to the CDN’s storage before users ask, so even the first request hits. Push suits large media libraries, software downloads and launches, when a wave of first misses at every edge could overload the origin, at the price of managing what is stored where and paying for storage. A strong answer adds that a shield tier gives a pull CDN much of push’s protection for the origin.

Key takeaways

  • A CDN edge is a shared cache near users: it answers hits itself and sends misses to the origin, often through a shield.
  • The cache key is the hit ratio: in the experiment, normalising the key raised hits from 17.4 % to 80 %, and dropping a parameter that mattered served 445 wrong pages.
  • Set lifetimes explicitly: max-age and s-maxage for freshness, no-cache to force revalidation, private or no-store for personal data, and never public on authenticated responses.
  • Release by versioned addresses with immutable, not purges: a purge cannot reach browser caches.
  • A shield kept origin load flat at about 17,000 requests while edges grew from 5 to 50; without it, origin load grew with every edge.
  • Edge compute suits stateless work such as redirects, token checks and key normalisation; shared mutable state stays at the origin.

Exercise

Exercise · Medium · Python

Normalise a URL into an edge cache key

Requests that should get the same response must get the same cache key, and requests that should not must not. Write cache_key(url, keep) in cache_key.py. keep is the set of query parameter names that change the response; every other parameter is dropped. Build the key from the URL like this:

- the scheme and the host in lower case; - no port when it is the scheme's default (:443 for https, :80 for http), and the port kept otherwise; - the path exactly as it is, upper case included (/Products and /products can be different pages), or / when it is empty; - only the parameters named in keep, sorted by name and then by value, with repeated parameters kept; - no ? when no parameter is left, and never the fragment (the part after #, which browsers do not send).

cache_key("https://Shop.Example:443/p/42?utm_source=mail&size=m&color=red", {"color", "size"})
# "https://shop.example/p/42?color=red&size=m"

Keep the parameters' values exactly as they appear in the URL. The sample tests import cache_key from cache_key.py and run in your browser.

Starter code · cache_key.py

def cache_key(url, keep):
    """The edge cache key of `url`, keeping only the query parameters named in `keep`."""
    # Replace this line with your code.
    return url
The sample tests · test_cache_key.py
from cache_key import cache_key


def test_tracking_parameters_dropped():
    """drops every parameter that is not in keep"""
    assert cache_key("https://shop.example/p/42?utm_source=mail&color=red&gclid=8812", {"color"}) == "https://shop.example/p/42?color=red"


def test_order_does_not_matter():
    """sorts the parameters, so their order makes no new key"""
    a = cache_key("https://shop.example/p/42?size=m&color=red", {"color", "size"})
    b = cache_key("https://shop.example/p/42?color=red&size=m", {"color", "size"})
    assert a == b == "https://shop.example/p/42?color=red&size=m"


def test_host_scheme_and_ports():
    """lower-cases the scheme and host, drops default ports and keeps others"""
    assert cache_key("HTTPS://Shop.Example:443/p/42", set()) == "https://shop.example/p/42"
    assert cache_key("http://shop.example:80/p/42", set()) == "http://shop.example/p/42"
    assert cache_key("https://shop.example:8443/p/42", set()) == "https://shop.example:8443/p/42"


def test_path_and_fragment():
    """keeps the path's case, uses / for an empty path and drops the fragment"""
    assert cache_key("https://shop.example/Products/42#reviews", set()) == "https://shop.example/Products/42"
    assert cache_key("https://shop.example?utm_source=ads", set()) == "https://shop.example/"


def test_repeated_parameters():
    """keeps repeated parameters, sorted by value"""
    assert cache_key("https://shop.example/s?tag=b&tag=a&page=2", {"tag"}) == "https://shop.example/s?tag=a&tag=b"
A hint

urllib.parse.urlsplit(url) gives you the parts you need: scheme, hostname (already in lower case), port, path and query. Split the query on & and each part on the first =, keep the pairs whose name is in keep, sort the pairs, and join them back with = and &.

The sample tests run on this device, in your browser (Pyodide): nothing is sent to mysmartcopilot.com. The first run downloads Python (about 13.5 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.

Check yourself

7 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.

  1. Question 1 of 7 A marketing team adds utm_source and utm_campaign parameters to every link to a product page. What happens to the CDN's hit ratio for that page, if the cache key is the full URL?

    Choose one answer.

    Show the answer to question 1

    Answer: It falls, because each combination of parameters is a new cache key that misses once at every edge

    A cache key is built from at least the method and the full target URI, query string included. In this lesson's example, keys built from the raw URLs gave a 17.4 % hit ratio for the same 200 responses that normalised keys served at 80 %.

  2. Question 2 of 7 Which response header suits a script published at a versioned address such as /assets/app.3f9c2b.js?

    Choose one answer.

    Show the answer to question 2

    Answer: Cache-Control: public, max-age=31536000, immutable

    A versioned address never changes its content, so it can be fresh for a year, and immutable tells browsers not to revalidate it even when the user reloads. A new version gets a new address instead of a purge.

  3. Question 3 of 7 What does Cache-Control no-cache tell a cache?

    Choose one answer.

    Show the answer to question 3

    Answer: It may store the response, but must check with the origin that it is still current before each reuse

    no-cache forbids reusing a stored response without validating it first, which costs a round trip but often only a short 304 response. no-store is the directive that forbids storing it at all.

  4. Question 4 of 7 A release changed /styles/site.css, which was served with max-age=86400. The team purged it from the CDN, yet some users still see the old styles. Why?

    Choose one answer.

    Show the answer to question 4

    Answer: Browsers that already hold the old file keep using it until its own max-age runs out, and a CDN purge cannot reach them

    A purge clears copies in the CDN's caches only. Each browser cached the file for a day and will not ask again before that. Publishing the new file at a new, versioned address avoids the problem entirely.

  5. Question 5 of 7 In this lesson's simulation with 20 edges, how many times fewer requests reached the origin with the shield than without it? Give one decimal.

    Type a number.

    Show the answer to question 5

    Answer: 2.8 times (anything from 2.75 to 2.85 counts)

    Without the shield the origin answered every edge miss, 48,432 requests; with it, only the shield's misses, 17,174: about 2.8 times fewer. With 50 edges the factor grew to 3.2, because more edges mean more first-time misses for the same files.

  6. Question 6 of 7 An API answers requests that carry an Authorization header with each user's own account data. A developer adds Cache-Control public, max-age=60 to speed it up. What can go wrong?

    Choose one answer.

    Show the answer to question 6

    Answer: public allows a shared cache to reuse the response for other requests, so one user's data can be served to another

    Shared caches leave responses to authenticated requests alone by default, and public, s-maxage and must-revalidate are the directives that switch that protection off. Personal responses need private or no-store, never public.

  7. Question 7 of 7 What does a shield tier change for the origin?

    Choose one answer.

    Show the answer to question 7

    Answer: Edges send their misses to the shield, so a file wanted at many edges costs the origin about one request instead of one per edge

    The shield is one more cache layer between the edges and the origin. Every edge's misses go through it, and simultaneous requests for the same file can be combined into a single origin request.

References

Related tools

Report a problem with this lesson

Quick answers and tool search

Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.