Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

System Design (High-Level Design)  Module 7 – Caching

Where caches live, from browser to database

The cache layers between a user and the disk, what each one is good at, how each goes stale, and how to plan for the empty cache after a flush or a deploy.

  • Beginner
  • 25 minutes
  • Examples run with Python 3.14.8, Pyodide 314.0.7, Node.js 24.21.0 and quickjs 0.32.0
  • By MySmartCoPilot

What you will learn

  • List the cache layers between a user and the database, and say who shares each one
  • Explain what each layer is good at and how its copy goes stale
  • Decide which layers a design needs from a staleness budget for each kind of data
  • Estimate the database load after a cache flush and compare three ways to warm a cache

Before you start

On this page

A cache is a copy of data kept closer to the place that needs it, so that a request can be answered without doing the original work again. Between a user and the disk that finally holds the data there are usually six places where such a copy can live: the browser or app, a CDN at the edge, a gateway or reverse proxy in front of the service, the memory of each application instance, a distributed cache such as Redis, Valkey or Memcached, and the database server’s own memory (its buffer pool and the operating system’s page cache).

Each layer is a trade. The closer a copy sits to the user, the more work and distance every hit saves, and the harder that copy is to correct when the data changes. Designing caching means choosing, for each kind of data, which of these copies may exist and how old each one may be, and then planning for the moment they are all empty.

The layers, from the user to the disk

Cache layers from the user to the disk: browser, CDN, gateway, in-process, distributed cache, then the buffer pool and OS page cache.UserBrowser or app cacheprivate to one userexample: a logo for a yearCDN edge cacheshared by everyone near itexample: a product photo for a dayGateway or reverse proxy cacheshared, in front of the serviceexample: a public page for 30 sIn-process cacheone copy in each app instanceexample: feature flags for 60 sDistributed cacheRedis, Valkey or Memcachedexample: a product record for 5 minDatabase serverBuffer poolpages the database managesOS page cachefile pages the kernel keepsDiskpage not in memorypage not cachedmissmissmissmissmiss: a query

Where a copy can live, from the user to the disk

Text description of the diagram

The diagram is a chain from top to bottom. Each layer passes a request it cannot answer, a miss, to the next one.

  1. The user's browser or app keeps a private cache for that one user, for example a logo kept for a year.
  2. A CDN edge cache is shared by everyone near it, for example a product photo kept for a day.
  3. A gateway or reverse proxy cache is shared and sits in front of the service, for example a public page kept for 30 seconds.
  4. An in-process cache lives inside each app instance, one copy per instance, for example feature flags kept for 60 seconds.
  5. A distributed cache such as Redis, Valkey or Memcached is shared by all app instances, for example a product record kept for 5 minutes.
  6. On a miss there, the app sends a query to the database server. Inside it, the database's own buffer pool holds the pages it manages; a page that is not in memory is read through the operating system's page cache, and a page the kernel has not cached is read from the disk.

A request falls through the layers until one of them has a usable copy. Each layer that answers saves the work of every layer below it, and each one that misses adds its own lookup to the time of the request. Each layer has its own strength and its own way of going stale:

  • Browser or app cache. One user on one device has the copy. It saves the whole trip, or turns it into a short check. It goes stale when the data changes before the copy’s lifetime ends, and the server cannot reach it until the browser asks again.
  • CDN edge cache. Everyone near one edge location shares the copy. It is good at public files and pages for users far from the origin. It goes stale when a file changes under the same URL, or when a purge reaches some locations later than others.
  • Gateway or reverse proxy cache. Everyone using the service shares it, in one place you control. It is good at public responses that many users request. It goes stale when the origin changes a response and nothing tells it.
  • In-process cache. Only the requests on one application instance see it. It is the fastest lookup there is, a read of the process’s own memory. It goes stale when another instance changes the data, because every instance keeps its own copy and its own clock.
  • Distributed cache. Every instance of every service that uses it shares one copy, which survives application restarts and deploys. It goes stale when the database changes and the code that should delete or update the entry fails or runs late.
  • Buffer pool and page cache. Every query on one database server shares them. They keep hot pages of tables and indexes in memory with no work for the application. The database keeps them correct, so they never serve stale rows, but they are empty after a restart.
HTTP Header Checker Look at the Cache-Control, Age and ETag headers a site sends, and work out which layer answered.

Close to the user: browser, app and CDN

The HTTP caching standard calls a cache that serves one user a private cache and one that stores responses for reuse by more than one user a shared cache (RFC 9111, section 1). A browser’s cache and a mobile app’s local store are private, so they may keep personal data such as the reader’s own order history. A CDN is shared, so a personal response must never be stored there: the server marks it with Cache-Control: private, because, as MDN puts it, a cookie alone does not make a response private.

The price of these layers is control. A CDN copy can be purged through its provider’s interface (see CDNs and edge caching), but a copy in a browser stays until that browser asks again. That is why static files carry a fingerprint of their content in the file name (app.3f9c2a.js): a new version is a new URL, so the old copy may live for a year without anyone seeing it.

In front of and inside the service: gateway, in-process and distributed

A reverse proxy or API gateway can cache public responses for the whole service, in one place your team runs. Behind it, every application instance can keep a small cache in its own memory. That in-process cache is the fastest lookup in the whole chain, but each instance has its own copy: the Amazon Builders’ Library article on caching warns that results become inconsistent from server to server across the fleet. With 40 instances and a 60-second lifetime, a changed value can be old on some instances and new on others for up to a minute.

A distributed cache is a separate service, such as Redis, Valkey or Memcached, that every instance reads over the network. It holds one copy for the whole fleet, so an update or a delete is seen by everyone, and it keeps its contents when the application restarts. Each hit costs a network round trip (see measured latency numbers), and the cache is now a system you size, monitor and keep alive. Redis can also tell an application which keys to drop from a local copy, which joins the two layers: the client-side caching reference describes how the server sends invalidation messages for the keys a client has read.

Inside the database: buffer pool and page cache

The database server caches too, without any code from you. PostgreSQL keeps table and index pages in its shared buffers (128 MB by default; the documentation suggests 25 % of the memory of a dedicated server as a starting point) and also relies on the operating system’s cache of file pages (PostgreSQL 18 documentation). These copies never serve stale rows, because the database manages them itself. They still cost something on a miss: a page that is in neither cache must be read from the disk, and both caches start empty after a restart.

A staleness budget for every kind of data

A cache is a copy, and a copy can be older than the data it copies. So before choosing layers, write down a staleness budget for each kind of data: how old may this value be, for this reader, before it causes a wrong decision or a complaint? The budget, not the layer, is the design decision; the layers follow from it.

Here is what that looks like for a news site:

  • Style sheet with a fingerprinted name: a year, in the browser and on the CDN. A new version gets a new name, so an old copy is never wrong.
  • Article text: 5 minutes, on the CDN or in a distributed cache. Corrections are rare and may take minutes to appear.
  • Live score ticker: 5 seconds, in a distributed cache or with a short CDN lifetime. Readers expect it current; a few seconds is invisible.
  • Comment count: 1 minute, in a distributed cache or in each instance. Nobody notices a count that is a minute behind.
  • The reader’s own subscription status: 0 right after a change, otherwise 10 minutes, in the reader’s own browser and in a distributed cache keyed by reader. A reader who has just paid must see the paid articles at once.
  • Payment status of a subscription order: 0, so no cache at all: read the database. A stale “failed” or “paid” decides money and access.

Two habits keep budgets honest. First, the same data can have different budgets for different readers: the person who just edited their profile needs to see the edit, while everyone else can wait a minute. Second, a budget of zero means “no cache for this read”, not “a very short lifetime”, because even a one-second copy can be the wrong one when it matters.

Choosing the layers a design needs

Work from the read path and the budgets, one kind of data at a time:

  1. Public and versioned (scripts, styles, images whose URL changes with their content): browser and CDN, with a long lifetime.
  2. Public and changing (product pages, articles, scores): a CDN or gateway with a lifetime inside the budget, or a distributed cache if the response is assembled per request.
  3. Personal (carts, profiles, order history): never a shared HTTP cache; a private copy in the browser, and a distributed cache with the user in the key.
  4. Read by every request on every instance (feature flags, configuration, permission rules): an in-process copy with a short lifetime, refreshed from the distributed cache or the database.
  5. Zero budget (balances, payment results, stock at the moment of purchase): the database, with indexes and replicas sized for it.

Each layer you add is one more copy to keep correct and one more cache that can be empty at the worst moment. A design rarely needs all six: many services run with a CDN for static files, a distributed cache for hot records and the database’s own caches.

Plan for the empty cache

Caches hide load. A database that serves a tenth of the reads while the cache is warm may be sized for that tenth, and then a flush, a restart or a deploy that changes the cache’s keys sends it every read at once. The Amazon Builders’ Library article describes such a service as addicted to its cache, and recommends load tests with caching switched off. This model shows what an empty cache does to a database sized for the warm case:

Database load after a cache flush, three ways to warm up Python · cold_start_sim.py
"""Database load after a product cache is emptied, for three ways of bringing the cache back into service.

The model: 20,000 reads a second reach a cache in front of a catalogue of 1,000,000 products, in four
popularity tiers. An entry stays in the cache for 5 minutes once it is loaded, and the database can serve
5,000 reads a second. A product that is not in the cache costs a database read every time it is asked for
until its first read loads it. Reads of one product arrive at random, so a product read r times a second is
still missing after x seconds of traffic with probability exp(-r * x). Every number below is an expected
value worked out from that model (no random draws), so the output is the same on every run.
"""
import math

READS = 20_000      # reads a second at the cache
TTL = 300           # seconds an entry is kept after it is loaded
DB_LIMIT = 5_000    # reads a second the database can serve
STEP = 10           # strategy 3 moves 10 % more of the traffic every 10 seconds
TIERS = [           # name, products, share of all reads
    ("bestsellers", 1_000, 0.35),
    ("popular", 9_000, 0.25),
    ("regular", 90_000, 0.25),
    ("long tail", 900_000, 0.15),
]


def per_product(share, products):
    return READS * share / products


def misses(exposure, tiers=TIERS):
    """Database reads a second from products not read yet, after `exposure` seconds of full traffic."""
    return sum(READS * share * math.exp(-per_product(share, n) * exposure) for _, n, share in tiers)


# Warm: a product read r times a second waits about 1 / r seconds for a read, then stays TTL seconds, so it
# misses once per cycle and its hit ratio is r * TTL / (1 + r * TTL).
WARM = sum(READS * share / (1 + per_product(share, n) * TTL) for _, n, share in TIERS)


def all_at_once(t):
    return misses(t)


def hottest_first(t):
    """The 10,000 bestsellers and popular products were loaded before the cache took any traffic."""
    return misses(t, TIERS[2:])


def share_moved(t):
    return min(1.0, 0.1 * (1 + int(t // STEP)))


def exposure(t):
    """Seconds of full traffic the new cache has seen by second t: the area under share_moved."""
    done = int(t // STEP)
    return sum(STEP * share_moved(STEP * j) for j in range(done)) + (t - STEP * done) * share_moved(t)


def in_steps(t):
    """The new cache misses on its share of the traffic; the old, warm cache still serves the rest."""
    moved = share_moved(t)
    return moved * misses(exposure(t)) + (1 - moved) * WARM


print(f"Products by tier, at {READS:,} reads a second")
print(f"{'tier':<12}{'products':>10}{'reads/s each':>14}{'warm hit ratio':>16}")
for name, n, share in TIERS:
    r = per_product(share, n)
    print(f"{name:<12}{n:>10,}{r:>14.3f}{r * TTL / (1 + r * TTL):>16.1%}")
print(f"Warm: {WARM:,.0f} database reads a second (limit {DB_LIMIT:,}); the cache answers {1 - WARM / READS:.1%} of reads.")
print()

strategies = {"all at once": all_at_once, "hottest 10,000 first": hottest_first, "10 % steps": in_steps}
print("Database reads a second after the cache is emptied")
print(f"{'second':<22}" + "".join(f"{name:>22}" for name in strategies))
for t in [0, 1, 2, 5, 10, 20, 30, 60, 90, 120]:
    print(f"{t:<22}" + "".join(f"{load(t):>22,.0f}" for load in strategies.values()))

# Peak, time over the limit and the reads beyond it, from 0.01-second steps over the first two minutes.
DT = 0.01
for label, measure in [
    ("peak", lambda load: f"{max(load(i * DT) for i in range(12_000)):,.0f}"),
    ("seconds over limit", lambda load: f"{sum(DT for i in range(12_000) if load(i * DT) > DB_LIMIT):.0f}"),
    ("reads beyond limit", lambda load: f"{sum(max(0, load(i * DT) - DB_LIMIT) * DT for i in range(12_000)):,.0f}"),
]:
    print(f"{label:<22}" + "".join(f"{measure(load):>22}" for load in strategies.values()))

Output

Products by tier, at 20,000 reads a second
tier          products  reads/s each  warm hit ratio
bestsellers      1,000         7.000          100.0%
popular          9,000         0.556           99.4%
regular         90,000         0.056           94.3%
long tail      900,000         0.003           50.0%
Warm: 1,816 database reads a second (limit 5,000); the cache answers 90.9% of reads.

Database reads a second after the cache is emptied
second                           all at once  hottest 10,000 first            10 % steps
0                                     20,000                 8,000                 3,635
1                                     10,595                 7,720                 3,252
2                                      9,100                 7,454                 3,049
5                                      7,049                 6,738                 2,820
10                                     5,790                 5,770                 3,572
20                                     4,453                 4,452                 3,715
30                                     3,659                 3,659                 3,770
60                                     2,635                 2,635                 3,593
90                                     2,256                 2,256                 2,993
120                                    2,017                 2,017                 2,414
peak                                  20,000                 8,000                 3,803
seconds over limit                        15                    15                     0
reads beyond limit                    29,856                19,797                     0

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 cold_start_sim.py

Warm, the cache answers 90.9 % of the reads, so the database serves about 1,800 reads a second, well inside its limit of 5,000. Look at the three strategies after a flush:

  • All at once. For an instant the database receives all 20,000 reads a second, four times its limit, and it stays over the limit for 15 seconds. About 30,000 reads arrive beyond what it can serve: in a real system they wait, time out or retry, and retries add even more load.
  • The hottest 10,000 first. Loading the bestsellers and popular products before taking traffic (1 % of the catalogue, 60 % of the reads) cuts the peak from 20,000 to 8,000 reads a second. It does not shorten the overload, because the 90,000 regular products, which nobody loaded, decide how long it lasts.
  • 10 % steps. Moving the traffic to the new cache in steps of 10 % every 10 seconds, while the old cache still serves the rest, keeps the database under 3,900 reads a second the whole time.

The model is kind to the database: it assumes every miss is answered at once. An overloaded database answers slowly, so products stay missing for longer and the reads that pile up on them make the overload worse. Read the peak as the number to design for, and the duration as a best case.

Interview questions

Warm-up (fresher to mid-level): What does a cache hit ratio measure, and what does it hide?

It is the share of lookups that the cache answers itself: hits divided by hits plus misses, over some period. A ratio of 95 % at 20,000 lookups a second means 1,000 a second still reach the next layer, and that number, not the ratio, is what the database must survive. The ratio hides three things. It averages over keys, so a few very hot keys can make it look healthy while the long tail misses. It says nothing about freshness: a cache full of stale values can score 100 %. And it describes the warm state only, so it tells you nothing about the load after a flush or a deploy.

Key takeaways

  • A copy can live in six places: the browser or app, a CDN, a gateway, each application instance, a distributed cache and the database server’s own memory.
  • Closer copies save more work and are harder to correct; in-process copies are fastest but differ between instances.
  • Give each kind of data a staleness budget for each reader. A budget of zero means no cache for that read.
  • Personal responses never belong in a shared cache: mark them private.
  • Size the database for the empty cache, or bring a cache into service gradually: an empty cache sends every read to the next layer at once.

Exercise

Exercise · Easy · Python

Plan where eight kinds of shop data may be cached

An online shop wants a caching plan. Write cache_plan() in plan.py: it returns a dictionary with one entry for each of the eight kinds of data below. Each entry is a pair (layers, budget):

- layers is a set of the layers that may hold a copy, chosen from LAYERS: "browser", "cdn", "gateway", "in-process" and "distributed". An empty set means no cache: every read goes to the database. - budget is the staleness budget in seconds, a whole number: how old a cached copy may be. It is 0 exactly when the set of layers is empty.

The eight kinds of data:

1. "app-bundle": the shop's JavaScript file, the same for every visitor; its name contains a hash of its content, so a new version always has a new URL. It must be cached in the browser and on the CDN, for at least a day. 2. "product-photo": public photos; a replaced photo keeps its URL and must show within a day. It must be on the CDN. 3. "product-price": public; it changes a few times a day. The checkout reads the price from the database again, so listing pages may show a price up to 60 seconds old. 4. "stock-badge": the "only 3 left" badge on listing pages; public, may be up to 30 seconds old. 5. "feature-flags": settings the servers read on every request; a change must apply within 60 seconds. They are never sent to browsers, so only server-side layers may hold them, and every instance needs its own copy. 6. "profile-header": the signed-in visitor's name in the page header; personal; may be up to 5 minutes old. 7. "order-history": the signed-in visitor's past orders; personal; may be up to 5 minutes old. 8. "wallet-balance": the money in the visitor's shop wallet; it must never be stale.

Personal data never goes into a shared HTTP cache: neither "cdn" nor "gateway". Every cached kind of data needs a budget greater than 0 and within its limit.

Starter code · plan.py

LAYERS = {"browser", "cdn", "gateway", "in-process", "distributed"}


def cache_plan():
    """Return {kind of data: (set of layers, staleness budget in seconds)} for the eight kinds in the prompt."""
    # Replace this line with your plan.
    return {}
The sample tests · test_plan.py
from plan import LAYERS, cache_plan

KINDS = {"app-bundle", "product-photo", "product-price", "stock-badge", "feature-flags", "profile-header", "order-history", "wallet-balance"}
PERSONAL = {"profile-header", "order-history", "wallet-balance"}
SHARED = {"cdn", "gateway"}


def test_every_kind_has_a_plan():
    """plans exactly the eight kinds of data, each as (set of layers, whole seconds)"""
    plan = cache_plan()
    assert set(plan) == KINDS
    for kind, (layers, budget) in plan.items():
        assert set(layers) <= LAYERS, kind
        assert isinstance(budget, int) and budget >= 0, kind


def test_zero_budget_means_no_cache():
    """uses a budget of 0 exactly for data with no cache"""
    for kind, (layers, budget) in cache_plan().items():
        assert (budget == 0) == (len(layers) == 0), kind


def test_personal_data_stays_out_of_shared_caches():
    """keeps personal data out of the CDN and the gateway"""
    plan = cache_plan()
    for kind in PERSONAL:
        assert not set(plan[kind][0]) & SHARED, kind


def test_money_is_never_cached():
    """reads the wallet balance from the database every time"""
    assert cache_plan()["wallet-balance"] == (set(), 0)


def test_versioned_bundle_is_cached_close_to_users():
    """caches the fingerprinted bundle in the browser and on the CDN for at least a day"""
    layers, budget = cache_plan()["app-bundle"]
    assert {"browser", "cdn"} <= set(layers)
    assert budget >= 86_400


def test_budgets_stay_within_their_limits():
    """keeps each cached kind of data within its staleness limit"""
    plan = cache_plan()
    limits = {"product-photo": 86_400, "product-price": 60, "stock-badge": 30, "feature-flags": 60, "profile-header": 300, "order-history": 300}
    for kind, limit in limits.items():
        assert 0 < plan[kind][1] <= limit, kind
    assert "cdn" in plan["product-photo"][0]


def test_flags_live_in_every_instance():
    """keeps feature flags in server-side layers, with a copy in each instance"""
    layers, _ = cache_plan()["feature-flags"]
    assert "in-process" in layers
    assert set(layers) <= {"in-process", "distributed"}
A hint

Start from the budget, then pick the layers. Write "wallet-balance": (set(), 0) first: money has a budget of 0, so it has no cache. For the personal data, choose only from "browser", "in-process" and "distributed"; for the feature flags, only server-side layers. The limits in the prompt are maximums, so any whole number from 1 up to the limit passes.

The sample tests run on this device, in your browser (Pyodide): nothing is sent to mysmartcopilot.com. The first run downloads Python (about 13.5 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.

Check yourself

5 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.

  1. Question 1 of 5 A service runs 40 instances behind a load balancer. Which cache layer can hold a different value for the same key on different instances?

    Choose one answer.

    Show the answer to question 1

    Answer: The in-process cache, because each instance keeps its own copy with its own lifetime

    Each instance fills and expires its in-process cache on its own, so after a change some instances can still serve the old value until their copy expires. A distributed cache holds one copy for the whole fleet, and the buffer pool is managed by the database, which never serves a stale row from it.

  2. Question 2 of 5 A response with a signed-in user's order history passes through a CDN on its way to the browser. Which header keeps it out of the CDN's cache while still letting the user's own browser keep it?

    Choose one answer.

    Show the answer to question 2

    Answer: Cache-Control: private

    private tells shared caches not to store the response, while a private cache such as the browser may keep it. public and max-age allow shared caches to store it, and a cookie alone does not make a response private.

  3. Question 3 of 5 A cache receives 20,000 lookups a second and its hit ratio is 95 %. How many lookups a second reach the database?

    Type a number.

    Show the answer to question 3

    Answer: 1000 a second

    The misses are 5 % of the lookups: 20,000 × 0.05 = 1,000 a second. That is the load the database must serve while the cache is warm; after a flush it would receive all 20,000.

  4. Question 4 of 5 In the cold-start model, loading the 10,000 hottest products before taking traffic cut the peak from 20,000 to 8,000 reads a second, but the database still stayed over its limit for 15 seconds. Why?

    Choose one answer.

    Show the answer to question 4

    Answer: The 90,000 regular products, which nobody loaded in advance, still missed on their first reads and kept the load high

    Pre-warming removes the misses of the products it loads, which were most of the reads at the very start. How long the overload lasts depends on the products that warm up slowly, here the regular tier, which still sent 5,000 reads a second at the start and needed about 15 seconds to fall below the limit together with the long tail.

  5. Question 5 of 5 A team gives the payment status of an order a staleness budget of 0. What does that mean for caching?

    Choose one answer.

    Show the answer to question 5

    Answer: That read is served from the database every time; no cache holds it

    A budget of zero means no copy may be older than the data at all, and any cache, even with a very short lifetime, can return the old value at the moment it matters. So the read goes to the database.

References

Related tools

Report a problem with this lesson

Quick answers and tool search

Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.