Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

System Design (High-Level Design)  Module 4 – Networking for system designers

Load balancing: L4 vs L7 and health checks

How layer-4 and layer-7 load balancers differ, how to design health checks that never empty the pool, and how to drain, warm up and pin servers safely.

  • Intermediate
  • 25 minutes
  • Examples run with Python 3.14.8, Pyodide 314.0.7, Node.js 24.21.0 and quickjs 0.32.0
  • By MySmartCoPilot

What you will learn

  • Compare layer-4 and layer-7 load balancing by what each one reads and decides
  • Design health checks, connection draining and slow start for a fleet of servers
  • Decide when sticky sessions are acceptable and what they cost

Before you start

On this page

A load balancer spreads incoming traffic over a group of servers and stops sending it to servers that fail. Its two kinds differ in how much of that traffic they read. A layer-4 (L4) balancer sees only addresses, ports and the transport protocol, so it chooses a server once for each TCP connection or UDP flow and forwards everything on it. A layer-7 (L7) balancer ends the client’s connection, reads each HTTP request and chooses a server for every request, by its path, a header or a cookie if you ask it to. L4 is cheaper and works for any protocol; L7 can route, retry and spread individual requests, for more work per request. Whichever you use, three habits decide whether users notice failures and deploys: health checks that judge each server on its own, draining a server before it leaves, and a slow start when it joins.

What each kind of balancer reads

The 4 and the 7 are layers of the OSI model: layer 4 is the transport layer (TCP and UDP), layer 7 the application layer (HTTP and the protocols built on it).

An L4 balancer never looks inside the byte stream. Amazon’s Network Load Balancer, which works at layer 4, picks a target for a TCP connection from a hash of the protocol, both addresses and ports, and the TCP sequence number, and keeps that connection on the same target for its whole life (Network Load Balancer). A Kubernetes Service in kube-proxy’s iptables mode is an L4 balancer too: by default it picks a backend pod at random for each new connection (Virtual IPs and Service Proxies). Because it reads nothing above the ports, an L4 balancer can carry any protocol (a database, a message broker, real-time game traffic over UDP) and can pass TLS through untouched, so only the server holds the certificate.

An L7 balancer is a reverse proxy that understands HTTP. It holds two connections, one to the client and one to the server, and makes a decision for every request: Amazon’s Application Load Balancer, which works at layer 7, evaluates its listener rules (path, host, headers, method, query string) and then picks a target from the matching group (Application Load Balancer). It also reuses its connections to each server for the requests of many clients, and adds X-Forwarded-For, X-Forwarded-Proto and X-Forwarded-Port headers so that servers still learn who the client was (How Elastic Load Balancing works). The price is work on every request: TLS, parsing and a second hop. gRPC’s own comparison of proxy designs lists higher latency and the proxy’s throughput as the costs of balancing in a proxy (gRPC Load Balancing).

An L4 balancer sends a whole TCP connection to one server; an L7 balancer reads each HTTP request and routes it by path.Layer 4: one choice per connectionLayer 7: one choice per requestClientL4 balancerreads addresses and ports,never the requestServer Bgets every requeston this connectionClientL7 balancerends TLS, readseach HTTP requestAPI serversImage serversone TCP connectionall of its packetsGET /api/cartGET /img/42.jpg/api/*/img/*

What each kind of balancer sees, and what it decides

Text description of the diagram

The figure has two panels, one above the other.

  1. Layer 4, one choice per connection. A client opens one TCP connection to an L4 balancer, which reads only addresses and ports, never the request inside. It sends all of that connection's packets to one server, server B, so server B gets every request sent on the connection.
  2. Layer 7, one choice per request. A client sends two requests, GET /api/cart and GET /img/42.jpg, to an L7 balancer, which ends the TLS connection and reads each HTTP request. It sends requests whose path starts with /api/ to the API servers and requests whose path starts with /img/ to the image servers.
Layer-4 and layer-7 load balancing
Criterion Layer 4 (L4)Layer 7 (L7)
What it reads Addresses, ports and the protocolThe whole HTTP request: method, path, headers, cookies
What it balances Connections or UDP flowsRequests, even several on one connection
Routes by path or header NoYes
Retries elsewhere Only a failed connection attemptWhole requests, when they are safe to repeat
TLS Can pass it through to the serverUsually ends it, so that it can read the request
Protocols Any protocol over TCP or UDPHTTP/1.1, HTTP/2, gRPC, WebSockets
Work per request Low: nothing to parseHigher: TLS, parsing and its own connections
When to choose Non-HTTP traffic, TLS kept to the server, huge packet ratesWebsites, HTTP APIs and gRPC that need routing or retries

Large sites often use both. An L4 tier spreads connections over a fleet of L7 proxies, and the proxies route each request to a service. Google’s Maglev is a well-documented example of the first tier: a software balancer on ordinary Linux servers, with the network’s routers spreading packets over the Maglev machines (Maglev, NSDI 2016).

Connections or requests: why the unit matters

Balancing connections works well while connections are short and similar. It breaks down when a few long-lived connections carry most of the traffic, which is exactly what HTTP/2 and gRPC clients do: they open one connection and send every request over it. This model gives four such clients to three servers and then adds a fourth server:

Four long-lived connections, balanced at L4 and at L7 JavaScript · units.mjs
// Connections or requests: what an L4 and an L7 balancer each spread over the servers.
// Four clients keep one long-lived HTTP/2 connection each (as gRPC clients do) and send all their requests on it.
const clients = [
  { name: "mobile gateway", perSecond: 600 },
  { name: "search service", perSecond: 100 },
  { name: "billing job", perSecond: 50 },
  { name: "admin panel", perSecond: 50 },
];
const total = clients.reduce((sum, c) => sum + c.perSecond, 0);

// One second of traffic: the clients' requests interleaved in time, as they arrive at the balancer.
function arrivals() {
  const list = [];
  for (const [i, c] of clients.entries()) {
    for (let k = 0; k < c.perSecond; k++) list.push({ client: i, at: (k + 0.5) / c.perSecond });
  }
  return list.sort((a, b) => a.at - b.at || a.client - b.client);
}

// L4: the balancer picks a server when a connection opens (round robin here) and every request on that
// connection follows it, because the balancer never looks inside the connection.
function layer4(servers, pinned) {
  const load = new Array(servers).fill(0);
  for (const r of arrivals()) load[pinned[r.client]]++;
  return load;
}

// L7: the balancer ends the client's connection, reads each request and picks a server per request.
function layer7(servers) {
  const load = new Array(servers).fill(0);
  let next = 0;
  for (const _ of arrivals()) {
    load[next]++;
    next = (next + 1) % servers;
  }
  return load;
}

const names = ["A", "B", "C", "D"];
function show(label, load) {
  const cells = names.map((_, i) => (i < load.length ? String(load[i]) : "-").padStart(5)).join("");
  const busiest = Math.max(...load) / (total / load.length);
  console.log(`${label.padEnd(14)}${cells}${`${busiest.toFixed(2)}x`.padStart(8)}`);
}

console.log("Four clients, one connection each,");
console.log(`sending ${clients.map((c) => c.perSecond).join(", ")} requests a second`);
console.log(`${"per second".padEnd(14)}${names.map((n) => n.padStart(5)).join("")}${"busiest".padStart(8)}`);
const pinnedTo3 = clients.map((_, i) => i % 3); // connections opened in turn: A, B, C, then A again
show("L4, 3 servers", layer4(3, pinnedTo3));
show("L7, 3 servers", layer7(3));
console.log("Server D joins; no client reconnects:");
show("L4, 4 servers", layer4(4, pinnedTo3));
show("L7, 4 servers", layer7(4));

Output

Four clients, one connection each,
sending 600, 100, 50, 50 requests a second
per second        A    B    C    D busiest
L4, 3 servers   650  100   50    -   2.44x
L7, 3 servers   267  267  266    -   1.00x
Server D joins; no client reconnects:
L4, 4 servers   650  100   50    0   3.25x
L7, 4 servers   200  200  200  200   1.00x

Recorded with Node.js 24.21.0 on macOS 26 arm64. To run it yourself: mise exec node@24.21.0 -- node units.mjs

At L4 the mobile gateway’s single connection lands on server A and takes 600 requests a second with it, so A carries 650 while C carries 50: the busiest server does 2.44 times its fair share. Adding server D changes nothing for the existing connections, so D sits idle until a client reconnects, and the imbalance grows to 3.25 times. At L7 every request is a new decision, the load is even, and the new server gets its quarter at once.

Three fixes are common. Put an HTTP/2-aware L7 proxy in front of the servers, so that it spreads individual requests. Let the clients balance requests themselves over a list of servers, which removes the extra hop but needs that logic in every client. Or make the servers close connections after a while: gRPC’s MAX_CONNECTION_AGE makes a server send GOAWAY once a connection reaches a maximum age, letting its calls finish, so that clients reconnect and an L4 balancer can spread them again (gRPC proposal A9).

Health checks: who stays in the pool

A balancer learns that a server is broken in two ways. Active checks send it a probe on a schedule and count the answers. Passive checks watch real traffic: open-source nginx marks a server failed when requests to it fail (max_fails is 1 within fail_timeout, 10 s by default), and Envoy’s outlier detection, which it calls a form of passive health checking, ejects a host after a run of server errors (nginx upstream module, Envoy outlier detection).

With active checks, the time to notice a dead server is about the check interval times the number of failures it takes to leave. HAProxy checks every 2 s by default and takes a server out after 3 failures in a row, so it notices in about 6 s (HAProxy manual, inter and fall). An Application Load Balancer’s defaults are a check every 30 s and 2 failures, about a minute (ALB health checks). At 200 requests a second per server, that minute sends about 12,000 requests to a server that cannot answer them, unless passive checks or retries catch them first. Faster checks shorten that window but make servers flap in and out of the pool during a brief pause, so pick the interval from how many failed requests you can afford.

What a check tests matters more than how often it runs. Kubernetes separates two questions. A liveness probe asks whether the process can make progress at all; when it fails often enough, the kubelet restarts the container. A readiness probe asks whether the server should get traffic right now; when it fails, the pod’s address is removed from the endpoints of every Service that selects it, and nothing restarts. The same page warns that a careless liveness probe can cause cascading failures, restarting containers under high load (Kubernetes probes).

The dangerous check is one that calls a dependency every server shares. This model runs six servers behind a balancer with HAProxy’s default check timing, and compares four check designs in two incidents: the shared database slows down, or one server’s disk fills up.

Four health-check designs in two incidents JavaScript · health_checks.mjs
// Four health-check designs, two incidents: which requests still get an answer?
// Six servers share one database. 70 % of requests are answered from each server's own cache; 30 % need the
// database. The balancer checks every server every 2 s; 3 failed checks in a row take a server out of the pool,
// 2 passes in a row bring it back.
const SERVERS = 6;
const PER_TICK = 120; // requests per 100 ms tick: 1,200 a second
const NEEDS_DATABASE = 0.3;
const CHECK_EVERY = 20; // ticks (2 s)
const FALL = 3;
const RISE = 2;

const incidents = [
  // From t = 30 s to t = 60 s (ticks 300 to 599).
  { name: "database slow 30 s", databaseSlow: (t) => t >= 300 && t < 600, broken: () => false },
  { name: "server 3 disk full 30 s", databaseSlow: () => false, broken: (t, s) => s === 2 && t >= 300 && t < 600 },
];

// What each design's check tests. A slow database makes a check that queries it time out.
const designs = [
  { name: "process up", passes: () => true, failOpen: false },
  { name: "deep check", passes: (inc, t, s) => !inc.broken(t, s) && !inc.databaseSlow(t), failOpen: false },
  { name: "deep, fail open", passes: (inc, t, s) => !inc.broken(t, s) && !inc.databaseSlow(t), failOpen: true },
  { name: "local checks", passes: (inc, t, s) => !inc.broken(t, s), failOpen: false },
];

function run(design, inc) {
  const inPool = new Array(SERVERS).fill(true);
  const streak = new Array(SERVERS).fill(0); // consecutive results that disagree with the current state
  let answered = 0;
  let offered = 0;
  let emptyTicks = 0;
  for (let t = 0; t < 900; t++) {
    if (t % CHECK_EVERY === 0) {
      for (let s = 0; s < SERVERS; s++) {
        const ok = design.passes(inc, t, s);
        if (ok === inPool[s]) streak[s] = 0;
        else if (++streak[s] >= (inPool[s] ? FALL : RISE)) {
          inPool[s] = !inPool[s];
          streak[s] = 0;
        }
      }
    }
    let targets = [...inPool.keys()].filter((s) => inPool[s]);
    if (targets.length === 0) {
      if (t >= 300 && t < 700) emptyTicks++;
      if (design.failOpen) targets = [...inPool.keys()]; // no healthy server: send to all of them
    }
    if (t < 300 || t >= 700) continue; // count the requests from t = 30 s to t = 70 s
    offered += PER_TICK;
    for (const s of targets) {
      const share = PER_TICK / targets.length;
      const success = inc.broken(t, s) ? 0 : inc.databaseSlow(t) ? 1 - NEEDS_DATABASE : 1;
      answered += share * success;
    }
  }
  return { answered: answered / offered, emptySeconds: emptyTicks / 10 };
}

console.log("Requests answered from t = 30 s to 70 s,");
console.log("and time with no healthy server");
for (const inc of incidents) {
  console.log(`\nIncident: ${inc.name}`);
  console.log(`${"check design".padEnd(17)}${"answered".padStart(9)}${"none healthy".padStart(14)}`);
  for (const d of designs) {
    const r = run(d, inc);
    console.log(`${d.name.padEnd(17)}${`${(100 * r.answered).toFixed(1)} %`.padStart(9)}${`${r.emptySeconds} s`.padStart(14)}`);
  }
}

Output

Requests answered from t = 30 s to 70 s,
and time with no healthy server

Incident: database slow 30 s
check design      answered  none healthy
process up          77.5 %           0 s
deep check          27.0 %          28 s
deep, fail open     77.5 %          28 s
local checks        77.5 %           0 s

Incident: server 3 disk full 30 s
check design      answered  none healthy
process up          87.5 %           0 s
deep check          98.3 %           0 s
deep, fail open     98.3 %           0 s
local checks        98.3 %           0 s

Recorded with Node.js 24.21.0 on macOS 26 arm64. To run it yourself: mise exec node@24.21.0 -- node health_checks.mjs

The deep check, which runs a database query, is the best design for the broken disk and the worst for the slow database. When the database slows, every server fails the same check at the same moment, the pool is empty for 28 s, and even the 70 % of requests that never needed the database go unanswered: 27 % answered, against 77.5 % for the other designs. A check that only asks whether the process is up keeps every server in the pool, including the broken one, which then fails one request in six. A check of each server’s own health (its process, its disk, its warm cache) gets both incidents right.

Amazon’s Builders’ Library describes the same trade. Checks that fail together on many servers do more harm than good, so fast-acting balancer checks are best kept to local health, and the balancer should fail open: when every server fails its checks, send traffic to all of them anyway (Implementing health checks). Real balancers do this. An Application Load Balancer whose targets are all unhealthy routes to all of them, and Envoy stops honouring health status once fewer than 50 % of a cluster’s hosts are healthy, its default panic threshold (Envoy panic threshold). A Kubernetes Service does not fail open on readiness: a pod that fails its readiness probe leaves the Service’s endpoints, so if every pod fails at once, the Service has no pod to send traffic to. The probes page suggests a readiness probe may check the back-end services a pod depends on, which helps when the fault is in one pod’s own connection; for a dependency all pods share, it is the 27 % row above.

Exercise · Medium · Python

Track health checks and fail open

Write the two decisions a load balancer makes from health checks, in health.py.

track(results, rise=2, fall=3) follows one server through its health checks. The server starts "up". results lists the checks in the order they ran (True for a pass), and the function returns the server's state after each check, as a list of "up" and "down". An up server goes down after fall failed checks in a row; a down server comes back up after rise passed checks in a row. A single result the other way starts the count again.

routable(states, fail_open=True) returns the names of the servers the balancer may send requests to, sorted. states maps each server's name to "up" or "down". Servers that are up get traffic. When no server is up, a balancer that fails open sends traffic to all of them anyway, and one that does not sends it to none.

For example, track([False, False, False, True, True]) is ["up", "up", "down", "down", "up"], and routable({"a": "down", "b": "down"}) is ["a", "b"].

Starter code · health.py

def track(results, rise=2, fall=3):
    """The server's state ("up" or "down") after each health check."""
    # Replace this line with your code.
    return []


def routable(states, fail_open=True):
    """Sorted names of the servers that may get requests."""
    # Replace this line with your code.
    return []
The sample tests · test_health.py
from health import routable, track


def test_goes_down_after_fall_failures():
    """goes down only after 3 failed checks in a row"""
    assert track([True, False, False, True, False, False, False]) == ["up", "up", "up", "up", "up", "up", "down"]


def test_comes_back_after_rise_passes():
    """comes back up only after 2 passed checks in a row"""
    assert track([False, False, False, True, False, True, True]) == ["up", "up", "down", "down", "down", "down", "up"]


def test_other_thresholds():
    """uses the rise and fall it is given"""
    assert track([False, True, False, False], rise=1, fall=1) == ["down", "up", "down", "down"]
    assert track([False, False, True, True, True], rise=3, fall=2) == ["up", "down", "down", "down", "up"]
    assert track([]) == []


def test_routable_up_servers():
    """sends only to up servers, sorted by name"""
    assert routable({"c": "up", "a": "up", "b": "down"}) == ["a", "c"]
    assert routable({"c": "up", "a": "up", "b": "down"}, fail_open=False) == ["a", "c"]


def test_fail_open():
    """fails open when no server is up, unless told not to"""
    assert routable({"b": "down", "a": "down"}) == ["a", "b"]
    assert routable({"b": "down", "a": "down"}, fail_open=False) == []
A hint

Keep two values while you walk through the results: the current state and how many results in a row disagreed with it. A result that agrees with the state resets the count to 0. When the count reaches fall (for an up server) or rise (for a down one), flip the state and reset the count. For routable, filter the up servers first, then decide what to return when that list is empty.

The sample tests run on this device, in your browser (Pyodide): nothing is sent to mysmartcopilot.com. The first run downloads Python (about 13.5 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.

Draining: let a server finish before it leaves

When a server leaves on purpose, for a deploy or a scale-in, the order of steps decides how many requests break. Take it out of the balancer’s pool first, so that no new requests arrive. Let the requests already in progress finish, up to a limit. Only then stop the process. An Application Load Balancer calls this the deregistration delay: a leaving target is draining for up to 300 s by default, finishes sooner once it has no requests in flight, and clients get a 5xx error if the target closes connections before then (target group attributes).

This model replaces four servers one at a time while 400 requests a second arrive, 2 % of them uploads that take 5 to 25 s:

Four ways to take a server out during a deploy Python · drain_demo.py
"""A rolling deploy of 4 servers, replaced one at a time, with four ways of taking each old server out.

400 requests a second arrive evenly and go round robin to the servers in the balancer's pool. 98 % take 50-300 ms;
2 % are uploads of 5-25 s. A replacement joins the pool 10 s after the old server stops.
"""
import random

RATE = 400  # requests a second
START = 60.0  # the deploy starts at t = 60 s
BOOT = 10.0  # seconds before a replacement takes traffic
CHECK_EVERY, FALL = 2.0, 3  # health checks every 2 s, 3 failures to leave the pool


def durations(seed=7):
    rng = random.Random(seed)
    while True:
        if rng.random() < 0.02:
            yield 5 + 20 * rng.random()  # an upload
        else:
            yield 0.05 + 0.25 * rng.random()


def deploy(strategy):
    """Returns (dropped requests, seconds the deploy took)."""
    pool = [0, 1, 2, 3]  # servers in the balancer's pool (ids); a replacement gets a new id
    ends = {s: [] for s in pool}  # when each server's in-flight requests finish
    dead = set()  # stopped, but the balancer may still send to them
    dropped, rr, served_by = 0, 0, durations()
    plan = [("leave", START, 0)]  # (event, time, server)
    t, n, next_id = 0.0, 0, 4
    while plan:
        plan.sort(key=lambda e: e[1])
        event, at, server = plan[0]
        if t >= at:  # handle the next deploy step before routing more requests
            plan.pop(0)
            if event == "leave":
                if strategy == "stop at once, checks find out":
                    dead.add(server)
                    first_check = (at // CHECK_EVERY + 1) * CHECK_EVERY
                    plan.append(("out of pool", first_check + (FALL - 1) * CHECK_EVERY, server))
                    plan.append(("stop", at, server))
                else:
                    pool.remove(server)
                    drain = {"stop at once": 0.0, "drain 2 s, then stop": 2.0}.get(strategy)
                    if drain is None:  # drain until idle, at most 30 s
                        busy_until = max([e for e in ends[server] if e > at], default=at)
                        drain = min(busy_until - at, 30.0)
                    plan.append(("stop", at + drain, server))
            elif event == "out of pool":
                pool.remove(server)
            elif event == "stop":
                dropped += sum(1 for e in ends[server] if e > at)  # cut off mid-request
                ends[server] = []
                dead.add(server)
                plan.append(("join", at + BOOT, next_id))
                next_id += 1
            elif event == "join":
                pool.append(server)
                ends[server] = []
                if server - 4 < 3:  # replace the next old server
                    plan.append(("leave", at, server - 3))
                else:
                    return dropped, at - START
            continue
        target = pool[rr % len(pool)]
        rr += 1
        if target in dead:
            dropped += 1  # sent to a server that has already stopped
        else:
            ends[target] = [e for e in ends[target] if e > t] + [t + next(served_by)]
        n += 1
        t = n / RATE


LABELS = {
    "stop at once, checks find out": "stop; checks find out",
    "stop at once": "stop at once",
    "drain 2 s, then stop": "drain 2 s, then stop",
    "drain until idle, 30 s max": "drain until idle, 30 s",
}
print(f"{'how a server leaves':23}{'dropped':>8}{'deploy':>9}")
for strategy, label in LABELS.items():
    lost, took = deploy(strategy)
    print(f"{label:23}{lost:>8}{took:>7.0f} s")

Output

how a server leaves     dropped   deploy
stop; checks find out      2638     40 s
stop at once                259     40 s
drain 2 s, then stop        144     48 s
drain until idle, 30 s        0    130 s

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 drain_demo.py

Stopping the process first and leaving the balancer to find out from its checks is the worst order: for about 6 s per server, new requests go to a server that is gone, and with the requests cut off in flight, 2,638 fail. Stopping at once after leaving the pool still cuts off everything in flight, 259 requests, most of them long uploads. A 2 s drain saves the short requests and loses 144 uploads. Draining until the server is idle, with a 30 s limit, loses nothing, and the deploy takes 130 s instead of 40 s. So set the drain limit above the longest request you are not willing to cut off, and expect it to set the pace of your deploys. Connections that last for hours, such as WebSockets, cannot be waited for; the server has to ask its clients to reconnect elsewhere, as gRPC’s GOAWAY does.

While one server of four is out, the other three take 133 requests a second each instead of 100, a third more. Keep that headroom, or deploys become overloads.

Slow start: let a new server warm up

A server that has just started is slower than its neighbours: its caches are empty, its connection pools have not filled and a just-in-time compiler may not have compiled its hot code yet. Given a full share of traffic at once, it answers slowly or times out. Slow start ramps its share up instead. HAProxy’s slowstart grows a returning server’s weight and connection limit linearly from 0 to 100 % over the time you set, and an Application Load Balancer increases the requests it sends to a new target linearly over its slow start duration (HAProxy manual, slowstart; target group attributes). With a 60 s linear ramp, a new server gets a quarter of a full share after 15 s and half after 30 s. Envoy’s slow start can also ramp non-linearly, through an aggression setting, and never starts below 10 % of the full weight by default (Envoy slow start).

Slow start and least-request balancing pull in opposite directions. A new server has no requests in progress, so a policy that picks the server with the fewest sees it as the best choice. The Application Load Balancer does not allow slow start together with its least outstanding requests algorithm, so with that algorithm the warm-up has to come from somewhere else, such as warming the caches before the server joins.

Sticky sessions: when they are acceptable

A sticky session sends all of one client’s requests to the same server. An L7 balancer does it with a cookie: the Application Load Balancer’s duration-based stickiness sets a cookie named AWSALB that names the chosen target (target group attributes). An L4 balancer can only use the client’s address. A Kubernetes Service with sessionAffinity: ClientIP keeps a client on one pod for 10,800 s (3 hours) by default, and nginx’s ip_hash hashes only the first three octets of an IPv4 address, so every client in the same /24 network lands on the same server (Virtual IPs and Service Proxies, nginx upstream module).

Stickiness costs three things. Load becomes uneven, because some clients are much busier than others and many users can share one address behind an office or carrier network. A failed server still loses its sessions, because stickiness decides where requests go, not where data lives. And new servers fill slowly, because existing sessions stay where they are: the Application Load Balancer’s guide warns of uneven load after a large scale-out and suggests expiring the cookies to rebalance.

So use stickiness only where losing it costs speed, not data. A per-user cache on the server is a fair case: a moved user gets a cache miss. A long-lived connection is sticky by nature: a WebSocket stays with the server that accepted the upgrade. A temporary bridge while you move in-memory sessions into a shared store is fine too. A cart or a login that exists only in one server’s memory is not: the Twelve-Factor App keeps processes stateless and calls sticky sessions a violation of its rules (Processes).

HTTP Header Checker Look at a site's response headers for the cookies and headers a load balancer adds.

Choosing in an interview

  • The protocol is not HTTP, TLS must reach the server untouched, or packet rates are extreme: L4.
  • You need routing by path or host, per-request balancing of HTTP/2 or gRPC, or retries: L7, retrying only requests that are safe to repeat. RFC 9110 calls PUT, DELETE and the safe methods idempotent and says a client should not retry other methods automatically unless it knows they are idempotent (RFC 9110, section 9.2.2). HAProxy’s default retry-on setting retries only a request that could not be sent at all.
  • Health checks: local and cheap, with a detection time you have calculated, and a balancer that fails open.
  • Leaving and joining: drain for longer than your longest normal request, and ramp new servers up.
  • Stickiness: an optimisation you can lose, never the only copy of state.

Interview questions

Warm-up (fresher to mid level): what is the difference between a layer-4 and a layer-7 load balancer, and when would you pick each? A layer-4 balancer works at the transport layer: it sees addresses, ports and the protocol, picks a server when a TCP connection or UDP flow starts, and keeps the whole connection on that server. It is fast, works for any protocol and can pass TLS through. A layer-7 balancer ends the client’s connection, reads every HTTP request and picks a server per request, so it can route by path or header, retry safe requests elsewhere and spread HTTP/2 or gRPC requests that share one connection. That costs TLS and parsing work on every request. Pick L4 for non-HTTP traffic, end-to-end TLS or very high packet rates, and L7 for websites and APIs that need routing or per-request balancing. Many systems use both: L4 in front, spreading connections over L7 proxies.

Key takeaways

  • An L4 balancer decides once per connection from addresses and ports; an L7 balancer reads and decides per request.
  • Long-lived connections defeat L4 balancing: in the run, one gateway made a server carry 2.44 times its share, and a new server stayed idle until clients reconnected.
  • Detection time is about the check interval times the failure threshold; size it from the failed requests you can afford.
  • Check what is local to each server, and let the balancer fail open: a shared-dependency check emptied the pool and answered 27 % of requests where local checks answered 77.5 %.
  • Leave the pool, drain, then stop: in the run that order took dropped requests from 2,638 to 0, for a slower deploy.
  • Ramp new servers up, and treat stickiness as an optimisation you can lose.

Check yourself

8 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.

  1. Question 1 of 8 Clients reach a pool of PostgreSQL read replicas over long-lived connections that speak PostgreSQL's own protocol. What should sit in front of the replicas?

    Choose one answer.

    Show the answer to question 1

    Answer: An L4 balancer that spreads connections, with a check that runs a cheap query on each replica itself

    The protocol is not HTTP, so an HTTP balancer has nothing to parse; an L4 balancer spreads the connections. The check asks each replica about itself (can it answer a query, is it too far behind), which is local to that replica.

  2. Question 2 of 8 Each client of a gRPC service keeps one HTTP/2 connection open for hours. After scaling from 4 to 8 servers behind an L4 balancer, the new servers stay idle. What fixes it?

    Choose one answer.

    Show the answer to question 2

    Answer: Balance each request, with an HTTP/2-aware L7 proxy or client-side balancing, or make servers close old connections now and then

    An L4 balancer chooses once per connection, and these connections never end, so new servers only get clients that reconnect. Balancing per request fixes it at the source; a maximum connection age (gRPC's MAX_CONNECTION_AGE) makes clients reconnect and spreads connections over time.

  3. Question 3 of 8 One domain serves /api from one service and /img from an image service, and the team wants a failed GET retried on another server. Which balancer?

    Choose one answer.

    Show the answer to question 3

    Answer: An L7 balancer with path rules, retrying only requests that are safe to repeat

    Routing by path needs the request, so it is layer 7, and so is resending a request to another server. GET is safe to repeat (RFC 9110 lists PUT, DELETE and the safe methods as idempotent); a POST is not, unless the API makes it so.

  4. Question 4 of 8 A multiplayer game sends its real-time state as UDP packets. What spreads players over the game servers?

    Choose one answer.

    Show the answer to question 4

    Answer: An L4 balancer that keeps each UDP flow on one server

    UDP has no connection, but each flow (the same addresses and ports) must stay on one game server, which is what an L4 balancer's flow hash does. Moving single packets between servers would split a game in two.

  5. Question 5 of 8 Every server's readiness check runs a query on the shared database. What happens when that database slows down for a minute?

    Choose one answer.

    Show the answer to question 5

    Answer: Every server fails its check at about the same time, and the pool empties unless the balancer fails open

    A shared dependency fails the same check on every server together. In the lesson's run, that design answered 27 % of requests during the incident against 77.5 % for local checks. Check what is local to each server, and let the balancer fail open when every server looks unhealthy.

  6. Question 6 of 8 Health checks run every 10 s and a server leaves the pool after 3 failed checks in a row. It crashes just after passing a check. About how many seconds pass before it leaves the pool?

    Type a number.

    Show the answer to question 6

    Answer: 30 seconds (anything from 29 to 31 counts)

    The next three checks fail at 10 s, 20 s and 30 s, and the third takes it out: interval × threshold, 10 × 3 = 30 s. At 200 requests a second, about 6,000 requests go to the dead server in that time unless retries or passive checks catch them.

  7. Question 7 of 8 Uploads to a service take up to 25 s. How should each old server leave during a deploy?

    Choose one answer.

    Show the answer to question 7

    Answer: Leave the balancer's pool first, let its requests finish with a drain limit above 25 s, then stop

    The drain must outlast the longest request you do not want to cut off. In the lesson's run, stopping first dropped 2,638 requests, a 2 s drain 144, and draining until idle (at most 30 s) none, at the price of a longer deploy.

  8. Question 8 of 8 When are sticky sessions a reasonable choice?

    Choose one answer.

    Show the answer to question 8

    Answer: When losing the stickiness costs only speed, such as a cache miss or a reconnect

    Stickiness is fine as an optimisation, for warm per-user caches or long-lived connections. It is not storage: when a server fails or leaves, its sticky users move and lose whatever lived only there. Shared addresses make source-address stickiness more uneven, not less.

References

Related tools

Report a problem with this lesson

Quick answers and tool search

Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.