System Design (High-Level Design) Module 4 – Networking for system designers
Load balancing: L4 vs L7 and health checks
How layer-4 and layer-7 load balancers differ, how to design health checks that never empty the pool, and how to drain, warm up and pin servers safely.
What you will learn
- Compare layer-4 and layer-7 load balancing by what each one reads and decides
- Design health checks, connection draining and slow start for a fleet of servers
- Decide when sticky sessions are acceptable and what they cost
Before you start
On this page
A load balancer spreads incoming traffic over a group of servers and stops sending it to servers that fail. Its two kinds differ in how much of that traffic they read. A layer-4 (L4) balancer sees only addresses, ports and the transport protocol, so it chooses a server once for each TCP connection or UDP flow and forwards everything on it. A layer-7 (L7) balancer ends the client’s connection, reads each HTTP request and chooses a server for every request, by its path, a header or a cookie if you ask it to. L4 is cheaper and works for any protocol; L7 can route, retry and spread individual requests, for more work per request. Whichever you use, three habits decide whether users notice failures and deploys: health checks that judge each server on its own, draining a server before it leaves, and a slow start when it joins.
What each kind of balancer reads
The 4 and the 7 are layers of the OSI model: layer 4 is the transport layer (TCP and UDP), layer 7 the application layer (HTTP and the protocols built on it).
An L4 balancer never looks inside the byte stream. Amazon’s Network Load Balancer, which works at layer 4, picks a target for a TCP connection from a hash of the protocol, both addresses and ports, and the TCP sequence number, and keeps that connection on the same target for its whole life (Network Load Balancer). A Kubernetes Service in kube-proxy’s iptables mode is an L4 balancer too: by default it picks a backend pod at random for each new connection (Virtual IPs and Service Proxies). Because it reads nothing above the ports, an L4 balancer can carry any protocol (a database, a message broker, real-time game traffic over UDP) and can pass TLS through untouched, so only the server holds the certificate.
An L7 balancer is a reverse proxy that understands HTTP. It holds two connections, one to the client and one to
the server, and makes a decision for every request: Amazon’s Application Load Balancer, which works at layer 7,
evaluates its listener rules (path, host, headers, method, query string) and then picks a target from the matching
group (Application Load Balancer).
It also reuses its connections to each server for the requests of many clients, and adds X-Forwarded-For,
X-Forwarded-Proto and X-Forwarded-Port headers so that servers still learn who the client was
(How Elastic Load Balancing works).
The price is work on every request: TLS, parsing and a second hop. gRPC’s own comparison of proxy designs lists
higher latency and the proxy’s throughput as the costs of balancing in a proxy
(gRPC Load Balancing).
What each kind of balancer sees, and what it decides
Text description of the diagram
The figure has two panels, one above the other.
- Layer 4, one choice per connection. A client opens one TCP connection to an L4 balancer, which reads only addresses and ports, never the request inside. It sends all of that connection's packets to one server, server B, so server B gets every request sent on the connection.
- Layer 7, one choice per request. A client sends two requests, GET /api/cart and GET /img/42.jpg, to an L7 balancer, which ends the TLS connection and reads each HTTP request. It sends requests whose path starts with /api/ to the API servers and requests whose path starts with /img/ to the image servers.
| Criterion | Layer 4 (L4) | Layer 7 (L7) |
|---|---|---|
| What it reads | Addresses, ports and the protocol | The whole HTTP request: method, path, headers, cookies |
| What it balances | Connections or UDP flows | Requests, even several on one connection |
| Routes by path or header | No | Yes |
| Retries elsewhere | Only a failed connection attempt | Whole requests, when they are safe to repeat |
| TLS | Can pass it through to the server | Usually ends it, so that it can read the request |
| Protocols | Any protocol over TCP or UDP | HTTP/1.1, HTTP/2, gRPC, WebSockets |
| Work per request | Low: nothing to parse | Higher: TLS, parsing and its own connections |
| When to choose | Non-HTTP traffic, TLS kept to the server, huge packet rates | Websites, HTTP APIs and gRPC that need routing or retries |
Large sites often use both. An L4 tier spreads connections over a fleet of L7 proxies, and the proxies route each request to a service. Google’s Maglev is a well-documented example of the first tier: a software balancer on ordinary Linux servers, with the network’s routers spreading packets over the Maglev machines (Maglev, NSDI 2016).
Connections or requests: why the unit matters
Balancing connections works well while connections are short and similar. It breaks down when a few long-lived connections carry most of the traffic, which is exactly what HTTP/2 and gRPC clients do: they open one connection and send every request over it. This model gives four such clients to three servers and then adds a fourth server:
// Connections or requests: what an L4 and an L7 balancer each spread over the servers.
// Four clients keep one long-lived HTTP/2 connection each (as gRPC clients do) and send all their requests on it.
const clients = [
{ name: "mobile gateway", perSecond: 600 },
{ name: "search service", perSecond: 100 },
{ name: "billing job", perSecond: 50 },
{ name: "admin panel", perSecond: 50 },
];
const total = clients.reduce((sum, c) => sum + c.perSecond, 0);
// One second of traffic: the clients' requests interleaved in time, as they arrive at the balancer.
function arrivals() {
const list = [];
for (const [i, c] of clients.entries()) {
for (let k = 0; k < c.perSecond; k++) list.push({ client: i, at: (k + 0.5) / c.perSecond });
}
return list.sort((a, b) => a.at - b.at || a.client - b.client);
}
// L4: the balancer picks a server when a connection opens (round robin here) and every request on that
// connection follows it, because the balancer never looks inside the connection.
function layer4(servers, pinned) {
const load = new Array(servers).fill(0);
for (const r of arrivals()) load[pinned[r.client]]++;
return load;
}
// L7: the balancer ends the client's connection, reads each request and picks a server per request.
function layer7(servers) {
const load = new Array(servers).fill(0);
let next = 0;
for (const _ of arrivals()) {
load[next]++;
next = (next + 1) % servers;
}
return load;
}
const names = ["A", "B", "C", "D"];
function show(label, load) {
const cells = names.map((_, i) => (i < load.length ? String(load[i]) : "-").padStart(5)).join("");
const busiest = Math.max(...load) / (total / load.length);
console.log(`${label.padEnd(14)}${cells}${`${busiest.toFixed(2)}x`.padStart(8)}`);
}
console.log("Four clients, one connection each,");
console.log(`sending ${clients.map((c) => c.perSecond).join(", ")} requests a second`);
console.log(`${"per second".padEnd(14)}${names.map((n) => n.padStart(5)).join("")}${"busiest".padStart(8)}`);
const pinnedTo3 = clients.map((_, i) => i % 3); // connections opened in turn: A, B, C, then A again
show("L4, 3 servers", layer4(3, pinnedTo3));
show("L7, 3 servers", layer7(3));
console.log("Server D joins; no client reconnects:");
show("L4, 4 servers", layer4(4, pinnedTo3));
show("L7, 4 servers", layer7(4)); Output
Four clients, one connection each, sending 600, 100, 50, 50 requests a second per second A B C D busiest L4, 3 servers 650 100 50 - 2.44x L7, 3 servers 267 267 266 - 1.00x Server D joins; no client reconnects: L4, 4 servers 650 100 50 0 3.25x L7, 4 servers 200 200 200 200 1.00x
Recorded with Node.js 24.21.0 on macOS 26 arm64. To run it yourself: mise exec node@24.21.0 -- node units.mjs
Runs on this device, in your browser. The first run downloads JavaScript (about 0.6 MB), which is kept for the next runs.
Your run, in this browser
At L4 the mobile gateway’s single connection lands on server A and takes 600 requests a second with it, so A carries 650 while C carries 50: the busiest server does 2.44 times its fair share. Adding server D changes nothing for the existing connections, so D sits idle until a client reconnects, and the imbalance grows to 3.25 times. At L7 every request is a new decision, the load is even, and the new server gets its quarter at once.
Three fixes are common. Put an HTTP/2-aware L7 proxy in front of the servers, so that it spreads individual requests.
Let the clients balance requests themselves over a list of servers, which removes the extra hop but needs that logic
in every client. Or make the servers close connections after a while: gRPC’s MAX_CONNECTION_AGE makes a server
send GOAWAY once a connection reaches a maximum age, letting its calls finish, so that clients reconnect and an L4
balancer can spread them again (gRPC proposal A9).
Health checks: who stays in the pool
A balancer learns that a server is broken in two ways. Active checks send it a probe on a schedule and count the
answers. Passive checks watch real traffic: open-source nginx marks a server failed when requests to it fail
(max_fails is 1 within fail_timeout, 10 s by default), and Envoy’s outlier detection, which it calls a form of
passive health checking, ejects a host after a run of server errors
(nginx upstream module,
Envoy outlier detection).
With active checks, the time to notice a dead server is about the check interval times the number of failures it
takes to leave. HAProxy checks every 2 s by default and takes a server out after 3 failures in a row, so it notices
in about 6 s (HAProxy manual, inter and fall). An Application
Load Balancer’s defaults are a check every 30 s and 2 failures, about a minute
(ALB health checks).
At 200 requests a second per server, that minute sends about 12,000 requests to a server that cannot answer them,
unless passive checks or retries catch them first. Faster checks shorten that window but make servers flap in and
out of the pool during a brief pause, so pick the interval from how many failed requests you can afford.
What a check tests matters more than how often it runs. Kubernetes separates two questions. A liveness probe asks whether the process can make progress at all; when it fails often enough, the kubelet restarts the container. A readiness probe asks whether the server should get traffic right now; when it fails, the pod’s address is removed from the endpoints of every Service that selects it, and nothing restarts. The same page warns that a careless liveness probe can cause cascading failures, restarting containers under high load (Kubernetes probes).
The dangerous check is one that calls a dependency every server shares. This model runs six servers behind a balancer with HAProxy’s default check timing, and compares four check designs in two incidents: the shared database slows down, or one server’s disk fills up.
// Four health-check designs, two incidents: which requests still get an answer?
// Six servers share one database. 70 % of requests are answered from each server's own cache; 30 % need the
// database. The balancer checks every server every 2 s; 3 failed checks in a row take a server out of the pool,
// 2 passes in a row bring it back.
const SERVERS = 6;
const PER_TICK = 120; // requests per 100 ms tick: 1,200 a second
const NEEDS_DATABASE = 0.3;
const CHECK_EVERY = 20; // ticks (2 s)
const FALL = 3;
const RISE = 2;
const incidents = [
// From t = 30 s to t = 60 s (ticks 300 to 599).
{ name: "database slow 30 s", databaseSlow: (t) => t >= 300 && t < 600, broken: () => false },
{ name: "server 3 disk full 30 s", databaseSlow: () => false, broken: (t, s) => s === 2 && t >= 300 && t < 600 },
];
// What each design's check tests. A slow database makes a check that queries it time out.
const designs = [
{ name: "process up", passes: () => true, failOpen: false },
{ name: "deep check", passes: (inc, t, s) => !inc.broken(t, s) && !inc.databaseSlow(t), failOpen: false },
{ name: "deep, fail open", passes: (inc, t, s) => !inc.broken(t, s) && !inc.databaseSlow(t), failOpen: true },
{ name: "local checks", passes: (inc, t, s) => !inc.broken(t, s), failOpen: false },
];
function run(design, inc) {
const inPool = new Array(SERVERS).fill(true);
const streak = new Array(SERVERS).fill(0); // consecutive results that disagree with the current state
let answered = 0;
let offered = 0;
let emptyTicks = 0;
for (let t = 0; t < 900; t++) {
if (t % CHECK_EVERY === 0) {
for (let s = 0; s < SERVERS; s++) {
const ok = design.passes(inc, t, s);
if (ok === inPool[s]) streak[s] = 0;
else if (++streak[s] >= (inPool[s] ? FALL : RISE)) {
inPool[s] = !inPool[s];
streak[s] = 0;
}
}
}
let targets = [...inPool.keys()].filter((s) => inPool[s]);
if (targets.length === 0) {
if (t >= 300 && t < 700) emptyTicks++;
if (design.failOpen) targets = [...inPool.keys()]; // no healthy server: send to all of them
}
if (t < 300 || t >= 700) continue; // count the requests from t = 30 s to t = 70 s
offered += PER_TICK;
for (const s of targets) {
const share = PER_TICK / targets.length;
const success = inc.broken(t, s) ? 0 : inc.databaseSlow(t) ? 1 - NEEDS_DATABASE : 1;
answered += share * success;
}
}
return { answered: answered / offered, emptySeconds: emptyTicks / 10 };
}
console.log("Requests answered from t = 30 s to 70 s,");
console.log("and time with no healthy server");
for (const inc of incidents) {
console.log(`\nIncident: ${inc.name}`);
console.log(`${"check design".padEnd(17)}${"answered".padStart(9)}${"none healthy".padStart(14)}`);
for (const d of designs) {
const r = run(d, inc);
console.log(`${d.name.padEnd(17)}${`${(100 * r.answered).toFixed(1)} %`.padStart(9)}${`${r.emptySeconds} s`.padStart(14)}`);
}
} Output
Requests answered from t = 30 s to 70 s, and time with no healthy server Incident: database slow 30 s check design answered none healthy process up 77.5 % 0 s deep check 27.0 % 28 s deep, fail open 77.5 % 28 s local checks 77.5 % 0 s Incident: server 3 disk full 30 s check design answered none healthy process up 87.5 % 0 s deep check 98.3 % 0 s deep, fail open 98.3 % 0 s local checks 98.3 % 0 s
Recorded with Node.js 24.21.0 on macOS 26 arm64. To run it yourself: mise exec node@24.21.0 -- node health_checks.mjs
Runs on this device, in your browser. The first run downloads JavaScript (about 0.6 MB), which is kept for the next runs.
Your run, in this browser
The deep check, which runs a database query, is the best design for the broken disk and the worst for the slow database. When the database slows, every server fails the same check at the same moment, the pool is empty for 28 s, and even the 70 % of requests that never needed the database go unanswered: 27 % answered, against 77.5 % for the other designs. A check that only asks whether the process is up keeps every server in the pool, including the broken one, which then fails one request in six. A check of each server’s own health (its process, its disk, its warm cache) gets both incidents right.
Amazon’s Builders’ Library describes the same trade. Checks that fail together on many servers do more harm than good, so fast-acting balancer checks are best kept to local health, and the balancer should fail open: when every server fails its checks, send traffic to all of them anyway (Implementing health checks). Real balancers do this. An Application Load Balancer whose targets are all unhealthy routes to all of them, and Envoy stops honouring health status once fewer than 50 % of a cluster’s hosts are healthy, its default panic threshold (Envoy panic threshold). A Kubernetes Service does not fail open on readiness: a pod that fails its readiness probe leaves the Service’s endpoints, so if every pod fails at once, the Service has no pod to send traffic to. The probes page suggests a readiness probe may check the back-end services a pod depends on, which helps when the fault is in one pod’s own connection; for a dependency all pods share, it is the 27 % row above.
Exercise · Medium · Python
Track health checks and fail open
Write the two decisions a load balancer makes from health checks, in health.py.
track(results, rise=2, fall=3) follows one server through its health checks. The server starts "up". results lists the checks in the order they ran (True for a pass), and the function returns the server's state after each check, as a list of "up" and "down". An up server goes down after fall failed checks in a row; a down server comes back up after rise passed checks in a row. A single result the other way starts the count again.
routable(states, fail_open=True) returns the names of the servers the balancer may send requests to, sorted. states maps each server's name to "up" or "down". Servers that are up get traffic. When no server is up, a balancer that fails open sends traffic to all of them anyway, and one that does not sends it to none.
For example, track([False, False, False, True, True]) is ["up", "up", "down", "down", "up"], and routable({"a": "down", "b": "down"}) is ["a", "b"].
Starter code · health.py
def track(results, rise=2, fall=3):
"""The server's state ("up" or "down") after each health check."""
# Replace this line with your code.
return []
def routable(states, fail_open=True):
"""Sorted names of the servers that may get requests."""
# Replace this line with your code.
return [] The sample tests · test_health.py
from health import routable, track
def test_goes_down_after_fall_failures():
"""goes down only after 3 failed checks in a row"""
assert track([True, False, False, True, False, False, False]) == ["up", "up", "up", "up", "up", "up", "down"]
def test_comes_back_after_rise_passes():
"""comes back up only after 2 passed checks in a row"""
assert track([False, False, False, True, False, True, True]) == ["up", "up", "down", "down", "down", "down", "up"]
def test_other_thresholds():
"""uses the rise and fall it is given"""
assert track([False, True, False, False], rise=1, fall=1) == ["down", "up", "down", "down"]
assert track([False, False, True, True, True], rise=3, fall=2) == ["up", "down", "down", "down", "up"]
assert track([]) == []
def test_routable_up_servers():
"""sends only to up servers, sorted by name"""
assert routable({"c": "up", "a": "up", "b": "down"}) == ["a", "c"]
assert routable({"c": "up", "a": "up", "b": "down"}, fail_open=False) == ["a", "c"]
def test_fail_open():
"""fails open when no server is up, unless told not to"""
assert routable({"b": "down", "a": "down"}) == ["a", "b"]
assert routable({"b": "down", "a": "down"}, fail_open=False) == [] A hint
Keep two values while you walk through the results: the current state and how many results in a row disagreed with it. A result that agrees with the state resets the count to 0. When the count reaches fall (for an up server) or rise (for a down one), flip the state and reset the count. For routable, filter the up servers first, then decide what to return when that list is empty.
Results of the sample tests
| Test | Result | Details |
|---|
What your code printed
The sample tests run on this device, in your browser (Pyodide): nothing is sent to mysmartcopilot.com. The first run downloads Python (about 13.5 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.
Draining: let a server finish before it leaves
When a server leaves on purpose, for a deploy or a scale-in, the order of steps decides how many requests break.
Take it out of the balancer’s pool first, so that no new requests arrive. Let the requests already in progress finish,
up to a limit. Only then stop the process. An Application Load Balancer calls this the deregistration delay: a leaving
target is draining for up to 300 s by default, finishes sooner once it has no requests in flight, and clients get a
5xx error if the target closes connections before then
(target group attributes).
This model replaces four servers one at a time while 400 requests a second arrive, 2 % of them uploads that take 5 to 25 s:
"""A rolling deploy of 4 servers, replaced one at a time, with four ways of taking each old server out.
400 requests a second arrive evenly and go round robin to the servers in the balancer's pool. 98 % take 50-300 ms;
2 % are uploads of 5-25 s. A replacement joins the pool 10 s after the old server stops.
"""
import random
RATE = 400 # requests a second
START = 60.0 # the deploy starts at t = 60 s
BOOT = 10.0 # seconds before a replacement takes traffic
CHECK_EVERY, FALL = 2.0, 3 # health checks every 2 s, 3 failures to leave the pool
def durations(seed=7):
rng = random.Random(seed)
while True:
if rng.random() < 0.02:
yield 5 + 20 * rng.random() # an upload
else:
yield 0.05 + 0.25 * rng.random()
def deploy(strategy):
"""Returns (dropped requests, seconds the deploy took)."""
pool = [0, 1, 2, 3] # servers in the balancer's pool (ids); a replacement gets a new id
ends = {s: [] for s in pool} # when each server's in-flight requests finish
dead = set() # stopped, but the balancer may still send to them
dropped, rr, served_by = 0, 0, durations()
plan = [("leave", START, 0)] # (event, time, server)
t, n, next_id = 0.0, 0, 4
while plan:
plan.sort(key=lambda e: e[1])
event, at, server = plan[0]
if t >= at: # handle the next deploy step before routing more requests
plan.pop(0)
if event == "leave":
if strategy == "stop at once, checks find out":
dead.add(server)
first_check = (at // CHECK_EVERY + 1) * CHECK_EVERY
plan.append(("out of pool", first_check + (FALL - 1) * CHECK_EVERY, server))
plan.append(("stop", at, server))
else:
pool.remove(server)
drain = {"stop at once": 0.0, "drain 2 s, then stop": 2.0}.get(strategy)
if drain is None: # drain until idle, at most 30 s
busy_until = max([e for e in ends[server] if e > at], default=at)
drain = min(busy_until - at, 30.0)
plan.append(("stop", at + drain, server))
elif event == "out of pool":
pool.remove(server)
elif event == "stop":
dropped += sum(1 for e in ends[server] if e > at) # cut off mid-request
ends[server] = []
dead.add(server)
plan.append(("join", at + BOOT, next_id))
next_id += 1
elif event == "join":
pool.append(server)
ends[server] = []
if server - 4 < 3: # replace the next old server
plan.append(("leave", at, server - 3))
else:
return dropped, at - START
continue
target = pool[rr % len(pool)]
rr += 1
if target in dead:
dropped += 1 # sent to a server that has already stopped
else:
ends[target] = [e for e in ends[target] if e > t] + [t + next(served_by)]
n += 1
t = n / RATE
LABELS = {
"stop at once, checks find out": "stop; checks find out",
"stop at once": "stop at once",
"drain 2 s, then stop": "drain 2 s, then stop",
"drain until idle, 30 s max": "drain until idle, 30 s",
}
print(f"{'how a server leaves':23}{'dropped':>8}{'deploy':>9}")
for strategy, label in LABELS.items():
lost, took = deploy(strategy)
print(f"{label:23}{lost:>8}{took:>7.0f} s") Output
how a server leaves dropped deploy stop; checks find out 2638 40 s stop at once 259 40 s drain 2 s, then stop 144 48 s drain until idle, 30 s 0 130 s
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 drain_demo.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
Stopping the process first and leaving the balancer to find out from its checks is the worst order: for about 6 s per
server, new requests go to a server that is gone, and with the requests cut off in flight, 2,638 fail. Stopping at
once after leaving the pool still cuts off everything in flight, 259 requests, most of them long uploads. A 2 s drain
saves the short requests and loses 144 uploads. Draining until the server is idle, with a 30 s limit, loses nothing, and the deploy takes 130 s instead of
40 s. So set the drain limit above the longest request you are not willing to cut off, and expect it to set the pace
of your deploys. Connections that last for hours, such as WebSockets, cannot be waited for; the server has to ask its
clients to reconnect elsewhere, as gRPC’s GOAWAY does.
While one server of four is out, the other three take 133 requests a second each instead of 100, a third more. Keep that headroom, or deploys become overloads.
Slow start: let a new server warm up
A server that has just started is slower than its neighbours: its caches are empty, its connection pools have not
filled and a just-in-time compiler may not have compiled its hot code yet. Given a full share of traffic at once, it
answers slowly or times out. Slow start ramps its share up instead. HAProxy’s slowstart grows a returning server’s
weight and connection limit linearly from 0 to 100 % over the time you set, and an Application Load Balancer increases
the requests it sends to a new target linearly over its slow start duration
(HAProxy manual, slowstart;
target group attributes).
With a 60 s linear ramp, a new server gets a quarter of a full share after 15 s and half after 30 s. Envoy’s slow start
can also ramp non-linearly, through an aggression setting, and never starts below 10 % of the full weight by
default (Envoy slow start).
Slow start and least-request balancing pull in opposite directions. A new server has no requests in progress, so a policy that picks the server with the fewest sees it as the best choice. The Application Load Balancer does not allow slow start together with its least outstanding requests algorithm, so with that algorithm the warm-up has to come from somewhere else, such as warming the caches before the server joins.
Sticky sessions: when they are acceptable
A sticky session sends all of one client’s requests to the same server. An L7 balancer does it with a cookie: the
Application Load Balancer’s duration-based stickiness sets a cookie named AWSALB that names the chosen target
(target group attributes).
An L4 balancer can only use the client’s address. A Kubernetes Service with sessionAffinity: ClientIP keeps a client
on one pod for 10,800 s (3 hours) by default, and nginx’s ip_hash hashes only the first three octets of an IPv4
address, so every client in the same /24 network lands on the same server
(Virtual IPs and Service Proxies,
nginx upstream module).
Stickiness costs three things. Load becomes uneven, because some clients are much busier than others and many users can share one address behind an office or carrier network. A failed server still loses its sessions, because stickiness decides where requests go, not where data lives. And new servers fill slowly, because existing sessions stay where they are: the Application Load Balancer’s guide warns of uneven load after a large scale-out and suggests expiring the cookies to rebalance.
So use stickiness only where losing it costs speed, not data. A per-user cache on the server is a fair case: a moved user gets a cache miss. A long-lived connection is sticky by nature: a WebSocket stays with the server that accepted the upgrade. A temporary bridge while you move in-memory sessions into a shared store is fine too. A cart or a login that exists only in one server’s memory is not: the Twelve-Factor App keeps processes stateless and calls sticky sessions a violation of its rules (Processes).
HTTP Header Checker Look at a site's response headers for the cookies and headers a load balancer adds.Choosing in an interview
- The protocol is not HTTP, TLS must reach the server untouched, or packet rates are extreme: L4.
- You need routing by path or host, per-request balancing of HTTP/2 or gRPC, or retries: L7, retrying only
requests that are safe to repeat. RFC 9110 calls PUT, DELETE and the safe methods idempotent and says a client
should not retry other methods automatically unless it knows they are idempotent
(RFC 9110, section 9.2.2). HAProxy’s default
retry-onsetting retries only a request that could not be sent at all. - Health checks: local and cheap, with a detection time you have calculated, and a balancer that fails open.
- Leaving and joining: drain for longer than your longest normal request, and ramp new servers up.
- Stickiness: an optimisation you can lose, never the only copy of state.
Interview questions
Warm-up (fresher to mid level): what is the difference between a layer-4 and a layer-7 load balancer, and when would you pick each? A layer-4 balancer works at the transport layer: it sees addresses, ports and the protocol, picks a server when a TCP connection or UDP flow starts, and keeps the whole connection on that server. It is fast, works for any protocol and can pass TLS through. A layer-7 balancer ends the client’s connection, reads every HTTP request and picks a server per request, so it can route by path or header, retry safe requests elsewhere and spread HTTP/2 or gRPC requests that share one connection. That costs TLS and parsing work on every request. Pick L4 for non-HTTP traffic, end-to-end TLS or very high packet rates, and L7 for websites and APIs that need routing or per-request balancing. Many systems use both: L4 in front, spreading connections over L7 proxies.
Key takeaways
- An L4 balancer decides once per connection from addresses and ports; an L7 balancer reads and decides per request.
- Long-lived connections defeat L4 balancing: in the run, one gateway made a server carry 2.44 times its share, and a new server stayed idle until clients reconnected.
- Detection time is about the check interval times the failure threshold; size it from the failed requests you can afford.
- Check what is local to each server, and let the balancer fail open: a shared-dependency check emptied the pool and answered 27 % of requests where local checks answered 77.5 %.
- Leave the pool, drain, then stop: in the run that order took dropped requests from 2,638 to 0, for a slower deploy.
- Ramp new servers up, and treat stickiness as an optimisation you can lose.
Check yourself
8 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.
References
- What is Elastic Load Balancing? (Amazon Web Services)
- How Elastic Load Balancing works (Amazon Web Services)
- What is a Network Load Balancer? (Amazon Web Services)
- What is an Application Load Balancer? (Amazon Web Services)
- Health checks for Application Load Balancer target groups (Amazon Web Services)
- Edit target group attributes for your Application Load Balancer (Amazon Web Services)
- Implementing health checks (Amazon Builders' Library) (Amazon Web Services)
- HAProxy 3.2 Configuration Manual (HAProxy Technologies)
- Liveness, Readiness, and Startup Probes (The Kubernetes Authors)
- Virtual IPs and Service Proxies (The Kubernetes Authors)
- Module ngx_http_upstream_module (nginx)
- Panic threshold (Envoy documentation) (Envoy Project)
- Outlier detection (Envoy documentation) (Envoy Project)
- Slow start mode (Envoy documentation) (Envoy Project)
- Maglev: A Fast and Reliable Software Network Load Balancer (NSDI 2016) (USENIX)
- RFC 9110: HTTP Semantics (section 9.2.2, Idempotent Methods) (IETF)
- gRPC Load Balancing (gRPC Authors)
- A9: Server-side Connection Management (gRPC proposal) (gRPC Authors)
- The Twelve-Factor App: VI. Processes (The Twelve-Factor App (Adam Wiggins))
Related tools
Report a problem with this lesson
Kept only in this browser. Your Learn progress