System Design (High-Level Design) Module 3 – Performance and reliability fundamentals
Circuit breakers, bulkheads and shuffle sharding
How circuit breakers fail fast, how bulkheads stop one slow dependency from taking every thread, and how shuffle sharding contains one bad client.
What you will learn
- Explain the closed, open and half-open states of a circuit breaker and what moves a breaker between them
- Size a bulkhead with Little's law so that one slow dependency can fill only its own share of threads
- Calculate how shuffle sharding shrinks the share of customers that one bad client can affect
- Choose between a breaker, a bulkhead, shuffle sharding and a retry budget for a given failure
Before you start
On this page
Three tools stop one failing part from taking everything around it down. A circuit breaker watches the calls to a dependency and, once too many of them fail, stops sending them for a while: callers fail fast instead of waiting, and the dependency gets room to recover. A bulkhead gives each dependency, or each tenant, its own limited share of threads, connections or queue slots, so a slow one can fill only its own share. Shuffle sharding gives every customer a small random set of servers out of a larger fleet, so a customer whose requests break servers reaches only a few of them, and almost no other customer depends on exactly those few. All three answer the same failure: a shared resource used up by one bad part. This lesson measures that failure first, then builds each tool with the numbers that size it.
One slow dependency can take every thread
Take a web server with 200 worker threads and three endpoints, each calling its own dependency: search, at 300 requests a second with calls of about 20 ms; profiles, at 200 a second and 30 ms; and recommendations, at 100 a second and 40 ms. By Little’s law, the calls in flight are the rate times the time each takes, so on a normal day they hold 6 + 6 + 4 = 16 threads, a twelfth of the pool. Now the recommendations dependency starts to hang, and every call to it waits for its 5-second timeout. The same law says it now wants 100 × 5 = 500 threads. There are 200, so within about two seconds every thread is waiting on recommendations, and requests for search and profiles, which never call that dependency, find no thread to run on.
The simulation below runs exactly this, with seeded random arrivals, for 30 seconds; the hang starts at 10 seconds. It compares the shared pool with two of this lesson’s tools.
# One web server with 200 worker threads serves three endpoints, each calling its own dependency. At 10 s the
# recommendations dependency starts hanging: every call waits for the 5 s timeout and then fails. What happens to the
# two endpoints that never call it? Three designs, the same seeded traffic, 30 simulated seconds in steps of 1 ms.
import heapq
import random
from collections import deque
POOL = 200
TIMEOUT_MS = 5_000
SLOW_FROM, END = 10_000, 30_000
# endpoint: requests per millisecond, mean latency of its dependency in ms
ENDPOINTS = {"search": (0.3, 20), "profile": (0.2, 30), "recommendations": (0.1, 40)}
class Breaker:
"""Opens when half of the last 20 calls failed; after 5 s it lets 3 trial calls through."""
def __init__(self):
self.state, self.results, self.opened_at, self.trials = "closed", deque(maxlen=20), 0, 0
def allow(self, now):
if self.state == "open" and now - self.opened_at >= 5_000:
self.state, self.trials = "half-open", 0
if self.state == "closed":
return True
if self.state == "half-open" and self.trials < 3:
self.trials += 1
return True
return False
def record(self, now, ok):
if self.state == "half-open":
if not ok:
self.state, self.opened_at = "open", now
elif self.trials == 3:
self.state = "closed"
self.results.clear()
return
self.results.append(ok)
if len(self.results) == 20 and self.results.count(False) >= 10:
self.state, self.opened_at = "open", now
def simulate(limit=None, breaker=False, seed=5):
"""limit: the most threads that recommendations calls may hold (a bulkhead); breaker: one on that dependency."""
rng = random.Random(seed)
busy = {name: 0 for name in ENDPOINTS}
done = [] # heap of (finish time, sequence number, endpoint, ok)
counts = {name: {"ok": 0, "fallback": 0, "failed": 0} for name in ENDPOINTS}
cb = Breaker() if breaker else None
seq = 0
for now in range(END):
while done and done[0][0] <= now:
_, _, name, ok = heapq.heappop(done)
busy[name] -= 1
if name == "recommendations" and cb:
cb.record(now, ok)
arrivals = [name for name, (rate, _) in ENDPOINTS.items() if rng.random() < rate]
rng.shuffle(arrivals)
for name in arrivals:
after = now >= SLOW_FROM
tally = counts[name] if after else {"ok": 0, "fallback": 0, "failed": 0}
if name == "recommendations" and cb and not cb.allow(now):
tally["fallback"] += 1 # open breaker: answer at once with a cached list of popular items
continue
if sum(busy.values()) >= POOL or (name == "recommendations" and limit is not None and busy[name] >= limit):
tally["failed"] += 1 # no thread for it: rejected with a 503
continue
hangs = name == "recommendations" and after
took = TIMEOUT_MS if hangs else max(1, round(rng.expovariate(1 / ENDPOINTS[name][1])))
busy[name] += 1
seq += 1
heapq.heappush(done, (now + took, seq, name, not hangs))
tally["failed" if hangs else "ok"] += 1
return counts
print("From 10 s to 30 s, while recommendations hang: share of each endpoint's requests")
print(f"{'design':<33}{'search ok':>10}{'profile ok':>11}{'recs ok':>9}{'recs fallback':>14}{'recs failed':>12}")
for label, kwargs in [
("one shared pool of 200 threads", {}),
("bulkhead: recommendations <= 20", {"limit": 20}),
("bulkhead and circuit breaker", {"limit": 20, "breaker": True}),
]:
c = simulate(**kwargs)
def share(name, key):
return c[name][key] / sum(c[name].values())
print(
f"{label:<33}{share('search', 'ok'):>10.1%}{share('profile', 'ok'):>11.1%}{share('recommendations', 'ok'):>9.1%}"
f"{share('recommendations', 'fallback'):>14.1%}{share('recommendations', 'failed'):>12.1%}"
) Output
From 10 s to 30 s, while recommendations hang: share of each endpoint's requests design search ok profile ok recs ok recs fallback recs failed one shared pool of 200 threads 40.8% 41.2% 0.0% 0.0% 100.0% bulkhead: recommendations <= 20 100.0% 100.0% 0.0% 0.0% 100.0% bulkhead and circuit breaker 100.0% 100.0% 0.0% 73.5% 26.5%
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 bulkhead_sim.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
With one shared pool, only about 41 % of search and profile requests succeed while recommendations hang, although those endpoints have nothing wrong with them. Limiting recommendations to 20 threads brings both back to 100 %: the bulkhead turns a failure of the whole server into a failure of one feature. Adding a circuit breaker changes what the recommendation users get. Once the breaker opens, their requests stop waiting 5 seconds for an error and get a cached list of popular items at once, which is why 73.5 % of them see a fallback instead of a failure. The rest are the requests from the first seconds of the hang, before enough calls had timed out for the breaker to notice.
Bulkheads: a separate share for each dependency
The name comes from the walls that divide a ship’s hull into compartments: if the hull is breached, only one compartment floods (Azure bulkhead pattern). In software the compartments are pools. A caller gives each dependency its own connection pool, thread pool or semaphore, so that a dependency that stops answering can exhaust only its own. A service can partition the other way too, giving important consumers their own instances or queues so that one heavy consumer cannot starve the rest. Resilience4j, a fault-tolerance library for Java, offers both common forms: a semaphore bulkhead that caps concurrent calls, 25 by default, and a thread-pool bulkhead with a bounded queue of 100 by default (Resilience4j bulkhead).
Size a bulkhead from the same law that showed the failure. Recommendations normally have 100 × 0.04 = 4 calls in flight, so a limit of 20 leaves five times the normal need for bursts and slow spells, and caps the damage of a hang at 20 of the 200 threads. A limit far above the normal need protects little; one at the normal need rejects good calls during every burst. The cost is that reserved capacity sits idle while its dependency is quiet, and the Azure pattern names that less efficient use of resources as the reason not to use bulkheads where they are not needed.
Proxies enforce the same limits outside the application. Envoy’s “circuit breakers” are in fact caps on connections,
pending requests, active requests and retries for each upstream cluster, 1,024 connections by default; requests over
a cap fail at once and carry an x-envoy-overloaded header
(Envoy circuit breaking).
Its documentation states the principle behind all of these limits: failing quickly and pushing back early is nearly
always better than waiting.
Circuit breakers: stop calling what is broken
A bulkhead limits how much a broken dependency can hold; a circuit breaker stops calling it. The Azure pattern describes the breaker as a proxy with three states (Azure circuit breaker pattern):
Closed, open and half-open
Text description of the diagram
The diagram shows the three states of a circuit breaker and the four moves between them.
- Closed: calls go through to the dependency, and the breaker counts the failures. When the failures pass the threshold, the breaker moves to open.
- Open: calls fail at once, without reaching the dependency, while a timer runs. When the timer runs out, the breaker moves to half-open.
- Half-open: a few trial calls go through. If the trials succeed, the breaker moves back to closed. If a trial fails, it moves back to open and the timer starts again.
- Closed: calls go through and the breaker counts the recent failures. When they pass a threshold within a time window, it opens and starts a timer.
- Open: calls fail at once, with an error or a fallback, and never reach the dependency.
- Half-open: when the timer runs out, a limited number of trial calls go through. If they succeed, the breaker closes and resets its counts; if one fails, it opens again and the timer restarts.
The half-open state exists because a service that has just recovered may cope with a trickle of requests and not with the full load at once. Libraries make the counting precise. Resilience4j keeps a sliding window of the last 100 calls by default, waits for at least 100 calls before it judges, opens at a failure rate of 50 %, can also count slow calls as failures, stays open for 60 seconds and then allows 10 trial calls (Resilience4j circuit breaker). The example below is a smaller breaker of the same kind, with a window of 10 calls and a clock passed in, so that time can be moved by a test instead of waited for. A dependency fails from 20 to 80 seconds and is called once a second.
# A circuit breaker that watches the failure rate of the last 10 calls, with a clock you pass in, so a test can move
# time forward instead of sleeping. Below: a dependency that fails from 20 s to 80 s, called once a second.
from collections import deque
class CircuitBreaker:
def __init__(self, clock, window=10, failure_rate=0.5, open_seconds=15, trial_calls=2, on_change=None):
self.clock, self.window, self.failure_rate = clock, window, failure_rate
self.open_seconds, self.trial_calls = open_seconds, trial_calls
self.on_change = on_change or (lambda old, new: None)
self.state = "closed"
self.results = deque(maxlen=window) # True for a success, False for a failure
self.opened_at = 0.0
self.trials_sent = self.trials_passed = 0
def allow(self):
"""May this call go to the dependency? Open: no, fail fast. Half-open: only the trial calls."""
if self.state == "open" and self.clock() - self.opened_at >= self.open_seconds:
self._move("half-open")
self.trials_sent = self.trials_passed = 0
if self.state == "half-open":
if self.trials_sent < self.trial_calls:
self.trials_sent += 1
return True
return False
return self.state == "closed"
def record(self, ok):
if self.state == "half-open":
if not ok:
self._open() # still broken: back to open, and the timer starts again
else:
self.trials_passed += 1
if self.trials_passed == self.trial_calls:
self._move("closed")
self.results.clear()
return
self.results.append(ok)
failures = self.results.count(False)
if len(self.results) == self.window and failures / self.window >= self.failure_rate:
self._open()
def _open(self):
self.opened_at = self.clock()
self._move("open")
def _move(self, new):
self.on_change(self.state, new)
self.state = new
now = 0.0
def show(old, new):
print(f"{now:>5.0f} s {old:>9} -> {new:<9} (the dependency is {'up' if not 20 <= now < 80 else 'down'})")
breaker = CircuitBreaker(clock=lambda: now, on_change=show)
sent = short = 0
for second in range(120):
now = float(second)
healthy = not (20 <= second < 80)
if breaker.allow():
sent += 1
breaker.record(healthy)
else:
short += 1 # failed fast: no thread waits, the dependency sees nothing
print(f"\n120 calls: {sent} reached the dependency, {short} were answered at once by the open breaker")
# Checks, as plain assertions: the same story told by a test with a fake clock.
t = [0.0]
b = CircuitBreaker(clock=lambda: t[0], window=4, failure_rate=0.5, open_seconds=10, trial_calls=1)
for ok in (True, True, False, False):
assert b.allow()
b.record(ok)
assert b.state == "open" and not b.allow()
t[0] = 10.0
assert b.allow() and b.state == "half-open" and not b.allow() # one trial call at a time
b.record(True)
assert b.state == "closed"
print("checks passed: opens at 2 failures in 4, fails fast, lets one trial through after 10 s, closes on success") Output
24 s closed -> open (the dependency is down) 39 s open -> half-open (the dependency is down) 39 s half-open -> open (the dependency is down) 54 s open -> half-open (the dependency is down) 54 s half-open -> open (the dependency is down) 69 s open -> half-open (the dependency is down) 69 s half-open -> open (the dependency is down) 84 s open -> half-open (the dependency is up) 85 s half-open -> closed (the dependency is up) 120 calls: 64 reached the dependency, 56 were answered at once by the open breaker checks passed: opens at 2 failures in 4, fails fast, lets one trial through after 10 s, closes on success
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 breaker.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
The breaker opens at 24 seconds, when 5 of the last 10 calls have failed, and from then on answers calls itself. Every 15 seconds it lets a trial call through, finds the dependency still down and opens again. Of 120 calls, 56 were answered at once instead of waiting on a broken dependency. The timeline also shows the breaker’s cost: the dependency recovered at 80 seconds, but callers kept failing until the next trial at 84 seconds. The Amazon Builders’ Library makes this criticism of breakers, that they add modes which are hard to test and can add time to recovery, and prefers a token-bucket limit on retries (Builders’ Library).
Three rules keep breakers honest:
- One breaker per thing that fails independently. A single breaker for a store with many shards blocks the healthy shards when one fails; the Azure pattern warns about this.
- Make retries respect the breaker. A call refused by an open breaker is not a transient fault, so the retry logic around it should stop, not try again.
- Test the open path. What the caller does while the breaker is open, a cached answer, a default or an error, runs only during failures. The Builders’ Library describes a fallback that turned a partial outage into a full one: when a cache on every web server failed at about the same time, every server fell back to the database directly, and the database could not take it (Avoiding fallback). A fallback that is not exercised regularly, or that sends more load to something else, can be worse than a fast, honest error.
Shuffle sharding: few customers share a bad customer’s servers
Bulkheads and breakers protect a caller from its dependencies. Shuffle sharding protects the customers of a service from each other. Suppose 8 workers serve many customers, and one customer sends a request that crashes any worker it reaches; its retries then crash the next worker too. With plain sharding into 4 fixed shards of 2 workers, the bad customer takes down its shard, and a quarter of all customers go down with it. With shuffle sharding, every customer gets its own pair of workers drawn from the 8, and there are 28 possible pairs, so only about 1 customer in 28 has exactly the bad customer’s pair; the Builders’ Library calls the result 7 times better than plain sharding (shuffle sharding).
# Shuffle sharding: give every customer its own random k workers out of n. How likely is another customer to share
# workers with a customer whose requests crash every worker they reach?
import random
from math import comb
def shares(n, k, j):
"""Chance that two random k-of-n shards have exactly j workers in common."""
return comb(k, j) * comb(n - k, k - j) / comb(n, k)
print(f"{'workers':>7}{'each':>5} {'possible shards':>15} {'share all workers':>21} {'share at least one':>18}")
for n, k in [(8, 2), (16, 4), (64, 4), (100, 5), (2048, 4)]:
same = f"1 in {comb(n, k):,}"
print(f"{n:>7}{k:>5} {comb(n, k):>15,} {same:>21} {1 - shares(n, k, 0):>18.1%}")
# A simulation with 8 workers and 1,000 customers. Customer 0 sends a request that crashes any worker it reaches,
# and it retries, so both of its workers go down. Clients retry on their other worker, so a customer loses service
# only if all of its workers are down.
rng = random.Random(11)
WORKERS, CUSTOMERS = 8, 1_000
# Plain sharding: 4 fixed shards of 2 workers each, customers spread evenly.
shard_of = [c % 4 for c in range(CUSTOMERS)]
plain_lost = sum(1 for c in range(1, CUSTOMERS) if shard_of[c] == shard_of[0])
# Shuffle sharding: every customer draws its own 2 of the 8 workers.
shards = [frozenset(rng.sample(range(WORKERS), 2)) for _ in range(CUSTOMERS)]
down = shards[0]
shuffle_lost = sum(1 for c in range(1, CUSTOMERS) if shards[c] <= down)
shuffle_touched = sum(1 for c in range(1, CUSTOMERS) if shards[c] & down)
print(f"\n{CUSTOMERS - 1} other customers, 8 workers, 2 per customer; the bad customer's workers are down:")
print(f" plain sharding: {plain_lost} lose service ({plain_lost / (CUSTOMERS - 1):.1%})")
print(f" shuffle sharding: {shuffle_lost} lose service ({shuffle_lost / (CUSTOMERS - 1):.1%}); "
f"{shuffle_touched} lose one of their two workers and carry on on the other") Output
workers each possible shards share all workers share at least one
8 2 28 1 in 28 46.4%
16 4 1,820 1 in 1,820 72.8%
64 4 635,376 1 in 635,376 23.3%
100 5 75,287,520 1 in 75,287,520 23.0%
2048 4 730,862,190,080 1 in 730,862,190,080 0.8%
999 other customers, 8 workers, 2 per customer; the bad customer's workers are down:
plain sharding: 249 lose service (24.9%)
shuffle sharding: 36 lose service (3.6%); 449 lose one of their two workers and carry on on the other
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 shuffle_shard.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
The table counts the possible shards, and the simulation plays the 8-worker case with 1,000 customers. Plain sharding loses 249 of them; shuffle sharding loses 36, because 449 customers lost one of their two workers but kept the other. That last condition matters: shuffle sharding works only if clients retry on another worker of their shard when one fails. The odds improve quickly with size. Amazon Route 53 assigns each customer domain 4 of 2,048 virtual name servers, which allows about 730 billion shards, and it chooses them so that no two domains share more than two name servers. Note what the last column says: sharing some workers becomes common in a bigger fleet, but sharing all of them becomes vanishingly rare, and only that takes a customer down.
Choosing the tool
Each tool answers a different shape of failure, and real systems combine them.
| The failure | The tool | What it does |
|---|---|---|
| A dependency is slow or hangs | Bulkhead, with a timeout | Caps the threads it can hold, so the rest keep working |
| A dependency is down for minutes | Circuit breaker | Fails fast, stops wasting calls and lets it recover |
| Brief, random errors | Retries with backoff and a budget | Turns blips into successes without multiplying load |
| One tenant’s requests crash or flood servers | Shuffle sharding, with per-tenant limits | Confines the tenant to a few servers few others share |
| Too much traffic overall | Load shedding | Rejects the excess early, the subject of the next lesson |
Practice
Exercise · Medium · Python
The three states of a circuit breaker
Finish the class Breaker in breaker.py: a circuit breaker that counts consecutive failures. Its state is "closed", "open" or "half-open", and clock is a function that returns the current time in seconds, so that the tests can move time forward.
- Closed:
allow()returnsTrue.on_failure()adds one to a count of failures in a row, andon_success()sets it back to 0. When the count reachesthreshold, the breaker opens and remembers the time. - Open:
allow()returnsFalse, untilcooldownseconds have passed since it opened. Then the breaker becomes half-open, and that call toallow()returnsTrue: it is the trial call. - Half-open: only one trial call at a time, so
allow()returnsFalsewhile the trial is out. If the trial succeeds (on_success()), the breaker closes and the count starts again from 0. If it fails (on_failure()), the breaker opens again, and the cooldown starts again from that moment.
For example, with threshold=3 and cooldown=10, three failures in a row open the breaker; ten seconds later one call is allowed through, and its success closes the breaker.
Starter code · breaker.py
class Breaker:
"""A circuit breaker that opens after `threshold` failures in a row and tries again after `cooldown` seconds."""
def __init__(self, threshold, cooldown, clock):
self.threshold = threshold
self.cooldown = cooldown
self.clock = clock
self.state = "closed"
# Add what else the breaker needs to remember.
def allow(self):
"""May the next call go to the dependency?"""
# Replace this line with your code.
return True
def on_success(self):
"""The call that allow() let through succeeded."""
# Replace this line with your code.
pass
def on_failure(self):
"""The call that allow() let through failed."""
# Replace this line with your code.
pass The sample tests · test_breaker.py
from breaker import Breaker
class Clock:
def __init__(self):
self.now = 0.0
def __call__(self):
return self.now
def test_opens_after_threshold():
"""three failures in a row open the breaker, and an open breaker fails fast"""
clock = Clock()
b = Breaker(threshold=3, cooldown=10, clock=clock)
for _ in range(3):
assert b.allow() is True
b.on_failure()
assert b.state == "open"
assert b.allow() is False
def test_success_resets_the_count():
"""a success between failures starts the count again"""
b = Breaker(threshold=3, cooldown=10, clock=Clock())
for ok in (False, False, True, False, False):
assert b.allow() is True
b.on_success() if ok else b.on_failure()
assert b.state == "closed"
b.on_failure()
assert b.state == "open"
def test_stays_open_during_cooldown():
"""an open breaker refuses calls until the cooldown has passed"""
clock = Clock()
b = Breaker(threshold=2, cooldown=10, clock=clock)
b.on_failure()
b.on_failure()
clock.now = 9.9
assert b.allow() is False
assert b.state == "open"
def test_one_trial_after_cooldown():
"""after the cooldown one trial call goes through, and only one at a time"""
clock = Clock()
b = Breaker(threshold=2, cooldown=10, clock=clock)
b.on_failure()
b.on_failure()
clock.now = 10.0
assert b.allow() is True
assert b.state == "half-open"
assert b.allow() is False
def test_trial_success_closes():
"""a successful trial closes the breaker, with the count back at zero"""
clock = Clock()
b = Breaker(threshold=2, cooldown=10, clock=clock)
b.on_failure()
b.on_failure()
clock.now = 12.0
assert b.allow() is True
b.on_success()
assert b.state == "closed"
assert b.allow() is True
b.on_failure()
assert b.state == "closed"
def test_trial_failure_reopens():
"""a failed trial opens the breaker again, and the cooldown starts from that moment"""
clock = Clock()
b = Breaker(threshold=2, cooldown=10, clock=clock)
b.on_failure()
b.on_failure()
clock.now = 10.0
assert b.allow() is True
b.on_failure()
assert b.state == "open"
clock.now = 19.0
assert b.allow() is False
clock.now = 20.0
assert b.allow() is True
assert b.state == "half-open" A hint
Keep four things: the state, the count of failures in a row, the time the breaker opened, and whether a trial call is out. Do the "has the cooldown passed?" check at the start of allow(): that is the moment an open breaker becomes half-open. In on_failure(), a failure while half-open opens the breaker at once, whatever the count.
Results of the sample tests
| Test | Result | Details |
|---|
What your code printed
The sample tests run on this device, in your browser (Pyodide): nothing is sent to mysmartcopilot.com. The first run downloads Python (about 13.5 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.
Check yourself
5 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.
Interview questions
Warm-up (fresher to mid level): what are the three states of a circuit breaker, and why does the half-open state exist? Closed is normal: calls go through and the breaker counts recent failures. When the failures pass a threshold, the breaker opens: calls fail at once, with an error or a fallback, so callers stop waiting on a broken dependency and the dependency gets less load. After a timer runs out, the breaker turns half-open and lets a few trial calls through. If they succeed it closes; if one fails it opens again and restarts the timer. The half-open state exists because a service that has just recovered may handle a trickle of requests and not the full load at once, so the breaker tests it gently before sending everything back.
Key takeaways
- One slow dependency can hold every thread: during a hang its calls in flight are its rate times its timeout. In the simulation, 200 threads and one hanging dependency left only 41 % of unrelated requests working.
- A bulkhead gives each dependency or tenant its own pool. Size it from Little’s law with headroom, for example 4 in flight normally and a limit of 20.
- A circuit breaker is closed, open or half-open. It fails fast while a dependency is broken and tests recovery with a few trial calls, at the price of a slower return to normal.
- Use one breaker per independent failure domain, never retry a call the breaker refused, and exercise the open path and its fallback regularly.
- Shuffle sharding gives each customer a random small set of servers: with 8 workers and pairs, 1 customer in 28 shares a bad customer’s servers, against 1 in 4 with plain sharding. Clients must retry on their other servers.
References
- Circuit Breaker pattern (Azure Architecture Center) (Microsoft)
- Bulkhead pattern (Azure Architecture Center) (Microsoft)
- Workload isolation using shuffle-sharding (Amazon Builders' Library) (Amazon Web Services)
- Avoiding fallback in distributed systems (Amazon Builders' Library) (Amazon Web Services)
- Timeouts, retries, and backoff with jitter (Amazon Builders' Library) (Amazon Web Services)
- CircuitBreaker (Resilience4j documentation) (Resilience4j project)
- Bulkhead (Resilience4j documentation) (Resilience4j project)
- Circuit breaking (Envoy documentation) (Envoy Project Authors)
Related tools
Report a problem with this lesson
Kept only in this browser. Your Learn progress