Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

System Design (High-Level Design)  Module 3 – Performance and reliability fundamentals

Circuit breakers, bulkheads and shuffle sharding

How circuit breakers fail fast, how bulkheads stop one slow dependency from taking every thread, and how shuffle sharding contains one bad client.

  • Intermediate
  • 30 minutes
  • Examples run with Python 3.14.8, Pyodide 314.0.7, Node.js 24.21.0 and quickjs 0.32.0
  • By MySmartCoPilot

What you will learn

  • Explain the closed, open and half-open states of a circuit breaker and what moves a breaker between them
  • Size a bulkhead with Little's law so that one slow dependency can fill only its own share of threads
  • Calculate how shuffle sharding shrinks the share of customers that one bad client can affect
  • Choose between a breaker, a bulkhead, shuffle sharding and a retry budget for a given failure

Before you start

On this page

Three tools stop one failing part from taking everything around it down. A circuit breaker watches the calls to a dependency and, once too many of them fail, stops sending them for a while: callers fail fast instead of waiting, and the dependency gets room to recover. A bulkhead gives each dependency, or each tenant, its own limited share of threads, connections or queue slots, so a slow one can fill only its own share. Shuffle sharding gives every customer a small random set of servers out of a larger fleet, so a customer whose requests break servers reaches only a few of them, and almost no other customer depends on exactly those few. All three answer the same failure: a shared resource used up by one bad part. This lesson measures that failure first, then builds each tool with the numbers that size it.

One slow dependency can take every thread

Take a web server with 200 worker threads and three endpoints, each calling its own dependency: search, at 300 requests a second with calls of about 20 ms; profiles, at 200 a second and 30 ms; and recommendations, at 100 a second and 40 ms. By Little’s law, the calls in flight are the rate times the time each takes, so on a normal day they hold 6 + 6 + 4 = 16 threads, a twelfth of the pool. Now the recommendations dependency starts to hang, and every call to it waits for its 5-second timeout. The same law says it now wants 100 × 5 = 500 threads. There are 200, so within about two seconds every thread is waiting on recommendations, and requests for search and profiles, which never call that dependency, find no thread to run on.

The simulation below runs exactly this, with seeded random arrivals, for 30 seconds; the hang starts at 10 seconds. It compares the shared pool with two of this lesson’s tools.

One hanging dependency, three designs Python · bulkhead_sim.py
# One web server with 200 worker threads serves three endpoints, each calling its own dependency. At 10 s the
# recommendations dependency starts hanging: every call waits for the 5 s timeout and then fails. What happens to the
# two endpoints that never call it? Three designs, the same seeded traffic, 30 simulated seconds in steps of 1 ms.
import heapq
import random
from collections import deque

POOL = 200
TIMEOUT_MS = 5_000
SLOW_FROM, END = 10_000, 30_000
# endpoint: requests per millisecond, mean latency of its dependency in ms
ENDPOINTS = {"search": (0.3, 20), "profile": (0.2, 30), "recommendations": (0.1, 40)}


class Breaker:
    """Opens when half of the last 20 calls failed; after 5 s it lets 3 trial calls through."""

    def __init__(self):
        self.state, self.results, self.opened_at, self.trials = "closed", deque(maxlen=20), 0, 0

    def allow(self, now):
        if self.state == "open" and now - self.opened_at >= 5_000:
            self.state, self.trials = "half-open", 0
        if self.state == "closed":
            return True
        if self.state == "half-open" and self.trials < 3:
            self.trials += 1
            return True
        return False

    def record(self, now, ok):
        if self.state == "half-open":
            if not ok:
                self.state, self.opened_at = "open", now
            elif self.trials == 3:
                self.state = "closed"
                self.results.clear()
            return
        self.results.append(ok)
        if len(self.results) == 20 and self.results.count(False) >= 10:
            self.state, self.opened_at = "open", now


def simulate(limit=None, breaker=False, seed=5):
    """limit: the most threads that recommendations calls may hold (a bulkhead); breaker: one on that dependency."""
    rng = random.Random(seed)
    busy = {name: 0 for name in ENDPOINTS}
    done = []  # heap of (finish time, sequence number, endpoint, ok)
    counts = {name: {"ok": 0, "fallback": 0, "failed": 0} for name in ENDPOINTS}
    cb = Breaker() if breaker else None
    seq = 0
    for now in range(END):
        while done and done[0][0] <= now:
            _, _, name, ok = heapq.heappop(done)
            busy[name] -= 1
            if name == "recommendations" and cb:
                cb.record(now, ok)
        arrivals = [name for name, (rate, _) in ENDPOINTS.items() if rng.random() < rate]
        rng.shuffle(arrivals)
        for name in arrivals:
            after = now >= SLOW_FROM
            tally = counts[name] if after else {"ok": 0, "fallback": 0, "failed": 0}
            if name == "recommendations" and cb and not cb.allow(now):
                tally["fallback"] += 1  # open breaker: answer at once with a cached list of popular items
                continue
            if sum(busy.values()) >= POOL or (name == "recommendations" and limit is not None and busy[name] >= limit):
                tally["failed"] += 1  # no thread for it: rejected with a 503
                continue
            hangs = name == "recommendations" and after
            took = TIMEOUT_MS if hangs else max(1, round(rng.expovariate(1 / ENDPOINTS[name][1])))
            busy[name] += 1
            seq += 1
            heapq.heappush(done, (now + took, seq, name, not hangs))
            tally["failed" if hangs else "ok"] += 1
    return counts


print("From 10 s to 30 s, while recommendations hang: share of each endpoint's requests")
print(f"{'design':<33}{'search ok':>10}{'profile ok':>11}{'recs ok':>9}{'recs fallback':>14}{'recs failed':>12}")
for label, kwargs in [
    ("one shared pool of 200 threads", {}),
    ("bulkhead: recommendations <= 20", {"limit": 20}),
    ("bulkhead and circuit breaker", {"limit": 20, "breaker": True}),
]:
    c = simulate(**kwargs)

    def share(name, key):
        return c[name][key] / sum(c[name].values())

    print(
        f"{label:<33}{share('search', 'ok'):>10.1%}{share('profile', 'ok'):>11.1%}{share('recommendations', 'ok'):>9.1%}"
        f"{share('recommendations', 'fallback'):>14.1%}{share('recommendations', 'failed'):>12.1%}"
    )

Output

From 10 s to 30 s, while recommendations hang: share of each endpoint's requests
design                            search ok profile ok  recs ok recs fallback recs failed
one shared pool of 200 threads        40.8%      41.2%     0.0%          0.0%      100.0%
bulkhead: recommendations <= 20      100.0%     100.0%     0.0%          0.0%      100.0%
bulkhead and circuit breaker         100.0%     100.0%     0.0%         73.5%       26.5%

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 bulkhead_sim.py

With one shared pool, only about 41 % of search and profile requests succeed while recommendations hang, although those endpoints have nothing wrong with them. Limiting recommendations to 20 threads brings both back to 100 %: the bulkhead turns a failure of the whole server into a failure of one feature. Adding a circuit breaker changes what the recommendation users get. Once the breaker opens, their requests stop waiting 5 seconds for an error and get a cached list of popular items at once, which is why 73.5 % of them see a fallback instead of a failure. The rest are the requests from the first seconds of the hang, before enough calls had timed out for the breaker to notice.

Bulkheads: a separate share for each dependency

The name comes from the walls that divide a ship’s hull into compartments: if the hull is breached, only one compartment floods (Azure bulkhead pattern). In software the compartments are pools. A caller gives each dependency its own connection pool, thread pool or semaphore, so that a dependency that stops answering can exhaust only its own. A service can partition the other way too, giving important consumers their own instances or queues so that one heavy consumer cannot starve the rest. Resilience4j, a fault-tolerance library for Java, offers both common forms: a semaphore bulkhead that caps concurrent calls, 25 by default, and a thread-pool bulkhead with a bounded queue of 100 by default (Resilience4j bulkhead).

Size a bulkhead from the same law that showed the failure. Recommendations normally have 100 × 0.04 = 4 calls in flight, so a limit of 20 leaves five times the normal need for bursts and slow spells, and caps the damage of a hang at 20 of the 200 threads. A limit far above the normal need protects little; one at the normal need rejects good calls during every burst. The cost is that reserved capacity sits idle while its dependency is quiet, and the Azure pattern names that less efficient use of resources as the reason not to use bulkheads where they are not needed.

Proxies enforce the same limits outside the application. Envoy’s “circuit breakers” are in fact caps on connections, pending requests, active requests and retries for each upstream cluster, 1,024 connections by default; requests over a cap fail at once and carry an x-envoy-overloaded header (Envoy circuit breaking). Its documentation states the principle behind all of these limits: failing quickly and pushing back early is nearly always better than waiting.

Circuit breakers: stop calling what is broken

A bulkhead limits how much a broken dependency can hold; a circuit breaker stops calling it. The Azure pattern describes the breaker as a proxy with three states (Azure circuit breaker pattern):

A circuit breaker moves from closed to open when failures pass a threshold, to half-open when a timer ends, then closes or opens again.Closedcalls go through;failures are countedOpencalls fail at once;a timer runsHalf-opena few trial callsgo throughfailures passthe thresholdthe timerruns outthe trials succeeda trial fails:the timer restarts

Closed, open and half-open

Text description of the diagram

The diagram shows the three states of a circuit breaker and the four moves between them.

  1. Closed: calls go through to the dependency, and the breaker counts the failures. When the failures pass the threshold, the breaker moves to open.
  2. Open: calls fail at once, without reaching the dependency, while a timer runs. When the timer runs out, the breaker moves to half-open.
  3. Half-open: a few trial calls go through. If the trials succeed, the breaker moves back to closed. If a trial fails, it moves back to open and the timer starts again.
  • Closed: calls go through and the breaker counts the recent failures. When they pass a threshold within a time window, it opens and starts a timer.
  • Open: calls fail at once, with an error or a fallback, and never reach the dependency.
  • Half-open: when the timer runs out, a limited number of trial calls go through. If they succeed, the breaker closes and resets its counts; if one fails, it opens again and the timer restarts.

The half-open state exists because a service that has just recovered may cope with a trickle of requests and not with the full load at once. Libraries make the counting precise. Resilience4j keeps a sliding window of the last 100 calls by default, waits for at least 100 calls before it judges, opens at a failure rate of 50 %, can also count slow calls as failures, stays open for 60 seconds and then allows 10 trial calls (Resilience4j circuit breaker). The example below is a smaller breaker of the same kind, with a window of 10 calls and a clock passed in, so that time can be moved by a test instead of waited for. A dependency fails from 20 to 80 seconds and is called once a second.

A failure-rate circuit breaker with a clock you control Python · breaker.py
# A circuit breaker that watches the failure rate of the last 10 calls, with a clock you pass in, so a test can move
# time forward instead of sleeping. Below: a dependency that fails from 20 s to 80 s, called once a second.
from collections import deque


class CircuitBreaker:
    def __init__(self, clock, window=10, failure_rate=0.5, open_seconds=15, trial_calls=2, on_change=None):
        self.clock, self.window, self.failure_rate = clock, window, failure_rate
        self.open_seconds, self.trial_calls = open_seconds, trial_calls
        self.on_change = on_change or (lambda old, new: None)
        self.state = "closed"
        self.results = deque(maxlen=window)  # True for a success, False for a failure
        self.opened_at = 0.0
        self.trials_sent = self.trials_passed = 0

    def allow(self):
        """May this call go to the dependency? Open: no, fail fast. Half-open: only the trial calls."""
        if self.state == "open" and self.clock() - self.opened_at >= self.open_seconds:
            self._move("half-open")
            self.trials_sent = self.trials_passed = 0
        if self.state == "half-open":
            if self.trials_sent < self.trial_calls:
                self.trials_sent += 1
                return True
            return False
        return self.state == "closed"

    def record(self, ok):
        if self.state == "half-open":
            if not ok:
                self._open()  # still broken: back to open, and the timer starts again
            else:
                self.trials_passed += 1
                if self.trials_passed == self.trial_calls:
                    self._move("closed")
                    self.results.clear()
            return
        self.results.append(ok)
        failures = self.results.count(False)
        if len(self.results) == self.window and failures / self.window >= self.failure_rate:
            self._open()

    def _open(self):
        self.opened_at = self.clock()
        self._move("open")

    def _move(self, new):
        self.on_change(self.state, new)
        self.state = new


now = 0.0


def show(old, new):
    print(f"{now:>5.0f} s  {old:>9} -> {new:<9}  (the dependency is {'up' if not 20 <= now < 80 else 'down'})")


breaker = CircuitBreaker(clock=lambda: now, on_change=show)
sent = short = 0
for second in range(120):
    now = float(second)
    healthy = not (20 <= second < 80)
    if breaker.allow():
        sent += 1
        breaker.record(healthy)
    else:
        short += 1  # failed fast: no thread waits, the dependency sees nothing
print(f"\n120 calls: {sent} reached the dependency, {short} were answered at once by the open breaker")

# Checks, as plain assertions: the same story told by a test with a fake clock.
t = [0.0]
b = CircuitBreaker(clock=lambda: t[0], window=4, failure_rate=0.5, open_seconds=10, trial_calls=1)
for ok in (True, True, False, False):
    assert b.allow()
    b.record(ok)
assert b.state == "open" and not b.allow()
t[0] = 10.0
assert b.allow() and b.state == "half-open" and not b.allow()  # one trial call at a time
b.record(True)
assert b.state == "closed"
print("checks passed: opens at 2 failures in 4, fails fast, lets one trial through after 10 s, closes on success")

Output

   24 s     closed -> open       (the dependency is down)
   39 s       open -> half-open  (the dependency is down)
   39 s  half-open -> open       (the dependency is down)
   54 s       open -> half-open  (the dependency is down)
   54 s  half-open -> open       (the dependency is down)
   69 s       open -> half-open  (the dependency is down)
   69 s  half-open -> open       (the dependency is down)
   84 s       open -> half-open  (the dependency is up)
   85 s  half-open -> closed     (the dependency is up)

120 calls: 64 reached the dependency, 56 were answered at once by the open breaker
checks passed: opens at 2 failures in 4, fails fast, lets one trial through after 10 s, closes on success

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 breaker.py

The breaker opens at 24 seconds, when 5 of the last 10 calls have failed, and from then on answers calls itself. Every 15 seconds it lets a trial call through, finds the dependency still down and opens again. Of 120 calls, 56 were answered at once instead of waiting on a broken dependency. The timeline also shows the breaker’s cost: the dependency recovered at 80 seconds, but callers kept failing until the next trial at 84 seconds. The Amazon Builders’ Library makes this criticism of breakers, that they add modes which are hard to test and can add time to recovery, and prefers a token-bucket limit on retries (Builders’ Library).

Three rules keep breakers honest:

  1. One breaker per thing that fails independently. A single breaker for a store with many shards blocks the healthy shards when one fails; the Azure pattern warns about this.
  2. Make retries respect the breaker. A call refused by an open breaker is not a transient fault, so the retry logic around it should stop, not try again.
  3. Test the open path. What the caller does while the breaker is open, a cached answer, a default or an error, runs only during failures. The Builders’ Library describes a fallback that turned a partial outage into a full one: when a cache on every web server failed at about the same time, every server fell back to the database directly, and the database could not take it (Avoiding fallback). A fallback that is not exercised regularly, or that sends more load to something else, can be worse than a fast, honest error.

Shuffle sharding: few customers share a bad customer’s servers

Bulkheads and breakers protect a caller from its dependencies. Shuffle sharding protects the customers of a service from each other. Suppose 8 workers serve many customers, and one customer sends a request that crashes any worker it reaches; its retries then crash the next worker too. With plain sharding into 4 fixed shards of 2 workers, the bad customer takes down its shard, and a quarter of all customers go down with it. With shuffle sharding, every customer gets its own pair of workers drawn from the 8, and there are 28 possible pairs, so only about 1 customer in 28 has exactly the bad customer’s pair; the Builders’ Library calls the result 7 times better than plain sharding (shuffle sharding).

The odds of sharing a bad customer's workers Python · shuffle_shard.py
# Shuffle sharding: give every customer its own random k workers out of n. How likely is another customer to share
# workers with a customer whose requests crash every worker they reach?
import random
from math import comb


def shares(n, k, j):
    """Chance that two random k-of-n shards have exactly j workers in common."""
    return comb(k, j) * comb(n - k, k - j) / comb(n, k)


print(f"{'workers':>7}{'each':>5}  {'possible shards':>15}  {'share all workers':>21}  {'share at least one':>18}")
for n, k in [(8, 2), (16, 4), (64, 4), (100, 5), (2048, 4)]:
    same = f"1 in {comb(n, k):,}"
    print(f"{n:>7}{k:>5}  {comb(n, k):>15,}  {same:>21}  {1 - shares(n, k, 0):>18.1%}")

# A simulation with 8 workers and 1,000 customers. Customer 0 sends a request that crashes any worker it reaches,
# and it retries, so both of its workers go down. Clients retry on their other worker, so a customer loses service
# only if all of its workers are down.
rng = random.Random(11)
WORKERS, CUSTOMERS = 8, 1_000

# Plain sharding: 4 fixed shards of 2 workers each, customers spread evenly.
shard_of = [c % 4 for c in range(CUSTOMERS)]
plain_lost = sum(1 for c in range(1, CUSTOMERS) if shard_of[c] == shard_of[0])

# Shuffle sharding: every customer draws its own 2 of the 8 workers.
shards = [frozenset(rng.sample(range(WORKERS), 2)) for _ in range(CUSTOMERS)]
down = shards[0]
shuffle_lost = sum(1 for c in range(1, CUSTOMERS) if shards[c] <= down)
shuffle_touched = sum(1 for c in range(1, CUSTOMERS) if shards[c] & down)
print(f"\n{CUSTOMERS - 1} other customers, 8 workers, 2 per customer; the bad customer's workers are down:")
print(f"  plain sharding:   {plain_lost} lose service ({plain_lost / (CUSTOMERS - 1):.1%})")
print(f"  shuffle sharding: {shuffle_lost} lose service ({shuffle_lost / (CUSTOMERS - 1):.1%}); "
      f"{shuffle_touched} lose one of their two workers and carry on on the other")

Output

workers each  possible shards      share all workers  share at least one
      8    2               28                1 in 28               46.4%
     16    4            1,820             1 in 1,820               72.8%
     64    4          635,376           1 in 635,376               23.3%
    100    5       75,287,520        1 in 75,287,520               23.0%
   2048    4  730,862,190,080   1 in 730,862,190,080                0.8%

999 other customers, 8 workers, 2 per customer; the bad customer's workers are down:
  plain sharding:   249 lose service (24.9%)
  shuffle sharding: 36 lose service (3.6%); 449 lose one of their two workers and carry on on the other

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 shuffle_shard.py

The table counts the possible shards, and the simulation plays the 8-worker case with 1,000 customers. Plain sharding loses 249 of them; shuffle sharding loses 36, because 449 customers lost one of their two workers but kept the other. That last condition matters: shuffle sharding works only if clients retry on another worker of their shard when one fails. The odds improve quickly with size. Amazon Route 53 assigns each customer domain 4 of 2,048 virtual name servers, which allows about 730 billion shards, and it chooses them so that no two domains share more than two name servers. Note what the last column says: sharing some workers becomes common in a bigger fleet, but sharing all of them becomes vanishingly rare, and only that takes a customer down.

Choosing the tool

Each tool answers a different shape of failure, and real systems combine them.

The failure The tool What it does
A dependency is slow or hangs Bulkhead, with a timeout Caps the threads it can hold, so the rest keep working
A dependency is down for minutes Circuit breaker Fails fast, stops wasting calls and lets it recover
Brief, random errors Retries with backoff and a budget Turns blips into successes without multiplying load
One tenant’s requests crash or flood servers Shuffle sharding, with per-tenant limits Confines the tenant to a few servers few others share
Too much traffic overall Load shedding Rejects the excess early, the subject of the next lesson

Practice

Exercise · Medium · Python

The three states of a circuit breaker

Finish the class Breaker in breaker.py: a circuit breaker that counts consecutive failures. Its state is "closed", "open" or "half-open", and clock is a function that returns the current time in seconds, so that the tests can move time forward.

  • Closed: allow() returns True. on_failure() adds one to a count of failures in a row, and on_success() sets it back to 0. When the count reaches threshold, the breaker opens and remembers the time.
  • Open: allow() returns False, until cooldown seconds have passed since it opened. Then the breaker becomes half-open, and that call to allow() returns True: it is the trial call.
  • Half-open: only one trial call at a time, so allow() returns False while the trial is out. If the trial succeeds (on_success()), the breaker closes and the count starts again from 0. If it fails (on_failure()), the breaker opens again, and the cooldown starts again from that moment.

For example, with threshold=3 and cooldown=10, three failures in a row open the breaker; ten seconds later one call is allowed through, and its success closes the breaker.

Starter code · breaker.py

class Breaker:
    """A circuit breaker that opens after `threshold` failures in a row and tries again after `cooldown` seconds."""

    def __init__(self, threshold, cooldown, clock):
        self.threshold = threshold
        self.cooldown = cooldown
        self.clock = clock
        self.state = "closed"
        # Add what else the breaker needs to remember.

    def allow(self):
        """May the next call go to the dependency?"""
        # Replace this line with your code.
        return True

    def on_success(self):
        """The call that allow() let through succeeded."""
        # Replace this line with your code.
        pass

    def on_failure(self):
        """The call that allow() let through failed."""
        # Replace this line with your code.
        pass
The sample tests · test_breaker.py
from breaker import Breaker


class Clock:
    def __init__(self):
        self.now = 0.0

    def __call__(self):
        return self.now


def test_opens_after_threshold():
    """three failures in a row open the breaker, and an open breaker fails fast"""
    clock = Clock()
    b = Breaker(threshold=3, cooldown=10, clock=clock)
    for _ in range(3):
        assert b.allow() is True
        b.on_failure()
    assert b.state == "open"
    assert b.allow() is False


def test_success_resets_the_count():
    """a success between failures starts the count again"""
    b = Breaker(threshold=3, cooldown=10, clock=Clock())
    for ok in (False, False, True, False, False):
        assert b.allow() is True
        b.on_success() if ok else b.on_failure()
    assert b.state == "closed"
    b.on_failure()
    assert b.state == "open"


def test_stays_open_during_cooldown():
    """an open breaker refuses calls until the cooldown has passed"""
    clock = Clock()
    b = Breaker(threshold=2, cooldown=10, clock=clock)
    b.on_failure()
    b.on_failure()
    clock.now = 9.9
    assert b.allow() is False
    assert b.state == "open"


def test_one_trial_after_cooldown():
    """after the cooldown one trial call goes through, and only one at a time"""
    clock = Clock()
    b = Breaker(threshold=2, cooldown=10, clock=clock)
    b.on_failure()
    b.on_failure()
    clock.now = 10.0
    assert b.allow() is True
    assert b.state == "half-open"
    assert b.allow() is False


def test_trial_success_closes():
    """a successful trial closes the breaker, with the count back at zero"""
    clock = Clock()
    b = Breaker(threshold=2, cooldown=10, clock=clock)
    b.on_failure()
    b.on_failure()
    clock.now = 12.0
    assert b.allow() is True
    b.on_success()
    assert b.state == "closed"
    assert b.allow() is True
    b.on_failure()
    assert b.state == "closed"


def test_trial_failure_reopens():
    """a failed trial opens the breaker again, and the cooldown starts from that moment"""
    clock = Clock()
    b = Breaker(threshold=2, cooldown=10, clock=clock)
    b.on_failure()
    b.on_failure()
    clock.now = 10.0
    assert b.allow() is True
    b.on_failure()
    assert b.state == "open"
    clock.now = 19.0
    assert b.allow() is False
    clock.now = 20.0
    assert b.allow() is True
    assert b.state == "half-open"
A hint

Keep four things: the state, the count of failures in a row, the time the breaker opened, and whether a trial call is out. Do the "has the cooldown passed?" check at the start of allow(): that is the moment an open breaker becomes half-open. In on_failure(), a failure while half-open opens the breaker at once, whatever the count.

The sample tests run on this device, in your browser (Pyodide): nothing is sent to mysmartcopilot.com. The first run downloads Python (about 13.5 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.

Check yourself

5 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.

  1. Question 1 of 5 A circuit breaker is open. What happens to the next call?

    Choose one answer.

    Show the answer to question 1

    Answer: It fails at once without reaching the dependency, until the breaker's timer runs out

    Open means fail fast: the caller gets an error or a fallback in microseconds and no thread waits on the broken dependency. Only when the timer runs out does the breaker turn half-open and let trial calls through.

  2. Question 2 of 5 Calls to a dependency arrive at 100 a second and normally take 40 ms. How many are in flight on average, the starting point for sizing its bulkhead?

    Type a number.

    Show the answer to question 2

    Answer: 4 calls

    Little's law: 100 a second × 0.04 s = 4 calls in flight. A bulkhead of a few times that, say 20, absorbs bursts and slow spells, and still leaves the rest of the pool to the other dependencies when this one hangs.

  3. Question 3 of 5 With 8 workers and every customer given its own 2 of them at random, in how many possible ways can a customer's pair be chosen? (A customer has exactly the bad customer's workers with a chance of 1 in this number.)

    Type a number.

    Show the answer to question 3

    Answer: 28 pairs

    There are 8 × 7 ÷ 2 = 28 pairs. With plain sharding into 4 fixed pairs, a quarter of the customers share the bad customer's pair; with shuffle sharding only about 1 in 28 does, and the rest keep at least one healthy worker.

  4. Question 4 of 5 One tenant of a shared API sends requests that crash any server they reach. Which tool limits how many other tenants lose service?

    Choose one answer.

    Show the answer to question 4

    Answer: Shuffle sharding, so that few tenants share all of their servers with the bad one

    The problem is a bad tenant reaching many servers. Shuffle sharding confines it to its own small set, and other tenants overlap that set rarely and partly. Longer timeouts and more retries make it worse, and breakers in the clients do not stop the crashing requests.

  5. Question 5 of 5 Why does the half-open state let only a few trial calls through instead of closing at once?

    Choose one answer.

    Show the answer to question 5

    Answer: A dependency that has just recovered may fall over again under the full load at once

    A recovering service may cope with a trickle and not with the full flood. A few trial calls test it gently; success closes the breaker, and a failure sends it back to open for another wait.

Permutation & Combination Calculator Count the possible shuffle shards, n choose k, for your own fleet and shard sizes. Python Online Compiler Change the bulkhead limit or the breaker's window in the simulations above and run them again.

Interview questions

Warm-up (fresher to mid level): what are the three states of a circuit breaker, and why does the half-open state exist? Closed is normal: calls go through and the breaker counts recent failures. When the failures pass a threshold, the breaker opens: calls fail at once, with an error or a fallback, so callers stop waiting on a broken dependency and the dependency gets less load. After a timer runs out, the breaker turns half-open and lets a few trial calls through. If they succeed it closes; if one fails it opens again and restarts the timer. The half-open state exists because a service that has just recovered may handle a trickle of requests and not the full load at once, so the breaker tests it gently before sending everything back.

Key takeaways

  • One slow dependency can hold every thread: during a hang its calls in flight are its rate times its timeout. In the simulation, 200 threads and one hanging dependency left only 41 % of unrelated requests working.
  • A bulkhead gives each dependency or tenant its own pool. Size it from Little’s law with headroom, for example 4 in flight normally and a limit of 20.
  • A circuit breaker is closed, open or half-open. It fails fast while a dependency is broken and tests recovery with a few trial calls, at the price of a slower return to normal.
  • Use one breaker per independent failure domain, never retry a call the breaker refused, and exercise the open path and its fallback regularly.
  • Shuffle sharding gives each customer a random small set of servers: with 8 workers and pairs, 1 customer in 28 shares a bad customer’s servers, against 1 in 4 with plain sharding. Clients must retry on their other servers.

References

Related tools

Report a problem with this lesson

Quick answers and tool search

Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.