Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

System Design (High-Level Design)  Module 3 – Performance and reliability fundamentals

Load shedding, backpressure and degradation

Reject excess work early and by priority, bound every queue and tell senders to slow down, and degrade features so goodput holds under overload.

  • Advanced
  • 30 minutes
  • Examples run with Python 3.14.8, Pyodide 314.0.7, Node.js 24.21.0 and quickjs 0.32.0
  • By MySmartCoPilot

What you will learn

  • Explain why goodput collapses when a server accepts more work than it can finish in time
  • Design admission control that rejects excess work early, cheaply, by priority and by deadline
  • Use bounded queues and explicit signals such as 429, 503 and Retry-After to apply backpressure
  • Plan graceful degradation that drops optional features before the journeys users pay for

Before you start

On this page

When more work arrives than a service can finish, something has to give, and these three techniques decide what. Load shedding rejects the excess quickly and cheaply, so that the requests the service does accept are still answered in time. Backpressure sends the problem back to where the work comes from: queues have limits, and a full queue tells its producer to slow down instead of growing until memory runs out. Graceful degradation makes each request cheaper by dropping optional features and serving cached or partial answers, so that the important journeys keep working. All three rest on one fact that surprises many engineers: a server that accepts everything during an overload ends up serving almost nobody, because every request waits behind the others until it is too late to matter. This lesson shows that collapse with numbers, then sheds by priority and by deadline, bounds the queues, and orders features for degradation.

Accepting everything serves almost no one

Throughput counts every request a server handles. Goodput counts only the useful ones: answers without an error that arrive soon enough for the caller to use them (Builders’ Library). The difference is all that matters during an overload. Suppose a server can finish 1,000 requests a second and is offered 1,500: its queue grows by 500 requests every second. By Little’s law, a new request then waits the queue length divided by the rate at which the server works through it: after two seconds, 1,000 requests stand ahead of it, one second of waiting, and a client that gives up after one second never sees its answer. The server stays fully busy, but from then on it works only for callers who have left.

The simulation offers a server with 50 workers and calls of 50 ms on average, about 1,000 requests a second, 1,500 a second for a minute. Its clients give up after 1 second. Requests are checkout (15 %), browsing (55 %) and prefetching (30 %), and the same seeded traffic meets four policies.

150 % load, four ways to handle it Python · overload_sim.py
# A server with 50 workers and calls of 50 ms on average can finish about 1,000 requests a second. For 60 seconds it is
# offered 1,500 a second, 150 % of what it can do. Clients give up after 1 second, so an answer that comes later is
# wasted work. Four ways to handle the excess, the same seeded traffic for each.
import heapq
import random
from collections import deque

WORKERS, SERVICE_MS = 50, 50.0
RATE_PER_MS, SECONDS, DEADLINE_MS = 1.5, 60, 1_000
CLASSES = [("checkout", 0.15), ("browse", 0.55), ("prefetch", 0.30)]  # share of the requests, most important first
# Priority admission: a request may join the queue only while fewer than this many requests are waiting.
PRIORITY_LIMIT = {"checkout": 300, "browse": 150, "prefetch": 20}


def traffic(seed=3):
    rng = random.Random(seed)
    t, out = 0.0, []
    while True:
        t += rng.expovariate(RATE_PER_MS)
        if t >= SECONDS * 1_000:
            return out
        kind = rng.choices([name for name, _ in CLASSES], weights=[share for _, share in CLASSES])[0]
        out.append((t, kind, rng.expovariate(1 / SERVICE_MS)))


def simulate(policy):
    queues = {name: deque() for name, _ in CLASSES}
    busy_until = []  # heap of times when each busy worker becomes free
    free = WORKERS
    good = {name: 0 for name, _ in CLASSES}
    total = {name: 0 for name, _ in CLASSES}
    latencies = []

    def waiting():
        return sum(len(q) for q in queues.values())

    def next_request(now):
        order = [name for name, _ in CLASSES] if policy == "priority" else None
        while True:
            if order:
                q = next((queues[name] for name in order if queues[name]), None)
            else:  # one FIFO queue: take the oldest request of any class
                q = min((q for q in queues.values() if q), key=lambda q: q[0][0], default=None)
            if q is None:
                return None
            req = q.popleft()
            if policy == "accept all" or DEADLINE_MS - (now - req[0]) > SERVICE_MS:
                return req
            # Too little time left for an average call: drop it rather than work for nobody.

    def start_work(now):
        nonlocal free
        while free:
            req = next_request(now)
            if req is None:
                return
            arrived, kind, work = req
            finish = now + work
            free -= 1
            heapq.heappush(busy_until, finish)
            if finish - arrived <= DEADLINE_MS:
                good[kind] += 1
                latencies.append(finish - arrived)

    for arrived, kind, work in traffic():
        while busy_until and busy_until[0] <= arrived:
            now = heapq.heappop(busy_until)
            free += 1
            start_work(now)
        total[kind] += 1
        if policy == "queue limit 100" and waiting() >= 100:
            continue  # rejected at once with a 503: costs almost nothing
        if policy == "priority" and waiting() >= PRIORITY_LIMIT[kind]:
            continue
        queues[kind].append((arrived, kind, work))
        start_work(arrived)
    latencies.sort()
    p99 = latencies[int(len(latencies) * 0.99)] if latencies else float("nan")
    return sum(good.values()) / SECONDS, p99, {name: good[name] / total[name] for name in good}


print("Offered 1,500 requests a second for 60 s to a server that finishes about 1,000 a second")
print(f"{'policy':<26}{'goodput/s':>10}{'p99 of good':>13}{'checkout':>10}{'browse':>8}{'prefetch':>10}")
for label, policy in [
    ("accept everything", "accept all"),
    ("drop doomed work", "drop expired"),
    ("queue limit of 100", "queue limit 100"),
    ("priority admission", "priority"),
]:
    goodput, p99, ok = simulate(policy)
    print(f"{label:<26}{goodput:>10,.0f}{p99:>10,.0f} ms{ok['checkout']:>10.1%}{ok['browse']:>8.1%}{ok['prefetch']:>10.1%}")

Output

Offered 1,500 requests a second for 60 s to a server that finishes about 1,000 a second
policy                     goodput/s  p99 of good  checkout  browse  prefetch
accept everything                 53       985 ms      3.5%    3.6%      3.6%
drop doomed work                 668       999 ms     45.4%   45.0%     44.8%
queue limit of 100             1,004       329 ms     67.4%   67.7%     67.7%
priority admission             1,007       356 ms    100.0%   95.9%      0.3%

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 overload_sim.py

Accepting everything yields a goodput of 53 requests a second, about 5 % of what the server can do: after the first seconds, every answer comes too late. Dropping doomed work, requests with too little time left to finish, helps (668 a second), but the queue still grows, so every request that is served waits nearly the whole second; the p99 is 999 ms. A queue limit of 100 restores full goodput, 1,004 a second with a p99 of 329 ms, by refusing a third of the requests at once. It refuses them blindly, though: checkout loses as much as prefetching. Priority admission gives each kind of request a different share of the queue, 300 places for checkout, 150 for browsing and 20 for prefetching, and the server then serves every checkout, 95.9 % of browsing and almost no prefetching. Same server, same traffic; only the policy changed.

Shed early, cheaply and by priority

A rejected request still costs something: reading it, deciding, answering. If that cost is close to the cost of serving it, the server is overloaded by its refusals alone. Google’s SRE book warns that a backend can end up spending most of its processor time just rejecting requests (SRE book, chapter 21), and the Azure throttling pattern says to make a rejection cheaper than the work it prevents, by rejecting as early in the pipeline as possible (Azure throttling pattern). So every layer can refuse work, and the earlier one is the cheaper:

Requests pass clients, a load balancer, admission control and a bounded queue; excess is refused with 429 or 503 and Retry-After.Clientsback off, keep a retry budget,throttle themselvesLoad balancerconnection limits; rejectsexcess instead of queueing itServer: admission controlby priority and by deadlineBounded queuefull means reject or waitWorkersa fixed number at a timerequestsrequests within the limitsadmitted worknext item429 or 503 with Retry-After

Every layer can refuse work, and earlier is cheaper

Text description of the diagram

The diagram shows the path of a request from top to bottom, and what each layer does when there is too much work.

  1. Clients send requests. They back off, keep a retry budget and throttle themselves when they are refused.
  2. The load balancer passes on requests within its connection limits and rejects the excess instead of queueing it.
  3. The server's admission control decides, by priority and by the request's deadline, which requests to accept.
  4. Accepted work waits in a bounded queue; when the queue is full, a new item is rejected or its sender waits.
  5. Workers take the next item from the queue, a fixed number at a time.

A dashed arrow goes from admission control back to the clients: refused requests get a 429 or 503 response with a Retry-After header, which tells the clients to slow down.

  • Clients can refuse their own requests. In the SRE book’s adaptive throttling, each client counts its requests and the backend’s accepts over the last two minutes, and once requests exceed K times the accepts it starts failing new requests locally, with a probability that grows with the gap. The book prefers K = 2.
P(reject locally)=max⁡(0, requests−K×acceptsrequests+1)P(\text{reject locally}) = \max\left(0,\ \frac{\text{requests} - K \times \text{accepts}}{\text{requests} + 1}\right)
  • Load balancers and proxies should reject the excess rather than queue it. A queue in a load balancer hides how long a request has already waited; the Builders’ Library calls a configuration that fails fast instead of queueing a safe default.
  • Servers run admission control: a limit on work in progress, checked before any expensive work starts.

Which requests to refuse is a business decision written into code. The SRE book gives every request a criticality, from the most critical down to work that may be shed freely, and a backend under load rejects lower criticalities first. The Builders’ Library adds three priorities that are easy to miss. Health checks from the load balancer come first, because a server that misses them is taken out of service and its load lands on the others. Work that has already started comes before new work, because failing step 9 of 10 wastes the first eight. And requests whose deadline has passed, or will pass before the work could finish, are dropped rather than served late.

Queues need the same attention to time. Measure how long items have waited, not only how many there are: CoDel, an algorithm for network queues described in RFC 8289, starts dropping when the time packets spend in a queue stays above a target of 5 ms for an interval of 100 ms (RFC 8289). A request queue can apply the same idea with its own target. Fixed limits also go stale as fleets grow or shrink, so some services adapt them. Netflix’s concurrency-limits library treats a service’s limit like a TCP congestion window, starts from Little’s law (the limit is the request rate times the latency) and adjusts it as latency changes, rejecting what is over the limit with HTTP 429 or gRPC’s UNAVAILABLE (concurrency-limits). Whatever sets the limit, keep the latency of refused requests out of the latency graphs, and alarm on how many are refused; a tiny rejection time mixed into the average can make an overloaded service look fast.

Backpressure: bound every queue, and say “slow down”

A queue without a limit does not remove an overload; it stores it. The work waits, memory grows, and every item behind the backlog waits longer, until the process runs out of memory or the waiting outlasts every client. The SRE book suggests queues of at most about half the size of the thread pool for steady traffic, and works an example: with a queue ten times the number of threads and 100 ms of work per request, a request that joins a full queue takes about 1.1 seconds, almost all of it waiting (SRE book, chapter 22).

A bounded queue has to do something when it is full, and there are two choices. It can shed, dropping or refusing the new item. Or it can push back, telling the producer to wait until there is room, which slows the producer to the consumer’s pace. The example connects a producer of 1,200 items a second to a consumer of 1,000 a second through three queues. The pushing-back one works like a Node.js writable stream, whose write() returns false when its buffer is full and emits drain when there is room again (Node.js backpressure guide).

No limit, drop when full, push back JavaScript · bounded_queue.mjs
// A producer makes 1,200 items a second and a consumer can handle 1,000. Three queues between them, on a simulated
// clock in steps of 1 ms: one with no limit, one that drops what does not fit, and one that pushes back. The
// pushing-back queue works like a Node.js writable stream: offer() returns false when it is full, and the producer
// pauses until the queue says "drain".
class BoundedQueue {
  constructor(capacity, onDrain) {
    this.capacity = capacity;
    this.items = [];
    this.onDrain = onDrain;
    this.wasFull = false;
  }
  offer(item) {
    if (this.items.length >= this.capacity) {
      this.wasFull = true;
      return false;
    }
    this.items.push(item);
    return true;
  }
  take() {
    const item = this.items.shift();
    if (this.wasFull && this.items.length <= this.capacity / 2) {
      this.wasFull = false;
      this.onDrain(); // there is room again: tell the producer it may continue
    }
    return item;
  }
}

// 12345 -> "12,345" (written out, because the browser runtime has no locale formatting).
const fmt = (n) => String(n).replace(/\B(?=(\d{3})+(?!\d))/g, ",");

function run(mode) {
  let paused = false;
  const queue = new BoundedQueue(mode === "unbounded" ? Infinity : 100, () => (paused = false));
  let owed = 0;
  let made = 0;
  let dropped = 0;
  let done = 0;
  let pausedMs = 0;
  const rows = [];
  for (let now = 1; now <= 60000; now++) {
    // The producer: 1.2 items a millisecond, unless the queue asked it to wait.
    if (paused) pausedMs++;
    else owed += 1.2;
    while (owed >= 1 && !paused) {
      owed -= 1;
      made++;
      if (!queue.offer({ madeAt: now })) {
        if (mode === "push back") {
          paused = true; // keep the item and wait for "drain"
          owed += 1;
          made--;
        } else dropped++; // "drop": the item is shed
      }
    }
    // The consumer: one item a millisecond.
    if (queue.items.length) {
      queue.take();
      done++;
    }
    if (now % 15000 === 0) {
      const oldest = queue.items.length ? now - queue.items[0].madeAt : 0;
      rows.push(`${String(now / 1000).padStart(4)} s ${fmt(queue.items.length).padStart(8)} ${fmt(oldest).padStart(8)} ms ${fmt(dropped).padStart(8)} ${String(Math.round(pausedMs / 1000)).padStart(6)} s`);
    }
  }
  console.log(`\n${mode}: ${fmt(made)} items made, ${fmt(done)} handled`);
  console.log("time   waiting     oldest   dropped  paused");
  for (const r of rows) console.log(r);
}

for (const mode of ["unbounded", "drop", "push back"]) run(mode);

Output


unbounded: 71,999 items made, 60,000 handled
time   waiting     oldest   dropped  paused
  15 s    2,999    2,499 ms        0      0 s
  30 s    5,999    4,999 ms        0      0 s
  45 s    8,999    7,499 ms        0      0 s
  60 s   11,999    9,999 ms        0      0 s

drop: 71,999 items made, 60,000 handled
time   waiting     oldest   dropped  paused
  15 s       99       98 ms    2,900      0 s
  30 s       99       98 ms    5,900      0 s
  45 s       99       98 ms    8,900      0 s
  60 s       99       98 ms   11,900      0 s

push back: 60,063 items made, 60,000 handled
time   waiting     oldest   dropped  paused
  15 s       59       98 ms        0      2 s
  30 s       61       49 ms        0      5 s
  45 s       62       51 ms        0      7 s
  60 s       63       52 ms        0     10 s

Recorded with Node.js 24.21.0 on macOS 26 arm64. To run it yourself: mise exec node@24.21.0 -- node bounded_queue.mjs

Without a limit, 11,999 items wait after a minute and the oldest has waited 10 seconds. Dropping keeps the queue at 99 items and every wait under 100 ms, at the cost of 11,900 lost items. Pushing back loses nothing: the producer spends 10 of the 60 seconds paused, so it runs at exactly the consumer’s pace, and items wait about 50 ms. The Node.js guide measured what ignoring the signal costs in a real stream: about 88 MB of memory with backpressure and about 1.5 GB without it.

Push back when the producer can wait: a batch job, a pipeline stage, a service that reads from a durable log. Shed when it cannot: users, devices and sensors keep producing whatever you do, so tell them clearly. Between services, the signals are HTTP status codes. 429 Too Many Requests says a client has sent too many requests in a given time (RFC 6585); 503 Service Unavailable says the server is temporarily overloaded; and either may carry Retry-After, the number of seconds to wait or a date (RFC 9110). Asynchronous libraries carry the same idea inside programs: Reactive Streams exists so that the receiving side is never forced to buffer an unlimited amount of data (Reactive Streams).

One rule ties the layers together: pass overload signals on. The Azure throttling pattern warns that a service which hides a 429 or 503 returned by its own dependency, by retrying silently or by answering 500, keeps its callers from slowing down, and the overload travels back up the chain.

HTTP Status Codes Reference Check what 429 and 503 promise before you choose which one your service sends when it sheds load.

Degrade gracefully, by priority

Shedding refuses whole requests; degradation makes requests cheaper. The SRE book describes degraded answers that are less accurate or hold less data but are easier to compute (chapter 21), such as searching only the data held in memory or ranking results with a cheaper method (chapter 22), and the Azure pattern’s example is a video service that drops to a lower resolution. Plan the order before the overload, feature by feature, from the first to go to the last:

  1. Prefetching and analytics events: stop them, and let the app keep events and send them later. Nobody notices for a while.
  2. Recommendations: serve a cached list of popular items. They are useful, not essential.
  3. Reviews and ratings: serve a slightly stale copy, or hide them. Users can still decide.
  4. Search: return fewer results, rank them more cheaply and serve cached queries. It is a core journey, but a cheaper answer still helps.
  5. Checkout and payment: keep them in full. They are the journey users pay through.

Tell users what changed (“recommendations are unavailable right now”) rather than failing silently. And exercise the degraded paths: the SRE book warns that a code path you never use is often a path that does not work, and suggests running a small set of servers near overload regularly so that the degraded mode is tested. Finally, test the whole system past its capacity on purpose. The Builders’ Library’s goal for such a test is a flat line: goodput that rises to the server’s capacity and stays there as the offered load keeps growing, instead of falling to zero.

Practice

Exercise · Medium · Python

Admit by priority, and say when to come back

Write the admission control of an overloaded server in admission.py: three functions that decide, for each new request, whether it may join the queue.

admit(priority, waiting, limit) returns whether a request may join a queue that already holds waiting requests, when the queue may hold limit requests in all. Important work may use more of the queue than optional work:

  • a "critical" request (checkout, payment) is admitted while fewer than limit requests wait;
  • a "normal" request (browsing) only while fewer than 80 % of limit wait;
  • a "sheddable" request (prefetching, analytics) only while fewer than 50 % of limit wait.

retry_after(waiting, drain_per_second) returns the whole number of seconds a refused client should wait before it tries again: the time the queue needs to drain at drain_per_second requests a second, rounded up, and never less than 1.

decide(priority, waiting, limit, drain_per_second) puts the two together. It returns (200, None) when the request is admitted, and (503, seconds) when it is refused, with seconds from retry_after, ready for a Retry-After header.

For example, with a limit of 100, a normal request is admitted while 79 requests wait but not when 80 do, and retry_after(250, 100) is 3.

Starter code · admission.py

import math

# The share of the queue each priority may fill.
SHARE = {"critical": 1.0, "normal": 0.8, "sheddable": 0.5}


def admit(priority, waiting, limit):
    """Whether a request of this priority may join a queue that already holds `waiting` of at most `limit`."""
    # Replace this line with your code.
    return True


def retry_after(waiting, drain_per_second):
    """Seconds a refused client should wait: the time the queue needs to drain, rounded up, at least 1."""
    # Replace this line with your code.
    return 0


def decide(priority, waiting, limit, drain_per_second):
    """(200, None) for an admitted request, (503, seconds to wait) for a refused one."""
    # Replace this line with your code.
    return (200, None)
The sample tests · test_admission.py
from admission import admit, decide, retry_after


def test_critical_uses_the_whole_queue():
    """a critical request is admitted until the queue is full"""
    assert admit("critical", 0, 100) is True
    assert admit("critical", 99, 100) is True
    assert admit("critical", 100, 100) is False


def test_normal_keeps_room_for_critical():
    """a normal request may use 80 % of the queue"""
    assert admit("normal", 79, 100) is True
    assert admit("normal", 80, 100) is False
    assert admit("normal", 159, 200) is True
    assert admit("normal", 160, 200) is False


def test_sheddable_goes_first():
    """a sheddable request may use only half of the queue"""
    assert admit("sheddable", 49, 100) is True
    assert admit("sheddable", 50, 100) is False
    assert admit("sheddable", 0, 2) is True
    assert admit("sheddable", 1, 2) is False


def test_retry_after_rounds_up():
    """the wait is the time the queue needs to drain, rounded up"""
    assert retry_after(250, 100) == 3
    assert retry_after(200, 100) == 2
    assert retry_after(201, 100) == 3


def test_retry_after_is_at_least_one_second():
    """even a nearly empty queue asks for at least a second"""
    assert retry_after(0, 100) == 1
    assert retry_after(30, 100) == 1


def test_decide():
    """admitted requests get 200; refused ones get 503 and the seconds for Retry-After"""
    assert decide("critical", 99, 100, 50) == (200, None)
    assert decide("sheddable", 60, 100, 50) == (503, 2)
    assert decide("normal", 90, 100, 20) == (503, 5)
A hint

Compare waiting with a share of the limit: limit, 0.8 * limit or 0.5 * limit, with < so that the boundary itself is refused. A dictionary from priority to share keeps admit short. For retry_after, math.ceil rounds up and max(1, …) keeps the answer at least 1.

The sample tests run on this device, in your browser (Pyodide): nothing is sent to mysmartcopilot.com. The first run downloads Python (about 13.5 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.

Check yourself

5 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.

  1. Question 1 of 5 A server finishes 1,000 requests a second and is offered 1,500 a second, and it queues everything. How many seconds of waiting does a new request face after 10 seconds?

    Type a number.

    Show the answer to question 1

    Answer: 5 seconds (anything from 4.8 to 5.2 counts)

    The queue grows by 500 requests a second, so after 10 seconds 5,000 are waiting. At 1,000 a second the server needs 5 seconds to reach a new request, far past a typical 1-second client timeout, which is why goodput collapses when nothing is shed.

  2. Question 2 of 5 What does goodput count?

    Choose one answer.

    Show the answer to question 2

    Answer: Requests answered without an error and soon enough for the caller to use the answer

    Throughput counts everything the server does. Goodput counts only the useful part: answers that are correct and arrive before the caller gives up. An overloaded server can be busy all the time and still have almost no goodput.

  3. Question 3 of 5 An online shop is overloaded. Put these kinds of work in the order they should be shed, the first to go at the top.

    Give each item its position, from 1 (first).

    Show the answer to question 3

    Answer:

    1. Prefetching the next page and sending analytics events
    2. Personalised recommendations
    3. Search results
    4. Checkout and payment

    Shed what users notice least and what can be retried or skipped: prefetches and analytics first, then optional features such as recommendations (a cached popular list will do), then degrade search, and keep checkout, the journey that earns money, for last.

  4. Question 4 of 5 Which of these tell a sender to slow down instead of letting work pile up?

    Choose every answer that is right.

    Show the answer to question 4

    Answer:

    • A 503 response with a Retry-After header
    • A Node.js writable stream whose write() returns false until it emits drain
    • A 429 Too Many Requests response

    The first three are explicit signals: the server, or the stream, says it has no room and when to come back. A queue without a limit says nothing until memory runs out, and instant retries add load instead of removing it.

  5. Question 5 of 5 Why must rejecting a request cost far less than serving it?

    Choose one answer.

    Show the answer to question 5

    Answer: Otherwise the server spends its capacity on rejections and still serves nobody

    During an overload most requests may be refused. If each refusal costs nearly as much as an answer, the server is overloaded by the refusals alone, which is why rejection belongs as early, and as cheap, as possible.

JavaScript & TypeScript Online Compiler Change the producer's rate or the queue's capacity in the backpressure example and run it again. Python Online Compiler Try other queue limits and priority shares in the overload simulation.

Interview questions

Warm-up (fresher to mid level): what is the difference between load shedding and rate limiting? Rate limiting enforces a quota per client, such as 100 requests a minute, whatever the server’s state; it is about fairness and abuse, and it answers with 429 Too Many Requests. Load shedding reacts to the server’s own state: when it is near its capacity it refuses work from anyone, ideally the least important work first, and answers with 503 and a Retry-After. They complement each other. A rate limit stops one client from taking more than its share; load shedding protects the server when the fair shares of all clients together are still more than it can finish.

Key takeaways

  • Goodput is the work that is correct and on time. Offered 150 % of its capacity, a server that queued everything kept only about 5 % of its possible goodput in the simulation.
  • Reject the excess early and cheaply: clients can throttle themselves, load balancers should refuse rather than queue, and servers admit work by priority and by deadline.
  • Give the most important work the largest share of capacity: with priority admission every checkout succeeded at 150 % load while prefetching was shed.
  • Bound every queue. When one is full, push back if the producer can wait and shed if it cannot, and signal clearly with 429 or 503 and Retry-After.
  • Pass overload signals on instead of hiding them behind retries or a 500.
  • Plan degradation feature by feature, tell users, and exercise the degraded paths and the overload itself in tests.

References

Related tools

Report a problem with this lesson

Quick answers and tool search

Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.