Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

System Design (High-Level Design)  Module 3 – Performance and reliability fundamentals

SLIs, SLOs, SLAs and error budgets

Define service level indicators that match what users feel, set objectives with an error budget, alert on burn rate, and use the budget to plan the work.

  • Intermediate
  • 25 minutes
  • Examples run with Python 3.14.8, Pyodide 314.0.7, Node.js 24.21.0 and quickjs 0.32.0
  • By MySmartCoPilot

What you will learn

  • Define service level indicators that reflect user experience, as good events over valid events
  • Set objectives and compute the error budget and the burn rate they allow
  • Use the error budget to decide between features and reliability work

Before you start

On this page

Reliability needs a number that everyone agrees on. A service level indicator (SLI) is a measurement of what users get, best written as a ratio: good events divided by valid events, such as successful checkouts out of all checkout attempts. A service level objective (SLO) is a target for that ratio over a window: 99.9% of checkouts succeed over 30 days. A service level agreement (SLA) is a contract that attaches a consequence, usually money, to missing an objective. The gap between the SLO and 100% is the error budget: a 99.9% objective allows 0.1% of requests to fail, 43.2 minutes of full outage a month. Teams spend the budget on releases and experiments, and when it runs out, reliability work comes first. This lesson shows how to write each of the four, measure them and act on them.

SLI, SLO and SLA

The three are easy to mix up, because all three are about the same numbers. Google’s SRE book separates them by what they are: the indicator is a carefully defined measurement, the objective is the value or range you aim for, and the agreement is a contract with your users that says what happens when the objectives are met or missed. Its test for telling an SLA from an SLO is to ask what happens if the target is missed: if nothing explicit does, it is an SLO.

A public example shows the shape of an SLA. Amazon’s compute SLA commits to a monthly uptime of at least 99.99% for instances spread across two or more Availability Zones of a region, and 99.5% for a single instance. Below the target, customers get service credits: 10% of the bill below 99.99%, 30% below 99%, and 100% below 95%. Monthly uptime is computed from the minutes the service was unavailable, and the credits are the agreement’s remedy.

Inside a company the objective should be stricter than any agreement built on it. The SRE book advises keeping a safety margin, an internal SLO tighter than the one users are promised, so a team can react to a chronic problem before customers feel it. It also warns against overachieving: a service that is far more reliable than its SLO teaches other teams to depend on that, which is why Google’s lock service, Chubby, was taken down on purpose when it had been too available for a quarter, to flush out dependencies that assumed it never failed.

Write SLIs as good events over valid events

The SRE workbook recommends treating every SLI as the ratio of good events to all events, so that it runs from 0%, nothing works, to 100%, nothing is broken. A ratio makes SLIs comparable, makes the error budget a count of bad events, and forces two decisions that matter: which events count at all, and which of them are good. The workbook groups SLIs by the kind of system:

  • Request-driven services: availability (the share of requests that succeed), latency (the share faster than a threshold) and quality (the share served without degradation, for example with every recommendation present).
  • Pipelines: freshness (the share of data updated recently enough), correctness (the share of outputs that are right) and coverage (the share of the data processed).
  • Storage: durability (the share of written records that can be read back).

Latency becomes a ratio by fixing a threshold: “99% of checkouts answer within 800 ms” counts requests over 800 ms as bad, which is easier to add up across servers than a percentile (the percentiles lesson shows why percentiles do not add up). With latency histograms, put a bucket boundary exactly at the threshold: Prometheus’s documentation shows that the share of requests within it can then be counted exactly instead of estimated. Pick a few journeys that users care about, such as signing in, paying and viewing an order, rather than every metric the monitoring system collects; the workbook suggests five SLI types or fewer per service. The definitions below are the ones the next program reads:

Error budgets and burn-rate alerts for two checkout SLIs

slo/slo_budget.py

"""Error budgets and burn-rate alerts for the SLIs in slis.toml, over 30 days of per-minute request counts.

The counts are generated, not measured: a daily traffic shape, a small background error rate, a 40-minute
incident on day 12 and a slow dependency for 12 hours on day 20. Bad events are spread over minutes deterministically, so the
program prints the same everywhere. To keep the sketch short, both SLIs count the same valid requests.
"""

import tomllib

with open("slis.toml", "rb") as f:
    config = tomllib.load(f)
MINUTES = config["window_days"] * 24 * 60


def traffic(minute):
    """Valid checkout requests in one minute: 120 at night, rising to 480 in the evening."""
    hour = minute // 60 % 24
    return 120 + 20 * hour if hour < 19 else 480 - 60 * (hour - 18)


def bad_share(name, minute):
    day, hour, in_day = minute // 1440, minute // 60 % 24, minute % 1440
    if name == "checkout availability":
        incident = day == 11 and 14 * 60 <= in_day < 14 * 60 + 40  # day 12, 14:00-14:40: 30% of requests fail
        return 0.30 if incident else 0.0002
    slow_day = day == 19 and 8 <= hour < 20  # day 20, 08:00-20:00: a slow dependency
    return 0.07 if slow_day else 0.004


def counts(name):
    """(valid, bad) per minute; fractions of a bad event carry over to the next minute."""
    out, carry = [], 0.0
    for m in range(MINUTES):
        valid = traffic(m)
        carry += valid * bad_share(name, m)
        bad = int(carry)
        carry -= bad
        out.append((valid, bad))
    return out


def first_alert(series, objective, rate, long_min, short_min):
    """First minute at which both windows burn faster than `rate`, as 'day D hh:mm'."""
    budget = 1 - objective / 100
    valid_sum, bad_sum = [0], [0]
    for valid, bad in series:
        valid_sum.append(valid_sum[-1] + valid)
        bad_sum.append(bad_sum[-1] + bad)

    def burn(end, length):
        start = max(0, end - length)
        return (bad_sum[end] - bad_sum[start]) / (valid_sum[end] - valid_sum[start]) / budget

    for end in range(long_min, MINUTES + 1):
        if burn(end, long_min) > rate and burn(end, short_min) > rate:
            m = end - 1
            return f"day {m // 1440 + 1} {m // 60 % 24:02}:{m % 60:02}"
    return "never"


ALERTS = [("page", 14.4, 60, 5), ("page", 6, 360, 30), ("ticket", 1, 3 * 1440, 360)]


def span(minutes):
    return f"{minutes // 1440} days" if minutes >= 1440 else f"{minutes // 60} h" if minutes >= 60 else f"{minutes} min"

for sli in config["sli"]:
    series = counts(sli["name"])
    valid = sum(v for v, _ in series)
    bad = sum(b for _, b in series)
    allowed = valid * (1 - sli["objective"] / 100)
    print(f"{sli['name']} (objective {sli['objective']}% over {config['window_days']} days)")
    print(f"  good: {sli['good']}")
    print(f"  SLI {100 * (1 - bad / valid):.3f}% of {valid:,} valid events")
    print(f"  budget {allowed:,.0f} bad events; spent {bad:,} ({bad / allowed:.0%}); left {1 - bad / allowed:.0%}")
    for kind, rate, long_min, short_min in ALERTS:
        rule = f"{rate:g}x over {span(long_min)} and {span(short_min)}"
        print(f"  {kind:6} {rule:25} first fires: {first_alert(series, sli['objective'], rate, long_min, short_min)}")

Output

checkout availability (objective 99.9% over 30 days)
  good: valid requests answered with a success, or with a card decline that the bank sent
  SLI 99.943% of 12,960,000 valid events
  budget 12,960 bad events; spent 7,388 (57%); left 43%
  page   14.4x over 1 h and 5 min  first fires: day 12 14:02
  page   6x over 6 h and 30 min    first fires: day 12 14:05
  ticket 1x over 3 days and 6 h    first fires: day 12 14:08
checkout latency (objective 99.0% over 30 days)
  good: answered within 800 ms, measured at the load balancer
  SLI 99.459% of 12,960,000 valid events
  budget 129,600 bad events; spent 70,056 (54%); left 46%
  page   14.4x over 1 h and 5 min  first fires: never
  page   6x over 6 h and 30 min    first fires: day 20 12:54
  ticket 1x over 3 days and 6 h    first fires: day 20 13:57

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 slo_budget.py

slo/slis.toml

# Service level indicators for a payments API: each one counts good events among valid events.
window_days = 30

[[sli]]
name = "checkout availability"
valid = "POST /checkout requests that reached the load balancer, except those rejected as malformed (4xx)"
good = "valid requests answered with a success, or with a card decline that the bank sent"
objective = 99.9

[[sli]]
name = "checkout latency"
valid = "POST /checkout requests answered with a success"
good = "answered within 800 ms, measured at the load balancer"
objective = 99.0

Each definition says precisely what is valid and what is good. A malformed request that the API rightly rejects is not valid, so a buggy client cannot spend the budget; a card that the bank declines is a correct answer, so it counts as good. Deciding these edge cases before an incident is most of the work of defining an SLI.

Where to measure

The workbook lists four sources: application server logs, load balancer monitoring, black-box probes that send requests from outside, and instrumentation in the client. Each sees a different set of failures:

Four places to measure an SLI, from the user inward: the client, a prober, the load balancer and the app servers.User's appor browser:sees whatthe user seesProber:test requestsfrom outsideLoad balancer:counts everyrequest thatarrivesApp servers:miss requeststhat neverarriveDatabase

Where to measure an SLI, from the user inward

Text description of the diagram

The diagram follows requests from top to bottom, and labels each place where a service level indicator can be measured.

  1. The user's app or browser sends the real requests. Measuring here sees everything the user sees, including slow networks and failing pages, but needs instrumentation in the client.
  2. A prober sends test requests from outside. It notices when nothing reaches the service at all, but its requests may not resemble real users'.
  3. Both reach the load balancer, which counts every request that arrives. It is often the best place to start.
  4. The app servers come next. Their logs are the cheapest source, but they miss requests that never reach them, for example during a crash.
  5. The database is at the bottom.

The closer a measurement is to the user, the more of the failures users notice it can see.

Server logs are the cheapest place to start, but they miss every request that never reached a server, which is exactly what happens when the servers are down. The load balancer sees every request that arrives and is closer to the user, which is why the workbook’s worked example uses it. Probes catch failures before requests reach your network but may miss problems that affect only some users. Client instrumentation sees what the user sees, including slow pages and networks, at the cost of code in every client. The SRE book makes the same trade: client-side latency is often the more user-relevant measure, while server-side measurement is easier to get.

The error budget, in requests and in minutes

The error budget is 100% minus the objective. Over a window it can be counted in bad events or, for a full outage, in minutes: a 99.9% objective over 3,000,000 requests allows 3,000 failures, and over 30 days allows 43.2 minutes of total outage. The workbook suggests a rolling four-week window as a general-purpose choice, because a window of whole weeks always contains the same number of weekends.

The program above generates 30 days of per-minute traffic for a checkout API, with a small background error rate, a 40-minute incident on day 12 in which 30% of checkouts fail, and a dependency that makes 7% of checkouts slow for 12 hours on day 20. Read the first lines of each block:

  • Availability ends at 99.943% of 12,960,000 checkouts. The budget is 12,960 failures; the incident and the background errors spent 7,388 of them, 57%, so 43% is left for the rest of the window.
  • Latency ends at 99.459% against an objective of 99%. Its budget of 129,600 slow responses is 54% spent, most of it by the slow day.

Burn rate: how fast the budget goes

A budget that is spent evenly lasts exactly the window. The burn rate compares the current pace with that: a burn rate of 1 uses the whole budget in 30 days, 2 in 15 days, and a total outage of a 99.9% service, 1,000 times the allowed failure rate, in 43 minutes. The share of the budget that a period uses is the burn rate times the period’s length divided by the window’s:

Burn rates, how long a 30-day budget lasts, and how much each window uses Python · burn_rates.py
"""Burn rate: how fast an error budget is spent, relative to spending it evenly over the window.

A burn rate of 1 spends exactly the whole budget in the window; 2 spends it in half the time. The share of the
budget that an alert window consumes is burn rate x alert window / SLO window. Window: 30 days.
"""

WINDOW_H = 30 * 24
print(f"{'burn rate':>9}{'error rate at 99.9%':>21}{'budget gone in':>16}{'1 h uses':>10}{'6 h uses':>10}{'3 days use':>12}")
for rate in (1, 2, 6, 10, 14.4, 36, 1000):
    hours = WINDOW_H / rate
    gone = f"{hours / 24:.1f} days" if hours >= 48 else f"{hours:.1f} h" if hours >= 2 else f"{hours * 60:.0f} min"
    shares = [min(1, rate * h / WINDOW_H) for h in (1, 6, 72)]
    print(f"{rate:>9g}{rate * 0.1:>20.2f}%{gone:>16}" + "".join(f"{s:>10.1%}" for s in shares[:2]) + f"{shares[2]:>12.1%}")

Output

burn rate  error rate at 99.9%  budget gone in  1 h uses  6 h uses  3 days use
        1                0.10%       30.0 days      0.1%      0.8%       10.0%
        2                0.20%       15.0 days      0.3%      1.7%       20.0%
        6                0.60%        5.0 days      0.8%      5.0%       60.0%
       10                1.00%        3.0 days      1.4%      8.3%      100.0%
     14.4                1.44%        2.1 days      2.0%     12.0%      100.0%
       36                3.60%          20.0 h      5.0%     30.0%      100.0%
     1000              100.00%          43 min    100.0%    100.0%      100.0%

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 burn_rates.py

The rows reproduce the starting points that the SRE workbook recommends for alerting on a 99.9% objective. Page when 2% of the budget goes in one hour, a burn rate of 14.4; page when 5% goes in six hours, a burn rate of 6; open a ticket when 10% goes in three days, a burn rate of 1. Each rule also checks a short window, a twelfth of the long one (5 minutes, 30 minutes and 6 hours), so that the alert stops soon after the problem does instead of firing for the rest of the hour.

The last lines of each block in the first program apply those three rules, minute by minute, to the generated traffic. The availability incident burns about 300 times faster than allowed, and the fast page fires at 14:02 on day 12, two minutes in. The slow day on the latency SLI burns about 7 times faster than allowed: it never triggers the one-hour page, the six-hour page fires at 12:54 on day 20, almost five hours in, and the three-day ticket at 13:57. That is the design: fast, severe burns wake someone at once, and slow ones that would still empty the budget are caught by the longer windows.

Spending the budget

The error budget turns “how reliable should we be?” from an argument into a rule. In the SRE book’s version, product and reliability engineers agree on the SLO; while budget remains, new releases go out, and when it is spent, releases pause and the effort moves to testing and hardening. Because both sides share the budget, developers who want to keep releasing have a reason to release carefully. The same chapter warns that 100% is almost never the right target: users rarely notice the difference between 100% and a little less, and every extra nine costs more than the one before.

The workbook’s example policy makes the rule concrete. If the service has exceeded its budget over the past four weeks, all changes and releases stop except urgent fixes and security fixes, until it is back within its SLO. A single incident that uses more than 20% of the budget over four weeks needs a postmortem with at least one top-priority action item. Disagreements about the numbers go to a senior decision-maker named in the policy, and the policy says explicitly that it is not a punishment.

HTTP Status Codes Reference Check which status codes an availability SLI should count as good, bad or not valid.

Practice

Exercise · Easy · Python

An error budget, what is left of it, and how fast it burns

Write three helpers in budget.py for an SLO given in percent, such as 99.9.

allowed_bad(slo_pct, valid) returns how many bad events the error budget allows among valid events. For example, a 99.9% objective over 2,000,000 valid requests allows 2,000 bad ones.

budget_left(slo_pct, valid, bad) returns the share of the budget that is still unspent after bad bad events: 1.0 when nothing is spent, 0.0 when exactly all of it is, and a negative number when the objective is already missed. For example, budget_left(99.9, 2_000_000, 500) is 0.75.

burn_rate(slo_pct, bad, valid) returns how fast a period spends the budget compared with spending it evenly: the period's share of bad events divided by the share the objective allows. For example, burn_rate(99.9, 50, 10_000) is 5.0, because 0.5% of events failed where 0.1% are allowed.

Starter code · budget.py

def allowed_bad(slo_pct, valid):
    """Bad events the error budget allows among `valid` events."""
    # Replace this line with your code.
    return 0


def budget_left(slo_pct, valid, bad):
    """Share of the error budget still unspent (negative once the objective is missed)."""
    # Replace this line with your code.
    return 0


def burn_rate(slo_pct, bad, valid):
    """How fast these events spend the budget, compared with spending it evenly."""
    # Replace this line with your code.
    return 0
The sample tests · test_budget.py
import math

from budget import allowed_bad, budget_left, burn_rate


def close(a, b):
    return math.isclose(a, b, rel_tol=1e-9, abs_tol=1e-9)


def test_allowed_bad():
    """is the valid events times the share that may fail"""
    assert close(allowed_bad(99.9, 2_000_000), 2_000)
    assert close(allowed_bad(99.0, 12_960_000), 129_600)
    assert close(allowed_bad(100, 5_000), 0)


def test_budget_left():
    """is one minus the share of the budget spent, and goes below zero when the objective is missed"""
    assert close(budget_left(99.9, 2_000_000, 500), 0.75)
    assert close(budget_left(99.9, 2_000_000, 0), 1.0)
    assert close(budget_left(99.9, 2_000_000, 2_000), 0.0)
    assert close(budget_left(99.9, 2_000_000, 3_000), -0.5)


def test_burn_rate():
    """divides the observed failure share by the share the objective allows"""
    assert close(burn_rate(99.9, 50, 10_000), 5.0)
    assert close(burn_rate(99.9, 10, 10_000), 1.0)
    assert close(burn_rate(99.0, 300, 10_000), 3.0)
    assert close(burn_rate(99.9, 0, 10_000), 0.0)
A hint

The budget is the share of events that may fail, 1 - slo_pct / 100. Multiply it by the valid events for the count, compare the bad events with that count for what is left, and divide the observed failure share (bad / valid) by it for the burn rate.

The sample tests run on this device, in your browser (Pyodide): nothing is sent to mysmartcopilot.com. The first run downloads Python (about 13.5 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.

Check yourself

6 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.

  1. Question 1 of 6 A provider publishes: "monthly uptime of at least 99.99%; below that, a credit of 10% of the monthly bill". What is that?

    Choose one answer.

    Show the answer to question 1

    Answer: An SLA, because it states a consequence for missing the objective

    A target becomes an SLA when a consequence comes with it, here a service credit. Without one it is an SLO. Providers usually keep their internal SLO tighter than the SLA, so that they notice and fix problems before the contract is broken.

  2. Question 2 of 6 Which is the best availability SLI for a checkout API?

    Choose one answer.

    Show the answer to question 2

    Answer: Successful checkout requests divided by valid checkout requests, counted at the load balancer

    An SLI should count what users experience: good events among valid ones, measured as close to the user as practical. CPU, health checks and server counts describe the machines and can look fine while checkouts fail.

  3. Question 3 of 6 A service has a 99.9% availability objective over four weeks and receives 3,000,000 valid requests in that time. How many requests may fail before the objective is missed?

    Type a number.

    Show the answer to question 3

    Answer: 3000 requests

    The error budget is 100% − 99.9% = 0.1% of valid events: 0.001 × 3,000,000 = 3,000 failed requests. An incident that fails 1,500 of them spends half the budget.

  4. Question 4 of 6 The objective is 99.9%. In the last hour, 0.5% of requests failed. What is the burn rate?

    Type a number.

    Show the answer to question 4

    Answer: 5

    The burn rate is the observed failure share divided by the allowed one: 0.5% ÷ 0.1% = 5. Kept up, that spends a 30-day budget in 6 days.

  5. Question 5 of 6 Why does a burn-rate page check a long window (one hour) and a short one (five minutes) together?

    Choose one answer.

    Show the answer to question 5

    Answer: The long window shows that enough budget is being spent to matter; the short one shows that it is still being spent, so the page stops soon after a fix

    A page at 14.4 times the allowed rate over one hour means 2% of a 30-day budget is gone. Requiring the last five minutes to burn that fast too means the alert fires only while the problem lasts, and clears minutes, not an hour, after it ends.

  6. Question 6 of 6 Halfway through the window the error budget is used up, because one bad release took checkout down. Under a typical error-budget policy, what happens next?

    Choose one answer.

    Show the answer to question 6

    Answer: Releases stop except urgent fixes until the service is back within its objective, and the release's postmortem gets an action item

    The budget is what makes the trade between features and reliability explicit: while it lasts, releases go ahead; once it is spent, reliability work comes first. An example policy in Google's SRE workbook freezes changes other than urgent and security fixes, and says the policy is not a punishment.

Interview questions

Warm-up (fresher to mid level): what is the difference between an SLI, an SLO and an SLA? An SLI is a measurement of the service as users experience it, usually a ratio such as successful requests out of valid requests, or requests answered within 300 ms. An SLO is the target for that SLI over a window, such as 99.9% over 30 days, and its complement is the error budget the team may spend. An SLA is a contract with customers that attaches a consequence, typically a credit or a refund, to missing an objective. The internal SLO is set tighter than the SLA, so problems are found and fixed before the contract is broken.

Key takeaways

  • An SLI is good events ÷ valid events; an SLO is a target for it over a window; an SLA adds a consequence for missing it, and should be looser than the internal SLO.
  • The error budget is 1 − SLO: 99.9% over 3,000,000 requests allows 3,000 failures, or 43.2 minutes of outage in 30 days.
  • Measure as close to the user as you can: server logs miss the requests that never arrive.
  • Alert on burn rate with a long and a short window: 14.4× over 1 hour and 6× over 6 hours page, 1× over 3 days opens a ticket.
  • While budget remains, ship; when it is spent, reliability work comes first, by a policy agreed in advance.

References

Related tools

Report a problem with this lesson

Quick answers and tool search

Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.