Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

System Design (High-Level Design)  Module 18 – Delivery, multi-region, disaster recovery and cost

Safe deployments: canaries, rollbacks and flags

Compare rolling, blue-green and canary deployments, judge a canary from its error counts, keep rollbacks safe and release features with flags.

  • Intermediate
  • 35 minutes
  • Examples run with Python 3.14.8, Pyodide 314.0.7, Node.js 24.21.0 and quickjs 0.32.0
  • By MySmartCoPilot

What you will learn

  • Compare rolling, blue-green and canary deployments by exposure, extra capacity and rollback speed
  • Decide from error counts whether a canary should be promoted, kept waiting or rolled back
  • Design a change to stored data that can be rolled back at every step
  • Use feature flags to separate deploying code from releasing a feature
  • Implement a deterministic percentage rollout that gives every user a stable answer
  • Track delivery with the five DORA metrics

Before you start

On this page

A safe deployment limits how many users a bad change can reach and how quickly it is undone. Three tools do most of the work. A canary sends a small share of traffic to the new version and compares it with the old one before going further. A rollback returns to the last version that worked, and it is safe only when the old and new versions can read each other’s data. A feature flag lets you deploy code switched off and release it later, to a percentage of users, with a switch that turns it off again in seconds. Below, each gets numbers: the requests a bad build fails under each strategy, how long a canary must run, and why some rollbacks make an incident worse.

Why a deployment needs a plan

Changes are where most outages start. Google’s SRE book reports that roughly 70% of outages come from changes to a live system, and names three habits that limit the damage: roll out progressively, detect problems quickly and accurately, and roll back safely (SRE book, introduction). The damage is roughly the share of traffic a bad build serves times the time until it is gone. The program below works it out for a build that fails 20% of its requests, deployed in five ways to a service taking 2,000 requests a second. Each strategy gets the same 10 minutes before someone notices, so only the exposure and the way back differ.

Failed requests for one bad build, five ways Python · blast_radius.py
"""How many requests a bad build fails under five ways of deploying it.

Assumptions (change them and run again): the service takes 2,000 requests a second, the bad build fails
20% of the requests it serves, and the problem is noticed 10 minutes after the first users reach the new
version, whatever the strategy. Undoing it then takes as long as each strategy needs.
"""

RPS = 2_000
BAD_BUILD_FAILS = 0.20
NOTICED_AFTER_S = 10 * 60
SLO = 0.999                      # a 99.9% monthly availability objective
MONTH_S = 30 * 24 * 3600


def failed_requests(steps):
    """steps: (seconds, share of traffic on the new version) in order; returns failed requests."""
    return sum(RPS * BAD_BUILD_FAILS * share * seconds for seconds, share in steps)


def until_noticed(schedule):
    """Cut a rollout schedule of (seconds, share) at the moment the problem is noticed."""
    steps, elapsed = [], 0
    for seconds, share in schedule:
        take = min(seconds, NOTICED_AFTER_S - elapsed)
        if take <= 0:
            break
        steps.append((take, share))
        elapsed += take
    if elapsed < NOTICED_AFTER_S:          # the rollout finished before anyone noticed
        steps.append((NOTICED_AFTER_S - elapsed, schedule[-1][1]))
    return steps


rolling = [(150, 0.25), (150, 0.50), (150, 0.75), (10**9, 1.0)]   # a batch of 25% every 2.5 minutes
strategies = {
    "all at once, redeploy the old version (5 min)": until_noticed([(10**9, 1.0)]) + [(300, 1.0)],
    "rolling, rolled back the same way": until_noticed(rolling) + [(150, 0.75), (150, 0.50), (150, 0.25)],
    "blue-green, switch back (1 min)": until_noticed([(10**9, 1.0)]) + [(60, 1.0)],
    "canary on 5%, weight back to 0 (1 min)": until_noticed([(10**9, 0.05)]) + [(60, 0.05)],
    "canary on 1%, weight back to 0 (1 min)": until_noticed([(10**9, 0.01)]) + [(60, 0.01)],
}

budget = (1 - SLO) * RPS * MONTH_S
print(f"Monthly error budget at {SLO:.1%}: {budget:,.0f} failed requests")
print(f"{'strategy':<46} {'failed requests':>15} {'budget used':>11}")
for name, steps in strategies.items():
    failed = failed_requests(steps)
    print(f"{name:<46} {failed:>15,.0f} {failed / budget:>11.2%}")

Output

Monthly error budget at 99.9%: 5,184,000 failed requests
strategy                                       failed requests budget used
all at once, redeploy the old version (5 min)          360,000       6.94%
rolling, rolled back the same way                      240,000       4.63%
blue-green, switch back (1 min)                        264,000       5.09%
canary on 5%, weight back to 0 (1 min)                  13,200       0.25%
canary on 1%, weight back to 0 (1 min)                   2,640       0.05%

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 blast_radius.py

Against a 99.9% monthly objective, whose error budget allows 5,184,000 failed requests at this traffic, a bad build deployed everywhere spends about 7% of the month’s budget in a quarter of an hour. Blue-green does better only because switching back takes a minute instead of a redeploy. A rolling update that nothing stops reaches every instance after 7.5 minutes, before anyone notices, and its way back takes as long. A 5% canary fails 13,200 requests, about a twenty-seventh of the damage: a twentieth of the traffic meets the bad build, for 11 minutes instead of 15.

Rolling, blue-green and canary deployments

The share of traffic on the new version, step by step, for a rolling update, a blue-green switch and a canary.Rolling updateBlue-greenCanary25%50%75%100%0%tested,no users100%blue keptfor rollback1%check5%check25%check100%

How traffic moves to a new version under each strategy

Text description of the diagram

Three rows, one per strategy, each a sequence of steps joined by arrows from left to right. Each step gives the share of the traffic on the new version.

- Rolling update: 25%, 50%, 75%, then 100%, as batches of instances are replaced. - Blue-green: 0% while the new environment is tested with no users, then 100% after one switch of the router, with the old environment, blue, kept for rollback. - Canary: 1%, 5% and 25%, each followed by a check against the old version, then 100%.

A rolling update replaces instances a few at a time until all of them run the new version. It is the default strategy of a Kubernetes Deployment, which by default keeps at least 75% of the desired Pods available and runs at most 125% of them during an update (maxUnavailable and maxSurge of 25% each) (Kubernetes documentation). It needs only a few spare instances, but old and new versions serve side by side for the whole rollout, and going back is another rollout in the other direction (kubectl rollout undo). Kubernetes reports a Deployment that has made no progress for 600 seconds (progressDeadlineSeconds) and keeps retrying it, but leaves any rollback to a higher-level tool. Recreate, the other built-in strategy, stops every old Pod before starting new ones: a gap in service, but never two versions at once.

Blue-green keeps two full environments. The new version is deployed to the idle one, green, and tested with no users; then the router sends everyone to green, and switching back to blue, kept running, is a rollback in seconds (Fowler, BlueGreenDeployment). It costs a second environment, and everyone moves at once, so a defect the tests missed reaches every user until the switch back. The two environments usually share one database, so change its schema first, in a form both versions can use.

A canary sends a small share of traffic, often 1% to 5%, to the new version and compares it with the old one before every step up. Argo Rollouts, a Kubernetes controller, writes the plan as steps, such as a weight of 10%, a pause of an hour, then 20%, with an analysis running beside them; a failed analysis aborts the rollout and sets the canary’s weight back to zero (canary, analysis). Without a router that splits traffic, the weight is approximated by counting Pods: 10 replicas at 10% means one canary Pod, and with 12 instances the smallest canary is about 8% of the traffic.

Three ways to move traffic to a new version
Criterion Rolling updateBlue-greenCanary
New version gets one batch of servers at a timeall traffic at oncea small share first
Extra capacity a few serversa second full copya few servers
Versions mixed for the whole rolloutonly in the shared databaseyes, on purpose
A bad build reaches every new batcheveryonethe canary's share
Going back a new rollout: minutesa switch: secondsweight to zero: seconds
Needs health checksa router switch, a schema both versions usea traffic split, metrics per version
When to choose low-risk stateless servicesa change all users get at oncebusy services where failures are costly

Judging a canary by its numbers

Compare the canary with the old version serving traffic during the same minutes, not with the whole fleet and not with the hour before the deploy. A fleet-wide error rate hides a small canary: one on 5% of traffic that fails 10% of its requests moves the fleet’s rate by half a percentage point. A before-and-after comparison mixes the release with everything else that changed, such as the time of day or a slow dependency. Google’s SRE workbook makes both points and advises a few metrics that show real problems, such as errors and latency (SRE workbook, canarying releases).

Whether the canary’s error rate is really higher is a question about two proportions, answered by the two-proportion z-test (NIST/SEMATECH e-Handbook, 7.3.3). With pcp_c and pbp_b the error rates of the canary and the baseline, ncn_c and nbn_b their requests, and pp the error rate of both groups together:

z=pc−pbp(1−p)(1nc+1nb)z = \frac{p_c - p_b}{\sqrt{p\,(1 - p)\left(\frac{1}{n_c} + \frac{1}{n_b}\right)}}

The program below turns z into a decision. It rolls back at z of 3 or more, a strict limit because an automated check looks again every minute and each look is another chance of a false alarm. It never promotes a canary that has served fewer than 10,000 requests, and an absolute ceiling of 2% errors, which would come from the SLO, rolls back a badly broken build without waiting for statistics.

Promote, wait or roll back Python · canary_analysis.py
"""Promote, wait or roll back: a canary's error rate against the baseline's over the same minutes."""
from math import sqrt

Z_LIMIT = 3.0             # roll back when the canary is worse by this many standard errors
MIN_REQUESTS = 10_000     # promote only after the canary has served at least this many requests
CEILING = 0.02            # roll back at once above 2% errors, whatever the comparison says


def z_score(base_errors, base_total, canary_errors, canary_total):
    """Two-proportion z statistic with the pooled error rate."""
    pooled = (base_errors + canary_errors) / (base_total + canary_total)
    if pooled in (0, 1):
        return 0.0
    standard_error = sqrt(pooled * (1 - pooled) * (1 / base_total + 1 / canary_total))
    return (canary_errors / canary_total - base_errors / base_total) / standard_error


def decide(base_errors, base_total, canary_errors, canary_total):
    if canary_errors / canary_total > CEILING:
        return "roll back: over the ceiling"
    if z_score(base_errors, base_total, canary_errors, canary_total) >= Z_LIMIT:
        return "roll back: worse than the baseline"
    if canary_total < MIN_REQUESTS:
        return "wait: too few requests to judge"
    return "promote to the next step"


# 2,000 requests a second with 5% on the canary: (baseline errors, baseline requests, canary errors, canary requests)
checks = {
    "healthy, after 10 min": (5_700, 1_140_000, 312, 60_000),
    "regressed, after 10 min": (5_700, 1_140_000, 480, 60_000),
    "regressed, after 30 s": (285, 57_000, 24, 3_000),
    "broken, after 30 s": (285, 57_000, 750, 3_000),
}
print(f"{'check':<24} {'baseline':>8} {'canary':>8} {'z':>6}  decision")
for name, (be, bt, ce, ct) in checks.items():
    z = z_score(be, bt, ce, ct)
    print(f"{name:<24} {be / bt:>8.2%} {ce / ct:>8.2%} {z:>6.1f}  {decide(be, bt, ce, ct)}")

Output

check                    baseline   canary      z  decision
healthy, after 10 min       0.50%    0.52%    0.7  promote to the next step
regressed, after 10 min     0.50%    0.80%   10.0  roll back: worse than the baseline
regressed, after 30 s       0.50%    0.80%    2.2  wait: too few requests to judge
broken, after 30 s          0.50%   25.00%  100.5  roll back: over the ceiling

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 canary_analysis.py

Compare the two “regressed” rows. The same 0.80% against 0.50% gives z = 10.0 after 10 minutes, but only 2.2 after 30 seconds, when the canary has served 3,000 requests and 24 errors: that difference could still be chance, so the answer is to wait, not to promote. How long, then? The next program finds the canary requests at which this rule catches a doubled error rate nine times in ten, and turns them into minutes at three traffic levels.

How long a 5% canary must run Python · canary_bake.py
"""How long must a canary run before the check in canary_analysis.py can see a real regression?

For a canary on 5% of traffic whose error rate is twice the baseline's, find the canary requests at which
the z >= 3 rule catches it 9 times in 10 (normal approximation), then the minutes that takes at three
traffic levels.
"""
from math import sqrt
from statistics import NormalDist

Z_LIMIT = 3.0
SHARE = 0.05          # the canary's share of traffic
POWER = 0.90          # the chance of catching the regression


def chance_to_catch(p_base, p_canary, canary_n):
    base_n = canary_n * (1 - SHARE) / SHARE
    pooled = (1 - SHARE) * p_base + SHARE * p_canary
    se_no_change = sqrt(pooled * (1 - pooled) * (1 / canary_n + 1 / base_n))
    se_change = sqrt(p_canary * (1 - p_canary) / canary_n + p_base * (1 - p_base) / base_n)
    return 1 - NormalDist().cdf((Z_LIMIT * se_no_change - (p_canary - p_base)) / se_change)


def requests_needed(p_base, p_canary):
    low, high = 1, 1
    while chance_to_catch(p_base, p_canary, high) < POWER:
        low, high = high, high * 2
    while low < high:                      # the smallest count that reaches POWER
        mid = (low + high) // 2
        if chance_to_catch(p_base, p_canary, mid) >= POWER:
            high = mid
        else:
            low = mid + 1
    return high


def duration(seconds):
    return f"{seconds / 60:.1f} min" if seconds >= 60 else f"{seconds:.0f} s"


print(f"Canary on {SHARE:.0%} of traffic, error rate doubled, caught {POWER:.0%} of the time at z >= {Z_LIMIT:g}")
print(f"{'baseline errors':<16} {'canary requests':>15} {'at 200/s':>10} {'at 2,000/s':>10} {'at 20,000/s':>11}")
for p_base in (0.001, 0.005, 0.02):
    n = requests_needed(p_base, 2 * p_base)
    times = [duration(n / (rps * SHARE)) for rps in (200, 2_000, 20_000)]
    print(f"{p_base:<16.1%} {n:>15,} {times[0]:>10} {times[1]:>10} {times[2]:>11}")

Output

Canary on 5% of traffic, error rate doubled, caught 90% of the time at z >= 3
baseline errors  canary requests   at 200/s at 2,000/s at 20,000/s
0.1%                      24,866   41.4 min    4.1 min        25 s
0.5%                       4,946    8.2 min       49 s         5 s
2.0%                       1,211    2.0 min       12 s         1 s

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 canary_bake.py

Rare errors need many requests: at a 0.1% baseline the canary needs about 25,000, which takes 25 seconds on a service with 20,000 requests a second and 41 minutes on one with 200. So size the bake by requests as well as by the clock. The hands-off pipelines of the Amazon Builders’ Library do both, waiting a minimum time (at least an hour after each single-instance stage) and a minimum amount of traffic before promoting (automating safe, hands-off deployments).

A/B Test Significance Calculator Enter the baseline's and the canary's requests as visitors and their errors as conversions: it runs the same two-proportion test.

Rollbacks that are safe and automatic

Rolling back is the quickest repair for a bad release: put back the last version that worked, then investigate. It needs three things. The old build must still be deployable: Kubernetes keeps 10 old revisions by default, and a Deployment whose revisionHistoryLimit is 0 cannot be rolled back at all. Nothing irreversible may have happened, such as a dropped column. And the old version must read everything the new one wrote. An article in the Amazon Builders’ Library names a change of protocol as the most common reason a rollback is impossible (ensuring rollback safety): a release that changes a stored format or a message looks fine until you try to leave it. The program below shows it with shopping carts: version 1 stores a cart as a list of product IDs, the new version stores lines with quantities and reads both formats.

Carts that a rollback cannot read Python · two_phase.py
"""Why a rollback can fail: shopping carts saved by one version and read by another.

The old format stores a cart as a list of product IDs; the new one stores each line with a quantity.
"""
import json


def write_old(cart):
    return json.dumps({"items": [pid for pid, qty in cart for _ in range(qty)]})


def write_new(cart):
    return json.dumps({"items": [{"id": pid, "qty": qty} for pid, qty in cart]})


def read_old_only(raw):
    """v1's reader: every item must be a product ID."""
    return ", ".join(json.loads(raw)["items"])


def read_both(raw):
    """The reader of the prepare and activate versions: understands both formats."""
    lines = {}
    for item in json.loads(raw)["items"]:
        pid, qty = (item, 1) if isinstance(item, str) else (item["id"], item["qty"])
        lines[pid] = lines.get(pid, 0) + qty
    return ", ".join(f"{pid} x{qty}" for pid, qty in lines.items())


VERSIONS = {  # name: (reader, writer)
    "v1": (read_old_only, write_old),
    "v1.5 (prepare)": (read_both, write_old),
    "v2 (activate)": (read_both, write_new),
    "v2 (one step)": (read_both, write_new),
}


def run(plan, rollback_to):
    """Each version in the plan saves its carts; then the fleet rolls back and reads every cart again."""
    stored = []
    for name, carts in plan:
        writer = VERSIONS[name][1]
        stored += [writer([(f"p{n % 7}", 1 + n % 3), ("p9", 1)]) for n in range(carts)]
    reader = VERSIONS[rollback_to][0]
    failures, first_error = 0, None
    for raw in stored:
        try:
            reader(raw)
        except TypeError as error:
            failures += 1
            first_error = first_error or error
    print(" -> ".join(name for name, _ in plan) + f", then roll back to {rollback_to}")
    print(f"  {failures:,} of {len(stored):,} carts cannot be read" + (f": TypeError: {first_error}" if first_error else ""))


run([("v1", 600), ("v2 (one step)", 400)], rollback_to="v1")
run([("v1", 600), ("v1.5 (prepare)", 200), ("v2 (activate)", 200)], rollback_to="v1.5 (prepare)")
run([("v1", 600), ("v1.5 (prepare)", 200), ("v2 (activate)", 200)], rollback_to="v1")

Output

v1 -> v2 (one step), then roll back to v1
  400 of 1,000 carts cannot be read: TypeError: sequence item 0: expected str instance, dict found
v1 -> v1.5 (prepare) -> v2 (activate), then roll back to v1.5 (prepare)
  0 of 1,000 carts cannot be read
v1 -> v1.5 (prepare) -> v2 (activate), then roll back to v1
  200 of 1,000 carts cannot be read: TypeError: sequence item 0: expected str instance, dict found

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 two_phase.py

In the one-step plan the rollback makes 400 carts unreadable, every cart the new version saved: the repair has become a second outage. The same article’s fix is a two-phase deployment. The prepare version reads both formats but keeps writing the old one, so rolling it back is harmless. Only when every server runs it does the activate version start writing the new format, and rolling that back is safe too, because the version you return to reads both. The third run shows the limit: two steps back at once fails on the 200 new carts, so each phase bakes before the next starts, and the old-format reader is removed only after every stored record has been rewritten.

Four versions of a two-phase deployment in order, with the rollbacks that are safe one step back and the one that is not.v1reads the old formatwrites the old formatv1.5, preparereads both formatswrites the old formatv2, activatereads both formatswrites the new formatv3, after the backfillreads and writesthe new format onlydeploy everywhere,check every serverafter a bakeafter every recordis rewrittensafe rollbacksafe rollbackunsafe: v1 cannotread new records

A two-phase deployment of a new stored format

Text description of the diagram

Four boxes from top to bottom, joined by arrows in the order the versions ship:

1. v1 reads and writes the old format. 2. v1.5, the prepare step, reads both formats and still writes the old one. The arrow to it says "deploy everywhere, check every server". 3. v2, the activate step, reads both formats and writes the new one. The arrow to it says "after a bake". 4. v3 reads and writes only the new format. The arrow to it says "after every record is rewritten".

Arrows labelled "safe rollback" lead back one step, from v2 to v1.5 and from v1.5 to v1. A dashed arrow from v2 straight back to v1 is labelled "unsafe: v1 cannot read new records".

Check that the prepare version reached every server, not a percentage of healthy ones: one server still running the old reader when the new format arrives is exactly the failure the phases exist to avoid. And test the way back before production: the article describes upgrade-downgrade testing, which deploys to about half of a production-like fleet, then all of it, then rolls back, each stage long enough for every API and batch job to run once. Schema changes need the same care: expand-and-contract migrations, the subject of the next lesson in this module, apply it to databases.

Make the rollback automatic. A person who must notice a graph, decide and act needs minutes; an alarm on the new version’s errors and latency that rolls back by itself needs seconds. The Builders’ Library pipelines deploy to a single instance first, then to waves of growing size, and roll back on alarms throughout each bake, so the rollback is often under way by the time the on-call engineer has been paged. Argo Rollouts aborts a canary whose analysis fails; a plain Kubernetes Deployment only reports that it is stuck.

Feature flags: deploying is not releasing

Deploying puts a new version of the code on the servers; releasing lets users use a new behaviour. A feature flag separates them: the new code ships switched off, and the release is a change of configuration that the code reads at run time, which can be undone in seconds without a deploy. Pete Hodgson’s article on martinfowler.com sorts flags by how long they live and how quickly their answer must change (feature toggles):

  • Release toggles hide unfinished or unreleased features and should be gone within a week or two.
  • Experiment toggles split users between variants until an A/B test has a significant result.
  • Ops toggles let operators turn costly features off under load; a few stay for good as kill switches.
  • Permissioning toggles open features to some users only, and can live for years.

A percentage rollout needs every server to give a user the same answer without storing who is in it. Hash the flag’s key with the user’s ID, keep the remainder after dividing by 10,000, and switch the feature on for buckets below the percentage times 100. flagd, an open-source engine that follows the OpenFeature specification, splits users between variants by hashing the user’s key with the flag key (MurmurHash3), and calls the assignment sticky (flagd). The example uses FNV-1a, a hash that fits in a few lines (RFC 9923), and checks what a rollout needs.

A percentage rollout from a hash JavaScript · flag_eval.mjs
// A percentage rollout without storing who is in it: each user's bucket is a hash of the flag key and the
// user ID, so every server computes the same answer for the same user.

/** FNV-1a, 32 bits, over the characters of `text` (all the IDs here are ASCII). */
function fnv1a32(text) {
  let hash = 0x811c9dc5;
  for (let i = 0; i < text.length; i++) {
    hash ^= text.charCodeAt(i);
    hash = Math.imul(hash, 0x01000193) >>> 0;
  }
  return hash;
}

/** A whole number from 0 to 9,999: the user's place in line for this flag. */
function bucket(flagKey, userId) {
  return fnv1a32(`${flagKey}:${userId}`) % 10_000;
}

/** On for `percent` per cent of users: the buckets below percent × 100. */
function isOn(flags, flagKey, userId) {
  const flag = flags[flagKey];
  if (!flag) return false;                       // an unknown flag gets the safe default: off
  return bucket(flagKey, userId) < flag.percent * 100;
}

function check(what, ok) {
  if (!ok) throw new Error(`check failed: ${what}`);
  console.log(`ok: ${what}`);
}

const users = Array.from({ length: 50_000 }, (_, i) => `user-${i}`);
const flags = { 'new-checkout': { percent: 10 }, 'dark-mode': { percent: 10 } };
const count = (pred) => users.filter(pred).length;
const share = (n) => `${((100 * n) / users.length).toFixed(2)}%`;
const grouped = (n) => String(n).replace(/\B(?=(\d{3})+(?!\d))/g, ',');

check('the same user gets the same answer on every call',
  users.slice(0, 1_000).every((u) => isOn(flags, 'new-checkout', u) === isOn(flags, 'new-checkout', u)));

const atTen = users.filter((u) => isOn(flags, 'new-checkout', u));
console.log(`new-checkout at 10%: ${grouped(atTen.length)} of 50,000 users (${share(atTen.length)})`);

flags['new-checkout'].percent = 25;
const atTwentyFive = count((u) => isOn(flags, 'new-checkout', u));
console.log(`raised to 25%: ${grouped(atTwentyFive)} users (${share(atTwentyFive)})`);
check('everyone who had it at 10% still has it at 25%', atTen.every((u) => isOn(flags, 'new-checkout', u)));

flags['new-checkout'].percent = 10;
const both = count((u) => isOn(flags, 'new-checkout', u) && isOn(flags, 'dark-mode', u));
console.log(`in both flags at 10%: ${both} users (${share(both)}; independent flags give about 1%)`);
check('an unknown flag is off', !isOn(flags, 'no-such-flag', 'user-1'));

Output

ok: the same user gets the same answer on every call
new-checkout at 10%: 5,076 of 50,000 users (10.15%)
raised to 25%: 12,491 users (24.98%)
ok: everyone who had it at 10% still has it at 25%
in both flags at 10%: 534 users (1.07%; independent flags give about 1%)
ok: an unknown flag is off

Recorded with Node.js 24.21.0 on macOS 26 arm64. To run it yourself: mise exec node@24.21.0 -- node flag_eval.mjs

The same user always gets the same answer. Raising the percentage only adds users, because buckets below 10% are also below 25%. And two flags at 10% share about 1% of users, as independent flags should, because the flag’s key changes every user’s bucket; without it, every flag at 10% would go to the same tenth of your users. Every service must also hash exactly the same text, or one user gets two answers from two services.

Choose the default as carefully as the rollout. OpenFeature, an open specification for a vendor-neutral flag API, has every evaluation take a flag key and a default value, plus an optional evaluation context whose targeting key identifies the user, and requires it to return the default value instead of throwing when anything goes wrong (flag evaluation, evaluation context). When the flag service is unreachable, users get the default, so make the default the old, safe behaviour. And a flag reverts only the code it guards: anything else the same deploy changed still needs a rollback.

Keeping flags under control

Every flag doubles the paths through the code it guards, and flags multiply because they are cheap to add. Hodgson recommends testing the production configuration with the flags you are about to turn on, and the same with them off, rather than every combination. To stop flags piling up, give each an owner and an expiry date when it is created, add the task to remove it at the same time, and let a test fail once a flag outlives its date.

The cost of a forgotten flag is on record in a regulator’s order. The US Securities and Exchange Commission found that when a trading firm, Knight Capital, deployed new order-routing code to its eight routing servers, a technician did not copy it to one of them, and no second person checked the deployment. The new code reused a flag that had once switched on an old function, Power Peg, which was never deleted. Orders carrying the flag reached the eighth server and woke the old code, which sent child orders without stopping: over 4 million executions in 154 stocks in about 45 minutes, and a loss of more than $460 million. While searching for the cause, the firm removed the new code from the seven correct servers, which made things worse, because the reused flag now woke the old code there too (SEC order). Each mistake breaks a rule of this lesson: verify that every server runs the version you expect, never give an old flag a new meaning, delete dead code with its flag, and roll back only to a version that handles the new inputs.

Measuring how well you deliver

DORA, a long-running research programme on software delivery, measures it with five metrics (DORA metrics). Three describe throughput: change lead time (from commit to production), deployment frequency and failed deployment recovery time. Two describe instability: change fail rate, the share of deployments that need immediate intervention, and deployment rework rate, the share of deployments that were unplanned and made because of an incident. DORA finds that speed and stability are not a trade-off: teams tend to be good, or poor, at all five. Automatic rollback shortens recovery time, canaries make each failure cheaper, and flags let small changes ship often. DORA also warns against turning the metrics into targets and against comparing teams whose systems differ: measure each service against its own past.

Interview questions

What is the difference between deploying and releasing?

Deploying puts a new version of the code on the servers; releasing lets users use a new behaviour. With feature flags they become separate steps: the code ships switched off, so a deploy changes nothing users see, and the release is a configuration change that can reach 1% of users, then everyone, and be switched off again in seconds without a deploy. Keeping them apart makes deploys small and routine, lets a release wait for the right moment and gives a switch that is faster than a rollback.

What is a canary deployment, and what is it compared with?

A canary sends a small share of real traffic, often 1% to 5%, to the new version and moves forward in steps only while the new version is as healthy as the old one. It is compared with the old version serving traffic during the same minutes, on a few metrics such as errors and latency: not with the whole fleet, whose average hides a small canary, and not with the hour before, because traffic and dependencies change over time. It also needs enough requests before a difference counts, and an absolute limit stops a badly broken build at once.

Key takeaways

  • The damage of a bad build is the share of traffic it serves times the time until it is gone; canaries cut the first factor, automatic rollbacks and flags the second.
  • Rolling updates are cheap but mix versions and roll back slowly; blue-green moves everyone at once and back in seconds; a canary exposes a small share and compares it with the old version before each step.
  • Judge a canary against a concurrent baseline with a test such as the two-proportion z-test, a minimum number of requests and an absolute ceiling; quiet services and rare errors need longer canaries.
  • A rollback is safe only when the old version reads what the new one wrote: ship a reader of both formats first, switch the writer second, and remove the old reader once every record is rewritten.
  • Flags separate deploying from releasing; bucket users by a hash of the flag key and user ID, make the default the safe behaviour, and give every flag an owner and an expiry date.
  • Track the five DORA metrics per service, never as targets.

Exercise

Exercise · Medium · JavaScript

Split users between the variants of a flag by hashing

A flag service has to give every user the same variant on every request and on every server, without storing who got what. Write two functions in rollout.mjs and export them; fnv1a32(text) is there already.

**bucketOf(flagKey, userId)** returns the user's bucket for that flag, a whole number from 0 to 9,999: hash the text flagKey + ':' + userId with fnv1a32 and take the remainder after dividing by 10,000. For example, bucketOf('new-checkout', 'user-42') is 9926.

**variantOf(flagKey, userId, split)** returns the name of the variant the user gets. split is a list of [variant, percent] pairs, such as [['on', 10], ['off', 90]]:

  • Each variant gets Math.round(percent * 100) buckets, and the variants take consecutive ranges in the order they are listed: with the split above, buckets 0 to 999 are 'on' and buckets 1,000 to 9,999 are 'off'.
  • A variant with 0 per cent gets no buckets, so it is never returned.
  • Percents may have fractions: [['on', 0.5], ['off', 99.5]] gives 'on' to buckets 0 to 49.
  • If a percent is negative, or the buckets do not add up to exactly 10,000, throw a RangeError.

Because the ranges follow the list, raising the first variant from 10 to 20 per cent keeps everyone who already had it. The sample tests check that, and the share of 20,000 users each variant gets; they import both functions from rollout.mjs and run in your browser.

Starter code · rollout.mjs

/** FNV-1a, 32 bits, over the characters of `text`. Complete: use it as it is. */
export function fnv1a32(text) {
  let hash = 0x811c9dc5;
  for (let i = 0; i < text.length; i++) {
    hash ^= text.charCodeAt(i);
    hash = Math.imul(hash, 0x01000193) >>> 0;
  }
  return hash;
}

/** The bucket of `userId` for the flag `flagKey`: a whole number from 0 to 9,999. */
export function bucketOf(flagKey, userId) {
  // Replace this line with your code.
  return 0;
}

/** The variant that `userId` gets from `split`, a list of [variant, percent] pairs that add up to 100. */
export function variantOf(flagKey, userId, split) {
  // Replace this line with your code.
  return split[0][0];
}
The sample tests · rollout.test.mjs
import { test, assert } from 'mysmartcopilot:test';
import { bucketOf, variantOf } from './rollout.mjs';

const USERS = Array.from({ length: 20000 }, (_, i) => `user-${i}`);
const share = (flagKey, split, variant) => USERS.filter((u) => variantOf(flagKey, u, split) === variant).length / USERS.length;

test('buckets are the hash of "flagKey:userId" modulo 10,000', () => {
  assert.equal(bucketOf('new-checkout', 'user-42'), 9926);
  assert.equal(bucketOf('new-checkout', 'user-7'), 8825);
  assert.equal(bucketOf('dark-mode', 'user-42'), 3953);
});

test('every bucket is a whole number from 0 to 9,999', () => {
  assert.ok(
    USERS.every((u) => {
      const b = bucketOf('search-v2', u);
      return Number.isInteger(b) && b >= 0 && b < 10000;
    }),
  );
});

test('ranges follow the order of the split, with the boundary in the next range', () => {
  assert.equal(variantOf('new-checkout', 'user-12785', [['on', 10], ['off', 90]]), 'on');   // bucket 999
  assert.equal(variantOf('new-checkout', 'user-15843', [['a', 25], ['b', 75]]), 'a');      // bucket 2499
  assert.equal(variantOf('new-checkout', 'user-5058', [['a', 25], ['b', 75]]), 'b');       // bucket 2500
});

test('0% and 100% mean nobody and everybody', () => {
  assert.equal(share('new-checkout', [['on', 0], ['off', 100]], 'on'), 0);
  assert.equal(share('new-checkout', [['on', 100]], 'on'), 1);
});

test('a 25/75 split gives the treatment to about a quarter of 20,000 users', () => {
  const s = share('price-test', [['control', 75], ['treatment', 25]], 'treatment');
  assert.ok(s > 0.24 && s < 0.26, `the treatment share is ${s}`);
});

test('raising "on" from 10% to 20% keeps everyone who had it', () => {
  const had = USERS.filter((u) => variantOf('new-checkout', u, [['on', 10], ['off', 90]]) === 'on');
  assert.ok(had.length > 0);
  assert.ok(had.every((u) => variantOf('new-checkout', u, [['on', 20], ['off', 80]]) === 'on'));
});

test('a fraction of a per cent is a whole number of buckets', () => {
  assert.equal(variantOf('new-checkout', 'user-880', [['on', 0.5], ['off', 99.5]]), 'on');  // bucket 12
  assert.equal(share('new-checkout', [['on', 0.5], ['off', 99.5]], 'on') * USERS.length, 95);
});

test('splits that do not add up to 100, or have a negative percent, throw a RangeError', () => {
  assert.throws(() => variantOf('new-checkout', 'user-1', [['on', 10], ['off', 80]]), RangeError);
  assert.throws(() => variantOf('new-checkout', 'user-1', [['on', 110], ['off', -10]]), RangeError);
});
A hint

bucketOf is one line: ` fnv1a32(${flagKey}:${userId}) % 10000 . For variantOf, first turn every percent into a number of buckets with Math.round(percent * 100), check that none is negative and that they add up to 10,000, then walk the list keeping the bucket where the current variant's range ends: end += size. The first variant whose end` is greater than the user's bucket is the answer.

The sample tests run on this device, in your browser (QuickJS): nothing is sent to mysmartcopilot.com. The first run downloads JavaScript (about 0.6 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.

Check yourself

9 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.

  1. Question 1 of 9 A payments API serves 4,000 requests a second from 40 instances behind a load balancer that can split traffic by percentage, and its metrics can be broken down by version. A bad release would be costly. Which way of deploying fits best?

    Choose one answer.

    Show the answer to question 1

    Answer: A canary on 1% of traffic, compared with the old version before every step

    The service is busy enough for a small canary to collect evidence in minutes, it has the two things a canary needs (a traffic split and per-version metrics), and the cost of a bad release is high. A canary limits a bad build to 1% of requests while the comparison runs; a rolling update without a gate would reach the whole fleet before most problems are noticed.

  2. Question 2 of 9 A nightly billing job runs as a single process. Two versions must never run at the same time, because both would charge the same customers, and a few minutes without the job are harmless. Which way of deploying fits?

    Choose one answer.

    Show the answer to question 2

    Answer: Recreate, stopping the old version before the new one starts

    Every other strategy has a moment when two versions run side by side, which is exactly what this job cannot allow. Recreate accepts a short gap instead, and for a job that runs once a night the gap costs nothing.

  3. Question 3 of 9 A stateless thumbnail service runs on 12 instances with readiness checks, and its load balancer cannot split traffic by percentage. A bad build would make some thumbnails slow, nothing worse. Which way of deploying fits?

    Choose one answer.

    Show the answer to question 3

    Answer: A rolling update that replaces a few instances at a time while readiness checks pass

    With 12 instances the smallest step is one instance, about 8% of traffic, so a 1% canary is not possible without a traffic router; a rolling update with health checks already gives a gradual swap. The risk is low, so a second full environment would cost more than it saves, and Recreate would take the service down for no reason.

  4. Question 4 of 9 A release changes a web page and the API behind it together, and every user must move to the new pair at the same moment. The team can afford a second full environment for an hour and wants to be back on the old pair within seconds if anything goes wrong. Which way of deploying fits?

    Choose one answer.

    Show the answer to question 4

    Answer: Blue-green, switching the router once both new parts are tested

    Blue-green moves everyone at once, and the old environment stays ready, so switching back takes seconds. A rolling update or a canary would mix old pages with the new API, or the reverse, for the length of the rollout.

  5. Question 5 of 9 A service takes 1,000 requests a second. A bad build that fails 10% of the requests it serves runs on a 5% canary for 10 minutes before the canary's share goes back to 0. How many requests fail?

    Type a number.

    Show the answer to question 5

    Answer: 3000 requests

    The canary serves 5% of 1,000, or 50 requests a second, for 600 seconds: 30,000 requests. A tenth of them fail: 3,000. Deployed to everyone for the same 10 minutes, the same build would fail 60,000.

  6. Question 6 of 9 Put the steps of a two-phase deployment that changes a stored format in order.

    Give each item its position, from 1 (first).

    Show the answer to question 6

    Answer:

    1. Deploy a version that reads both formats but still writes the old one
    2. Check that every server runs that version, then let it bake
    3. Deploy the version that writes the new format
    4. Rewrite every stored record in the new format
    5. Remove the code that reads the old format

    Readers go before writers. Once every server can read both formats, writing the new one is safe to roll back, because the version you would return to reads it. The old reader goes last, and only after no record in the old format is left.

  7. Question 7 of 9 Which of these make rolling back to the previous version unsafe? Choose every one that does.

    Choose every answer that is right.

    Show the answer to question 7

    Answer:

    • The release ran a migration that dropped a column the previous version reads
    • The new version writes records in a format the previous version cannot parse

    A rollback is safe when the version you return to can work with everything the new one left behind. New formats it cannot read and columns it needs that are gone both break it. A kept build is what makes a rollback possible, and code behind a flag that was never turned on has not written anything new.

  8. Question 8 of 9 Why compare a canary with the old version serving traffic during the same minutes, rather than with the service's error rate before the deploy?

    Choose one answer.

    Show the answer to question 8

    Answer: Because traffic, time of day and dependencies change over time, and both groups see the same changes

    A comparison with the past mixes the release's effect with everything else that changed since: the time of day, the traffic mix, a slow dependency. A baseline running at the same time sees those changes too, so a difference between the two groups points at the release.

  9. Question 9 of 9 A service evaluates the flag new-checkout through an OpenFeature client with the default value false, and the flag provider is unreachable. What does the evaluation return?

    Choose one answer.

    Show the answer to question 9

    Answer: false, the default value the code passed in

    The OpenFeature specification requires evaluation calls not to throw and to return the default value when anything goes wrong, with an error code in the evaluation details. That is why the default should be the safe behaviour, usually the old code path.

References

Related tools

Report a problem with this lesson

Quick answers and tool search

Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.