System Design (High-Level Design) Module 18 – Delivery, multi-region, disaster recovery and cost
Safe deployments: canaries, rollbacks and flags
Compare rolling, blue-green and canary deployments, judge a canary from its error counts, keep rollbacks safe and release features with flags.
What you will learn
- Compare rolling, blue-green and canary deployments by exposure, extra capacity and rollback speed
- Decide from error counts whether a canary should be promoted, kept waiting or rolled back
- Design a change to stored data that can be rolled back at every step
- Use feature flags to separate deploying code from releasing a feature
- Implement a deterministic percentage rollout that gives every user a stable answer
- Track delivery with the five DORA metrics
Before you start
On this page
A safe deployment limits how many users a bad change can reach and how quickly it is undone. Three tools do most of the work. A canary sends a small share of traffic to the new version and compares it with the old one before going further. A rollback returns to the last version that worked, and it is safe only when the old and new versions can read each other’s data. A feature flag lets you deploy code switched off and release it later, to a percentage of users, with a switch that turns it off again in seconds. Below, each gets numbers: the requests a bad build fails under each strategy, how long a canary must run, and why some rollbacks make an incident worse.
Why a deployment needs a plan
Changes are where most outages start. Google’s SRE book reports that roughly 70% of outages come from changes to a live system, and names three habits that limit the damage: roll out progressively, detect problems quickly and accurately, and roll back safely (SRE book, introduction). The damage is roughly the share of traffic a bad build serves times the time until it is gone. The program below works it out for a build that fails 20% of its requests, deployed in five ways to a service taking 2,000 requests a second. Each strategy gets the same 10 minutes before someone notices, so only the exposure and the way back differ.
"""How many requests a bad build fails under five ways of deploying it.
Assumptions (change them and run again): the service takes 2,000 requests a second, the bad build fails
20% of the requests it serves, and the problem is noticed 10 minutes after the first users reach the new
version, whatever the strategy. Undoing it then takes as long as each strategy needs.
"""
RPS = 2_000
BAD_BUILD_FAILS = 0.20
NOTICED_AFTER_S = 10 * 60
SLO = 0.999 # a 99.9% monthly availability objective
MONTH_S = 30 * 24 * 3600
def failed_requests(steps):
"""steps: (seconds, share of traffic on the new version) in order; returns failed requests."""
return sum(RPS * BAD_BUILD_FAILS * share * seconds for seconds, share in steps)
def until_noticed(schedule):
"""Cut a rollout schedule of (seconds, share) at the moment the problem is noticed."""
steps, elapsed = [], 0
for seconds, share in schedule:
take = min(seconds, NOTICED_AFTER_S - elapsed)
if take <= 0:
break
steps.append((take, share))
elapsed += take
if elapsed < NOTICED_AFTER_S: # the rollout finished before anyone noticed
steps.append((NOTICED_AFTER_S - elapsed, schedule[-1][1]))
return steps
rolling = [(150, 0.25), (150, 0.50), (150, 0.75), (10**9, 1.0)] # a batch of 25% every 2.5 minutes
strategies = {
"all at once, redeploy the old version (5 min)": until_noticed([(10**9, 1.0)]) + [(300, 1.0)],
"rolling, rolled back the same way": until_noticed(rolling) + [(150, 0.75), (150, 0.50), (150, 0.25)],
"blue-green, switch back (1 min)": until_noticed([(10**9, 1.0)]) + [(60, 1.0)],
"canary on 5%, weight back to 0 (1 min)": until_noticed([(10**9, 0.05)]) + [(60, 0.05)],
"canary on 1%, weight back to 0 (1 min)": until_noticed([(10**9, 0.01)]) + [(60, 0.01)],
}
budget = (1 - SLO) * RPS * MONTH_S
print(f"Monthly error budget at {SLO:.1%}: {budget:,.0f} failed requests")
print(f"{'strategy':<46} {'failed requests':>15} {'budget used':>11}")
for name, steps in strategies.items():
failed = failed_requests(steps)
print(f"{name:<46} {failed:>15,.0f} {failed / budget:>11.2%}") Output
Monthly error budget at 99.9%: 5,184,000 failed requests strategy failed requests budget used all at once, redeploy the old version (5 min) 360,000 6.94% rolling, rolled back the same way 240,000 4.63% blue-green, switch back (1 min) 264,000 5.09% canary on 5%, weight back to 0 (1 min) 13,200 0.25% canary on 1%, weight back to 0 (1 min) 2,640 0.05%
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 blast_radius.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
Against a 99.9% monthly objective, whose error budget allows 5,184,000 failed requests at this traffic, a bad build deployed everywhere spends about 7% of the month’s budget in a quarter of an hour. Blue-green does better only because switching back takes a minute instead of a redeploy. A rolling update that nothing stops reaches every instance after 7.5 minutes, before anyone notices, and its way back takes as long. A 5% canary fails 13,200 requests, about a twenty-seventh of the damage: a twentieth of the traffic meets the bad build, for 11 minutes instead of 15.
Rolling, blue-green and canary deployments
How traffic moves to a new version under each strategy
Text description of the diagram
Three rows, one per strategy, each a sequence of steps joined by arrows from left to right. Each step gives the share of the traffic on the new version.
- Rolling update: 25%, 50%, 75%, then 100%, as batches of instances are replaced. - Blue-green: 0% while the new environment is tested with no users, then 100% after one switch of the router, with the old environment, blue, kept for rollback. - Canary: 1%, 5% and 25%, each followed by a check against the old version, then 100%.
A rolling update replaces instances a few at a time until all of them run the new version. It is the default
strategy of a Kubernetes Deployment, which by default keeps at least 75% of the desired Pods available and runs at
most 125% of them during an update (maxUnavailable and maxSurge of 25% each)
(Kubernetes documentation). It needs only a
few spare instances, but old and new versions serve side by side for the whole rollout, and going back is another
rollout in the other direction (kubectl rollout undo). Kubernetes reports a Deployment that has made no progress
for 600 seconds (progressDeadlineSeconds) and keeps retrying it, but leaves any rollback to a higher-level tool.
Recreate, the other built-in strategy, stops every old Pod before starting new ones: a gap in service, but never
two versions at once.
Blue-green keeps two full environments. The new version is deployed to the idle one, green, and tested with no users; then the router sends everyone to green, and switching back to blue, kept running, is a rollback in seconds (Fowler, BlueGreenDeployment). It costs a second environment, and everyone moves at once, so a defect the tests missed reaches every user until the switch back. The two environments usually share one database, so change its schema first, in a form both versions can use.
A canary sends a small share of traffic, often 1% to 5%, to the new version and compares it with the old one before every step up. Argo Rollouts, a Kubernetes controller, writes the plan as steps, such as a weight of 10%, a pause of an hour, then 20%, with an analysis running beside them; a failed analysis aborts the rollout and sets the canary’s weight back to zero (canary, analysis). Without a router that splits traffic, the weight is approximated by counting Pods: 10 replicas at 10% means one canary Pod, and with 12 instances the smallest canary is about 8% of the traffic.
| Criterion | Rolling update | Blue-green | Canary |
|---|---|---|---|
| New version gets | one batch of servers at a time | all traffic at once | a small share first |
| Extra capacity | a few servers | a second full copy | a few servers |
| Versions mixed | for the whole rollout | only in the shared database | yes, on purpose |
| A bad build reaches | every new batch | everyone | the canary's share |
| Going back | a new rollout: minutes | a switch: seconds | weight to zero: seconds |
| Needs | health checks | a router switch, a schema both versions use | a traffic split, metrics per version |
| When to choose | low-risk stateless services | a change all users get at once | busy services where failures are costly |
Judging a canary by its numbers
Compare the canary with the old version serving traffic during the same minutes, not with the whole fleet and not with the hour before the deploy. A fleet-wide error rate hides a small canary: one on 5% of traffic that fails 10% of its requests moves the fleet’s rate by half a percentage point. A before-and-after comparison mixes the release with everything else that changed, such as the time of day or a slow dependency. Google’s SRE workbook makes both points and advises a few metrics that show real problems, such as errors and latency (SRE workbook, canarying releases).
Whether the canary’s error rate is really higher is a question about two proportions, answered by the two-proportion z-test (NIST/SEMATECH e-Handbook, 7.3.3). With and the error rates of the canary and the baseline, and their requests, and the error rate of both groups together:
The program below turns z into a decision. It rolls back at z of 3 or more, a strict limit because an automated check looks again every minute and each look is another chance of a false alarm. It never promotes a canary that has served fewer than 10,000 requests, and an absolute ceiling of 2% errors, which would come from the SLO, rolls back a badly broken build without waiting for statistics.
"""Promote, wait or roll back: a canary's error rate against the baseline's over the same minutes."""
from math import sqrt
Z_LIMIT = 3.0 # roll back when the canary is worse by this many standard errors
MIN_REQUESTS = 10_000 # promote only after the canary has served at least this many requests
CEILING = 0.02 # roll back at once above 2% errors, whatever the comparison says
def z_score(base_errors, base_total, canary_errors, canary_total):
"""Two-proportion z statistic with the pooled error rate."""
pooled = (base_errors + canary_errors) / (base_total + canary_total)
if pooled in (0, 1):
return 0.0
standard_error = sqrt(pooled * (1 - pooled) * (1 / base_total + 1 / canary_total))
return (canary_errors / canary_total - base_errors / base_total) / standard_error
def decide(base_errors, base_total, canary_errors, canary_total):
if canary_errors / canary_total > CEILING:
return "roll back: over the ceiling"
if z_score(base_errors, base_total, canary_errors, canary_total) >= Z_LIMIT:
return "roll back: worse than the baseline"
if canary_total < MIN_REQUESTS:
return "wait: too few requests to judge"
return "promote to the next step"
# 2,000 requests a second with 5% on the canary: (baseline errors, baseline requests, canary errors, canary requests)
checks = {
"healthy, after 10 min": (5_700, 1_140_000, 312, 60_000),
"regressed, after 10 min": (5_700, 1_140_000, 480, 60_000),
"regressed, after 30 s": (285, 57_000, 24, 3_000),
"broken, after 30 s": (285, 57_000, 750, 3_000),
}
print(f"{'check':<24} {'baseline':>8} {'canary':>8} {'z':>6} decision")
for name, (be, bt, ce, ct) in checks.items():
z = z_score(be, bt, ce, ct)
print(f"{name:<24} {be / bt:>8.2%} {ce / ct:>8.2%} {z:>6.1f} {decide(be, bt, ce, ct)}") Output
check baseline canary z decision healthy, after 10 min 0.50% 0.52% 0.7 promote to the next step regressed, after 10 min 0.50% 0.80% 10.0 roll back: worse than the baseline regressed, after 30 s 0.50% 0.80% 2.2 wait: too few requests to judge broken, after 30 s 0.50% 25.00% 100.5 roll back: over the ceiling
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 canary_analysis.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
Compare the two “regressed” rows. The same 0.80% against 0.50% gives z = 10.0 after 10 minutes, but only 2.2 after 30 seconds, when the canary has served 3,000 requests and 24 errors: that difference could still be chance, so the answer is to wait, not to promote. How long, then? The next program finds the canary requests at which this rule catches a doubled error rate nine times in ten, and turns them into minutes at three traffic levels.
"""How long must a canary run before the check in canary_analysis.py can see a real regression?
For a canary on 5% of traffic whose error rate is twice the baseline's, find the canary requests at which
the z >= 3 rule catches it 9 times in 10 (normal approximation), then the minutes that takes at three
traffic levels.
"""
from math import sqrt
from statistics import NormalDist
Z_LIMIT = 3.0
SHARE = 0.05 # the canary's share of traffic
POWER = 0.90 # the chance of catching the regression
def chance_to_catch(p_base, p_canary, canary_n):
base_n = canary_n * (1 - SHARE) / SHARE
pooled = (1 - SHARE) * p_base + SHARE * p_canary
se_no_change = sqrt(pooled * (1 - pooled) * (1 / canary_n + 1 / base_n))
se_change = sqrt(p_canary * (1 - p_canary) / canary_n + p_base * (1 - p_base) / base_n)
return 1 - NormalDist().cdf((Z_LIMIT * se_no_change - (p_canary - p_base)) / se_change)
def requests_needed(p_base, p_canary):
low, high = 1, 1
while chance_to_catch(p_base, p_canary, high) < POWER:
low, high = high, high * 2
while low < high: # the smallest count that reaches POWER
mid = (low + high) // 2
if chance_to_catch(p_base, p_canary, mid) >= POWER:
high = mid
else:
low = mid + 1
return high
def duration(seconds):
return f"{seconds / 60:.1f} min" if seconds >= 60 else f"{seconds:.0f} s"
print(f"Canary on {SHARE:.0%} of traffic, error rate doubled, caught {POWER:.0%} of the time at z >= {Z_LIMIT:g}")
print(f"{'baseline errors':<16} {'canary requests':>15} {'at 200/s':>10} {'at 2,000/s':>10} {'at 20,000/s':>11}")
for p_base in (0.001, 0.005, 0.02):
n = requests_needed(p_base, 2 * p_base)
times = [duration(n / (rps * SHARE)) for rps in (200, 2_000, 20_000)]
print(f"{p_base:<16.1%} {n:>15,} {times[0]:>10} {times[1]:>10} {times[2]:>11}") Output
Canary on 5% of traffic, error rate doubled, caught 90% of the time at z >= 3 baseline errors canary requests at 200/s at 2,000/s at 20,000/s 0.1% 24,866 41.4 min 4.1 min 25 s 0.5% 4,946 8.2 min 49 s 5 s 2.0% 1,211 2.0 min 12 s 1 s
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 canary_bake.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
Rare errors need many requests: at a 0.1% baseline the canary needs about 25,000, which takes 25 seconds on a service with 20,000 requests a second and 41 minutes on one with 200. So size the bake by requests as well as by the clock. The hands-off pipelines of the Amazon Builders’ Library do both, waiting a minimum time (at least an hour after each single-instance stage) and a minimum amount of traffic before promoting (automating safe, hands-off deployments).
A/B Test Significance Calculator Enter the baseline's and the canary's requests as visitors and their errors as conversions: it runs the same two-proportion test.Rollbacks that are safe and automatic
Rolling back is the quickest repair for a bad release: put back the last version that worked, then investigate.
It needs three things. The old build must still be deployable: Kubernetes keeps 10 old revisions by default, and a
Deployment whose revisionHistoryLimit is 0 cannot be rolled back at all. Nothing irreversible may have happened,
such as a dropped column. And the old version must read everything the new one wrote. An article in the Amazon
Builders’ Library names a change of protocol as the most common reason a rollback is impossible
(ensuring rollback safety):
a release that changes a stored format or a message looks fine until you try to leave it. The program below shows
it with shopping carts: version 1 stores a cart as a list of product IDs, the new version stores lines with
quantities and reads both formats.
"""Why a rollback can fail: shopping carts saved by one version and read by another.
The old format stores a cart as a list of product IDs; the new one stores each line with a quantity.
"""
import json
def write_old(cart):
return json.dumps({"items": [pid for pid, qty in cart for _ in range(qty)]})
def write_new(cart):
return json.dumps({"items": [{"id": pid, "qty": qty} for pid, qty in cart]})
def read_old_only(raw):
"""v1's reader: every item must be a product ID."""
return ", ".join(json.loads(raw)["items"])
def read_both(raw):
"""The reader of the prepare and activate versions: understands both formats."""
lines = {}
for item in json.loads(raw)["items"]:
pid, qty = (item, 1) if isinstance(item, str) else (item["id"], item["qty"])
lines[pid] = lines.get(pid, 0) + qty
return ", ".join(f"{pid} x{qty}" for pid, qty in lines.items())
VERSIONS = { # name: (reader, writer)
"v1": (read_old_only, write_old),
"v1.5 (prepare)": (read_both, write_old),
"v2 (activate)": (read_both, write_new),
"v2 (one step)": (read_both, write_new),
}
def run(plan, rollback_to):
"""Each version in the plan saves its carts; then the fleet rolls back and reads every cart again."""
stored = []
for name, carts in plan:
writer = VERSIONS[name][1]
stored += [writer([(f"p{n % 7}", 1 + n % 3), ("p9", 1)]) for n in range(carts)]
reader = VERSIONS[rollback_to][0]
failures, first_error = 0, None
for raw in stored:
try:
reader(raw)
except TypeError as error:
failures += 1
first_error = first_error or error
print(" -> ".join(name for name, _ in plan) + f", then roll back to {rollback_to}")
print(f" {failures:,} of {len(stored):,} carts cannot be read" + (f": TypeError: {first_error}" if first_error else ""))
run([("v1", 600), ("v2 (one step)", 400)], rollback_to="v1")
run([("v1", 600), ("v1.5 (prepare)", 200), ("v2 (activate)", 200)], rollback_to="v1.5 (prepare)")
run([("v1", 600), ("v1.5 (prepare)", 200), ("v2 (activate)", 200)], rollback_to="v1") Output
v1 -> v2 (one step), then roll back to v1 400 of 1,000 carts cannot be read: TypeError: sequence item 0: expected str instance, dict found v1 -> v1.5 (prepare) -> v2 (activate), then roll back to v1.5 (prepare) 0 of 1,000 carts cannot be read v1 -> v1.5 (prepare) -> v2 (activate), then roll back to v1 200 of 1,000 carts cannot be read: TypeError: sequence item 0: expected str instance, dict found
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 two_phase.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
In the one-step plan the rollback makes 400 carts unreadable, every cart the new version saved: the repair has become a second outage. The same article’s fix is a two-phase deployment. The prepare version reads both formats but keeps writing the old one, so rolling it back is harmless. Only when every server runs it does the activate version start writing the new format, and rolling that back is safe too, because the version you return to reads both. The third run shows the limit: two steps back at once fails on the 200 new carts, so each phase bakes before the next starts, and the old-format reader is removed only after every stored record has been rewritten.
A two-phase deployment of a new stored format
Text description of the diagram
Four boxes from top to bottom, joined by arrows in the order the versions ship:
1. v1 reads and writes the old format. 2. v1.5, the prepare step, reads both formats and still writes the old one. The arrow to it says "deploy everywhere, check every server". 3. v2, the activate step, reads both formats and writes the new one. The arrow to it says "after a bake". 4. v3 reads and writes only the new format. The arrow to it says "after every record is rewritten".
Arrows labelled "safe rollback" lead back one step, from v2 to v1.5 and from v1.5 to v1. A dashed arrow from v2 straight back to v1 is labelled "unsafe: v1 cannot read new records".
Check that the prepare version reached every server, not a percentage of healthy ones: one server still running the old reader when the new format arrives is exactly the failure the phases exist to avoid. And test the way back before production: the article describes upgrade-downgrade testing, which deploys to about half of a production-like fleet, then all of it, then rolls back, each stage long enough for every API and batch job to run once. Schema changes need the same care: expand-and-contract migrations, the subject of the next lesson in this module, apply it to databases.
Make the rollback automatic. A person who must notice a graph, decide and act needs minutes; an alarm on the new version’s errors and latency that rolls back by itself needs seconds. The Builders’ Library pipelines deploy to a single instance first, then to waves of growing size, and roll back on alarms throughout each bake, so the rollback is often under way by the time the on-call engineer has been paged. Argo Rollouts aborts a canary whose analysis fails; a plain Kubernetes Deployment only reports that it is stuck.
Feature flags: deploying is not releasing
Deploying puts a new version of the code on the servers; releasing lets users use a new behaviour. A feature flag separates them: the new code ships switched off, and the release is a change of configuration that the code reads at run time, which can be undone in seconds without a deploy. Pete Hodgson’s article on martinfowler.com sorts flags by how long they live and how quickly their answer must change (feature toggles):
- Release toggles hide unfinished or unreleased features and should be gone within a week or two.
- Experiment toggles split users between variants until an A/B test has a significant result.
- Ops toggles let operators turn costly features off under load; a few stay for good as kill switches.
- Permissioning toggles open features to some users only, and can live for years.
A percentage rollout needs every server to give a user the same answer without storing who is in it. Hash the flag’s key with the user’s ID, keep the remainder after dividing by 10,000, and switch the feature on for buckets below the percentage times 100. flagd, an open-source engine that follows the OpenFeature specification, splits users between variants by hashing the user’s key with the flag key (MurmurHash3), and calls the assignment sticky (flagd). The example uses FNV-1a, a hash that fits in a few lines (RFC 9923), and checks what a rollout needs.
// A percentage rollout without storing who is in it: each user's bucket is a hash of the flag key and the
// user ID, so every server computes the same answer for the same user.
/** FNV-1a, 32 bits, over the characters of `text` (all the IDs here are ASCII). */
function fnv1a32(text) {
let hash = 0x811c9dc5;
for (let i = 0; i < text.length; i++) {
hash ^= text.charCodeAt(i);
hash = Math.imul(hash, 0x01000193) >>> 0;
}
return hash;
}
/** A whole number from 0 to 9,999: the user's place in line for this flag. */
function bucket(flagKey, userId) {
return fnv1a32(`${flagKey}:${userId}`) % 10_000;
}
/** On for `percent` per cent of users: the buckets below percent × 100. */
function isOn(flags, flagKey, userId) {
const flag = flags[flagKey];
if (!flag) return false; // an unknown flag gets the safe default: off
return bucket(flagKey, userId) < flag.percent * 100;
}
function check(what, ok) {
if (!ok) throw new Error(`check failed: ${what}`);
console.log(`ok: ${what}`);
}
const users = Array.from({ length: 50_000 }, (_, i) => `user-${i}`);
const flags = { 'new-checkout': { percent: 10 }, 'dark-mode': { percent: 10 } };
const count = (pred) => users.filter(pred).length;
const share = (n) => `${((100 * n) / users.length).toFixed(2)}%`;
const grouped = (n) => String(n).replace(/\B(?=(\d{3})+(?!\d))/g, ',');
check('the same user gets the same answer on every call',
users.slice(0, 1_000).every((u) => isOn(flags, 'new-checkout', u) === isOn(flags, 'new-checkout', u)));
const atTen = users.filter((u) => isOn(flags, 'new-checkout', u));
console.log(`new-checkout at 10%: ${grouped(atTen.length)} of 50,000 users (${share(atTen.length)})`);
flags['new-checkout'].percent = 25;
const atTwentyFive = count((u) => isOn(flags, 'new-checkout', u));
console.log(`raised to 25%: ${grouped(atTwentyFive)} users (${share(atTwentyFive)})`);
check('everyone who had it at 10% still has it at 25%', atTen.every((u) => isOn(flags, 'new-checkout', u)));
flags['new-checkout'].percent = 10;
const both = count((u) => isOn(flags, 'new-checkout', u) && isOn(flags, 'dark-mode', u));
console.log(`in both flags at 10%: ${both} users (${share(both)}; independent flags give about 1%)`);
check('an unknown flag is off', !isOn(flags, 'no-such-flag', 'user-1')); Output
ok: the same user gets the same answer on every call new-checkout at 10%: 5,076 of 50,000 users (10.15%) raised to 25%: 12,491 users (24.98%) ok: everyone who had it at 10% still has it at 25% in both flags at 10%: 534 users (1.07%; independent flags give about 1%) ok: an unknown flag is off
Recorded with Node.js 24.21.0 on macOS 26 arm64. To run it yourself: mise exec node@24.21.0 -- node flag_eval.mjs
Runs on this device, in your browser. The first run downloads JavaScript (about 0.6 MB), which is kept for the next runs.
Your run, in this browser
The same user always gets the same answer. Raising the percentage only adds users, because buckets below 10% are also below 25%. And two flags at 10% share about 1% of users, as independent flags should, because the flag’s key changes every user’s bucket; without it, every flag at 10% would go to the same tenth of your users. Every service must also hash exactly the same text, or one user gets two answers from two services.
Choose the default as carefully as the rollout. OpenFeature, an open specification for a vendor-neutral flag API, has every evaluation take a flag key and a default value, plus an optional evaluation context whose targeting key identifies the user, and requires it to return the default value instead of throwing when anything goes wrong (flag evaluation, evaluation context). When the flag service is unreachable, users get the default, so make the default the old, safe behaviour. And a flag reverts only the code it guards: anything else the same deploy changed still needs a rollback.
Keeping flags under control
Every flag doubles the paths through the code it guards, and flags multiply because they are cheap to add. Hodgson recommends testing the production configuration with the flags you are about to turn on, and the same with them off, rather than every combination. To stop flags piling up, give each an owner and an expiry date when it is created, add the task to remove it at the same time, and let a test fail once a flag outlives its date.
The cost of a forgotten flag is on record in a regulator’s order. The US Securities and Exchange Commission found that when a trading firm, Knight Capital, deployed new order-routing code to its eight routing servers, a technician did not copy it to one of them, and no second person checked the deployment. The new code reused a flag that had once switched on an old function, Power Peg, which was never deleted. Orders carrying the flag reached the eighth server and woke the old code, which sent child orders without stopping: over 4 million executions in 154 stocks in about 45 minutes, and a loss of more than $460 million. While searching for the cause, the firm removed the new code from the seven correct servers, which made things worse, because the reused flag now woke the old code there too (SEC order). Each mistake breaks a rule of this lesson: verify that every server runs the version you expect, never give an old flag a new meaning, delete dead code with its flag, and roll back only to a version that handles the new inputs.
Measuring how well you deliver
DORA, a long-running research programme on software delivery, measures it with five metrics (DORA metrics). Three describe throughput: change lead time (from commit to production), deployment frequency and failed deployment recovery time. Two describe instability: change fail rate, the share of deployments that need immediate intervention, and deployment rework rate, the share of deployments that were unplanned and made because of an incident. DORA finds that speed and stability are not a trade-off: teams tend to be good, or poor, at all five. Automatic rollback shortens recovery time, canaries make each failure cheaper, and flags let small changes ship often. DORA also warns against turning the metrics into targets and against comparing teams whose systems differ: measure each service against its own past.
Interview questions
What is the difference between deploying and releasing?
Deploying puts a new version of the code on the servers; releasing lets users use a new behaviour. With feature flags they become separate steps: the code ships switched off, so a deploy changes nothing users see, and the release is a configuration change that can reach 1% of users, then everyone, and be switched off again in seconds without a deploy. Keeping them apart makes deploys small and routine, lets a release wait for the right moment and gives a switch that is faster than a rollback.
What is a canary deployment, and what is it compared with?
A canary sends a small share of real traffic, often 1% to 5%, to the new version and moves forward in steps only while the new version is as healthy as the old one. It is compared with the old version serving traffic during the same minutes, on a few metrics such as errors and latency: not with the whole fleet, whose average hides a small canary, and not with the hour before, because traffic and dependencies change over time. It also needs enough requests before a difference counts, and an absolute limit stops a badly broken build at once.
Key takeaways
- The damage of a bad build is the share of traffic it serves times the time until it is gone; canaries cut the first factor, automatic rollbacks and flags the second.
- Rolling updates are cheap but mix versions and roll back slowly; blue-green moves everyone at once and back in seconds; a canary exposes a small share and compares it with the old version before each step.
- Judge a canary against a concurrent baseline with a test such as the two-proportion z-test, a minimum number of requests and an absolute ceiling; quiet services and rare errors need longer canaries.
- A rollback is safe only when the old version reads what the new one wrote: ship a reader of both formats first, switch the writer second, and remove the old reader once every record is rewritten.
- Flags separate deploying from releasing; bucket users by a hash of the flag key and user ID, make the default the safe behaviour, and give every flag an owner and an expiry date.
- Track the five DORA metrics per service, never as targets.
Exercise
Exercise · Medium · JavaScript
Split users between the variants of a flag by hashing
A flag service has to give every user the same variant on every request and on every server, without storing who got what. Write two functions in rollout.mjs and export them; fnv1a32(text) is there already.
**bucketOf(flagKey, userId)** returns the user's bucket for that flag, a whole number from 0 to 9,999: hash the text flagKey + ':' + userId with fnv1a32 and take the remainder after dividing by 10,000. For example, bucketOf('new-checkout', 'user-42') is 9926.
**variantOf(flagKey, userId, split)** returns the name of the variant the user gets. split is a list of [variant, percent] pairs, such as [['on', 10], ['off', 90]]:
- Each variant gets
Math.round(percent * 100)buckets, and the variants take consecutive ranges in the order they are listed: with the split above, buckets 0 to 999 are'on'and buckets 1,000 to 9,999 are'off'. - A variant with 0 per cent gets no buckets, so it is never returned.
- Percents may have fractions:
[['on', 0.5], ['off', 99.5]]gives'on'to buckets 0 to 49. - If a percent is negative, or the buckets do not add up to exactly 10,000, throw a
RangeError.
Because the ranges follow the list, raising the first variant from 10 to 20 per cent keeps everyone who already had it. The sample tests check that, and the share of 20,000 users each variant gets; they import both functions from rollout.mjs and run in your browser.
Starter code · rollout.mjs
/** FNV-1a, 32 bits, over the characters of `text`. Complete: use it as it is. */
export function fnv1a32(text) {
let hash = 0x811c9dc5;
for (let i = 0; i < text.length; i++) {
hash ^= text.charCodeAt(i);
hash = Math.imul(hash, 0x01000193) >>> 0;
}
return hash;
}
/** The bucket of `userId` for the flag `flagKey`: a whole number from 0 to 9,999. */
export function bucketOf(flagKey, userId) {
// Replace this line with your code.
return 0;
}
/** The variant that `userId` gets from `split`, a list of [variant, percent] pairs that add up to 100. */
export function variantOf(flagKey, userId, split) {
// Replace this line with your code.
return split[0][0];
} The sample tests · rollout.test.mjs
import { test, assert } from 'mysmartcopilot:test';
import { bucketOf, variantOf } from './rollout.mjs';
const USERS = Array.from({ length: 20000 }, (_, i) => `user-${i}`);
const share = (flagKey, split, variant) => USERS.filter((u) => variantOf(flagKey, u, split) === variant).length / USERS.length;
test('buckets are the hash of "flagKey:userId" modulo 10,000', () => {
assert.equal(bucketOf('new-checkout', 'user-42'), 9926);
assert.equal(bucketOf('new-checkout', 'user-7'), 8825);
assert.equal(bucketOf('dark-mode', 'user-42'), 3953);
});
test('every bucket is a whole number from 0 to 9,999', () => {
assert.ok(
USERS.every((u) => {
const b = bucketOf('search-v2', u);
return Number.isInteger(b) && b >= 0 && b < 10000;
}),
);
});
test('ranges follow the order of the split, with the boundary in the next range', () => {
assert.equal(variantOf('new-checkout', 'user-12785', [['on', 10], ['off', 90]]), 'on'); // bucket 999
assert.equal(variantOf('new-checkout', 'user-15843', [['a', 25], ['b', 75]]), 'a'); // bucket 2499
assert.equal(variantOf('new-checkout', 'user-5058', [['a', 25], ['b', 75]]), 'b'); // bucket 2500
});
test('0% and 100% mean nobody and everybody', () => {
assert.equal(share('new-checkout', [['on', 0], ['off', 100]], 'on'), 0);
assert.equal(share('new-checkout', [['on', 100]], 'on'), 1);
});
test('a 25/75 split gives the treatment to about a quarter of 20,000 users', () => {
const s = share('price-test', [['control', 75], ['treatment', 25]], 'treatment');
assert.ok(s > 0.24 && s < 0.26, `the treatment share is ${s}`);
});
test('raising "on" from 10% to 20% keeps everyone who had it', () => {
const had = USERS.filter((u) => variantOf('new-checkout', u, [['on', 10], ['off', 90]]) === 'on');
assert.ok(had.length > 0);
assert.ok(had.every((u) => variantOf('new-checkout', u, [['on', 20], ['off', 80]]) === 'on'));
});
test('a fraction of a per cent is a whole number of buckets', () => {
assert.equal(variantOf('new-checkout', 'user-880', [['on', 0.5], ['off', 99.5]]), 'on'); // bucket 12
assert.equal(share('new-checkout', [['on', 0.5], ['off', 99.5]], 'on') * USERS.length, 95);
});
test('splits that do not add up to 100, or have a negative percent, throw a RangeError', () => {
assert.throws(() => variantOf('new-checkout', 'user-1', [['on', 10], ['off', 80]]), RangeError);
assert.throws(() => variantOf('new-checkout', 'user-1', [['on', 110], ['off', -10]]), RangeError);
}); A hint
bucketOf is one line: ` fnv1a32(${flagKey}:${userId}) % 10000 . For variantOf, first turn every percent into a number of buckets with Math.round(percent * 100), check that none is negative and that they add up to 10,000, then walk the list keeping the bucket where the current variant's range ends: end += size. The first variant whose end` is greater than the user's bucket is the answer.
Results of the sample tests
| Test | Result | Details |
|---|
What your code printed
The sample tests run on this device, in your browser (QuickJS): nothing is sent to mysmartcopilot.com. The first run downloads JavaScript (about 0.6 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.
Check yourself
9 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.
References
- Kubernetes documentation: Deployments (rolling updates, rollback, progress deadline) (The Kubernetes Authors)
- Argo Rollouts: Canary deployment strategy (Argo Project)
- Argo Rollouts: Analysis and progressive delivery (Argo Project)
- BlueGreenDeployment (Martin Fowler)
- The Site Reliability Workbook, chapter 16: Canarying Releases (Google (O'Reilly Media))
- Site Reliability Engineering, chapter 1: Introduction (change management) (Google (O'Reilly Media))
- NIST/SEMATECH e-Handbook of Statistical Methods, 7.3.3: comparing two proportions (National Institute of Standards and Technology)
- Ensuring rollback safety during deployments (Amazon Builders' Library) (Amazon Web Services)
- Automating safe, hands-off deployments (Amazon Builders' Library) (Amazon Web Services)
- Feature Toggles (aka Feature Flags), by Pete Hodgson (martinfowler.com)
- OpenFeature specification: Flag Evaluation API (OpenFeature (CNCF))
- OpenFeature specification: Evaluation Context (OpenFeature (CNCF))
- flagd: the fractional operation (OpenFeature (flagd))
- RFC 9923: The FNV Non-Cryptographic Hash Algorithm (RFC Editor (Independent Submission))
- DORA's software delivery performance metrics (DORA)
- In the Matter of Knight Capital Americas LLC, Release No. 70694 (order) (U.S. Securities and Exchange Commission)
Related tools
Report a problem with this lesson
Kept only in this browser. Your Learn progress