System Design (High-Level Design) Module 4 – Networking for system designers
DNS, GeoDNS and anycast for system design
How DNS resolution and TTL caching work, how GeoDNS and health-checked records steer traffic, and how anycast compares, with a failover model that runs.
What you will learn
- Explain recursive resolution, caching and TTLs
- Use GeoDNS and health-checked records for traffic steering
- Calculate how much traffic still reaches a failed region after a DNS change
- Compare DNS-based steering with anycast routing
Before you start
On this page
DNS turns a name into an address through a chain of servers: the device asks a recursive resolver, which follows referrals from the root to the servers that are authoritative for the name, and every answer is cached for its TTL along the way. Because the authoritative server chooses each answer, it can answer differently by where the question came from (GeoDNS), by weights, or by which regions pass their health checks. That makes DNS a traffic-steering tool, with one weakness: a change reaches users only as caches expire, so DNS failover takes minutes, not seconds. Anycast steers differently: many sites announce the same IP address, Internet routing delivers each packet to the nearest one, and a failed site simply withdraws its announcement.
This lesson resolves names in a small simulated resolver, measures what a TTL costs during a failover, and compares the two ways to steer.
How a name is resolved
Resolving a name on a cold cache, step by step
Text description of the diagram
A device's stub resolver asks a recursive resolver for www.shop.example (1). The recursive resolver has nothing cached, so it walks the hierarchy: - it asks a root server (2), which answers with a referral to the servers of the example. zone (3); - it asks those servers (4), which refer it to the servers of shop.example. (5); - it asks the authoritative servers of shop.example. (6), which return the answer with its TTL (7).
The recursive resolver caches the referrals and the answer, and returns the answer to the device (8). Until the TTL runs out, the next users of the same resolver get the answer from its cache without any of steps 2 to 7.
A device does not walk the DNS itself. Its stub resolver sends the whole question to a recursive resolver, run by the network, an ISP or a public service. The recursive resolver starts at the root, which refers it to the servers of the top-level domain, which refer it to the servers that hold the zone, and those give the answer (RFC 1034). Every referral and every answer carries a TTL, and the resolver keeps each for that long, so the next user of the same resolver skips most or all of the walk.
This resolver runs on invented zones under .example, a name reserved for documentation
(RFC 6761). The address is an alias, a CNAME, that points into a
second zone, as names served through a CDN often do:
"""A tiny recursive resolver on invented zones. The clock moves only when the script says so, so every run prints the
same thing. It counts the queries it sends, keeps answers for their TTL and keeps "no such name" answers for the
time the zone's SOA record allows."""
# Each zone: delegations to child zones (name -> TTL of the referral), its records ((name, type) -> (value, TTL)),
# and the negative-caching TTL of its SOA record. Names under .example are reserved for documentation.
ZONES = {
".": {"children": {"example.": 172800}, "records": {}, "negative_ttl": 86400},
"example.": {"children": {"shop.example.": 86400, "cdn.example.": 86400}, "records": {}, "negative_ttl": 900},
"shop.example.": {"children": {}, "records": {("www.shop.example.", "CNAME"): ("shop.cdn.example.", 3600)}, "negative_ttl": 300},
"cdn.example.": {"children": {}, "records": {("shop.cdn.example.", "A"): ("192.0.2.10", 60)}, "negative_ttl": 60},
}
cache = {} # (name, type) -> (value, expires at)
referrals = {} # zone -> expires at: which zones' servers the resolver knows
negative = {} # name -> expires at
def fresh(expires, now):
return expires > now
def closest_known_zone(name, now):
"""The most specific zone the resolver already knows the servers of (the root's are built in)."""
labels = name.split(".")
for i in range(len(labels) - 1):
zone = ".".join(labels[i:])
if zone in referrals and fresh(referrals[zone], now):
return zone
return "."
def resolve(name, now, asked):
"""The address of `name`, following CNAMEs. Appends every zone it queries to `asked`."""
if (name, "A") in cache and fresh(cache[name, "A"][1], now):
return cache[name, "A"][0]
if (name, "CNAME") in cache and fresh(cache[name, "CNAME"][1], now):
return resolve(cache[name, "CNAME"][0], now, asked)
if name in negative and fresh(negative[name], now):
return "NXDOMAIN (cached)"
zone = closest_known_zone(name, now)
while True:
asked.append(zone)
data = ZONES[zone]
child = next((c for c in data["children"] if name == c or name.endswith("." + c)), None)
if child: # a referral: "ask the servers of this child zone"
referrals[child] = now + data["children"][child]
zone = child
continue
for rtype in ("A", "CNAME"):
if (name, rtype) in data["records"]:
value, ttl = data["records"][name, rtype]
cache[name, rtype] = (value, now + ttl)
return value if rtype == "A" else resolve(value, now, asked)
negative[name] = now + data["negative_ttl"]
return "NXDOMAIN"
def left(name, rtype, now):
entry = cache.get((name, rtype))
return max(0, entry[1] - now) if entry else 0
print(f"{'time':>6} {'name':<19} {'answer':<19}{'queries':>8} zones asked")
for now, name in [(0, "www.shop.example."), (30, "www.shop.example."), (90, "www.shop.example."),
(100, "wwww.shop.example."), (200, "wwww.shop.example."), (500, "wwww.shop.example.")]:
asked = []
answer = resolve(name, now, asked)
print(f"{now:>5}s {name:<19} {answer:<19}{len(asked):>8} {', '.join('root' if z == '.' else z for z in asked) or '-'}")
if name.startswith("www."):
print(f"{'':>8}cache: CNAME {left('www.shop.example.', 'CNAME', now)} s left, A {left('shop.cdn.example.', 'A', now)} s left") Output
time name answer queries zones asked
0s www.shop.example. 192.0.2.10 5 root, example., shop.example., example., cdn.example.
cache: CNAME 3600 s left, A 60 s left
30s www.shop.example. 192.0.2.10 0 -
cache: CNAME 3570 s left, A 30 s left
90s www.shop.example. 192.0.2.10 1 cdn.example.
cache: CNAME 3510 s left, A 60 s left
100s wwww.shop.example. NXDOMAIN 1 shop.example.
200s wwww.shop.example. NXDOMAIN (cached) 0 -
500s wwww.shop.example. NXDOMAIN 1 shop.example.
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 resolve_sim.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
Read it line by line:
- The first lookup asks five servers. Two of them, the root and the
example.servers, only refer it onward. The alias costs a second walk, but a shorter one, because the referral toexample.is already cached. - Thirty seconds later nothing is asked. Both records come from the cache, with their TTLs counting down.
- At 90 seconds one server is asked. The address had a TTL of 60 seconds and has expired, while the alias, with a TTL of 3,600, has not. Each link of a chain expires on its own clock.
- A typo is cached too. A name that does not exist gets an NXDOMAIN answer, and the zone’s SOA record says how long that “no” may be cached (RFC 2308), here 300 seconds. The second typo is answered from the cache; the third, after the 300 seconds, asks again.
Negative caching has a practical edge: if someone looks up a record before you create it, resolvers may keep answering “no such name” for the negative TTL after you have added it.
DNS Lookup Look up a real name's records and see each record's TTL as a public resolver reports it.TTLs trade the speed of change against load
The TTL is a limit on how long a record may be kept in a cache, and the DNS specification suggests that for a typical host it can be days. It also gives the standard trick for planned changes: lower the TTL before a change you can anticipate, then raise it again afterwards (RFC 1034). Lower it at least one old TTL ahead of time, so that the long-lived copies are gone when the change happens.
A short TTL is not free:
- Load on the DNS servers rises in proportion. Every resolver that keeps the name warm asks again once per TTL, so going from 3,600 seconds to 30 multiplies those queries by 120.
- Users wait more often. Each lookup that misses the cache costs round trips, which the previous lessons counted.
- Not every client obeys. Some applications and devices keep an address longer than its TTL. Resolvers may also keep serving an expired answer, with a short TTL attached, when they cannot reach the authoritative servers to refresh it; the standard for this suggests keeping such stale data for 1 to 3 days at most (RFC 8767). That makes DNS more resilient when its servers are under attack, and it means an old answer can outlive its TTL.
Steering traffic with DNS
An authoritative server can apply a policy to every query instead of returning one fixed answer. Managed DNS services sell these policies by name; one provider’s documentation lists simple, failover, geolocation, geoproximity, latency-based, IP-based, multivalue and weighted routing (Route 53 Developer Guide). For a design, four of them matter most:
- GeoDNS answers by the location the query appears to come from, so users reach a nearby region.
- Latency-based answers with the region that measured fastest for that part of the network.
- Weighted splits answers by percentages, which is how traffic moves gradually to a new region or a canary.
- Failover answers with a standby only while health checks find the primary unhealthy.
GeoDNS has a blind spot: the authoritative server sees the address of the recursive resolver, not of the user. A user whose resolver is far away, or whose resolver’s address is placed elsewhere, gets an answer for the wrong place. EDNS Client Subnet lets a resolver pass on part of the user’s address, recommended at 24 bits for IPv4 and 56 for IPv6, to get a tailored answer. It has costs: answers must be cached per subnet, and the user’s network becomes visible to every server involved, which is why its own specification recommends shipping it switched off and enabling it only where it clearly helps (RFC 7871).
Health-checked failover adds its own delay. A checker has to see several failed probes in a row before it changes the answer, and only then do the caches start to expire.
Failover by DNS, simulated
Two regions serve users; one fails at the evening peak. How many requests still go to it, and for how long? This model needs only the TTL, the time the health check needs to notice, and a small share of clients that ignore the TTL. Each of those numbers is an assumption, and the model says so:
"""How long requests keep going to a failed region after DNS fails over, for three TTLs. A model: every number at
the top is an assumption, and each one says so."""
DETECT_S = 30 # assumption: the health check probes every 10 s and needs 3 failures in a row
STICKY_SHARE = 0.02 # assumption: 2 % of clients keep an address for at least 30 minutes, whatever the TTL
STICKY_S = 1800
REQUESTS_PER_S = 2000 # assumption: the requests a second that the failed region was serving
RESOLVERS = 20000 # assumption: resolvers that keep the name in their caches all day
TTLS = (30, 300, 3600)
def still_sent(t, ttl):
"""Share of requests still sent to the failed region t seconds after it failed."""
if t < DETECT_S:
return 1.0 # the record still names the failed region
since = t - DETECT_S
# A resolver's cached answer was fetched at a random moment during the last TTL, so after the change the share
# of caches still holding the old answer falls in a straight line to zero over one TTL.
obeying = max(0.0, 1 - since / ttl)
sticky = max(0.0, 1 - since / max(ttl, STICKY_S))
return (1 - STICKY_SHARE) * obeying + STICKY_SHARE * sticky
def label(seconds):
return f"{seconds // 60} min" if seconds % 60 == 0 else f"{seconds} s"
print(f"Share of requests still sent to the failed region ({DETECT_S} s until the health check notices):")
print(f"{'after':>8}" + "".join(f"{'TTL ' + str(ttl) + ' s':>13}" for ttl in TTLS))
for t in (15, 30, 45, 60, 120, 300, 600, 1800, 3600):
print(f"{label(t):>8}" + "".join(f"{still_sent(t, ttl):>13.1%}" for ttl in TTLS))
failed = [sum(still_sent(t, ttl) for t in range(0, 4 * 3600)) * REQUESTS_PER_S for ttl in TTLS]
print()
print(f"{'requests sent to the failed region in all':<44}" + "".join(f"{n:>13,.0f}" for n in failed))
print(f"{'queries a second at the DNS servers':<44}" + "".join(f"{RESOLVERS / ttl:>13,.0f}" for ttl in TTLS)) Output
Share of requests still sent to the failed region (30 s until the health check notices):
after TTL 30 s TTL 300 s TTL 3600 s
15 s 100.0% 100.0% 100.0%
30 s 100.0% 100.0% 100.0%
45 s 51.0% 95.1% 99.6%
1 min 2.0% 90.2% 99.2%
2 min 1.9% 70.5% 97.5%
5 min 1.7% 11.5% 92.5%
10 min 1.4% 1.4% 84.2%
30 min 0.0% 0.0% 50.8%
60 min 0.0% 0.0% 0.8%
requests sent to the failed region in all 126,400 391,000 3,661,000
queries a second at the DNS servers 667 67 6
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 ttl_failover_sim.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
- Nothing moves until the health check notices. For the first 30 seconds every request goes to the failed region, whatever the TTL. Detection time is part of every failover budget.
- A 30-second TTL leaves only the stragglers after a minute: the 2 % of clients that ignore the TTL, which now dominate the tail.
- A 300-second TTL still sends 90 % of requests to the dead region after a minute and 11.5 % after five.
- A one-hour TTL sends half the requests there for half an hour: 3.66 million failed requests at 2,000 a second, against 391,000 at 300 seconds and 126,400 at 30.
- The price of the short TTL is the last line: 667 queries a second at the DNS servers instead of 67 or 6.
So DNS failover is a tool for minutes. When a design needs a region to drop out within seconds, the steering has to happen below DNS: in anycast routing, or in a load balancer or edge that retries a healthy region itself.
DNS Propagation Checker Watch a DNS change spread: compare a record and its remaining TTL across public resolvers in many places.Anycast: one address, many sites
Anycast: one address, announced from two sites
Text description of the diagram
Site A and site B both announce the same IP address, 203.0.113.10, to the Internet's routing system. Users in the west and users in the east send their packets to that one address, and routing delivers each packet to the announcement that is nearest by its own measure: - traffic from the west reaches site A; - traffic from the east reaches site B while it is up.
When site B fails, it withdraws its announcement. Once routers have learned of the withdrawal, traffic from the east also goes to site A. No DNS answer changes, so no cache has to expire.
With anycast, a service gets a stable set of IP addresses, and several independent sites announce reachability to those addresses into the routing system. Routing then makes the choice for each client: packets go to the announcement that is closest by the routing system’s own measures, which is not always the closest in kilometres or the fastest (RFC 4786, RFC 7094). DNS servers are its classic use, including many of the root servers, and the same architecture notes record its growing use by content delivery networks, despite the caveats about connections described below.
Failover needs no DNS change at all. The operational guidance is to tie a site’s route announcement to the health of the service there, so that a failed site withdraws its route and traffic moves to the remaining sites as routing converges. It also warns against announcing and withdrawing in rapid alternation, which routers may punish by ignoring the route for a while (RFC 4786).
The catch is state. Routing picks the site again for every packet, so a connection survives only while routing stays stable for longer than the connection lasts. A DNS query is a single packet each way and does not care. A long TCP download can break if the route changes halfway, because the new site has never heard of the connection; the architecture notes on anycast state plainly that unmodified stateful transports can fail when used with anycast under normal routing changes (RFC 7094). One answer in the guidance is to let anycast handle only the start of an exchange and to hand the long part to a unicast address (RFC 4786).
IP Address Lookup See which network announces an address and its prefix: the routing facts behind anycast.DNS steering or anycast?
| Criterion | DNS steering (GeoDNS, weights, failover) | Anycast |
|---|---|---|
| Who chooses the site | Your authoritative DNS servers, per query | Internet routing, per packet |
| What it knows about the user | The resolver's address, or a client subnet | Nothing: only the routing topology |
| Time to move users after a failure | Detection plus up to one TTL, and a tail that ignores TTLs | Detection plus routing convergence |
| Fine control | Yes: weights, percentages, per-country answers | Little: routing decides |
| What you need | A DNS service with policies and health checks | Your own address space and routing relationships, or a provider that has them |
| Weak spot | Caches keep old answers | Long connections can break when routes change |
| When to choose | Choosing an application region, gradual migrations and canaries | DNS servers, CDN edges and front doors that must fail over fast |
Large systems often use both: an anycast front door, a CDN or a global load balancer, keeps the first hop close and fails over fast, while DNS policies decide which application region sits behind it.
Interview questions
Warm-up (fresher to mid level): what does a DNS TTL control, and what happens when you lower it? The TTL says how many seconds resolvers and clients may keep a record in their caches before they must ask the authoritative servers again. Lowering it makes changes, such as a failover or a move to new servers, reach users sooner, because cached copies expire sooner. The cost is load and latency: every resolver asks again once per TTL, so going from 3,600 to 30 seconds multiplies those queries by 120, and more user lookups miss the cache and wait for round trips. A lower TTL also does not help clients that ignore it, and it takes effect only after the old, longer TTL has run out, so for a planned change it should be lowered at least one old TTL in advance.
Key takeaways
- A recursive resolver walks from the root to the authoritative servers once, then serves the cached answer until its TTL runs out; every link of a CNAME chain and every “no such name” has its own TTL.
- A TTL trades the speed of change against load: in the model, 30 seconds cost 667 queries a second against 6 for an hour.
- DNS failover takes detection time plus up to a TTL, plus a tail of clients that ignore TTLs: minutes, not seconds.
- GeoDNS answers by the resolver’s location; EDNS Client Subnet improves it at a cost in caching and privacy.
- Anycast lets routing pick the nearest site and fails over by withdrawing a route, but long-lived connections need routing to stay stable.
- Use DNS policies to choose regions and move traffic gradually; use anycast where the first hop must stay close and fail over fast.
Exercise
Exercise · Medium · Python
Choose the longest TTL that still meets a failover goal
A failover goal reads like this: "deadline_s seconds after a region fails, at most target of the clients may still be sent to it." A long TTL keeps load on the DNS servers low, so you want the longest TTL that still meets the goal.
Use this lesson's model for clients that respect the TTL. Nothing changes until the health check notices the failure, detect_s seconds after it happens. From then on, the share of clients still holding the old answer falls in a straight line from 1 to 0 over one TTL: t seconds after the failure it is 1 - (t - detect_s) / ttl.
Write longest_ttl(deadline_s, detect_s, target) in ttl.py. Return the longest whole number of seconds that meets the goal. Raise ValueError when no TTL can meet it, because the deadline is not later than the detection time, and when target is not at least 0 and below 1.
longest_ttl(300, 30, 0.05) # 284: after 5 minutes at most 5 % of clients on the failed region
The sample tests import longest_ttl from ttl.py and run in your browser.
Starter code · ttl.py
def longest_ttl(deadline_s, detect_s, target):
"""The longest TTL in whole seconds that leaves at most `target` of clients on a failed region at the deadline."""
# Replace this line with your code.
return 0 The sample tests · test_ttl.py
from ttl import longest_ttl
def raises_value_error(call):
try:
call()
except ValueError:
return True
return False
def test_five_minutes_five_percent():
"""after 5 minutes, at most 5 % of clients on the failed region"""
assert longest_ttl(300, 30, 0.05) == 284
def test_everyone_moved():
"""with a target of 0, every cache must have expired by the deadline"""
assert longest_ttl(120, 30, 0.0) == 90
def test_half_after_an_hour():
"""a lenient goal allows a long TTL"""
assert longest_ttl(3600, 60, 0.5) == 7080
def test_impossible_goals():
"""raises ValueError when no TTL can meet the goal, or the target is out of range"""
assert raises_value_error(lambda: longest_ttl(20, 30, 0.05))
assert raises_value_error(lambda: longest_ttl(30, 30, 0.05))
assert raises_value_error(lambda: longest_ttl(300, 30, 1.0))
assert raises_value_error(lambda: longest_ttl(300, 30, -0.1)) A hint
Write the goal as an inequality and solve it for the TTL: 1 - (deadline_s - detect_s) / ttl <= target. Moving ttl to one side gives ttl <= (deadline_s - detect_s) / (1 - target). The longest whole number of seconds that fits is the floor of the right-hand side (math.floor).
Results of the sample tests
| Test | Result | Details |
|---|
What your code printed
The sample tests run on this device, in your browser (Pyodide): nothing is sent to mysmartcopilot.com. The first run downloads Python (about 13.5 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.
Check yourself
7 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.
References
- RFC 1034: Domain Names, Concepts and Facilities (IETF)
- RFC 1035: Domain Names, Implementation and Specification (IETF)
- RFC 2308: Negative Caching of DNS Queries (DNS NCACHE) (IETF)
- RFC 8767: Serving Stale Data to Improve DNS Resiliency (IETF)
- RFC 7871: Client Subnet in DNS Queries (IETF)
- RFC 4786: Operation of Anycast Services (IETF)
- RFC 7094: Architectural Considerations of IP Anycast (IAB)
- Choosing a routing policy (Amazon Route 53 Developer Guide) (Amazon Web Services)
- RFC 6761: Special-Use Domain Names (IETF)
Related tools
Report a problem with this lesson
Kept only in this browser. Your Learn progress