Skip to main content
Network

Why your DNS change is live for some people and not for you

A customer says the new site is down. You check on your phone: it loads. They check again: still the old certificate warning. Nobody is lying. DNS propagation is not a broadcast, it is a stack of independent caches that expire on their own schedules, and understanding the stack turns "wait 48 hours" from folklore into a prediction.

The packet takes its own path

When you update a record at your registrar, the change lands on the authoritative nameservers within minutes. That part is fast. What is slow is every resolver between a user and those nameservers realizing their cached answer has gone stale.

A typical chain looks like this: your laptop asks your home router, the router asks the ISP's resolver, that resolver may ask a forwarding provider, and only that final layer asks your authoritative server. Each hop caches your record independently. Some are configured to respect the TTL you set. Some ignore it and use a minimum of their own. A resolver that cached your record 30 seconds before your change will happily keep serving it for the full TTL of the old value.

This is also why two people at the same desk can see different answers. Your browser might be using a DNS-over-HTTPS service while your colleague's uses the ISP resolver. Different resolvers, different cache clocks, different views of the same record.

TTL is a promise others may not keep

The TTL on a DNS record says "you may reuse this answer for this many seconds." It is a suggestion that well-behaved resolvers follow. Three things make real propagation longer than the TTL:

  • Negative caching. If you queried a name that did not exist, that NXDOMAIN answer is also cached, sometimes for the full negative TTL. Add the record later and users who looked it up during the gap still see nothing until that cache expires. A DNS propagation checker shows you per-resolver whether the answer is old, new, or NXDOMAIN, which is the fastest way to tell a cache problem from a configuration problem.
  • Minimum-value overrides. Some resolvers clamp TTLs to their own floor, so a 60-second record can live in memory for 5 minutes regardless.
  • Your own OS and browser. Local stub caches and internal browser DNS caches refresh on their own timers. Restarting the browser clears one of them; flushing the OS cache clears another; neither touches the ISP.

The practical rule: when you know a record will change, lower its TTL to 300 seconds a day in advance, make the change, verify with a checker from several networks, then raise the TTL back for stability.

CNAME, A, or ALIAS: picking under pressure

Half the confusion around migrations is record-type choice. An A record points a name at an IP. A CNAME points it at another name, and the name is then resolved separately. A hostname at the zone apex cannot legally carry a CNAME, which is why registrars invented vendor-specific "ALIAS" or "flatten" types for the root.

Two failure modes eat whole evenings:

  • A record pointing at an IP that a load balancer later changed. The site dies for everyone at once, which at least makes it obvious.
  • A CNAME chain two or three links long. Every link adds a resolver round trip and its own cache clock, so propagation gets slower and harder to reason about. Keep chains to one hop.

Checking without fooling yourself

Testing from your own machine mostly tests your own cache. To see what a random visitor would see, query resolvers directly, from outside your network:

dig @8.8.8.8 example.com A +short dig @1.1.1.1 example.com A +short

Different answers at different resolvers are the normal mid-migration state, and they converge as each TTL expires. If every public resolver agrees but your office still sees the old value, the culprit is a corporate resolver that clamps TTLs hard; escalation there is an IT tickets problem, not a DNS one.

And when a lookup returns NXDOMAIN for a name you just created, do not add the record twice. The gap between creating it and first query got cached as "does not exist." It clears on the negative TTL, then everything works, and the record you panicked and duplicated will be a cleanup task.

The part people forget

Mail is DNS with a longer memory. MX records, SPF, DKIM selectors, and DMARC all resolve through the same caching chain, and a mail provider migration means old MX answers keep flowing for hours. Moving mail without dropping TTLs in advance is how a Friday migration becomes a lost-Monday-morning incident. A DNS lookup of every record type in the zone before touching anything costs two minutes and catches the CNAME-under-apex mistake, the stale selector, and the MX pointing at decommissioned servers while nothing is actually broken yet.

Caching is doing exactly what it was designed to do. It just has no idea that today, of all days, you are in a hurry.