SSL Certificate Expiry: The Outage Nobody Watches For
Certificate expiry is the most predictable outage in the industry. The end date is visible months in advance, printed inside the certificate itself. Yet every year, major services go dark because nobody looked. Microsoft Teams, in February 2020. Ericsson's network in the UK, in December 2018 — 33 million O2 customers offline. Spotify, in 2020. All expired certificates. The fix takes four minutes. The detection takes one scheduled check.
Why renewals still fail
The failure is almost never the renewal itself. It is the chain of assumptions around it:
- The cert lives on one box, the knowledge lives in a person. The engineer who set up Let's Encrypt three years ago left. The auto-renew cron died with their account, and nobody noticed because renewals only fail loudly twice a year.
- Everything is behind a CDN. Your origin cert can expire while the edge serves a valid one — until the CDN cache rules change or someone bypasses the edge in a debug session.
- Multiple certs per domain. The wildcard covers
*.example.com, but the load balancer in front ofapi.example.comholds its own leaf cert with its own expiry date. Checking the public endpoint validates one of them. - Test and staging environments. They expire, nobody cares, and then a monitoring integration or a mobile app's staging build starts failing TLS handshakes. Staging outages regularly leak into production incidents through shared dependencies.
Auto-renewal reduces the work but not the risk. Let's Encrypt rate limits, failed ACME challenges after a DNS change, and port 80 blocked by a new firewall rule all silently stop renewals. An automation that succeeds 200 times and fails once on the 201st looks identical to one that always worked — until the expiry date.
What actually breaks at expiry
The blast radius is wider than "the site shows a warning":
- Browsers hard-block. Chrome and Firefox show a full-page interstitial, and users must click through multiple warnings. Most leave instead.
- Mobile apps break harder. Native HTTP stacks reject expired certs with no user-facing bypass. If your app talks to
api.example.com, an expired cert means an app-wide outage until users update — except there is no update to ship, because the app is fine and the server is not. - Machine traffic fails first. Webhooks stop delivering, payment providers reject callbacks, SSO logins break. API clients rarely show warnings; they just fail. Often the browser still works while half your integrations are down, which makes the incident look like a mysterious partial outage.
- Monitoring is blind too. If your monitoring agents validate TLS like normal clients do, they fail at the same moment — so the alert that was supposed to catch the outage is itself failing to report.
Set a deadline that precedes the real one
Treat expiry like a hard deadline with a buffer. A workable policy:
- Alert at 30 days (time to renew normally, without urgency),
- Alert again at 14 days (time to notice the auto-renewal failed and fix it by hand),
- Page someone at 7 days (renewal now involves debugging why automation broke).
Run the check against every public endpoint, not just the apex: example.com, www, api, mail, staging, and every load balancer or CDN origin that terminates TLS. You can inspect any endpoint's current certificate chain and expiry with a SSL certificate checker — it shows the dates, issuer, and remaining validity in one view, which is exactly what you want during an incident triage.
Renewal is not the whole job
A renewed cert can still produce an outage. After every renewal, verify three things that fail silently:
- Chain completeness. A missing intermediate certificate works in browsers that cached it and fails on mobile devices and API clients that did not. This is the classic "works on my laptop" TLS bug. Curl against the endpoint and look for
unable to get local issuer certificate. - The full chain was deployed. Tools that copy only the leaf cert leave the intermediates stale. The endpoint then serves a valid cert nobody can validate.
- Old protocols and ciphers. A renewal sometimes rides along with a server config change. If TLS 1.0 got disabled at the same time, old clients break and the expiry gets the blame.
Check the response headers while you are there — Strict-Transport-Security max-age matters, because HSTS pinning means a broken cert cannot even be bypassed by clicking through the warning. An HTTP headers inspector shows the security headers on the live endpoint in a single request.
Automation with a human backstop
The setup that actually survives staff turnover has three layers:
- Auto-renewal (certbot, Caddy, or your platform's managed TLS) — the workhorse, trusted but never assumed to be working.
- Independent expiry monitoring that does not depend on the same credentials or network path as the renewal. An external check is the honest one: it sees what the world sees. A simple daily check per endpoint is enough.
- A runbook. One page: where the cert lives, which command renews it, which services need a reload, who to call. Write it during a calm day; an incident is the worst possible time to discover the renewal runs on a server nobody remembers.
The monitoring layer must be separate from the renewal layer. When both run on the same server, the same disk-full event kills the renewal and the alert. A cheap external probe — even just an HTTP reachability check on endpoint availability plus a scheduled cert date check — breaks the dependency.
A five-minute audit you can do today
List every domain your users and machines touch, including staging. Check each cert's not-after date and the chain. Write the dates down with owners' names. If any cert expires within 30 days, renew it now, not at the next deploy. Add the 30/14/7-day alerting. That list is the difference between a four-minute renewal and an incident retro.
Expiry dates are known years in advance. There is no excuse for an unwatched one — only the absence of a scheduled check.