- 15:40 UTC · one backbone command, every link down · audit tool: bug
- name servers withdraw their own BGP routes, by design · SERVFAIL worldwide
- ~6 hours · tools, out-of-band and badges on the same network
+ git blame: Facebook 55 · DNS design 25 · one network 15 · the racks 5Facebook removes its own address from the internet, and when its engineers arrive to put it back, the badge readers are down too, because they run on Facebook. Meta's own postmortem, next day: a routine maintenance command takes down every backbone link, and the tool built to block commands like that has a bug. Cloudflare, from outside: at 15:40 UTC the routes to Facebook's name servers vanish, and dig facebook.com returns SERVFAIL everywhere on the planet.
How it happens, why a safety feature makes it worse, and who gets the blame. Early afternoon, UTC. A command meant to check free backbone capacity instead takes down every link between Facebook's data centres, and the audit tool built to catch exactly this waves it through. 15:40. Cloudflare's BGP feed fills with withdrawals; Facebook's DNS prefixes leave the routing table. Within a minute Cloudflare's engineers are in a room wondering whether they broke 1.1.1.1.
15:45, Hacker News, twenty-six hundred points. 16:07, Facebook's spokesman, on Twitter of all places: some people are having trouble accessing our apps. Some people is three and a half billion. 18:51, the New York Times: employees locked out of their buildings, badges dead. 19:52, the CTO apologises. 21:00, five hours in, routes return; 21:20, facebook.com resolves. Why does a backbone fault delete Facebook from the internet? Its name servers have a rule: can't reach the data centres, declare yourself unhealthy, stop advertising your address. Sensible when one site goes bad.
When the whole backbone goes, every name server does it at once. The servers are up. Nobody can find them. A safety feature turns a network fault into an existence fault. Then the internet piles on: apps retry, users reload, and Cloudflare's resolver sees thirty times its normal load, for a site that isn't there. And the doors. Tools, out-of-band access, badges: same network. Engineers drive to the data centres, can't get in, then meet racks hardened against anyone with physical access. Built to slow an attacker, it slows the owner too.
The backbone returns; they ramp traffic slowly, on purpose. Each data centre has dropped tens of megawatts, and reversing that in one step is a second outage. git blame. Facebook, fifty-five percent: the command was audited, the auditor had the bug, one command reached every router. The DNS design, twenty-five: a health check with no sense of proportion. One network, fifteen: tools, out-of-band, badges. The racks, five, for doing their job. Blast radius: three and a half billion users, nearly six hours. The stock closes down five percent, six billion off Zuckerberg, on paper. Telegram claims seventy million sign-ups that day.
Verdict, postmortem: needs review. Same-evening statement, named-author postmortem within a day, audit-tool bug admitted in plain English. The fix list: strengthen testing and drills. Nothing about taking the doors off the production network. Monday: put your out-of-band access, and your door, on a network you don't operate. Send me the incident you're still not allowed to talk about, in the comments, or at the daily diff dot dev.
Verdict: NEEDS REVIEW — honest postmortem in 24 h · fix list: "strengthen testing and drills" · doors still on the backbone
Sources
Primary
Meta Engineering, "More details about the October 4 outage" (Santosh Janardhan, Oct 5, 2021): https://engineering.fb.com/2021/10/05/networking-traffic/outage-details/
Meta Engineering, first note, Oct 4, 2021: https://engineering.fb.com/2021/10/04/networking-traffic/outage/
Cloudflare, "Understanding how Facebook disappeared from the Internet": https://blog.cloudflare.com/october-2021-facebook-outage/
Mark Zuckerberg, note to employees, Oct 5, 2021: https://www.facebook.com/zuck/posts/10113961365418581
Posts shown
Andy Stone (Facebook), 16:07 UTC: https://twitter.com/andymstone/status/1445058088436908045
Sheera Frenkel (NYT), the badges, 18:51 UTC: https://twitter.com/sheeraf/status/1445099150316503057
Mike Schroepfer (CTO), 19:52 UTC: https://twitter.com/schrep/status/1445114730151043073
Press
Hacker News, 2,589 points: https://news.ycombinator.com/item?id=28748203
The New York Times, "Gone in Minutes, Out for Hours": https://www.nytimes.com/2021/10/04/technology/facebook-down.html
Forbes, "Zuckerberg Loses $5.9B in a Day" (estimate): https://www.forbes.com/sites/abrambrown/2021/10/04/zuckerberg-net-worth-billionaire-facebook-stock-outage/
Reuters, Telegram's 70 million sign-ups (Durov's figure): https://www.reuters.com/technology/telegram-founder-says-over-70-mln-new-users-joined-during-facebook-outage-2021-10-05/
Krebs on Security: https://krebsonsecurity.com/2021/10/what-happened-to-facebook-instagram-whatsapp/
And that's the diff for today. I'm Niko from Axrisi. Merge responsibly.
YouTube · thedailydiff.dev · forward this to the intern who deployed on Friday.

