- Jan 31, 2017, 23:27 UTC: rm -Rvf on the PRIMARY's data directory; ~300 GB removed, 4.5 GB left
- 5 of 5 backups fail: empty S3 bucket, no Azure snapshots on the DB, wiped replica, daily copy without webhooks
+ Feb 1, 18:00 UTC: GitLab.com back from a 6-hour-old manual snapshot; live doc, live stream, blameless postmortemAn engineer at GitLab runs rm -rf on the wrong database server, and three hundred gigabytes of GitLab.com vanish in a second or two, about how long it takes to read a hostname. January 31st, 2017, 11:27 p.m. UTC. GitLab tweets that it accidentally deleted production data, opens its incident notes to the internet, and streams the recovery on YouTube, the number two live stream on the platform. Next day, in writing: out of five backup techniques, none are working reliably. How it happens, why it is possible, and who actually gets the blame.
5:20 p.m.: an engineer snapshots production to test a load balancer in staging. 7 p.m.: spam hammers the database, plus a job hard-deleting a GitLab employee a troll reported for abuse. 11 p.m.: the replica falls so far behind that the primary has already discarded the log it needs; the only fix is to wipe the replica and copy the primary again. pg_basebackup hangs with no output. It is actually waiting, silently, for the primary; nobody knows that, and the runbook does not say. The engineer, who meant to sign off at eleven, decides the empty data directory is the problem and removes it. On db1. The primary. He notices a second or two later; of roughly three hundred gigabytes, 4.5 remain. The backups. One: pg_dump to S3, daily. The bucket is empty. The cron job runs on an app server with no database, so the package picks PostgreSQL 9.2 binaries for a 9.6 database, fails, and emails the failure, which bounces for missing DMARC. Two: Azure disk snapshots, enabled for the file servers, not the databases. Three: the replica, wiped on purpose an hour ago. Four: the daily snapshot, 24 hours old, every webhook stripped out by the staging sync. Five: the manual snapshot from 5:20, for an unrelated test. That one wins.
Restoring means copying the staging disk back to production over Azure's cheap storage at sixty megabits per second: eighteen hours. GitLab.com is back February 1st at six p.m. UTC, six hours of data older. git blame: two hostnames one character apart, and five backup systems nobody has ever restored from. Not the engineer. The postmortem, signed by the CEO, keeps him anonymous, colours the production prompt red, and gives data durability an owner, because until now it had none. Blast radius: eighteen hours down, six hours of data gone, roughly five thousand projects, five thousand comments, seven hundred new users, and five thousand people watching a progress bar. Hacker News gives the live doc 1,162 points and quotes one line back at them: out of five backups, none.
Verdict, postmortem: ship it, on the response. They run the incident in public, blame the process, and publish the fix list with issue numbers. Monday action: restore a backup. If you have never restored it, you do not have one. Send me the incident you are still not allowed to talk about, in the comments, or at thedailydiff.dev.
Verdict: SHIP IT — the response: live doc, live stream, blameless postmortem with issue numbers
Sources
GitLab, "Postmortem of database outage of January 31" (Feb 10, 2017) — https://about.gitlab.com/blog/postmortem-of-database-outage-of-january-31/
GitLab, "GitLab.com database incident" (Feb 1, 2017) — https://about.gitlab.com/blog/gitlab-dot-com-database-incident/
The live incident doc (Wayback copy) — http://web.archive.org/web/20170202034719/https://docs.google.com/document/d/1GCK53YDcBWQveod9kfzW-VCxIABGiryG7_z_6jHdVik/pub
@gitlabstatus: "We accidentally deleted production data…" — https://twitter.com/gitlabstatus/status/826591961444384768
@gitlabstatus: emergency maintenance notice — https://twitter.com/gitlabstatus/status/826572933304827904
Hacker News, "GitLab Database Incident – Live Report" (1,162 points) — https://news.ycombinator.com/item?id=13537052
Hacker News, the postmortem thread (377 points) — https://news.ycombinator.com/item?id=13619714
And that's the diff for today. I'm Niko from Axrisi. Merge responsibly.
YouTube · thedailydiff.dev · forward this to the intern who deployed on Friday.

