- 13:31–13:42 UTC, Jul 2, 2019: PR merged, CI green (no CPU test), one WAF rule pushed to 180+ cities in seconds — no staging
- 13:45–14:07: CPU 100 % on every core, traffic −82 %, 502s everywhere; the kill switch sits behind Cloudflare Access, which is behind Cloudflare
- 14:07–14:09: global WAF terminate; traffic normal after 27 minutes; 14:52 the WAF is back minus one rule
+ Jul 2 (15:50 UTC) and Jul 12: same-day note, then the full postmortem — CPU guard re-added, 3,868 rules re-read, staged rollouts, a linear-time regex engineOne regular expression goes live on every Cloudflare server at once, and for the next twenty-seven minutes the sites behind it are a 502 page, which, for a firewall, is the strictest possible setting. July 2nd, 2019, 13:42 UTC. Cloudflare posts within two hours: not an attack, a bad deploy, traffic down 82 percent. Ten days later CTO John Graham-Cumming publishes the full postmortem, regex included, and Hacker News gives it 698 points, which for an outage is a standing ovation. How it happens, why it is possible, and who actually gets the blame.
Timeline. 13:31. A pull request merges: one new firewall rule against cross-site scripting, in simulate mode, so it blocks nothing. 13:37, tests pass; none measures CPU. 13:42, the rule ships to 180 cities in two seconds, because WAF rules skip the DOG, PIG and Canary stages other releases get. 13:45, the first page. 13:49, Hacker News has a thread on the status page, which still says all systems operational.
Mechanism. The rule ends in .*(?:.*=.*) — .*.*=.*: anything, then anything, then an equals sign. PCRE guesses greedily, fails, and backtracks through every other split. x=x takes 23 steps. Twenty x's after the equals: 555. Twenty x's, no equals sign: 4,067 steps to find nothing. Run that on every request and every core is at one hundred percent, doing nothing, thoroughly.
Two guards should have caught it. The CPU limit on rules was removed by mistake weeks earlier, in a refactor meant to make the WAF use less CPU. And the procedure lets any rule skip staging, because rules exist to stop live attacks; this one was not an emergency, and it went global anyway.
14:00, the WAF is identified; no attack. 14:02, someone proposes the global terminate: one component, off, worldwide. The switch is behind Cloudflare Access. Cloudflare Access is behind Cloudflare. Some credentials have expired from disuse, so the fastest network on the internet spends five minutes on a bypass nobody drilled. 14:07, kill. 14:09, traffic normal.
git blame: a rollout with one speed, global; a CPU guard refactored away by accident; a regex engine with no upper bound. Not the engineer who wrote the rule: the postmortem lists eleven causes and names nobody.
Blast radius: 27 minutes, 82 percent of traffic, 100 percent CPU on every core, in every city. The dashboard and API sit behind the same edge, so customers cannot even switch it off. Hacker News, under the postmortem: they had one problem, used a regular expression, now they have two. Old joke. Still compiles.
Verdict, postmortem: ship it. The CPU guard is back, all 3,868 rules get read by hand, rules go through staging, and the engine moves to one with linear-time guarantees, published by Ken Thompson in 1968. Monday: no .*.* in anything that runs per request, and keep the kill switch off the thing it kills. Send me the incident you are still not allowed to talk about, in the comments, or at thedailydiff.dev.
Verdict: SHIP IT — guard back · 3,868 rules re-read · staged rollouts · linear-time engine
Primary sources
John Graham-Cumming, "Details of the Cloudflare outage on July 2, 2019" (Jul 12, 2019) — https://blog.cloudflare.com/details-of-the-cloudflare-outage-on-july-2-2019/
Matthew Prince, "Cloudflare outage caused by bad software deploy (updated)" (Jul 2, 2019) — https://blog.cloudflare.com/cloudflare-outage/
Matthew Prince on X, Jul 2, 2019, 14:22 UTC — https://x.com/eastdakota/status/1146061591143538688
Matthew Prince on X, Jul 2, 2019, 14:36 UTC ("No evidence yet attack related") — https://x.com/eastdakota/status/1146065231270907907
Press and reactions
Hacker News, Jul 2, 2019, 13:49 UTC — "Cloudflare Network Performance Issues" (631 points, 310 comments) — https://news.ycombinator.com/item?id=20334924
Hacker News, Jul 2, 2019 — "Cloudflare outage caused by bad software deploy" (348 points) — https://news.ycombinator.com/item?id=20336332
Hacker News, Jul 12, 2019 — "Details of the Cloudflare outage on July 2, 2019" (698 points, 149 comments) — https://news.ycombinator.com/item?id=20421538
TechCrunch, Jul 2, 2019 — https://techcrunch.com/2019/07/02/a-cloudflare-outage-is-impacting-sites-everywhere/
And that's the diff for today. I'm Niko from Axrisi. Merge responsibly.
YouTube · thedailydiff.dev · forward this to whoever owns your WAF rules.

