+ red-button to turn off the pathA policy row with blank fields hits a null pointer, and Google Cloud crashes in every region at once — and Cloudflare falls with it, then calls Google "a third-party provider". June 12, 2025, 17:49 UTC. Google's report: Service Control, the binary that approves every API request, in a crash loop for three hours. Cloudflare's, the same evening: Workers KV, ninety percent failing, two hours twenty-eight. How it happened, why one blank field went global in seconds, and who gets the blame.
May 29. Service Control gets a new quota check, shipped region by region with a red button — but the new code path never runs during rollout; nothing triggers it yet. No null check. No feature flag. 17:45. A policy change with blank fields lands in the Spanner table. Quota is global, so the row replicates everywhere within seconds. Every Service Control binary reads it, hits the null pointer, crashes, restarts, reads it again. Every API call: 503. Google is fast: triage in two minutes, root cause in ten. The red button is out in forty, and small regions recover first. The status page posts its first update an hour in, because it runs on Google Cloud.
17:52. Cloudflare's WARP team sees new devices fail to register. Workers KV, the store half of Cloudflare uses for config and identity, lives on a third-party cloud. Access fails every login — by design, it fails closed. Workers AI fails every inference. 19:11, Gergely Orosz: two independent clouds down at once, never seen before. 19:32, Google support tells a user there are no known disruptions — try clearing your cookies. Everywhere else recovers by 19:48. us-central1 doesn't: restarting tasks stampede the same Spanner table, no randomized backoff, so Google throttles them by hand. Two hours forty.
One: the policy check sits inside the request path; when the check dies, the API dies. Two: quota data goes global with no staging; a bad row is a global row. Three: Cloudflare knew. Workers KV was mid-migration to its own R2, down to one provider — the postmortem calls it a gap in coverage. git blame. Google, fifty-five percent: no null check, no feature flag, no backoff — their report says a flag would have caught it in staging. Global replication, twenty: one row, every region, seconds. Cloudflare, twenty: half a product line on one vendor's store, and a postmortem that never says Google. The status page, five, for living on what it reports on. Blast radius: seventy Google Cloud products, ten Workspace apps, every region. Three hours. At Cloudflare: Access, WARP, Workers AI, Pages, Stream — two and a half. Hacker News, fourteen hundred points; top comment: the status page is green.
Verdict, postmortem: ship it. Both postmortems land within thirty hours, both blame themselves, and Google's fixes are concrete: fail open, flags off by default, exponential backoff. The Monday line: new code behind a flag that ships off, and a null check where the data comes in. Send me the incident you're still not allowed to talk about, in the comments, or at the daily diff dot dev.
Verdict: SHIP IT — two postmortems in 30 h, both self-blaming · fixes: fail open, flags off by default, backoff
Sources
https://status.cloud.google.com/incidents/ow5i3PPK96RduMcb1SsW
https://blog.cloudflare.com/cloudflare-service-outage-june-12-2025/
https://www.cloudflarestatus.com/incidents/25r9t0vz99rp
https://x.com/Google/status/1933246051512644069
https://x.com/GergelyOrosz/status/1933240698716729511
https://x.com/ThomasOrTK/status/1933337436970709493
https://news.ycombinator.com/item?id=44260810
https://news.ycombinator.com/item?id=44261064
https://news.ycombinator.com/item?id=44274563
https://news.ycombinator.com/item?id=44264580
https://forgecode.dev/blog/gcp-cloudflare-anthropic-outage/
And that's the diff for today. I'm Niko from Axrisi. Merge responsibly.
YouTube · thedailydiff.dev · forward this to the intern who deployed on Friday.

