+ livenerf: Opus 5.5 on a 30-day clock
- EU data centers keep power use secret
+ Backblaze: 354k drives, 1.73% AFREveryone knows Claude gets dumber a few weeks after launch. So one developer put Opus 5.5 on a clock from launch week, and the clock says nobody could have known yet. And one line in the repo's credits made me laugh out loud. I'll get to it at the end. Three stories today, and the big one is a GitHub repo called livenerf. For months, people have said Anthropic quietly nerfs its models after launch. The usual theories are a squeezed model, a smaller model under the same name or simply less thinking. The repo calls every argument so far vibes versus vibes.
Opus 5.5 shipped on September 22, and a week ago I told you it topped the independent board while writing three times more words than the median model. Two and a half days in, a developer called ninjahawk started the clock, once a day for thirty days, with logs nobody may edit. The Hacker News thread hit almost eight hundred points and split like a team chat. One commenter says nerfing models isn't real in the vast majority of reported cases. Another says people are just getting used to the new level of intelligence. And one Claude Code power user now dismisses the rate-this-session pop-up every time, because rating Claude seemed to make it worse. Peer review. So, question one, how do you catch a nerf? You start with about 2,300 hard exam questions, and Opus gets 93% right on the first try, which is useless for a detector, since a question it always gets right can't drop. 97% were always right or always wrong, and the 78 it only sometimes gets right became the panel.
Same prompt, no tools, and exact-match grading with no AI judge, because the judge would drift too. It runs on a Max subscription through headless Claude Code, since the API version was priced at about $1,600 a month. And the Claude Code version is pinned, because, in the repo's words, a changed harness looks exactly like a changed model. The statistics come from a paper called Adding Error Bars to Evals, by Evan Miller at Anthropic. So Anthropic's own math is now auditing Anthropic, which is the academic version of quoting the terms of service back at customer support. Question two, what can it see? Per ten-day window, a drop of about 7.5 points. To prove the rig works, the author turned Claude's effort down on purpose. Medium effort cut the output by about a quarter, and the score by only four points. Low effort cut the output by almost two thirds, for eight points.
That's the part to steal. Less thinking shows up in the token count before it shows up in the answers. So pin your model version, log output tokens per task, and keep a few frozen questions of your own. If your agent suddenly writes shorter, check your logs before you check Reddit. And here's the honest part. Quietly swapping in the older Opus five was too small a change for the rig to call, and the readme says so up front. The audit also found eight answer keys that look wrong, so the nerf detector found bugs in the benchmark before it found a nerf. Question three, what does Anthropic say? Last year, after a wave of complaints, Anthropic published a postmortem. It says, quote, we never reduce model quality due to demand, time of day or server load. The same post admits three bugs had degraded Claude, and almost a third of Claude Code users hit the wrong servers at least once.
So the complaints were real and the cause was bugs, which is why the repo warns that launch week could be the worst week. Six of thirty days are in. The first possible call lands around October 24, so the honest answer is that nobody knows, including everyone who was sure. Speaking of numbers nobody publishes. Lighthouse Reports and the Dutch paper Trouw found that the vast majority of European data centers keep their power and water use secret, although an EU rule has required reporting for three years. In the Netherlands, fewer than a quarter publish. One Microsoft data center alone uses about one percent of Dutch electricity. Meanwhile, Backblaze did the boring thing again. Its quarterly drive stats cover about 350,000 hard drives, and the failure rate rose to 1.73%, the highest in a while. Three Seagate models had zero failures. That's what a public baseline looks like, quarter after quarter, for thirteen years.
Now, that credits line. The readme admits a lot of the repo was written with the help of Claude, which is the model being measured. So Claude helped build the meter that checks whether Claude got dumber. That's why the graders are plain functions. In the author's words, you shouldn't have to trust the author, human or otherwise.
Verdict: SHIP IT — a timestamp, a public rulebook, a stated blind spot
Sources
https://github.com/ninjahawk/livenerf
https://news.ycombinator.com/item?id=49901736
https://github.com/ninjahawk/livenerf/blob/main/README.md
https://github.com/ninjahawk/livenerf/blob/main/docs/VALIDATION.md
https://github.com/ninjahawk/livenerf/blob/main/PLAN.md
https://inspect.aisi.org.uk/
https://arxiv.org/abs/2411.00640
https://news.ycombinator.com/item?id=49902336
https://news.ycombinator.com/item?id=49902182
https://news.ycombinator.com/item?id=49902054
https://news.ycombinator.com/item?id=49905447
https://www.anthropic.com/engineering/a-postmortem-of-three-recent-issues
https://nltimes.nl/2026/09/30/data-centers-refusing-say-much-water-electricity-use
https://news.ycombinator.com/item?id=49907057
https://www.backblaze.com/blog/backblaze-drive-stats-for-q2-2026/
https://news.ycombinator.com/item?id=49893002
And that's the diff for today. I'm Niko from Axrisi. Merge responsibly.
YouTube · thedailydiff.dev · forward this to the intern who deployed on Friday.

