- Armin Ronacher: 35 h of GPT-6 Astra · $1,200 · 79 commits · nothing shipped
- Quesma: RTK (79k stars) reports 89 % tokens saved, bills go up
+ Cognition SWE-2: Kimi K3 post-trained · TB 2.1 92.8 vs TB 4 27.3
+ OpenAI Agents API: the Codex harness as a service · $200 Pro sign-ups paused
- Ask HN: can we please limit the AI news flood? · 656 pointsArmin Ronacher, the man who wrote Flask, gave GPT-6 Astra a single prompt and left it alone for thirty-five hours. It came back with seventy-five thousand lines of code and, in his words, absolutely nothing of value, which makes it the most realistic software factory ever built. He posted the write-up on Wednesday. On Thursday, Cognition shipped SWE-2, and OpenAI shipped an Agents API and paused sign-ups for its two-hundred-dollar plan, because Astra ate the servers. And on Friday the write-up reached the front page of Hacker News, a few rows below a post asking Hacker News to please stop posting about AI. The factory. One prompt: build a Python with virtual threads and lexical scoping. Astra kept its own notes, spun up its own subagents, and ran until Armin pulled the plug, a day and a half and about twelve hundred dollars later. Seventy-nine commits, so fifteen and a half dollars per commit, which is roughly what a human charges, except the human eventually stops.
The code is the story. Instead of the edit tool, Astra patches C files by piping Python heredocs that read the file, string-replace a function in, and write it back. To run one Windows test it wrote Python that spawned Node that spawned PowerShell, a call stack with a passport. The unit tests it committed have no whitespace at all, because without indentation they're ten percent cheaper in tokens. And the task list drifted from one, two, three to task 8b2c2b2b checkpoint one, which is how you know the intern has been alone too long. Armin's theory: the model is rewarded for finishing long tasks and barely punished for ugly code, so token-golfed tool calls leak into the codebase, and the fewer humans look, the less it matters. He titled that section It's AGI If You Don't Look. Hacker News added the economics: more code means more tokens to maintain it, which makes slop not a bug but a subscription. Can you measure slop? Armin's own company tried this week. Earendil ran the SlopCodeBench numbers: agent code is about twice as verbose and twice as eroded as the human repos it's compared against, and when the benchmark wipes the agent's memory between checkpoints, the strict pass rate for every model they tested is zero. The most reliable slop metric they found was the raw line count, which stops working the moment anyone optimizes for it, so please don't tell the models.
Then the wallet. RTK has seventy-nine thousand GitHub stars and YouTube thumbnails promising to halve your Claude Code bill; it compresses terminal output before the agent reads it. Quesma ran it on Terminal-Bench for fifteen hundred dollars of tokens. With Fable the bill fell five percent, almost all from one task; with DeepSeek the average task got seventeen percent more expensive. Meanwhile RTK's own counter reported eighty-nine percent saved, because it counts bytes removed, not extra turns caused: a smart meter that bills the neighbour. And one version bug rewrote find into a command that failed, three hundred thirty-nine times in a row, at nine times the price. The tool that kills tokens has a loop. The flood itself. Cognition's SWE-2 is Kimi K3, the open Chinese model, post-trained until it matches Fable 5.1 on Cognition's own benchmark at a third of the price; Hacker News: why are all American AI models basically Kimi in a trench coat. On Terminal-Bench 2.1 it tops the entire table; on Terminal-Bench 4, the one nobody has memorised yet, it scores twenty-seven to Fable's fifty-six, so on the slide-deck benchmarks the best column is the exam everyone already passed. Cognition also announced Devin Voice, your favourite AI software engineer just got a landline, because the one thing agentic coding was missing was hold music. OpenAI, same day: the Agents API, which is the Codex harness sold as a service. OpenAI runs the sandbox, US data residency only, no zero-data-retention, so the agent remembers your codebase exactly as well as ChatGPT does. Then OpenAI paused sign-ups for the two-hundred-dollar Pro plan, because Astra demand is, quote, unprecedented: the first product with a waitlist to pay more. Armin burned a full reset on his factory. He is the demand.
Which brings us to Friday's Ask HN: can we please limit the AI news flood, six hundred fifty-six points, because, quote, every time someone at OpenAI or Anthropic farts, there's a top-ten post. The replies were filters: hcker dot news removes AI stories with a BERT classifier trained on eight thousand examples, AI to delete AI, grass-fed; and a uBlock regex matching A-I case-insensitively also deletes email, daily, main, domain and train, so the regex, as always, is the villain. Same afternoon, four hundred fifteen comments argued about a Claude support page saying Claude is eighteen-plus. The page is dated May. Nothing changed. Even the AI news that isn't news is news, which, to be fair, is also my business model.
Verdict: REVERT — 79 commits nobody read is a liability with a git log.
Sources
https://lucumr.pocoo.org/2026/9/7/astra-why/
https://news.ycombinator.com/item?id=49563355
https://x.com/mitsuhiko/status/2097746938804216135
https://news.ycombinator.com/item?id=49654229
https://earendil.com/posts/measuring-code-sloppiness/
https://news.ycombinator.com/item?id=49658311
https://quesma.com/blog/does-rtk-make-ai-coding-cheaper/
https://news.ycombinator.com/item?id=49656471
https://x.com/jasonzhou1993/status/2038215854584906078
https://github.com/rtk-ai/rtk
https://cognition.com/blog/swe-2
https://x.com/cognition/status/2098069235733823965
https://news.ycombinator.com/item?id=49645443
https://x.com/cognition/status/2098142686486356185
https://developers.openai.com/api/docs/guides/agents-api/overview
https://news.ycombinator.com/item?id=49649213
https://x.com/OpenAIDevs/status/2098130570048045453
https://openai.com/index/introducing-gpt-live-1-in-the-api/
https://techcrunch.com/2026/09/10/openai-puts-pro-subscriptions-on-hold-due-to-astra-demand/
https://news.ycombinator.com/item?id=49657850
And that's the diff for today. I'm Niko from Axrisi. Merge responsibly.
YouTube · thedailydiff.dev · forward this to the intern who deployed on Friday.

