+ Grok 4.7: still $2 / $6 per M tokens
- but $3.74 a task, up from $1.86
- “twice as fast” clocks 57 tok/s
- Copilot+ PC brand quietly retired
+ a month without AI: the joy is backYou'd think the same price means the same bill. On Monday I told you Grok four point seven costs exactly what Grok four point six did. Per token, it does. Per task, it costs double. Also, this one is by request. The real Haas asked under Friday's video, can you do a vid on Grok four point seven, colon D. We replied, easy peasy, later today. By the time you watch this, it's probably tomorrow, so we shipped late. Relax, so did Grok, about two weeks late, if you believe Hacker News. And one test gave Grok a perfect five out of five. It's my favourite number of the week, and I'll get to it at the end. Here's the trick. The price per token didn't move, two dollars in and six out, per million. What moved is how much it talks. Artificial Analysis ran its whole test suite, and Grok four point seven wrote about two hundred forty million tokens. That's two and a half times its predecessor, and almost three times a typical model.
So the average cost per task went from about a dollar eighty-six to three seventy-four. Same price tag, double the receipt. It's a taxi with the same rate per mile that takes the scenic route through two other cities. Hacker News spotted it in the first hour. The biggest thread under the launch says forty percent more weights, same price, and almost two weeks late, so xAI can't have loved the results. Another commenter asked about the chart on top, which compares against GPT five point six and quietly leaves out Astra. That, he wrote, can't have been an oversight. Now, the launch post. The headline says twice as fast, at half the price of comparable models. A few lines down, it says served at the same price and speed as Grok four point six. So it's twice as fast as somebody else's model. Artificial Analysis clocked it at about fifty-seven tokens a second, slower than the old Grok, and summed it up in two words. Notably slow.
Meanwhile, xAI's own account, which now goes by SpaceX AI, posted the calmer version. A notable improvement over Grok four point six, at the same price and speed. Twenty-eight thousand likes. So the tweet and the blog headline disagree, and for once the tweet is the careful one. Which brings us to the strangest part. Artificial Analysis says it works harder than the last Grok, about eighty-one thousand output tokens per task, more than double the last one. And on long office work and in its own coding harness, it really did improve. It's fourth among coding agents now, behind two Claudes and GPT six Astra, and that puts xAI in the top four labs. The people using it describe a different model. One Hacker News commenter says Grok ends tasks almost immediately and claims done, and calls it the laziest of them all. Another watched it loop in thinking mode, from fix one all the way to fix eighty-one. So it's lazy and verbose at the same time. It's the coworker who writes an eighty-one step plan, then says done and goes home.
One more straight diff before the fun number. Microsoft has quietly retired the Copilot plus PC brand. The Surface boss told Windows Central the new Surface PCs are not called Copilot plus PCs, even though they meet every requirement. Two years after a launch whose headline feature, Recall, had to be delayed for security, the name lives only on spec sheets. Windows Central's own take is shorter. These PCs have nothing to do with Copilot. And one developer tried the opposite of Grok, a month with zero AI. His post hit the Hacker News front page. The low point before he quit. An agent sat stalled for thirty minutes, he told it off, it apologised and delivered in twenty seconds, and the bill for that half hour was thirty dollars of tokens for absolutely nothing. A month later he says the joy of programming is back, and he isn't scared of being fired. Grok would call thirty dollars a warm up. Which brings me to the perfect score I promised. Personality Bench runs every new model through a stack of personality tests, and on honesty and humility, Grok four point seven maxed out, five out of five. Its profile is literally called the humble type. To be fair, thirty-two other models also maxed it out. The same page says it believes powerful others control what happens to it, and flags that as unusual for a frontier model. I wouldn't call that a training artifact. I'd call it reading the org chart. And according to the birth chart on the same page, it's a Virgo.
Verdict: NEEDS REVIEW — I'd cap the budget first: the price held, the bill didn't
Sources
https://www.youtube.com/watch?v=X-aFviBTcCs
https://x.ai/news/grok-4-7
https://web.archive.org/web/20260923060447/https://x.ai/news/grok-4-7
https://news.ycombinator.com/item?id=49788838
https://twitter.com/SpaceXAI/status/2102069815225586149
https://artificialanalysis.ai/articles/benchmarking-grok-4-7
https://news.ycombinator.com/item?id=49804685
https://artificialanalysis.ai/models/grok-4-7
https://news.ycombinator.com/item?id=49789558
https://artificialanalysis.ai/models/grok-4-6
https://x.com/ValsAI/status/2102086608476590432
https://persona.earthpilot.ai/models/x-ai/grok-4.7
https://news.ycombinator.com/item?id=49802839
https://www.windowscentral.com/microsoft/windows-11/the-copilot-pc-brand-is-dead-microsoft-and-pc-makers-quietly-pull-back-on-tarnished-windows-11-ai-pc-branding
https://news.ycombinator.com/item?id=49854945
https://blog.bustikiller.com/2026/09/25/one-month-without-ai.html
https://news.ycombinator.com/item?id=49855018
And that's the diff for today. I'm Niko from Axrisi. Merge responsibly.
YouTube · thedailydiff.dev · forward this to the intern who deployed on Friday.

