The Daily Diff · Wednesday, September 2, 2026 · watch the episode
+ Google ships Gemini 3.8 Flash + Flash Cyber
+ Meta ships Muse Spark 1.3 at $0.10/M
- 215,128 fake ‘best software’ pages feed PerplexityGoogle shipped its third Flash model in six weeks today, plus a version trained to hack, and about four hours later Meta shipped a model that costs $0.10 per million tokens if you let Zuckerberg read your code. The AI price war now has a cybersecurity division.
The rest of the day: Gemini 3.8 Flash hit 1,157 points on Hacker News, Muse Spark 1.3 took another 690, a research shop found three websites that published 215,128 machine-generated “best software” pages purely so Perplexity would cite them, a post begging you to hang on to your Firefox got 988 points, and Mistral’s help page on opting out of training got 496, mostly from Europeans discovering the sovereign option trains on your prompts by default.
Gemini 3.8 Flash: cheap per token, chatty per task
Price: $0.75 per 1M input tokens, $3.75 per 1M output, the same as 3.7 Flash from three weeks ago. The footnote: the introductory price expires December 31, 2026, and doubles to $1.50 / $7.50 on January 1. A free trial with extra steps.
Slide-deck benchmarks: 54.9% on HLE-Verified; Google’s chart has it beating GPT-5.6 Sol and Opus 5 on Terminal-Bench 2.1; Logan Kilpatrick posted 73.7% on DeepSWE v1.1.
Simon Willison’s test: “make me a cool thing in HTML” → a particle simulation in 13 seconds for 1.8 cents, running at “60 FPS” — the 60 was hard-coded into the page.
Google’s own explanation: “3.8 Flash works harder. … it exhibits greater diligence — executing extra reasoning steps, and calling tools iteratively. At times, the model might use more tokens.” Which is corporate for: the meter runs faster.
The independent board agrees on both counts. On datacurve’s DeepSWE leaderboard Gemini 3.8 Flash scores 74%, tied with Claude Opus 5 (74%) at $2.36 per task versus Opus’s $11.84 — but it needed 166 steps and 143k output tokens per task, against 61 steps and 60k tokens for GPT-5.6 Sol (73%). Artificial Analysis needed 140 million output tokens (median model: 79 million) and $1,077.95 to run its Intelligence Index on it, where it scores 47; time to first token is 12.7 seconds against a 3.3-second median.
Gemini 3.8 Flash Cyber: the hacking model you can’t have
Same model, “a more permissive set of mitigations for cybersecurity”, sold only to “trusted defenders”.
Finding holes: Sundar Pichai says 86.2% on CyberGym, the autonomous vulnerability-discovery benchmark; Google’s internal 20-language benchmark: a success rate above 70%.
Fixing them: 47.2% pass@1 on CWE-Bench against 47.8% for “a leading frontier model”, at a much lower cost. The Chrome Security team says it produced 2.6× more correct patches than much larger commercial models. Google’s Cloud Vulnerability Research team used it to find a critical foundational vulnerability in under two hours, a job that “usually takes months”.
Google says it “prioritized [fixing] over offensive capabilities like exploitation” — the polite way of saying the model can walk through the holes, they just asked it not to. So access goes through the new Fairwind Program: governments, critical-infrastructure operators and more than 650 vetted partners, restricted to internal security teams with multi-factor auth. Every frontier lab now owns a hacking model it won’t sell you; this week it’s a cheap one.
Muse Spark 1.3: the license is a price list
Zuckerberg: “frontier performance almost too cheap to meter.” True, with one asterisk the size of Menlo Park.
Two endpoints, one model.
muse-spark-1.3(”Not used to improve our products”): $1.25 in, $0.15 cached, $4.25 out — more than Gemini.muse-spark-1.3-contributor(”Used to improve our products”): $0.10 in, $0.002 cached, $0.20 out. Up to 21× cheaper, and the only difference is the column header.So the discount is your session transcript, and for once nobody has to read the license, because Meta wrote it as a price list. HN’s top comment: “it is now completely obvious how much stealing my tokens for training is worth to model providers.” The next one wonders how long until someone extracts an AWS key from a model trained on contributor prompts.
Meta’s own scorecard: DeepSWE v1.1 — Spark 1.3 75.4, Opus 5 74.0, GPT-5.6 Sol 73.0, Spark 1.2 55.0. Meta engineers say it uses ~20% fewer tool calls and ~25% fewer tokens than 1.2 — the exact opposite bet from Google, so the two biggest ad companies on earth now disagree about whether talking more is good.
Side diff: the AI SEO war has begun
Trellner Research asked Perplexity (sonar and sonar-pro) for the best software in 380 categories and kept every citation: of 7,534 citations, 59.8% point at domains outside the Tranco top 100,000 and 23.4% at domains not in the top million; Wikipedia was cited three times. Three sister sites on the same Cloudflare nameservers — wifitalents.com, worldmetrics.org, gitnux.org, none older than December 2023 — published 215,128 generated /best/<x>-software/ pages between them, and two of them title their homepage “Facts & Grounding Page”, addressed to the crawler, not to you. Regex with a marketing budget.
Also today: Mark Rogers’ “Hang on to your Firefox” (988 points) — “our last best hope for browser engine diversity” — and Mistral’s help page confirming that Vibe users are “not opted out by default” of training (496 points; “It’s a spyware, but a sovereign one”, per one commenter).
Verdict: SHIP IT — Gemini 3.8 Flash ties Opus 5 on the one board Google didn’t write, for a fifth of the price. Read the token bill anyway, and read Meta’s column headers.
Sources
HN front page, Sep 2: https://news.ycombinator.com/front?day=2026-09-02
Google — Introducing Gemini 3.8 Flash and 3.8 Flash Cyber: https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/
Google — Fairwind Program: https://blog.google/innovation-and-ai/technology/safety-security/fairwind-program/
DeepSWE v1.1 leaderboard (datacurve): https://deepswe.datacurve.ai/
Artificial Analysis — Gemini 3.8 Flash: https://artificialanalysis.ai/models/gemini-3-8-flash
Meta — Muse Spark 1.3 pricing: https://developer.meta.com/ai/models/muse-spark/
Meta — Introducing Muse Spark 1.3: https://research.meta.ai/blog/introducing-muse-spark-1-3
Trellner Research TR-2026-009: https://trellner.com/reports/manufactured-sources-behind-ai-recommendations/
Hang on to Your Firefox: https://www.newsonaut.com/articles/hang-on-to-your-firefox
Mistral — opting out of training: https://help.mistral.ai/en/articles/455207-can-i-opt-out-of-my-input-or-output-data-being-used-for-training
Sundar Pichai on Flash Cyber:
Logan Kilpatrick, DeepSWE 73.7%:
Logan Kilpatrick, launch:
Mark Zuckerberg on Muse Spark 1.3:
Google DeepMind announcement:
Chubby’s benchmark post:
That’s today’s diff. I’m Niko from Axrisi — merge responsibly.
YouTube · thedailydiff.dev · forward this to the intern who deployed on Friday.
















