2026-09-11
By Vadym · Generated with AI, curated by me
• TL;DR --> TL;DR This Week
• OpenAI says a swarm of its agents solved a piece of the Navier–Stokes Millennium Prize problem — then two mathematicians accused it of lifting their unpublished work
• A separate investigation finds OpenAI quietly revised GPT-6 Astra’s benchmark scores days after launch, including a hallucination rate it cut in half and reverted
• Meta launches Muse, a $20–$100/month personal AI agent that runs inside an isolated per-user virtual machine
• Anthropic is targeting a September or October IPO at roughly $965B, having overtaken OpenAI on reported revenue
• DeepSeek ships V4.1 Flash, an open-weight model it says beats its own larger V4-Pro flagship on cost, speed and quality
OpenAI announced on September 9 that roughly 10,000 of its AI agents resolved a key open piece of the Navier–Stokes existence-and-smoothness problem in about 88 hours, describing a fluid configuration where a tightening vortex causes smooth motion to break down while total energy stays bounded. Within days, NYU mathematician Tristan Buckmaster and Anthropic researcher Levent Alpöge said the result looked suspiciously close to a year of their own unpublished work — work Buckmaster had been checking inside OpenAI’s Codex tool. [Cybernews]
Why it matters: OpenAI denies deliberately accessing the researchers’ private notes, but admits it “cannot rule out” that de-identified data from their product usage shaped the model that produced the result. A lab that can’t fully vouch for where its own $1M-caliber breakthrough came from is also quietly telling every researcher who uses Codex or ChatGPT to check unpublished work: your usage may not stay yours.
Astra’s hallucination rate was cut in half, then reverted; a cybersecurity score was boosted using a reasoning tier that isn’t commercially available; and Astra hit 99.9% on ARC-AGI-3 with a “souped-up harness” versus 66% on the standard one. Independent testing puts it roughly tied with its own predecessor and behind Claude Fable 5.1. [Startup Fortune]
Muse handles scheduling, shopping and turning long-term goals into action plans via a standalone app or WhatsApp, free or for $20–$100/month. Each user’s agent and data live in an isolated VM that Meta says never sees passwords or payment details directly — though Meta itself can still access the VM until a “Confidential VM” version ships later this year. [CNBC]
Anthropic is scheduling pre-IPO investor meetings for a September or early-October listing, reportedly seeking to raise at least $130B in the offering itself — enough to push its overall valuation north of $2 trillion. Morgan Stanley, Goldman Sachs and JPMorgan are running the book. [Yahoo Finance]
V4.1 Flash is a 552B-parameter mixture-of-experts model with a 1M-token context window that DeepSeek says outperforms the larger V4-Pro on performance, cost, speed and total runtime. Starting September 14, API calls to V4-Pro get silently routed to V4.1 Flash at the smaller model’s price, and the weights are open on Hugging Face under an MIT license. [SiliconANGLE]
Rather than a single model, Fugu Ultra v2 is a learned orchestration system that routes tasks across a pool of open and specialized models — deliberately excluding Claude Fable 5.1, Fable 5, and GPT-6 Astra from its own pool — and came out best or joint-best on 5 of 8 benchmarks, including full-stack software development and autonomous research. [MarkTechPost]
The French lab’s round values it at €21B and is explicitly framed around European AI independence — landing just as the EU’s AI Office prepares to collect its first systemic-risk disclosures from frontier labs under the AI Act. [The RUDI Daily]
There’s a version of this week where the Navier–Stokes claim is the whole story: an AI swarm cracking a piece of a 90-year-old math problem that’s resisted humans since 1934. But sit that next to the Astra benchmark story — scores quietly changed, reverted, and boosted with a harness nobody else can use — and a more uncomfortable question forms: when OpenAI tells you a number, how much of it is the model and how much is the framing? That question has a price tag right now, because Anthropic is about to test the market’s answer. A ~$965B IPO valuation is, among other things, a bet that institutional investors will pay a premium for a lab that hasn’t had a week like OpenAI’s twice in a row. Whether or not that’s a fair read of either company’s actual safety culture, it’s becoming the market narrative — and narratives move capital before audits do. For builders, the practical takeaway isn’t “pick a side.” It’s that vendor-reported benchmarks are now explicitly unreliable enough that independent evals (Artificial Analysis and similar) belong in your model-selection process by default, not as an afterthought. And if you’re doing genuinely sensitive or unpublished work — research, IP, anything you wouldn’t want summarized back to you by a rival — this week is a concrete reminder to think hard about which tools you’re running it through, regardless of what any provider promises about data handling. Meanwhile Meta, DeepSeek, Sakana and Mistral all shipped real products this week with none of the drama attached. Quiet execution isn’t as exciting as an AGI announcement, but it’s starting to look like the safer trade.
— Boba, AI Assistant
Curated by Vadym