← Back to newsletters

AI Daily Brief — Sat Jun 13

2026-06-13

By Vadym · Generated with AI, curated by me


Listen to this issue

The model race did not pause for the weekend. A new frontier-class public model, a 1-trillion-parameter open-weight coding release, a $12 billion industrial AI round, and a major developer platform cutting 14% of staff all landed in the same short window. Meanwhile, a new benchmark quietly exposed how much air was in the previous coding scores.


Headlines & News
Research

Anthropic Ships Fable 5 — Its Most Capable Public Model, Four Days After Warning AI May Be Getting Too Dangerous

Fable 5, released June 9, is the first Mythos-class model Anthropic has made available to the general public. It tops nearly every major benchmark by a wide margin: 95% on SWE-bench Verified, 10+ percentage points above Opus 4.8 on several evals, and state-of-the-art on vision and scientific reasoning tasks. Access is free through June 22 on Pro, Max, Team, and Enterprise plans; after that, pricing is $10 per million input tokens and $50 per million output tokens. The release came four days after Anthropic published a report warning that its models were approaching capability levels that could pose serious biosecurity risks — the same week it quietly applied for a classified government contract. Building the most dangerous model and publishing a paper about the danger on the same day is a new kind of disclosure strategy.

Source: TechCrunch

Funding

Bezos’s Industrial AI Startup Prometheus Raises $12 Billion at a $41 Billion Valuation

Prometheus, the industrial AI startup co-led by Jeff Bezos and former Google executive Vik Bajaj, closed a $12 billion Series B on June 11, taking its total raised to over $18 billion since launching in late 2025. The round was backed by JPMorgan, BlackRock, Goldman Sachs, DST Global, and Arch Venture Partners — notably not Amazon or Blue Origin, which Bezos has kept at arm’s length. Prometheus builds AI for engineering and manufacturing workflows, describing its long-term goal as building an “artificial general engineer.” It currently employs roughly 150 people. At $41 billion, Prometheus is worth more per employee than nearly any company in history. The bet is that the physical world is the next major surface for AI — and that it requires purpose-built models, not general-purpose LLMs deployed sideways.

Source: Axios

Workforce

GitLab Cuts 14% of Staff and Exits 22 Countries to Rebuild Around Agentic AI

GitLab announced on June 2 that it is eliminating approximately 350 positions — 14% of its global workforce — while simultaneously flattening management by up to three layers and reorganizing R&D into 60 autonomous teams. CEO Bill Staples said the driver is not cost-cutting but infrastructure: agentic workloads are pushing dramatically more traffic through GitLab’s platform than it was designed to handle, and the current organization is not fast enough to adapt. The company is wiring AI agents into its internal review, approval, and handoff workflows, then right-sizing headcount where agents take over. GitLab is still growing revenue and is not cutting from a position of weakness — which makes this a more legible signal than most: agentic automation is now changing the internal structure of dev-tooling companies, not just the products they ship.

Source: TechCrunch

Benchmarks

SWE-bench Pro Exposes a 23-Point Score Gap — AI Coding Was Better at Taking Tests Than Fixing Code

SWE-bench Pro, built by Scale AI specifically to resist benchmark contamination, has become the new standard for measuring AI coding ability — and the numbers are sobering. The best result on the new benchmark as of June 9 is 59.1%, held by Fable 5. The same model scores 95% on the older SWE-bench Verified. The gap is consistent: every model tested drops 19 to 26 percentage points moving from Verified to Pro, with an average drop of 23 points. The original SWE-bench was validated on a fixed set of 500 GitHub issues that are now widely known and partially in training data. Pro rotates in new, unseen issues in a way that makes memorization useless. The practical implication is that AI coding tools are still genuinely useful — a 59% solve rate on hard real-world bugs is meaningful — but the “superhuman at code” narrative was running ahead of what the benchmarks actually measured.

Source: MorphLLM


Analysis

Takeaway

Today’s stories share an uncomfortable subtext: the AI that exists now is both more capable and more honestly measured than it was six months ago. Fable 5 is the strongest public model yet. SWE-bench Pro shows the prior scoring system was flattering. Kimi K2.7-Code is open and commercial-friendly. And GitLab, a growing company, is cutting people anyway — not because of revenue pressure but because agents are handling workflows faster than humans can. The honest version of the story is not that AI is coming for jobs someday. It’s that it already changed the unit economics of software companies in a way that is now showing up in headcount.

— Boba


Curated by Vadym