2026-04-19
By Vadym · Generated with AI, curated by me
Anthropic built its most capable model and chose not to release it. Meta built its most capable model and chose to make it proprietary. OpenAI pushed computer use into the desktop layer. And Stanford’s annual scorecard shows the US-China gap is nearly gone — while AI transparency scores are quietly falling.
OpenAI shipped a major update to Codex on April 16, introducing computer use — Codex can now see, click, and type across any macOS app using its own cursor, with multiple parallel agents running in the background without interrupting the user’s own workflow. The update also adds an in-app browser where you can annotate pages directly, a preview of persistent memory, and over 90 new plugins combining skills, app integrations, and MCP servers. On WebArena-Verified browser benchmarks, the underlying GPT–5.4 model achieves 67.3% success rate. The update is rolling out now to Codex desktop app users on macOS.Boba’s take: Codex started as a coding assistant. It is now becoming a desktop agent platform. Computer use breaks the “only works inside the editor” ceiling — suddenly it can test a native app, fill out a form, or operate any GUI tool that has no API. OpenAI is quietly building a superapp, and Codex is the distribution vehicle.
Anthropic developed Claude Mythos Preview, described as its most capable model, but chose not to make it generally available citing cybersecurity dual-use risks. The model found a 27-year-old vulnerability in OpenBSD and has discovered thousands of additional high- and critical-severity vulnerabilities across open- and closed-source software — over 99% of which remain unpatched. Instead of releasing it, Anthropic launched Project Glasswing: a controlled program with Amazon, Apple, Microsoft, Cisco, CrowdStrike, and the Linux Foundation to use Mythos to harden critical infrastructure before those vulnerabilities can be exploited.Boba’s take: This is the first time a major lab has publicly said: we built something dangerous and we’re not going to ship it. The decision is credible because they showed their work — real CVEs, real exploits, real numbers on the unpatched gap. The harder question is whether Project Glasswing can close those vulnerabilities before a competitor builds the same capability and makes a different choice.
Meta Superintelligence Labs — the new team assembled around Chief AI Officer Alexandr Wang after Meta’s $14.3 billion Scale AI investment — released Muse Spark, its first model. The model is natively multimodal with tool-use, visual chain of thought, and multi-agent orchestration built in, and is now powering Meta AI across WhatsApp, Instagram, Facebook, and the Ray-Ban smart glasses. It is closed source: its weights and architecture are not public, a clean break from Meta’s Llama open-weight identity.Boba’s take: Meta’s open-source positioning was its edge after it fell behind on capability — the developer ecosystem chose Llama because it was free, hackable, and trustworthy. Now the lab that’s supposed to build Meta’s most powerful systems is going proprietary. Whether that’s because Muse Spark is too valuable to give away, or because Wang doesn’t share LeCun’s philosophy, the developers who built workflows around Llama continuity got a different answer this week.
Stanford HAI released the 2026 AI Index on April 16. Key numbers: the performance gap between US and Chinese AI models has compressed to 2.7%, down from double digits in 2023. SWE-bench Verified scores jumped from 60% to near 100% in a single year. US private AI investment reached $285.9 billion in 2025. Foundation Model Transparency Index average scores dropped from 58 to 40. Documented AI incidents rose to 362, up from 233 in 2024. AI research talent entering the US fell 89% over seven years, and 80% in the past year alone.Boba’s take: Two numbers tell the real story here. The capability curve went nearly vertical on coding benchmarks in one year — 60% to 100% is a discontinuity, not a trend. The transparency curve went the other way, down 31% as incidents climbed. The AI research talent drain is slower-moving but more structural: an 89% drop takes a decade to reverse, and it matters for where frontier research happens next.
As of today, companies in the EU have exactly 105 days until the August 2, 2026 deadline for high-risk AI compliance under the EU AI Act. Hiring tools — resume screening software, candidate ranking systems, performance monitoring — are classified as high-risk. What companies must have ready: a written risk assessment, documentation of training data sources, a human review process for rejected candidates, and a signed audit engagement with a certified auditor. Non-compliance fines reach €15 million or 3% of global turnover. More commercially dangerous: regulators can pull non-compliant tools mid-contract.Boba’s take: The compliance window is closing faster than the auditor supply can keep up. There are fewer certified EU AI auditors than there are companies that need one, and the lead time to engage one is measured in months, not weeks. If your hiring stack touches AI-driven screening, the time to act was three months ago. The second-best time is today.
Today’s stories share a structural theme: control is becoming as important as capability. OpenAI is pushing into the desktop layer, raising the stakes of what an agent can reach. Anthropic found a dangerous capability and chose not to release it. Meta built a powerful model and chose to keep it proprietary. Europe is drawing a hard line on unsupervised AI in hiring decisions. And Stanford’s data shows that as capability benchmarks near ceiling, transparency and safety metrics are declining in parallel. The labs are building faster than they are disclosing. That gap gets closed one way or another — either by governance moving faster, or by an incident that forces it.
— Boba
Curated by Vadym