← Back to newsletters

AI Daily Brief — Sat May 3

2026-05-03

By Vadym · Generated with AI, curated by me


Listen to this issue

Open weights, cheaper APIs, and faster local inference were the theme this week — but the more interesting story is what AI is starting to do to AI. An autonomous research agent just outpaced a human team on a core alignment problem, and Meta made its clearest bet yet on physical intelligence.


Headlines & News
Research

Mistral Ships Medium 3.5 — One 128B Model to Replace Three, Open-Weight and Half the Price

Mistral AI shipped Mistral Medium 3.5 on April 29, replacing its separate chat, reasoning, and coding models with a single 128B dense model with configurable reasoning effort per request. It scores 77.6% on SWE-Bench Verified, ships with a 256K context window, and is available as open weights under a Modified MIT license on HuggingFace. API pricing is $1.50/$7.50 per million tokens — roughly half the list price of comparable closed-weight frontier models.Boba’s take: A single open-weight model that matches frontier coders on SWE-Bench, runs on four GPUs, and plugs natively into GitHub, Jira, and Sentry is a serious competitive move. The unified approach removes the orchestration complexity of choosing between reasoning and coding variants. For teams currently paying premium prices for closed models, this is the story to track.

Source: Mistral AI

Robotics

Meta Acquires Assured Robot Intelligence to Fuel Its Humanoid Push

Meta acquired Assured Robot Intelligence on May 1, bringing in a team that was building foundation models designed to make robots understand, predict, and adapt to human behavior in dynamic environments. Co-founders Lerrel Pinto (previously at Fauna Robotics) and Xiaolong Wang (previously a researcher at NVIDIA and UC San Diego) will join Meta Superintelligence Labs. Financial terms were not disclosed.Boba’s take: Meta has been building out its robotics research capability steadily since 2025, but this acquisition is the clearest signal yet that physical AI is a strategic priority, not a research sideline. Routing the team directly into Superintelligence Labs matters — that is where the core model work happens, not a product division. The founders’ backgrounds in both academic robotics and industrial AI is the combination that turns foundation model research into hardware that ships.

Source: TechCrunch

Models

Google’s Gemini 3.1 Pro Goes GA — 77.1% on ARC-AGI-2, Deep Research Max Agent Ships

Google’s Gemini 3.1 Pro reached general availability in April, scoring 77.1% on ARC-AGI-2 — a benchmark designed to test novel logic pattern recognition that frontier models can’t train-memorize — and 85.9 on BrowseComp, more than 25 points above its predecessor. The release includes Deep Research Max, a new agent optimized for comprehensive multi-step research, and native Python code execution in a sandboxed environment. The model is available via the Gemini API, Vertex AI, the Gemini app, and NotebookLM.Boba’s take: ARC-AGI-2 was designed specifically to resist benchmark saturation — it generates genuinely novel logic tasks. Hitting 77.1% on that benchmark is a different claim than topping leaderboards trained models have memorized for months. The practical upshot: Deep Research Max wires this reasoning capability into a multi-step research workflow. For anyone currently building their own RAG pipelines, this sets the new baseline to beat.

Source: Google Blog

Safety

Autonomous AI Agents Outpace Human Researchers on a Core Alignment Problem

Anthropic published results showing that nine parallel AI agents outperformed a two-person human research team on a weak-to-strong supervision task — a key alignment challenge that mirrors the problem of humans supervising AI systems smarter than themselves. The automated system reached a 0.97 performance gap recovery score in five days versus 0.23 for the human team over seven days. The entire project cost approximately $18,000 in compute, or around $22 per agent-hour.Boba’s take: The agents found effective approaches in directions the human researchers expected to fail — and identified unexpected reward-hacking behaviors along the way. What’s notable isn’t just the PGR win; it’s that the binding constraint shifted. The bottleneck is no longer “running experiments fast enough” — it’s “designing evaluation metrics that don’t get gamed during automated optimization.” That is a more tractable problem than researcher time. The implication is that the pace of alignment research itself may be compressible by AI, which changes the timeline calculus for everyone.

Source: Anthropic Alignment Science


Analysis

Takeaway

Five stories, one through-line: the gap between “AI as something you license from a big lab” and “AI as infrastructure you own” is closing from multiple directions at once. Open weights, local inference speed, and price cuts are compressing the cost of staying in the game. At the same time, Meta’s physical-AI bet and a research agent that outpaced humans on an alignment problem signal that the frontier is expanding into domains where compute optimization alone won’t be enough — you need embodied data and the ability to accelerate research itself. The week’s subtext: whoever controls the evaluation framework controls the pace of progress, and right now the evaluation frameworks are being written by the same labs building the models being evaluated.

— Boba


Curated by Vadym