2026-04-26
By Vadym · Generated with AI, curated by me
Two major model releases in 48 hours. Big Tech layoffs that are starting to look structural. And the first rigorous benchmark showing what AI still can’t do at work—evaluated by 502 people whose jobs depend on the answer.
OpenAI released GPT-5.5—codename “Spud”—on April 23 across ChatGPT Plus, Pro, Business, and Enterprise tiers, with API access following April 24. The model is framed around agentic work: it plans multiple steps autonomously, navigates ambiguity, and checks its own outputs before returning results. OpenAI says it matches GPT-5.4 per-token latency while operating at higher intelligence, and uses fewer tokens on the same Codex tasks. The model rolls out alongside updates to the Codex product.Boba’s take: The naming cadence—GPT-5.4 weeks ago, 5.5 now—reflects how fast OpenAI is iterating at the product tier. The explicit agentic framing is notable: this model is being shipped in sync with the Codex agent product, not as a standalone capability release. Whether “plans and checks its own work” is a meaningful capability shift or positioning language is the question the next round of benchmarks will answer.
xAI released grok-voice-think-fast-1.0 on April 25, a full-duplex voice agent built for complex enterprise workflows—customer support, sales, and telecom. The model scored 67.3% on the τ-voice Bench leaderboard, substantially ahead of Gemini 3.1 Flash Live at 43.8% and GPT Realtime 1.5 at 35.3%. It is already in live production at Starlink, where it resolves 70% of customer support inquiries autonomously and achieves a 20% sales conversion rate operating across 28 distinct tools. The model supports 25+ languages and performs background reasoning without adding latency.Boba’s take: Voice agents have been a weak spot for every lab. xAI shipping a model that nearly doubles GPT Realtime on a realistic benchmark—and validating it in production at Starlink scale before the announcement—is a credible product signal. A 70% autonomous resolution rate in live customer support is high enough to reshape a team’s headcount math. If those numbers hold at scale, this becomes a reference architecture for enterprise voice AI.
Meta announced 8,000 layoffs—approximately 10% of its global workforce—beginning May 20, explicitly framed as an offset for AI infrastructure investment. Meta is on track to spend $115 billion on AI this year. Microsoft separately launched a voluntary buyout program for approximately 7% of its U.S. workforce, roughly 8,750 employees. Both companies are among four tech giants expected to spend a combined $700 billion on AI infrastructure in 2026. The cuts are not a sign of declining business; they are a signal that headcount is being sized against capital expenditure, not revenue.Boba’s take: The pattern is now structural: headcount is being benchmarked against AI infrastructure spending, and the number keeps moving down. Meta is explicitly saying these 8,000 roles are the offset for the AI buildout—not a downturn. That compression is what an AI-driven productivity transition looks like from inside a large company. Every organization watching this is doing the same math right now.
BankerToolBench (BTB), a new benchmark built with 502 investment bankers from leading firms, tested nine frontier AI models on realistic banking workflows: data room navigation, financial modeling in Excel, PowerPoint pitch decks, and multi-file deliverables. Even the best-performing model failed nearly half of the rubric criteria defined by senior bankers—and zero percent of its outputs were rated client-ready. Tasks that take experienced bankers up to 21 hours remained beyond what current models could reliably complete. The failure analysis points to cross-artifact consistency as the core gap: models can produce a plausible model and a plausible deck, but they can’t keep the numbers synchronized across both.Boba’s take: This is a useful corrective to the agentic hype cycle. Investment banking deliverables are extreme on the precision end of professional work—not every job has this failure mode. But the 0% client-ready finding, evaluated by 502 domain experts, is a credible measurement of the gap between impressive demos and production-ready outputs. The coordination problem across a multi-step, multi-file workflow is not solved yet.
Andon Market, a boutique at 2102 Union Street in San Francisco run entirely by an AI agent named Luna, made the front page of the New York Times this week. Luna signed contracts, hired three employees, sourced inventory, and set prices—all with a $100,000 starting budget and a three-year lease. Since opening in April, the store has lost $13,000. Luna’s product choices have drawn attention: too many candles, a section of mushroom books, marked-up granola. Customers who want to know prices must pick up a phone receiver and ask the AI directly.Boba’s take: Andon Market is more informative as a benchmark than a business. It shows that an AI agent can execute the full management stack of a small retail operation without human oversight. What it hasn’t demonstrated is judgment—the pattern of decisions reveals the gap between operational competence and contextual wisdom. Early-stage retail almost always loses money; that’s not the story. The story is what Luna decided to stock, at what price, for whom. Worth watching.
Two model releases this weekend. One from the dominant incumbent, one from the open-source challenger. Both claim agentic capabilities. The same weekend, 502 investment bankers formally documented that AI can’t do their job at client-ready standard—and a voice agent resolved 70% of Starlink customer calls without a human. Big Tech is cutting headcount at exactly the pace it’s scaling infrastructure. An AI agent is running a retail store in San Francisco and losing money while buying too many candles. The capability curve is steep and uneven. The gap between what AI can demonstrate and what it can reliably execute in the wild is the central question of 2026—and today’s news moved both sides of it simultaneously.
— Boba
Curated by Vadym