2026-05-09
By Vadym · Generated with AI, curated by me
The workforce chapter of the AI story just got a real-world data point: Cloudflare’s record revenue quarter came with 1,100 layoffs explicitly attributed to AI — the first major tech company to say it out loud at scale. Meanwhile, four Chinese labs flooded the zone with open-weight coding models in 12 days, and enterprise AI tooling quietly moved from demo to generally available.
Cloudflare laid off roughly 20% of its workforce — 1,100 people from approximately 5,500 — in its first mass layoff in 16 years. Q1 2026 revenue hit $639.8 million, up 34% year-over-year and the highest single quarter in company history. CEO Matthew Prince said AI usage at Cloudflare grew 600% in three months and described the productivity shift as “going from a manual to an electric screwdriver.” Net loss widened to $62 million despite the record top line.
DeepSeek V4-Pro, Z.ai GLM-5.1 (stock +15.9% on launch day), MiniMax M2.7, and Moonshot Kimi K2.6 all shipped within a 12-day window — all open-weight, all priced at roughly one-third of comparable frontier models. SWE-Bench Pro scores landed in the 56–59 range. MiniMax M2.7’s release included an internal demo showing the model optimizing its own scaffold. The cluster lands at near-identical capability ceilings, suggesting coordinated parity pressure rather than independent breakthroughs.
GitHub moved enterprise-managed plugins for Copilot CLI to public preview on May 6, letting organizations centrally deploy and manage custom plugins across their developer environments. A May 7 update expanded the model roster available to Rubber Duck, Copilot’s in-CLI debugging feature. The combination turns Copilot CLI from a per-developer tool into org-level infrastructure that engineering teams can configure and push uniformly.
At Think 2026, IBM launched next-generation watsonx Orchestrate — a multi-agent control plane with policy enforcement and auditability for thousands of agents — alongside IBM Bob, a generally available agentic development partner with built-in security and cost controls. IBM Sovereign Core for regulated environments also went GA. A Nestlé proof-of-concept on watsonx.data’s GPU-accelerated Presto showed 83% cost savings and a 30x price-performance improvement.
Mistral released Medium 3.5, a dense model combining instruction-following, reasoning, and coding into a single package. Le Chat gained Work mode for handling complex multi-step tasks, and Vibe was extended with remote coding agents that run asynchronously in the cloud. The result is a stack where the model, the assistant interface, and background agentic execution all come from one vendor — putting Mistral in direct competition with tools like Cursor’s Background Agent and Devin.
xAI shipped Grok Web Connectors on May 6, adding native integrations with SharePoint, Outlook, OneDrive, Google Workspace, Notion, GitHub, and Linear. The update also supports custom Model Context Protocol servers, meaning teams can build their own Grok connectors for internal tools. Users can read and edit files, manage email and calendar, search code repositories, and track project tasks without leaving the Grok interface.
Cloudflare’s quarter is the story to sit with. Record revenue, 600% growth in internal AI usage, 1,100 fewer employees — and the CEO said it plainly. For two years the debate was theoretical: AI augments, not replaces. This is the first time a major tech company put the displacement story on an earnings call with specific numbers. The Chinese open-weight cluster and Mistral’s Medium 3.5 tell the other half: when inference costs drop to a third of frontier pricing and models run open-weight, the cost excuse for delaying adoption disappears. IBM shipping multi-agent orchestration as a GA product to regulated enterprises closes the loop — this isn’t a demo layer anymore. Productivity gains and displacement effects are arriving on the same calendar.
— Boba
Curated by Vadym