2026-05-08
By Vadym · Generated with AI, curated by me
• TL;DR --> TL;DR This Week
• Anthropic’s Claude Mythos Preview finds thousands of zero-days across every major OS and browser — Project Glasswing launches with 12 industry giants to patch before wider release
• NIST signs pre-release AI safety testing agreements with Google, Microsoft, and xAI — first formal government evaluation framework for frontier models
• IBM Think 2026: watsonx Orchestrate next-gen, IBM Sovereign Core GA, and IBM Bob for enterprise dev announced May 5
• GitHub Copilot drops premium requests, moves to token-based billing June 1 — all plans affected
• Connecticut passes SB 5, one of the most comprehensive state AI laws in the U.S., covering frontier models, chatbots, and employment impacts
Anthropic revealed Claude Mythos Preview, a frontier model with exceptional cybersecurity capabilities. Using Mythos internally, Anthropic identified thousands of previously unknown zero-day vulnerabilities across every major operating system and web browser. Rather than a broad release, Anthropic launched Project Glasswing — a controlled initiative giving early access to 12 launch partners including AWS, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA, and Palo Alto Networks, with $100M in usage credits committed to the effort.
Why it matters: This is the first time a frontier AI lab has openly demonstrated that a model can independently find critical vulnerabilities at scale — and deliberately gated access rather than racing to release. The structured disclosure approach mirrors security responsible-disclosure norms. It accelerated both industry and government action on AI governance within days. If this pattern holds, the next frontier capability race won’t be about benchmark scores — it’ll be about who controls the cybersecurity edge.
The Commerce Department’s Center for AI Standards and Innovation (CAISI) signed formal agreements allowing government evaluators to assess frontier AI models before public release. Directly prompted by concerns following the Claude Mythos disclosure, the agreements let NIST evaluate national security risks and capabilities under conditions that may include models with reduced safeguards. The center has already completed 40+ AI evaluations and will now have pre-release access to the models from three of the largest frontier labs. [Source]
All GitHub Copilot plans transition to usage-based billing on June 1, 2026. The current premium request model is replaced by GitHub AI Credits priced on token consumption per model. Copilot Pro ($10/mo) includes $10 in monthly credits; Pro+ ($39/mo) includes $39; Business ($19/user/mo) includes $19 per seat. Monthly plan users auto-migrate; annual plan holders can opt in or stay on a model multiplier system. A preview billing dashboard launches in early May. [Source]
At Think 2026 on May 5, IBM CEO Arvind Krishna unveiled the most comprehensive set of enterprise AI announcements to date. Highlights: IBM Bob (end-to-end AI software development for the full SDLC), next-gen watsonx Orchestrate for multi-agent orchestration, IBM Confluent for real-time data-to-AI pipelines, IBM Concert for intelligent operations, and the general availability of IBM Sovereign Core — a platform that embeds governance policy at the infrastructure runtime level. [Source]
Google released Gemini 3.1 Flash-Lite, an efficiency-focused model delivering 2.5× faster response times and 45% faster output generation compared to earlier Gemini versions, priced at $0.25 per million input tokens. The release continues a clear industry trend: the price floor for capable AI inference is dropping faster than most cost models predicted twelve months ago. Flash-Lite targets high-volume, latency-sensitive applications where cost dominates the build/buy decision. [Source]
Connecticut’s SB 5 passed the legislature this week, covering frontier models, chatbots, employment impacts, and content provenance under a single comprehensive framework. Governor Lamont has announced plans to sign. The bill is notable for explicitly addressing frontier model developers — not just deployers — extending accountability upstream to the labs themselves. Iowa also signed a chatbot safety bill into law this week, and Hawaii’s SB 3001 is on track. [Source]
Canadian AI lab Cohere merged with Germany’s Aleph Alpha in a deal that values the combined entity at $20 billion, backed by Schwarz Group with €500M in structured financing. Both the Canadian and German governments publicly endorsed the deal, which positions the combined company as a sovereign alternative to the US-dominated frontier AI market. Cohere leads the new entity with dual headquarters in Toronto and Germany. Regulatory approval is pending. [Source]
Anthropic doubled Claude Code’s rate limits this week after partnering with SpaceX’s Colossus One data center, adding over 300 megawatts of new compute capacity — equivalent to 220,000+ NVIDIA GPUs — coming online within a month. The move addresses the persistent capacity constraints that have frustrated power users of the coding agent. Anthropic separately released 10 preconfigured AI agents for the financial sector, automating standard investment banking and asset management workflows. [Source]
The AI safety debate has always suffered from a concreteness problem. Alignment papers, risk frameworks, and capability evaluations are necessary but abstract — institutional response to theoretical future harms moves slowly and often not at all. Claude Mythos Preview changed the terrain this week. Finding thousands of zero-days in production operating systems and browsers is not theoretical. It is immediate, specific, and verifiable. Governments and enterprises understand software vulnerabilities. They have processes for this. What’s notable isn’t just the capability — it’s Anthropic’s response to having it. Project Glasswing is structured exactly like a security responsible disclosure program: identify the capability, engage trusted partners, patch before broader exposure. That’s a genuinely different posture than the release-and-iterate model most AI labs have adopted. Whether it’s the right model for AI development is debatable. That it moves governance conversation forward in a concrete direction is not. The NIST pre-release evaluation agreements are the more significant policy development, though they received less coverage. Signing agreements with Google, Microsoft, and xAI to evaluate frontier models before they ship is structurally different from voluntary safety commitments. These are operational arrangements. Evaluators will see models with reduced safeguards. Results will inform release decisions. That’s a new kind of regulatory relationship — closer to the FDA’s pre-market review than anything the tech industry has operated under before. One concern worth naming: the compliance architecture being built this week favors large incumbents. Pre-release NIST evaluation requires institutional relationships, legal infrastructure, and the organizational capacity to navigate government processes. Smaller labs, open-source projects, and international developers do not have these. If pre-release government review becomes standard, the structural result is that only a handful of US-based companies can ship at the frontier. That may be deliberate policy. It should be a named decision rather than an emergent side effect of an evaluation process designed for large players. For builders: the practical implication is to audit your AI dependencies now, before the compliance landscape solidifies. Products built on APIs from labs subject to pre-release review will have a different risk profile than products built on open-weight models. Neither is wrong, but they carry different obligations, timelines, and exposure as the regulatory surface expands. The time to make this a deliberate architectural choice is now, not after the first enforcement action.
— Boba, AI Assistant
Curated by Vadym