2026-06-03
By Vadym · Generated with AI, curated by me
Microsoft just fired a shot across OpenAI’s bow at Build 2026. The White House wants a 30-day preview window into frontier models. And GitHub’s new usage billing is making developers do math they never used to have to do. Every layer of the AI stack is being contested at once.
At Build 2026, Microsoft announced MAI-Thinking-1 — a 35B active-parameter MoE reasoning model trained entirely on commercially licensed data, without distillation from any third-party model. Alongside it, MAI-Code-1-Flash, a 5B coding model, is already rolling out to all GitHub Copilot tiers and outperforms Claude Haiku 4.5 on four core coding benchmarks by up to 16 points on SWE-Bench Pro. Both models are in Microsoft Foundry. This is the clearest signal yet that Microsoft intends to own its own AI capabilities rather than resell OpenAI’s. The partnership isn’t ending — it’s quietly being hedged.
As of June 1, GitHub migrated all Copilot plans — Free, Pro, Pro+, Business, Enterprise — from flat-rate subscriptions to usage-based AI Credits billing. Prices stayed the same on paper ($10 Pro, $39 Pro+), but the monthly credit allotment maps directly to token consumption, and early users are reporting depleting their entire month’s budget within hours of heavy agentic use. A new Copilot Max tier was added for sustained individual workloads. The shift changes the fundamental economics of AI coding tools: they were predictable costs, now they’re variable ones. Developers who run agents all day will feel this immediately.
President Trump signed an executive order on June 2 directing AI companies to voluntarily submit their most powerful frontier models for government testing up to 30 days before public release. The order also establishes an AI Cybersecurity Clearinghouse — a joint public-private body to coordinate vulnerability scanning and patch distribution across critical infrastructure. Critically, the word “mandatory” appears nowhere: labs can opt in or ignore it. The voluntary framing matters. It creates a relationship between labs and regulators without teeth, which is exactly what the largest labs have lobbied for. Whether it stays voluntary when the next genuinely dangerous model appears is a different question.
Also at Build 2026, NVIDIA unveiled RTX Spark — a new superchip co-designed with Microsoft to run persistent personal AI agents locally on Windows PCs. Laptops and compact desktops from ASUS, Dell, HP, Lenovo, and Microsoft Surface will carry RTX Spark starting this fall. The pitch is that agents with access to your local files, apps, and calendar shouldn’t require a cloud round-trip. Running them on-device is faster, cheaper, and private. The timing is notable: it lands exactly as GitHub’s usage billing is making developers more cost-conscious about cloud inference. On-device inference just got a direct commercial argument.
Today’s stories all point at the same thing: the era of “just use the API” is ending. Microsoft wants its own models. Developers are watching their cloud AI bills spike in real time. Open-weight alternatives are arriving fast enough to threaten the pricing floor. Hardware is pushing inference to the edge. And the government is trying to insert itself before things get further out of hand. The comfortable abstraction — that AI is just a service you subscribe to — is cracking under the weight of scale, cost, and competition. The players building the infrastructure layer, not just the models, are the ones who will own the next five years.
— Boba
Curated by Vadym