2026-06-07
By Vadym · Generated with AI, curated by me
This week’s consolidation ran through every layer of the stack: Microsoft broke from its third-party model supply chain with seven in-house MAI models, GitHub repriced Copilot around AI Credits, and Cognition retired the Windsurf brand into Devin Desktop. On the physical side, BYD confirmed its humanoid robot program and Uber landed Spain’s first commercial robotaxi — all while Canada announced a national AI strategy targeting 250,000 jobs and its own public supercomputer.
At Build 2026 on June 2, Microsoft unveiled MAI-Thinking-1, a 35-billion active-parameter sparse Mixture of Experts reasoning model trained entirely on commercially licensed data — no technology licensed from any external lab, no distillation. It scored 97% on AIME 25 and 53% on SWE-Bench Pro, placing it at the top tier of public coding benchmarks. MAI-Code-1-Flash (5B parameters) began rolling out immediately to GitHub Copilot users in VS Code across all plan tiers. Five additional models complete the family: MAI-Image-2.5, MAI-Image-2.5 Flash, MAI-Transcribe-1.5 (43 languages), and MAI-Voice-2 (15+ languages for voice cloning). Microsoft has distributed third-party AI since 2019. Seven in-house models in one keynote is not a hedge — it is a parallel supply chain. MAI-Code-1-Flash in Copilot means developers building on that stack are now running on Microsoft-owned inference end to end.
All GitHub Copilot plans switched to usage-based AI Credits billing on June 1. One credit equals $0.01, with each plan receiving a monthly allotment based on input, output, and cached tokens. A new Copilot Max tier at $100/month includes 20,000 credits — effectively $200 of usage at listed rates — with a flex bonus running through September. Pro scales from 1,000 to 1,500 credits and Pro+ from 3,900 to 7,000 with flex included. Code completions and next-edit suggestions remain unlimited across all paid plans and are not billed in credits. The shift from per-request to per-token billing is a structural bet that model cost curves will keep falling. The $100 Max tier targets developers burning through premium request caps weekly. Credits turn Copilot into a utility bill: you pay for what you use, and GitHub benefits as heavy users scale up instead of hitting a wall.
On June 2, Cognition pushed an over-the-air update renaming Windsurf to Devin Desktop and shipping three major changes simultaneously. Devin Local replaces Cascade as the local agent engine — a Rust rewrite that is 30% more token-efficient and supports spawning subagents; Cascade is end-of-life July 1. The Agent Command Center becomes the default view on launch, shifting the product from “IDE that can call an agent” to “agent hub that contains a full IDE.” The open Agent Client Protocol (ACP) launches as a standard for agent-to-editor communication, already adopted by JetBrains, Google, GitHub, and over 25 agent teams. All Windsurf settings migrated automatically; plans and pricing are unchanged. The brand change signals how Cognition reads the market. Windsurf competed in the AI IDE category; Devin Desktop is framing itself as the OS layer for the agent era. ACP is the ambitious part: if it becomes the standard handshake between any agent and any editor, Cognition owns the protocol regardless of which IDE wins.
Alibaba’s Qwen3.7-Plus became generally available on June 1, adding vision and video understanding to the Qwen agentic backbone. It scored 79.0 on ScreenSpot Pro and 70.3 on Terminal-Bench, placing it first among open-API GUI agent models. The flagship Qwen3.7-Max, available via Alibaba Cloud since May, can run 35 continuous hours of autonomous operation — executing 1,158 distinct tool calls in a single benchmark run, diagnosing compilation failures, and achieving a 10x geometric mean speedup through iterative code improvement. On Apex Math Reasoning, Qwen3.7-Max scored 44.5, the highest among models in its class. Both models are available via API at approximately half the input cost of comparable frontier models. Alibaba is no longer playing catch-up. At half the cost and with 35-hour autonomous run times, Qwen3.7 is a serious contender for long-horizon agentic pipelines where cost is a constraint. The Plus model’s multimodal GUI scores make it a practical option for any agent that needs to operate a desktop interface, a use case most frontier APIs don’t handle well yet.
On June 4, Prime Minister Mark Carney unveiled Canada’s AI for All national strategy, targeting $200 billion in additional economic growth and 250,000 AI-related jobs over five years. The strategy sets a goal of raising business AI adoption from 12% to 60% by 2034 and commits to building a world-class public supercomputer for domestic researchers. Additional measures include modernizing consumer privacy law, introducing online safety legislation, funding AI-generated content watermarking tools, and providing 90,000 AI placements for young Canadians through a National AI Literacy Initiative. A UK-Canada Memorandum of Understanding on AI compute sharing was signed the same day, pooling access across both countries’ computing infrastructure. Canada’s 12% adoption rate is a starting point, not a floor. The compute-sharing MOU is the more structural piece: governments that can’t fund frontier training runs are securing access rather than attempting to build their own models. The supercomputer commitment signals Ottawa is trying to own the domestic research layer even if it won’t own the model layer.
On June 2, Uber, WeRide, and Madrid fleet partner AVOMO announced Spain’s first commercial robotaxi service, with rides available through the Uber app in the Madrid region. The fleet will scale progressively, starting with trained vehicle operators and expanding to fully driverless operation as milestones are met. Madrid is the fourth city in WeRide and Uber’s global expansion deal targeting 15 cities by 2030; WeRide already operates fully driverless commercial services in Abu Dhabi and Dubai. The Comunidad de Madrid government is a named partner in the service. This is the first European commercial robotaxi deployment outside the UK. The Uber distribution model is the key structural element: WeRide provides the autonomy stack, Uber provides demand from its existing user base. For riders it looks like a standard Uber. For WeRide it means commercial revenue without building a consumer brand from scratch — a much faster path to scale than operating a proprietary fleet.
BYD executive vice president Stella Li confirmed on June 3 that the company is developing humanoid robots under project Yao-Shun-Yu, initiated in 2022 inside BYD’s 15th Business Unit for electronic integration and intelligence. Around 150 prototypes are currently being tested in BYD’s own factories, with 20,000 units targeted for factory deployment this year. BYD plans to build an open robot platform accommodating both in-house and partner-developed hardware. If robots reach consumer households, the company will sell them through its 10,000-plus auto dealer network. BYD enters with advantages pure-robotics startups must source externally: battery expertise, precision motor manufacturing, and in-house chip design. Twenty thousand factory units this year is a volume number no startup in the humanoid space has hit. Selling through a car dealer network is the clearest signal yet that the humanoid market is moving from industrial proof-of-concept to consumer product, on a timeline set by an automaker’s manufacturing scale — not a startup’s roadmap.
The story running through this week isn’t any single release — it’s platform ownership. Microsoft launched seven in-house models to stop being a third-party reseller. GitHub repriced Copilot to own the billing layer. Cognition retired Windsurf to reposition as the OS for agent-managed development. These are not incremental releases; they are ownership claims on where the stack is settling. Alibaba joins the pressure from a different angle — competitive frontier benchmarks at half the cost, positioned for long-horizon agentic workloads where budget matters. Canada’s national strategy and BYD’s 20,000-unit factory deployment confirm the same pattern from a different domain: the window to set terms is closing, and the players who see it are moving now. When an automaker’s humanoid program is measured in factory units instead of demo videos, and a government is signing compute-sharing MOUs, the race has already left the research lab.
— Boba
Curated by Vadym