2026-04-07
By Vadym · Generated with AI, curated by me
Today’s news belongs to the infrastructure and open layer. Gemma 4 arrived multimodal. Ollama went MLX-native on Apple Silicon. Cursor rebuilt what a coding IDE looks like when agents are the primary unit of work. And Kubernetes got its first AI conformance standard — developed jointly by the three companies that run the enterprise cloud.
Google DeepMind released Gemma 4 on April 2 — a family of four open-weight models spanning E2B, E4B, a 26B Mixture-of-Experts variant, and a 31B Dense model. All four accept text, images, audio, and video input. Framework support across Hugging Face Transformers, Ollama, MLX, vLLM, and NVIDIA NIM was live at the time of announcement — not promised later, live at launch.
Cursor released version 3 with a redesigned Agents Window that manages multiple AI coding agents running in parallel across different branches and workspaces. The release also includes Design Mode for UI-level editing, Composer 2 for faster iteration, cloud handoff for long-running tasks, and new /worktree and /best-of-n slash commands for evaluating parallel agent outputs.
Ollama released a version 0.19 preview with a native MLX backend replacing the previous GGML/llama.cpp path on Apple Silicon. On M5, M5 Pro, and M5 Max chips, the MLX runtime taps the new GPU Neural Accelerators to improve both time-to-first-token and token throughput. Performance gains are tiered by chip variant and apply across the full model catalog.
X Square Robot hosted the inaugural Embodied AI Developers Conference (EAIDC 2026) on April 2, focused on moving robotics from lab demonstrations to production deployment. The event featured live robot demonstrations, a national hackathon, and sessions on the simulation-to-real pipeline, hardware standardization, and commercial integration for AI-powered physical systems.
Google announced the Certified Kubernetes AI Conformance program on April 6, developed jointly with Microsoft, Red Hat, and Kubermatic. The standard defines portability, reliability, and efficiency requirements for AI workloads running on Kubernetes, addressing the hardware demands, networking latency, and stateful characteristics unique to AI systems. Google Kubernetes Engine and Azure Kubernetes Service have already adopted the standard.
The through-line today is infrastructure maturing to meet the frontier. Open-weight models are now multimodal. Local inference on consumer Apple Silicon hardware is faster. Coding environments are being rebuilt around parallel agents rather than single completions. Enterprise Kubernetes has an AI portability standard. And the embodied AI field organized its first developer conference to tackle the sim-to-real deployment gap. None of these are frontier model breakthroughs — they’re the scaffolding that makes the breakthroughs usable. The most meaningful developments in any technology cycle tend to happen not at the research edge but one level below it, where the tools, runtimes, and standards get built that let everyone else catch up. Today was that layer.
— Boba
Curated by Vadym