← Back to newsletters

AI Daily Brief — Fri Jul 10

2026-07-10

By Vadym · Generated with AI, curated by me


Listen to this issue

The capability curve keeps bending up — a cheaper model, a cheaper chip, robots doing real surgery — while the accountability layer underneath it keeps getting thinner. Today’s clearest evidence: the same four labs graded on AI safety this week all quietly walked back the exact commitments that mattered most.


Headlines & News
Models

SpaceXAI and Cursor Ship Grok 4.5, Trained on Real Coding Sessions

xAI — folded into SpaceX after its $60B acquisition of Cursor in June — released Grok 4.5 on July 8 at $2/$6 per million input/output tokens, undercutting Anthropic’s Opus 4.8 ($5/$25). Artificial Analysis ranks it #4 overall (score 54, behind Fable 5, Opus 4.8, and GPT-5.5) but #1 on agentic tool use at 33%, needing roughly 4x fewer tokens per SWE-Bench Pro task than Opus 4.8.

Source: MarkTechPost

Hardware

Nvidia Rival Positron in Talks to Triple Its Valuation to $5B in Five Months

Positron, a Reno-based startup building energy-efficient inference chips, is negotiating a two-phase raise of roughly $750M that would value it at $3.5B in the first tranche and up to $5B in the second — up from just over $1B in February. The round lands alongside SambaNova’s $1B raise at an $11B valuation and Etched’s $800M round, while Cerebras, the first inference chipmaker to go public, now trades below its IPO price.

Source: Yahoo Finance

Workforce

Microsoft Cuts 4,800 Jobs, Guts Xbox to Fund Its AI Buildout

Microsoft eliminated 4,800 roles — 2.1% of its workforce — on July 6, with about 1,600 coming from Xbox as the gaming unit spins off four studios; further cuts are expected to push total gaming losses toward 3,200. The company framed the move as part of an ongoing “transformation” rather than a response to weak revenue.

Source: CNBC

Safety

Nobody Passes: FLI’s New AI Safety Index Gives Every Lab a C+ or Worse

The Future of Life Institute’s Summer 2026 AI Safety Index graded nine companies on risk assessment, governance, and transparency. Anthropic topped the field with a C+, OpenAI and Google DeepMind scored a C, Meta got a D+, and xAI, DeepSeek, and Mistral all failed outright. Reviewers singled out Anthropic, OpenAI, Google DeepMind, and Meta for quietly weakening or dropping earlier pledges to pause development if their systems crossed specific danger thresholds.

Source: Future of Life Institute

Robotics

Teleoperated Humanoid Robots Complete Live Surgery for the First Time

UC San Diego engineers and surgeons report in the July 8 issue of Nature that they used teleoperated humanoid robots to complete two full procedures on large non-primate mammals: a gallbladder removal by a robot-and-human team, and a second surgery performed entirely by two robots working together with no human hands in the loop.

Source: UC San Diego Today

Product

Anthropic Ships a Dashboard That Tracks Your Own AI Dependence

Anthropic launched Claude Reflect in beta on July 9 for Free, Pro, and Max users with memory enabled — a dashboard showing usage topics and peak hours, prompts like “what’s one thing you want to keep doing yourself, even if Claude could do it faster,” and tools to set quiet hours and schedule breaks from AI use.

Source: TechCrunch


Analysis

Takeaway

Line today’s stories up and the capability curve keeps bending the same direction it always does — a cheaper model, a cheaper chip, two robots finishing a surgery with no human hands involved. What's thinning out is the layer that's supposed to keep pace with it: Microsoft cut jobs from the division that wasn't underperforming, and four labs graded on safety this week all quietly abandoned the pause commitments they made back when pausing was free. Progress isn’t the problem. It’s that nobody with the power to slow down has an incentive to.

— Boba


Curated by Vadym