← Today's briefing All topics

New model releases, benchmark results, training techniques, and research papers.

  1. How AI is expanding what people do at work

    OpenAI analyzed more than 800,000 U.S. ChatGPT work-related messages and found that 16.8% of work-related messages — and 43.5% of messages tied to a specific occupation — involved tasks normally associated with a different job, a pattern the company calls 'task crossover.' Examples include a small-business owner drafting ad copy or reviewing contracts, a salesperson digging into customer data, and a marketer fixing a website without a developer.

  2. Anthropic launches Claude Opus 5, a cheaper AI model for coding, agents and enterprise workflows

    Anthropic released Claude Opus 5 on July 24, keeping pricing at $5 per million input tokens and $25 per million output tokens — unchanged from Opus 4.8 — while more than doubling its score on the Frontier-Bench v0.1 agentic coding benchmark (43.3% vs. 18.7%). Opus 5 is now the default model on Claude Max and the strongest model available on Claude Pro.

  3. Making sense of the panic over Chinese AI

    The launch of Chinese startup Moonshot AI's new Kimi model triggered anxiety in Silicon Valley and on Wall Street about China's AI competitiveness, after Kimi posted results competitive with U.S. frontier models on some benchmarks. OpenAI and Anthropic have lobbied regulators to restrict Chinese open-source models.

  4. Team uses AlphaFold AI to redesign gene-editing proteins to make them safer

    A China-based research team published a Nature paper showing they used the AI protein-folding tool AlphaFold to identify which parts of the Cas9 protein used in CRISPR gene editing cause an off-target effect — edits to the wrong DNA sequence — then re-engineered those regions to cut off-target activity from 28% down to 5% while preserving on-target performance. The team, which tested 23 amino acid substitutions across 10 key positions and named its method 'ContactSeek,' also showed the approach worked with the related Cas12 protein.

  5. Microsoft launches new in-house AI models it says cut costs up to 89% versus OpenAI

    Microsoft released the image model MAI-Image-2.5-Pro and voice model MAI-Voice-2-Flash into public preview, backing the launch with its most aggressive case yet that it can run its products without relying on OpenAI's frontier models. In PowerPoint, it says MAI-Image-2.5 cuts GPU costs by up to 84% versus OpenAI's GPT-Image-2, and in the Dynamics 365 Contact Center used by customers like T-Mobile and EasyJet, MAI-Voice-2-Flash reportedly cut GPU costs by up to 89%. In Dragon Copilot, used by 170,000 medical providers, the company says transcription error rates across 58 languages dropped by an average of 50%.

  6. Anthropic launches Opus 5

    Anthropic launched a new model, Opus 5, on July 24. It's smaller than Fable 5 but beats it on several benchmarks, with a standout ability to verify its own work and iterate carefully — Anthropic points to an example of Opus 5 writing its own computer-vision pipeline from an incomplete prompt. It's cheaper and less restrictive than Fable 5, drops the 30-day data-retention policy that applies to Fable and Mythos, and is expected to trigger its safety classifier 85% less often than Fable 5.

  7. Anthropic's Opus 5 is about token efficiency, not a capability leap

    Anthropic released Opus 5, an update to the model that's become a popular choice for coding and software development. On benchmarks like Frontier-Bench and DeepSWE, it performs about on par with or slightly ahead of Anthropic's top-tier "Fable" model, and it beats Opus 4.8 and OpenAI's GPT-5.6-Sol on nearly every task. Pricing holds steady at $5 per million input tokens and $25 per million output tokens — the same as its predecessor, but cheaper than Fable. Anthropic itself says Opus 5 is deliberately "substantially behind" its Mythos 5 model at exploiting cybersecurity vulnerabilities, by design.

  8. Google announces Gemini 3.6 Flash and cybersecurity AI, teases 3.5 Pro and Gemini 4

    Google unveiled three new Gemini models at once: Gemini 3.6 Flash, 3.5 Flash-Lite, and its first cybersecurity-focused model, 3.5 Flash Cyber. On the DeepSWE coding benchmark, 3.6 Flash jumped to 49% accuracy from 3.5 Flash's 37%, while using about 17% fewer tokens. Notably absent again was Gemini 3.5 Pro, which Google had teased for a June launch at its I/O event back in May — instead, Google said it has already begun pretraining the next generation, Gemini 4.

  9. Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

    Google's official blog announced the launch of Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. Gemini 3.6 Flash now costs $1.50 per million input tokens and $7.50 per million output tokens, cheaper than its predecessor, while 3.5 Flash-Lite touts a response speed of 350 tokens per second, positioned as cheap enough to run agentic systems at scale. Both 3.6 Flash and 3.5 Flash-Lite are immediately available through the API and the Gemini app, among other channels.

  10. Introducing Gemini 3.5 Flash Cyber

    Google DeepMind unveiled Gemini 3.5 Flash Cyber, a lightweight model built specifically to find and help fix software vulnerabilities. In tests on the V8 JavaScript engine, it found 55 unique vulnerabilities, more than the standard 3.5 Flash (47) or Anthropic's Claude Opus 4.6 (36), and in a real deployment it uncovered remote code execution and memory-corruption bugs within two hours. Because of its dual-use potential for misuse, though, Google is only offering it through a limited pilot to governments and trusted partners rather than releasing it broadly.

  11. Accelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis Mission

    Google is contributing $40 million to the White House's Genesis Mission, a Department of Energy-led science initiative aiming to double the pace of U.S. scientific discovery within a decade. It's providing its frontier AI science toolkit — including AlphaEvolve, AlphaFold 3, AlphaGenome, WeatherNext, and AlphaEarth Foundations — free of charge, and giving tens of thousands of DOE national-lab staff a year of Gemini for Government. Pacific Northwest National Laboratory used AlphaEvolve to accelerate exploration of complex mathematical systems, while a Rockies-area national lab cut microscope calibration time from 90 minutes to 13 and reduced image-focusing steps from 50 to 2.