2026-09-19 · ~9 min read Past Edition View today's briefing →

10 picked from 90 candidates · ordered by significance

Today's Insight

The same Claude is now both the weapon that breached OpenAI and the tool auditing Anthropic itself, but Anthropic is still the one hiring and paying its auditor

Three security researchers at Hacktron AI used Claude Opus 5 to take over OpenAI employee accounts and reach a private code repository in under 72 hours. It started with a flaw they found in how Discourse, which runs OpenAI's community forum, handles HEIF images, the format iPhones use for photos, and OpenAI paid them $6,500 for reporting it. On the same day, Anthropic named consulting firm Accenture as its first "embedded evaluator" to vet its own models, agreeing to pay it at least $1 billion over five years. Since Anthropic is the one hiring and paying the auditor of its own models, it remains to be seen whether that review will ever be allowed to reach a conclusion that embarrasses the company.

Continue reading (2) Show less

Newly unsealed documents in the New York Times' lawsuit against OpenAI and Microsoft show both companies knew exactly what they were doing. Microsoft's director of applied science, Brent Hecht, called the companies' data scraping the "largest theft of labor in human history," and another internal document admitted the company's AI content strategy had started a "doom loop" that was eating away at the web itself. OpenAI's head of ChatGPT, Nick Turley, said that once a chatbot gives an answer, there is "no good reason to click" through to the original source, and the company's own estimates put the resulting drop in referral traffic to sites like the Times at as much as 60 percent. The documents make clear both companies saw this damage coming and pushed ahead anyway.

AI's reach kept expanding into homes and labs. Google launched a household agent called CC that up to six family members can share to manage schedules, shopping lists, and meal planning, while Anthropic opened a in the San Francisco Bay Area where robots run experiments so Claude can be put to work on drug discovery and clinical research. Google also hired a new slate of economists, including Nobel laureate Philippe Aghion, to study how AI is reshaping jobs and productivity. There was progress at the frontier of raw capability too: OpenAI reportedly moved closer to solving the Hodge Conjecture, one of math's that has resisted proof for over a century, and its new Astra for Law model scored more than 15 percentage points higher than a plain web search on the same legal-research benchmark.

Signal to watch Worth watching whether Accenture's first evaluation ever surfaces a finding that's actually unflattering to Anthropic.

New experts join Google’s AI & Economy team

Summary

Google has added four new experts to its AI & Economy research team, which studies how AI is reshaping the economy. Philippe Aghion, an INSEAD professor and 2025 Nobel laureate in economics, joins as an academic adviser, alongside directors Anu Madgavkar (formerly of the McKinsey Global Institute) and Daniel Rock (Wharton), plus visiting fellow Ajay Agrawal of the University of Toronto. They join a bench that already includes MIT's David Autor and Nobel laureates Michael Spence and Diane Coyle.

Why it matters

Why It Matters

Google is now both a maker of the large language models under study and the employer of the economists studying their labor impact, which raises the question of how independently that research can be published if the findings turn out unflattering to the company.

Read the original

Co-creating the future of fashion with Google

Summary

Google used its video-generation tool Flow to build two custom AI tools for New York Fashion Week designers Jane Wade and Sergio Hudson. Wade's "Styling Suite" lets her virtually dress digital models in clothes, accessories, hair, and makeup, while Hudson's "Runway Visualization" tool simulates stage lighting and props within a fixed budget. Both tools were built using only natural-language prompts, with no coding involved.

Why it matters

Why It Matters

Since Wade's stated goal was cutting down on in-person casting and fittings that used to eat up three full days, tools like this could end up narrowing the gap for smaller designers with limited time and staff more than for big-budget brands.

Read the original

Anthropic’s first embedded evaluator is … Accenture?

Summary

Anthropic announced consulting giant Accenture as its first "embedded evaluator" to vet its own models. The work will be handled by Faculty, the AI unit Accenture acquired in January, which will run model evaluations, , alignment assessments, and security checks. Anthropic will pay Accenture at least $1 billion over the next five years, and Accenture's stock jumped 8% in after-hours trading following the news. Anthropic said it plans to announce additional evaluators, including nonprofit group METR, in the coming weeks.

Why it matters

Why It Matters

Anthropic argued the deal makes it more verifiable rather than less accountable, but the conflict-of-interest concern remains since Anthropic itself is still the one choosing and paying its evaluator. Whether Accenture's first report ever surfaces a finding that's actually unflattering to Anthropic will be the first real test of how meaningful this arrangement is.

Read the original
See it drawn
A consulting firm is now paid billions to vet Anthropic's own models

OpenAI and Microsoft knew they were starting a ‘doom loop’ for the web

Summary

Newly unsealed court documents from the New York Times' copyright lawsuit against OpenAI and Microsoft, running to 92 pages, reveal that both companies' own staff warned about the damage they were causing. Microsoft's director of applied science, Brent Hecht, called ChatGPT and Copilot's data harvesting the "largest theft of labor in human history," while another internal Microsoft document admitted the company's AI content strategy had started a "doom loop" eating away at the web itself. OpenAI's head of ChatGPT, Nick Turley, is quoted saying there's "no good reason to click" through to a source once a chatbot answers a question, and the company's own estimates put the resulting drop in referral traffic to sites like the Times at as much as 60 percent. Microsoft pushed back, saying Hecht's comments "reflect one employee's individual perspective" rather than the company's position.

Why it matters

Why It Matters

Having internal admissions on record that both companies foresaw the harm but pressed on anyway makes an "we didn't know" defense much harder to sustain in this case. The acknowledgment of a steep drop in referral traffic could also become a reference point for other publishers weighing similar lawsuits.

Read the original

Security researchers used Claude to help them hack into OpenAI

Summary

A team of three independent security researchers at Hacktron used Anthropic's Claude Opus 4.8 and 5 to take over OpenAI employee accounts in under 72 hours, The Wall Street Journal reported. They achieved by exploiting a flaw in how Discourse, the third-party service hosting OpenAI's community forums, handles HEIF images, the format iPhones use for photos, gaining access to OpenAI's GitHub repository known as "Monorepo." They finished the exploit within a day of Claude Opus 5's release, and adapting the technique to other companies, including Slack, Meta, and GitHub Enterprise, took only a day or two and less than $3,000 in tokens. OpenAI paid them a $6,500 .

Why it matters

Why It Matters

Hacktron CTO Mohan Pedhapati said, "I don't think we are as strong as Chinese threat actors… We're just three guys with Claude and Codex subscriptions," which cuts the other way too: breaching a major company's systems no longer requires state backing, just a couple hundred dollars a month in AI subscriptions. As models keep improving, the cost of this kind of intrusion is likely to keep falling.

Read the original

Also covered by The Decoder

OpenAI takes aim at the legal market with Astra for Law

Summary

OpenAI unveiled Astra for Law, a version of its latest GPT-6 Astra model tuned for legal work. It can search more than 230 million URLs of US case law, statutes, and regulations, drawing on Free Law Project data covering more than 99.9% of US precedent. On Vals AI's Legal Research Bench, it answered 54% of 200 questions correctly, well above the 38.7% scored by plain GPT-6 Astra with ordinary web search. OpenAI also released 26 plugins connecting to legal software like Relativity and Clio, plus a data-retention-free "Trusted Access" program.

Why it matters

Why It Matters

A gap of more than 15 percentage points between the general-purpose version and the law-tuned one shows there's still a real difference between bolting web search onto a generic model and retraining it on industry-specific data. That gap matters most in accuracy-sensitive work like legal research.

Read the original
See it drawn
39 Web search 54 Astra Law +15pp
The law-tuned version lifted accuracy by more than 15 points on the same benchmark

Google’s new ‘CC’ is an AI agent that helps families run their households

Summary

Google unveiled CC, an AI agent designed to help families run their households together. Up to six family members can share email, calendars, chat, and task lists, letting CC track school schedules, bills, and sports activities to manage calendars, fill out forms, and even build shopping lists and weekly meal plans. It runs on Gemini and Google's coding agent Antigravity, with each household getting its own dedicated cloud computer. The test is opening to US Gmail users 18 and older, with existing users receiving email invitations.

Why it matters

Why It Matters

Letting several family members open their email and calendars to a single shared agent also means that if one person's account is compromised, the whole family's schedule and document access could be exposed along with it. How much visibility and control users get over that expanding access is the real question.

Read the original

OpenAI reportedly nears solving another Millennium Prize math problem, the Hodge Conjecture

오픈AI, 또 다른 밀레니엄 수학 문제 '호지 추측' 해결에 근접

Summary

OpenAI is reportedly nearing a solution to the Hodge Conjecture, another of math's , following its earlier work on the Navier-Stokes existence and smoothness problem, The Information reported on September 17. The Hodge Conjecture asks whether certain geometric features of shapes defined by polynomial equations can be expressed using simpler algebraic building blocks, one of seven Millennium Prize Problems named by the Clay Mathematics Institute in 2000, each carrying a $1 million reward. OpenAI used a next-generation pretrained model variant called "Doug" and expects to be able to automate a substantial share of mathematical research within six to nine months. Given the potential for controversy in academia, OpenAI plans to coordinate carefully with mathematicians before any announcement.

Why it matters

Why It Matters

Touching two problems that had resisted proof for over a century in succession suggests this isn't a fluke but a repeatable capability. Still, the plan to coordinate with mathematicians before announcing anything is itself a sign that even OpenAI is being cautious about how confident it is in the result before outside verification.

Read the original

Anthropic opens AI-powered biology research lab

Summary

Anthropic has opened a in the San Francisco Bay Area where Claude-powered robots run biology experiments in its place, Reuters reported. The company is developing a "Model Hardware Standard" for controlling lab equipment with medical research institute HHMI, and said its latest model, Claude Mythos 5.1, achieved a 50% success rate designing high-affinity binding molecules, well above the 10-15% typical of ordinary protein-design projects. A spokesperson said the company isn't specialized in drug discovery and is exploring licensing its findings to pharmaceutical companies instead.

Why it matters

Why It Matters

A success rate several times higher than conventional methods for designing candidate proteins could sharply cut the time and cost of drug discovery. But since Anthropic itself says it isn't a drug-discovery specialist, turning this into an actual treatment still depends on the next step of licensing deals with pharmaceutical companies.

Read the original

Anthropic wants you to know Claude leads a quarter of its research, but "lead" doesn't mean what you think

Summary

Anthropic released its first set of metrics on how it builds its own AI. As of August, Claude reached "Autonomy Level 4" (AL4) on 26% of the company's research work, up sharply from under 1% in February. The company's internal platform runs roughly 30,000 AI agents at once, blocks just 0.002% of more than a billion decisions, and allocated about 6% of July's AI research compute to safety work. But the word "leads" doesn't mean full autonomy (AL5); AL4 means Claude completes tasks like bug fixes under human supervision, and human raters only agreed with each other about 33% of the time when scoring the same work.

Why it matters

Why It Matters

Since the company generated these metrics itself, with Claude involved in setting the scoring criteria, and human raters only agreed with each other about a third of the time, it's too early to read "26% automated" as a reliable measure of real autonomy. Anthropic bringing in Accenture as an outside evaluator on the same day looks connected to exactly this credibility gap in its self-reported numbers.

Read the original

Terms in this briefing

Past Briefings

Monthly Reports

Weekly Recaps