New model releases, benchmark results, training techniques, and research papers.
Current topic Models & Research · 77 Switch topic
Latest stories
-
OpenAI just wants to win
OpenAI claims to have solved the Navier-Stokes equation, one of the Millennium Prize Problems, but the way it got there has infuriated mathematicians. The company says it took about 10,000 agents and tens of millions of dollars in compute over 88 hours, and once it learned NYU professor Tristan Buckmaster and Anthropic researcher Levent Alpöge were closing in on the same problem, it offered Buckmaster near-unlimited compute and sole authorship of its paper, but not Alpöge. Buckmaster called the offer a "bribe," rejected it, and went public with concerns that his own Codex prompts from working the problem may have been fed back into OpenAI's training.
-
OpenAI's feud with mathematicians is only escalating
NYU professor Tristan Buckmaster publicly questioned whether OpenAI took research he'd co-developed using Codex and published it as its own proof of the Navier-Stokes equations. He said OpenAI pressured him over how collaborator credit would be listed, and OpenAI subsequently withdrew its sponsorship of a Caltech math event. In an open letter, 25 Fields Medal winners warned that proofs are being announced in a rush with no time for proper verification, and that AI-generated ideas never become part of the mathematical canon without mathematicians willing to refine them.
-
Mathematicians want proof OpenAI didn’t use their work
German mathematician Andreas Thom accused OpenAI of unethical conduct, saying the company's recent breakthrough on non-sofic groups may have drawn on unpublished work by him and a colleague. He says OpenAI researchers only answered whether his ChatGPT conversations were directly accessible, not whether they entered the vast pools of training data, calling it the same kind of evasive denial OpenAI gave when announcing its Navier-Stokes solution.
-
OpenAI's GPT-Live-1 API lets developers build apps that talk and listen at the same time
OpenAI released GPT-Live-1, an API version of the voice model already used in ChatGPT, for developers. It scored 80.1% on a full-duplex interaction test, meaning it can listen and speak at once, far above the previous model's 45.4%, and cut turn-taking latency from 1.4 seconds to 0.8. It costs 5 cents per minute, and Yelp is already using it to handle phone reservations.
-
What OpenAI’s latest controversy tells us about the future of math
OpenAI said on September 9 that an unreleased model solved the Navier-Stokes existence and smoothness problem, a Millennium Prize Problem that had stood unsolved for 87 years, in just 88 hours using a swarm of 10,000 AI agents. A day earlier, NYU mathematician Tristan Buckmaster and Anthropic researcher Levent Alpöge had already published a proof of a simplified version of the same problem. Buckmaster says an OpenAI researcher warned he would ruin his career by going public with what happened, while OpenAI denies training its model on Alpöge's unpublished work.
-
OpenAI Just Claimed a Huge Math Discovery. Some Academics Are Crying Foul
OpenAI announced on September 8 that an internal AI model had solved the existence and smoothness problem for the Navier-Stokes equations, a 200-year-old fluid-dynamics problem and one of the seven Millennium Prize Problems, each worth $1 million. The company says it began training the model on August 28 and used as many as 10,000 agents over more than 50 hours before landing on a Lean-formalized proof. NYU mathematician Tristan Buckmaster says he and Anthropic researcher Levent Alpoge had been working toward a related result first, and that OpenAI rushed to the same approach after learning of their progress.
-
Update to Google’s AI weather model improves forecast accuracy
Google unveiled WeatherNext 3, an update to its AI weather forecasting model. Unlike prior versions that relied solely on six-hourly reanalysis snapshots, the new model ingests raw satellite data directly and updates forecasts hourly. Google's white paper reports roughly a 5 percent accuracy gain for upper-atmosphere conditions, equivalent to about six more hours of reliable forecast lead time, and up to 30 percent better accuracy for surface temperature and dew point at specific locations.
-
AlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genome
Google DeepMind released AlphaGenome Atlas on September 8, predicting the molecular impact of roughly 9 billion possible single-nucleotide variants across the human genome. The dataset spans 1 petabyte, more than 30 times larger than the AlphaFold Database. Its 'AVI score' reduces each variant's impact to a single number, covering not just the 2 percent of the genome that codes for genes but the remaining 98 percent of non-coding regions too.
-
OpenAI's GPT-6 Astra clears 3D puzzle game Portal without human help
OpenAI's GPT-6 Astra completed the 3D puzzle game Portal from start to finish without any human controlling it. Game creator cozyblaze posted video of the run on X on September 5; it took about $571 in API costs, roughly 24 hours, and 3,336 tool calls. It's the first time a general-purpose AI agent has autonomously finished a 3D game that requires spatial reasoning.
-
Opaque recurrence, and other AI terms that you should probably know
TechCrunch published a glossary of 36 frequently used AI industry terms, including AGI, hallucination, chain-of-thought, and mixture of experts. It noted that OpenAI's new GPT-6 Astra model uses a technique called "opaque recurrence" that has raised concern among safety researchers. The glossary is maintained as a living document that gets updated over time.
September 202613
- OpenAI developer claims Astra boosted productivity so much it pulled some plans forward by six months ↗
- OpenAI's Astra Tops Web Dev Arena, Overtaking Anthropic ↗
- Google's WeatherNext 3 ditches physics simulations and learns weather directly from live satellite data ↗
- Deepmind put 100 AI agents in a room and they sorted into cheaters, converts, and whistleblowers ↗
- Why OpenAI Led With Safety, Not Performance, in Its Astra Launch ↗
- OpenAI's GPT-6 Astra hallucinates less but remains vulnerable to hidden prompt injections ↗
- GPT-6 Astra Is Here—and OpenAI Thinks It May Kick Off the AGI Era ↗
- Introducing WeatherNext 3, our most advanced and accurate global weather AI model ↗
- Introducing Gemini 3.8 Flash and 3.8 Flash Cyber ↗
- Google: LLM Hallucination Stems From Failed Memory Retrieval, Not Missing Knowledge ↗
- Meta's Muse Spark 1.3 Launches, Overtakes OpenAI to Rank Third Worldwide ↗
- OpenAI Is About to Release Its First AI Model With 'Critical' Cyber Abilities ↗
- Anthropic launches Claude Fable 5.1 and says it's up to 45 percent cheaper for agentic work ↗
August 202642
- Google's WikiSkill gives AI agents a persistent memory of past mistakes to sharpen future performance ↗
- An Anthropic researcher just gave us a peek at self-improving AI ↗
- Meta researchers taught an 8B AI model to match Claude Opus 4.5 — without the frontier price tag ↗
- Google Deepmind's AI Co-Scientist now plans experiments, runs lab equipment, and writes scientific papers ↗
- Gemini Omni 1.1 Flash lets you build with more control ↗
- Jensen Huang says Nvidia achieved AGI, again — not that it matters ↗
- Intelligent transcription with Gemini 3.5 Transcribe ↗
- Sam Altman says OpenAI will have AGI by the end of 2026 if you accept his definition ↗
- Is Google skipping Gemini 3.5 Pro to prepare next-gen Gemini 4? ↗
- Thomson Reuters bets $40M on owning its AI instead of renting from OpenAI or Anthropic ↗
- Nvidia Cracks the Model-Switching Bottleneck With Linear KV Cache Transfer ↗
- Kids outlearn AI—and we still don't know why ↗
- Who’s behind the new ‘stealth model’ Ox Alpha? ↗
- 'The harness matters more than the LLM': Nvidia's AVO proves long-horizon agents work ↗
- Inherent, founded by DeepMind alumni, says its AI ‘teammate’ just outperformed Anthropic and OpenAI at replicating research ↗
- Study explains why AI agents benefit from "skills" and when they fail ↗
- World models that ignore human beliefs predict the wrong actions, new research shows ↗
- Nvidia just showed that the harness, not the AI model, is now the real hero ↗
- From Atari to EVE Online: Building on 15 Years of AI Research in Games ↗
- Nvidia finds that simple linear math can replace costly AI model handoffs ↗
- AI’s recursive self-improvement might not come so quickly after all ↗
- Upstage's Solar LLM Now Powers 100% of Daum's AI Search Summaries ↗
- Mistral AI ships 'OCR 4.1' with improved document layout preservation ↗
- Alibaba releases 'Qwen3.8-Max' weights, requiring a license for commercial use ↗
- Alibaba releases 'Qwen3.8-27B' weights, claiming Claude Opus 4.6-level performance ↗
- OpenAI and Anthropic in price war as Chinese AI rivals gain ground ↗
- Introducing Gemini 3.7 Flash ↗
- AMIE, our research medical AI system, demonstrates real-time clinical video consultation capabilities in a first-of-its-kind study. ↗
- OpenAI launches GPT-5.6-Cyber with reduced refusals, 95% completion on advanced cybersecurity tasks ↗
- A New Trick Reveals AI Models' Inner Thoughts ↗
- AI Is Dead. Organoids Are Alive ↗
- 1 Million Virtual Humans Test AI Before Launch: Enter 'MatrAIx' ↗
- These startups are chasing the next big thing in LLMs ↗
- DeepMind's hurricane breakthrough has surprised weather scientists ↗
- Responding to the next frontier of critical cyber capabilities ↗
- ByteDance trains massive AI model in bid to rival Anthropic ↗
- Scientists Used AI to Create 16 New Viruses ↗
- Improving GPT‑5.6 Sol in ChatGPT—and expanding access to GPT-5.6 Luna for free users ↗
- Google's Top AI Brains Are Leaving to Launch Discovery Loop ↗
- The latest AI news we announced in July 2026 ↗
- China’s Alibaba takes another swipe at America’s AI supremacy ↗
- A Tutorial on GeoAI: Designing Footprint Extraction from NAIP Imagery Using U-Net, Grounding DINO, SAM, and Mask R-CNN ↗
July 202612
- Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration ↗
- How AI is expanding what people do at work ↗
- Anthropic launches Claude Opus 5, a cheaper AI model for coding, agents and enterprise workflows ↗
- Making sense of the panic over Chinese AI ↗
- Team uses AlphaFold AI to redesign gene-editing proteins to make them safer ↗
- Microsoft launches new in-house AI models it says cut costs up to 89% versus OpenAI ↗
- Anthropic launches Opus 5 ↗
- Anthropic's Opus 5 is about token efficiency, not a capability leap ↗
- Google announces Gemini 3.6 Flash and cybersecurity AI, teases 3.5 Pro and Gemini 4 ↗
- Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber ↗
- Introducing Gemini 3.5 Flash Cyber ↗
- Accelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis Mission ↗