AI September 9, 2026 mixed ⇧ 3621 pts across 4 threads

OpenAI's Navier-Stokes claim ignites AI credibility debate

OpenAI published a blog post claiming its internal AI agents made meaningful progress on one of the six unsolved Millennium Prize problems in math, the Navier-Stokes equations. The agents sent 4.9 million messages and consumed roughly 300 billion output tokens, a sum that would cost an eye-watering amount at normal API rates. Tristan Buckmaster, a mathematician working on concurrent related work, posted a sharp response on Mastodon, and HN threads on both sides ran hot all day.

The pattern here is trust. Multiple commenters noted that OpenAI framed this as a breakthrough before peer review, and that the concurrent work section of their post points to independent math that complicates the claim. One commenter called it hard to see how this wouldn't qualify as misrepresentation of model capabilities, a line you don't often see directed at a major lab so bluntly. The Terry Tao thread on open math problems being 'mined' by AI added fuel, with people debating whether AI is genuinely advancing math or just finding the low-hanging fruit that human mathematicians left documented online.

There is a real counterpoint: even if the specific Navier-Stokes claim is overstated, the ability to run million-message agent swarms on hard problems is genuinely new. The question is whether OpenAI earned the right to announce it the way they did.


So what?

If the math community successfully pushes back on this claim, it will make it harder for AI labs to use scientific breakthroughs as marketing. Founders building on top of AI capabilities should track how quickly lab announcements get stress-tested, because the gap between announcement and scrutiny is shrinking fast. Don't build roadmaps around headlines.

Read these