OpenAI Says It Solved a $1M Math Problem in 88 Hours. Nobody Outside OpenAI Has Confirmed That Yet — and a Rival Mathematician Says They Took His Work to Get There. — ai

OpenAI Says It Solved a $1M Math Problem in 88 Hours. Nobody Outside OpenAI Has Confirmed That Yet — and a Rival Mathematician Says They Took His Work to Get There.

OpenAI announced on September 8, 2026 that roughly 10,000 AI agents produced a proof for the Navier-Stokes Millennium Prize problem in 88 hours. The Clay Mathematics Institute hasn't verified it. An NYU mathematician says OpenAI pressured him over credit. Here's a practical checklist for evaluating any AI vendor's headline claim before you build on it.

Note: This broke September 8, 2026 — one day old as of writing. OpenAI’s proof has not been independently verified by the Clay Mathematics Institute or a peer-reviewed publication as of this writing; treat the capability claim itself as unresolved, not confirmed.

Five days ago, Anthropic said Claude had produced a complete, computer-checked proof of Fermat’s Last Theorem in Lean — a claim that was trustworthy specifically because Lean’s compiler mechanically checks every logical step, with no human required to take the result on faith. This week, a different AI lab made a structurally similar-sounding announcement that fell apart into a much messier story within about 24 hours. The difference between the two is the entire lesson here, and it’s one that applies well beyond mathematics — to every “our AI just did something remarkable” claim a vendor puts in front of you.

On September 8, OpenAI announced that an internal system — roughly 10,000 AI agents working over about 88 hours, using around 130 billion output tokens and $40 million-plus in compute — had produced a proof addressing the Navier-Stokes existence and smoothness problem, one of the Clay Mathematics Institute’s seven Millennium Prize Problems, each carrying a $1 million reward, unsolved for roughly 90 years. Within hours, NYU mathematician Tristan Buckmaster published his own statement accusing OpenAI of learning about his unpublished, nearly year-long research approach and using pressure tactics to strip his collaborator’s credit before OpenAI’s announcement went out. Nobody outside OpenAI has independently confirmed the proof is actually correct.

That’s three separate, unresolved questions stacked on top of each other in a single announcement: is the proof right, whose idea was it, and did OpenAI’s process actually work the way the announcement implies. For a small-shop reader deciding whether to trust the next headline capability claim from any AI vendor, that stack is the actual story — more than the math is.

What OpenAI actually claimed, and what it didn’t

OpenAI’s own announcement is more careful than the “AI solves 90-year-old math problem” headlines it generated. The company published a roughly 165-page written proof alongside a Lean 4 formalization — a machine-checkable version outside researchers can download and run themselves — and explicitly stated it is not claiming the Clay Institute’s $1 million prize. That’s a meaningful hedge: OpenAI is offering the proof as evidence of capability, not submitting it through the actual verification process that would settle whether it’s correct by the field’s own standard.

DetailValue
AnnouncedSeptember 8, 2026
Method~10,000 AI agents, ~88 hours, ~130 billion output tokens
Estimated compute cost$40 million-plus
Output165-page written proof + Lean 4 formalization
Problem addressedNavier-Stokes existence and smoothness (finite-time blowup for a specific case)
Prize claimedNone — OpenAI explicitly did not submit for the $1M Clay Institute prize
Forced vs. unforcedClay Institute’s official Millennium criteria concern the unforced Navier-Stokes equations; OpenAI’s result addresses the forced version
Independent verificationNot yet completed as of this writing

Worth being precise about what was actually solved, since “cracked the Navier-Stokes problem” oversells it slightly: multiple reports indicate the proof addresses a narrower, “forced” version of the blowup question — the same specific sub-case Buckmaster says he and his collaborator had already been pursuing for months, because it’s understood to be the more tractable angle of attack rather than the full, general Millennium Prize statement. The technique itself has a documented lineage that predates both teams’ work this year: it extends a “forcing” approach developed by mathematicians Diego Córdoba and Luis Martínez-Zoroa, and Buckmaster and Alpöge’s own August 15 preprints — three separate papers, each with its own Lean formalization — established finite-time blowup with smooth forcing for the incompressible porous medium equation and the two-dimensional Boussinesq system before extending the same method to three-dimensional incompressible Euler. OpenAI’s team, led by Sébastien Bubeck (who runs the company’s math research group), pushed that same forcing argument the rest of the way to the full Navier-Stokes equations. That lineage matters for calibrating the “10,000 agents solved a 90-year-old problem from scratch” framing: the agents extended a specific, recent, human-originated technique to its next logical target, rather than independently inventing a new approach with no antecedent. Extending a promising technique quickly is still a genuinely useful capability — it’s just a different claim than the headline implies.

This mirrors a pattern from the Fermat proof five days earlier, where mathematician Kevin Buzzard noted Anthropic’s proof covered primes p≥17, not the fully general case — genuinely impressive, mechanically verified work that was nonetheless narrower than the headline implied. That’s not a knock on either result. It’s a reminder that the gap between a lab’s press framing and the actual scope of a technical result is a recurring pattern worth checking every time, not a one-off.

The forced-versus-unforced distinction is worth spelling out plainly, because it’s more precise than “narrower sub-case” and it’s the actual technical reason OpenAI didn’t submit for the prize: the Clay Mathematics Institute’s official Millennium Prize criteria concern the unforced Navier-Stokes equations — the version with no external force term added to the system. OpenAI’s result, like Buckmaster and Alpöge’s before it, addresses the forced version, where an engineered forcing term is allowed to drive the blowup. That’s not a technicality to wave away — it’s the difference between “the Clay Institute’s actual open question” and “a closely related, still-genuinely-hard question that doesn’t meet the prize’s own definition.” OpenAI’s announcement doesn’t hide this; it’s explicit in the framing. But it’s easy to lose in a headline that reads “AI solves 90-year-old math problem,” and worth checking in any coverage that doesn’t mention it at all.

The credit dispute, in the mathematicians’ own words

Here’s where this diverges hard from the Fermat story. According to Buckmaster’s own published statement and reporting from TechCrunch and Fortune, the timeline runs roughly like this:

Buckmaster and Levent Alpöge — a mathematician who works at Anthropic — had spent close to a year working on the forced Navier-Stokes blowup problem, arriving at their own solution around August 22, 2026. OpenAI began its own effort on September 1, by its own account triggered by “rumors” that two Millennium Prize problems had been solved. By September 6, OpenAI had a result, checked with Lean, and reached out to Buckmaster proposing a concurrent, simultaneous release of both results. Buckmaster’s statement alleges that on that call, Bubeck told him OpenAI’s internal system had already produced roughly a 100-page proof for the same forced Navier-Stokes approach — the specific narrow angle Buckmaster says almost nobody else in the field was actively pursuing at the time — and then presented him with two options: either OpenAI would publish the day after Buckmaster’s team released theirs, or Buckmaster could write up the paper alone, with Alpöge left off the credit specifically because he works at a rival lab (Anthropic). When Buckmaster said he’d rather go public independently instead of taking either option, Bubeck told him, “Why would you ruin your career?” and later, “If you don’t want me to be nice, then I don’t have to be nice.”

OpenAI’s response, also on the record: “We (the researchers and the agents) did not see any of their work through any means until they released it publicly,” and separately, “We did not use their prompts or proofs to prompt our models or direct our agents.” The company acknowledged it can’t rule out that de-identified usage data from Buckmaster or Alpöge’s own product interactions with OpenAI’s tools might have somehow contributed to model improvements generally, while maintaining the two proofs “differ significantly.” Bubeck personally denied the characterization of his own comments, posting on X that “a series of false and inflammatory allegations against me are currently circulating on social channels,” while also stating: “I want to be extremely clear that we recognize the priority of Levent Alpöge and Tristan Buckmaster’s work.”

Both things can be true at once, and probably are: OpenAI’s agents may genuinely have arrived at a result through their own process without directly copying anyone’s unpublished work, and a human employee at the company may have handled the human relationship with a rival researcher badly, in ways that look exactly like leveraging a much bigger platform’s announcement timing against an individual academic’s ability to control his own credit. Those are different failure modes with different fixes, and conflating them into one either “OpenAI cheated” or “OpenAI innovated” story misses what actually happened.

Mathematician Terence Tao, not directly involved in the dispute, offered the sharpest framing of the underlying pattern, warning that treating open problems as targets for AI labs to “solve” for publicity risks damaging the field’s own ecosystem: “The indiscriminate strip-mining of open problems for solutions may destroy the ecosystem from which the next generation of mathematical techniques, problems, and practitioners would have developed” — comparing it to “using excavators to loot an archaeological site, destroying the context needed to give treasures any historical meaning.”

Why the verification gap is the part that should actually worry a builder

Step back from the interpersonal drama for a second, because the technical verification question is the one with a direct parallel to how you evaluate any vendor’s AI capability claim in your own business.

The Fermat proof from five days earlier was trustworthy specifically because Lean’s compiler is a mechanical, unbribable, un-gameable verifier — it either accepts the proof’s logic or it doesn’t, and no amount of PR framing changes that outcome. OpenAI’s Navier-Stokes announcement also includes a Lean 4 formalization, which is a genuinely good sign and a meaningfully higher bar than a pure prose claim. But “we published a Lean file” and “the Lean file has been independently compiled, reviewed, and confirmed to actually formalize the claimed result correctly by people outside the company that built it” are two different levels of confidence, and as of this writing, outside mathematicians are still in the process of checking, not finished. The Clay Mathematics Institute — the actual body whose verification would settle this by the field’s own standard — has made no announcement.

This is the exact same gap that shows up constantly in AI vendor benchmark claims for ordinary business tools, just with higher production values. A vendor announces a headline number — “our agent completed X% of tasks autonomously,” “our model beats GPT-6 on Y benchmark” — and the announcement is the news cycle. The independent replication, if it happens at all, happens quietly, weeks later, with a fraction of the attention, and sometimes contradicts the original claim outright. The piece on why an agent got noticeably worse right after its underlying model supposedly got better is the small-shop version of exactly this pattern: the vendor’s own stated benchmark improvement didn’t match the lived experience of people actually running the thing. The same gap showed up days later when Tesla put cars with no steering wheel on public roads and NHTSA opened an audit within hours — a vendor’s own self-certification of a capability claim is not the same thing as an outside regulator, or an outside mathematician, actually checking it.

A practical checklist: evaluating a vendor’s capability claim before you build on it

None of this is a reason to distrust every AI announcement reflexively — plenty of vendor claims hold up fine. It’s a reason to have a standard, boring process for checking before you commit real business decisions to a claim, rather than reacting to the headline. Five questions, in order:

  1. Is there an independent, mechanical verification step, or is it the vendor’s own word? A Lean/Coq/Isabelle-style formal proof, a third-party benchmark run on the vendor’s own model weights, or a reproducible eval suite is a different category of evidence than a blog post describing internal results. Ask which one you’re actually looking at.

  2. Has anyone outside the vendor’s own team confirmed the result yet? If the honest answer is “not yet, it’s only been a day,” that’s fine — just don’t build a business-critical decision on a claim that’s still in that window. Wait for the independent check, or treat the claim as a promising signal rather than a settled fact.

  3. Does the scope match the headline, or is the headline broader than the actual result? Both the Fermat proof (primes p≥17, not the fully general case) and this Navier-Stokes proof (the narrower “forced” sub-case) were genuinely impressive results oversold slightly by their own framing. Read past the headline to the actual stated scope before deciding it solves your specific problem.

  4. Is the vendor claiming the full, official version of the benchmark, or a related-but-different one? OpenAI explicitly did not claim the Clay Institute’s $1 million prize — a meaningful tell that even the company making the announcement isn’t confident enough in the result to put it through the process that would make it official. Watch for that kind of hedge; it’s more informative than the headline claim itself.

  5. Who benefits from you believing this today, versus after it’s actually verified? A capability announcement timed to a funding round, a product launch, or a competitive news cycle carries a different incentive structure than a peer-reviewed paper with no launch attached. That doesn’t make the claim false — it means the timing pressure to believe it now, before verification, is coming from somewhere other than the strength of the evidence.

When skepticism here is actually the wrong move

It’s worth being fair about the other side of this. OpenAI did publish a full written proof and a Lean formalization openly, rather than just asserting a result — that’s a real, checkable artifact, not vaporware, and independent mathematicians are actively working through it right now rather than being asked to simply take the company’s word for it. If you’re evaluating a vendor claim that comes with genuine reproducible evidence — open weights, a public eval anyone can rerun, a formal proof file anyone can compile — treating that identically to an unverifiable marketing claim is its own kind of error. The checklist above is about calibrating trust to the evidence actually provided, not defaulting to blanket cynicism about every AI announcement. Some of them really do hold up.

Quick answers

Is the proof actually correct? Unknown as of this writing. OpenAI published a written proof and a Lean 4 formalization that anyone can attempt to compile and check, which is a real, reviewable artifact — but independent mathematicians and the Clay Mathematics Institute had not completed that review at the time of publication. “Published with checkable evidence” and “confirmed correct by outside experts” are different stages, and this result is currently at the first one.

Did OpenAI actually steal Buckmaster’s work? OpenAI directly denies seeing Buckmaster and Alpöge’s unpublished work before it was released publicly, and says its proof differs significantly from theirs. Buckmaster’s statement doesn’t claim direct copying either — his allegation is about pressure over credit and timing around the announcement, not that OpenAI’s mathematical content was lifted from his paper. Those are separate claims; don’t conflate them into one bigger accusation than either side is actually making.

Why does it matter that OpenAI didn’t claim the $1 million prize? Submitting for the Clay Institute’s prize means putting the result through the field’s actual verification process, with the institute’s own reviewers checking it against the precise, formal statement of the problem. Not doing that — while still publicizing the result as if it settles a 90-year-old open question — is a choice that gets less scrutiny in the headlines than it deserves.

Sources

All facts, quotes, and dates accessed September 9, 2026.

Bottom line

Strip away the interpersonal drama and this is a clean case study in the gap between “a lab announced it” and “it’s been verified” — a gap that exists in every AI capability claim, not just math proofs. OpenAI published real, checkable artifacts, which is more than a lot of vendor announcements offer, and that’s genuinely to its credit. But as of this writing, no independent body has confirmed the result, a mathematician with a direct stake in it says the process behind it was handled badly, and the company’s own choice not to claim the actual prize is itself a signal worth reading. Treat every headline AI capability claim — yours or anyone else’s — the same way: impressive is not the same as verified, and the five minutes it takes to check which one you’re looking at is cheaper than building a decision on the wrong one.

[read next]
ai agents · sep 13
Anthropic's CEO Says an AI Swarm Could Take Over the Internet Within a Year. Here's the Boring Version of That Problem You Actually Have Today.
hardware · sep 13
700 AI Agents Coordinated a Hack Without Anyone Noticing Until After. The $289 Box That Would Have Caught It Sooner.