OpenAI says it pursued the Navier-Stokes existence and smoothness problem with 10,000 AI agents that ran for 88 hours. The company then used Astra for 17 hours of formal verification in Lean. The claim matters because it shifts the story from a single-model stunt to a much larger test-time compute pipeline.
Henry Yin on test-time compute
Henry Yin, the founding partner of MoE Capital, described the approach as a new phase of test-time compute. On Late Talk, he said, “Previously, we might have one model think for a relatively long time to solve harder problems, but for particularly difficult problems, if a single model were to think it through, it might take a very, very long time — like decades or even a century.”
He also said, “They gave these agents some basic communication tools, and then trained the model itself to learn when sending a message is good, when to message others, when to ask for help, and when to change course.” That is the operational shift here: the system was not just thinking longer, it was coordinating across many agents and deciding when to hand work around.
Tristan Buckmaster and Codex
OpenAI said it decided in early September to attempt the problem after hearing that Tristan Buckmaster had made major progress on an important math problem. Buckmaster, the NYU mathematician, and his collaborator had previously given unpublished drafts to Codex for use. OpenAI then reached out to Buckmaster proactively and proposed excluding his collaborator from authorship because that person works at Anthropic.
Buckmaster made the communication process public online, turning the exchange into a public authorship dispute as well as a technical claim. OpenAI also said the two sides solved slightly different problems, one with viscosity and one without, and used different proof methods.
Lean verification in Astra
After obtaining the proof, OpenAI said Astra spent 17 hours completing formal verification in Lean. In practice, that means the result had to survive a machine-checked proof step rather than resting only on informal mathematical reasoning. The reported 130 billion tokens show how much text and intermediate reasoning the system burned through before that last check.
The unanswered question is the proof itself. OpenAI has made a strong compute claim, but the source does not show the full argument, so readers cannot yet judge whether the result will hold up outside the company’s own framing.







