ByteBulletin

[research] · · 4 min read

OpenAI claims Navier-Stokes solution, sparking data privacy dispute with NYU professor

OpenAI announced a solution to a 90-year-old math problem using 10,000 agents, but the claim is complicated by allegations that the model may have accessed private data from a competing research team.

By ByteBulletin Editors · Editorial Team


OpenAI has announced that it has discovered a solution to the Navier-Stokes existence and smoothness problem, a mathematical challenge that has remained unsolved for approximately 90 years. The announcement, made via a blog post on Tuesday, states that the company utilized an internal AI model—described as more powerful than its recently released GPT-6 Astra—alongside 10,000 concurrent agents to derive the proof. The Navier-Stokes problem is one of the seven Millennium Prize Problems, each carrying a $1 million reward from the Clay Mathematics Institute, though OpenAI has stated it does not intend to claim the prize. This development marks a significant milestone for the mathematics community, yet it arrives amidst immediate controversy regarding the integrity of the data used to train the model.

The controversy centers on the timeline of the discovery and the potential overlap with the work of New York University mathematics professor Tristan Buckmaster. Just one day before OpenAI’s public announcement, Buckmaster, in partnership with Anthropic researcher Levent Alpöge, published findings on a related problem. Buckmaster claims that he contacted OpenAI after learning the company had heard about their progress, only to discover that OpenAI had already produced a proof using a similar route. This proximity in timing has raised questions about whether OpenAI’s internal model had access to the private drafts and sessions that Buckmaster and Alpöge had been actively working on.

In a detailed statement, Buckmaster expressed concern that OpenAI may have accessed their Codex data to refine its solution. “I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project,” Buckmaster wrote. He noted that he was initially told the model did not look up user data, but when he pressed for clarification regarding training data, he did not receive a definitive answer. This lack of transparency has led Buckmaster to accuse OpenAI of using training data from a period after the researchers had already found their result, a claim he posted on Mastodon in response to OpenAI’s official statement.

OpenAI has moved to address these suspicions directly in its Tuesday announcement. The company stated that “no specific user data was accessed in order to solve this problem.” However, it added a caveat that “while unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.” Sébastien Bubeck, a member of technical staff at OpenAI, further clarified that the team did not see Buckmaster and Alpöge’s work until it was released publicly the previous night. Bubeck emphasized that the proofs differ significantly, stating, “One can in hindsight see that our proofs differ significantly and even the precise results proved are different.”

The Navier-Stokes equations describe the motion of liquids and gases, and the existence and smoothness problem asks whether solutions to these equations always exist and remain smooth, or if they can develop singularities. Solving this problem is considered one of the most difficult challenges in mathematics, with implications for fluid dynamics, aerodynamics, and weather prediction. OpenAI’s internal model, which began training on August 28th, reportedly exhibited “unprecedented performance in our benchmarks, including mathematics.” The use of 10,000 concurrent agents suggests a massive parallelization of reasoning tasks, a technique that has become increasingly common in advanced AI research.

For developers and researchers, this event highlights the growing intersection of AI capabilities and high-level mathematical reasoning. The ability of large language models to assist in proving complex theorems or solving long-standing problems represents a shift in how scientific discovery might be conducted. However, the controversy surrounding the Navier-Stokes solution also underscores the importance of data provenance and privacy in AI training. As AI models become more capable of handling specialized tasks, the question of how they are trained and what data they access becomes critical to maintaining trust in the scientific community.

What this means for developers is a reminder to be cautious when using AI tools for sensitive or proprietary work. If a model is trained on de-identified data from user sessions, there is a risk that proprietary information could inadvertently influence the model’s outputs. This is particularly relevant for researchers and companies working on cutting-edge problems where the integrity of the solution is paramount. Developers should be aware of the terms of service and data usage policies of the AI tools they employ, especially when working on projects that could have significant academic or commercial implications.

The situation also raises questions about the future of mathematical research. If AI models can solve problems that have stumped human mathematicians for decades, it could accelerate progress in fields such as physics, engineering, and computer science. However, it also challenges the traditional peer review process, as the validity of AI-generated proofs may be harder to verify. The mathematics community will likely need to develop new standards for evaluating AI-assisted research to ensure that the results are both correct and original.

What to watch in the coming weeks is the response from the Clay Mathematics Institute, which oversees the Millennium Prize Problems. The institute may need to determine whether OpenAI’s solution meets the criteria for the prize, despite the company’s decision not to claim it. Additionally, the mathematical community will likely scrutinize the proofs provided by both OpenAI and Buckmaster to determine their validity and originality. The outcome of this dispute could set a precedent for how AI-generated research is recognized and validated in the future.

In the end, the Navier-Stokes controversy is a reminder that while AI has the potential to revolutionize scientific discovery, it also introduces new challenges related to data privacy, intellectual property, and the integrity of the research process. As AI continues to advance, the scientific community will need to navigate these challenges carefully to ensure that the benefits of AI are realized without compromising the principles of open and transparent research.

SHARE

RELATED

← All stories