OpenAI’s 10,000-Agent AI System Produces Proposed Solution to Navier-Stokes Millennium Problem

OpenAI recently announced a remarkable development in the realm of theoretical mathematics, reporting that an internal artificial intelligence system has produced a proposed solution to the Navier-Stokes existence and smoothness problem. This legendary challenge stands as one of the seven Millennium Prize Problems designated by the Clay Mathematics Institute in the year 2000, each carrying a one-million-dollar bounty for a verified resolution.

More specifically, the internal AI system constructed a finite-time singularity for the three-dimensional Navier-Stokes equations accompanied by a smooth external force. This route is explicitly permitted by the official formulation of the problem provided by the Clay Mathematics Institute. It is important to note that this breakthrough does not settle the separate, widely known question regarding whether unforced equations always remain smooth.

According to OpenAI’s official report, the monumental effort required roughly 10,000 concurrent AI agents working in tandem. Over the course of the experiment, these agents exchanged approximately 2.7 million messages and generated around 130 billion output tokens. The distributed system reached its proposed solution roughly 88 hours after the experiment commenced. Following this generation phase, the system underwent an additional 17 hours of formalization and verification using Lean, an interactive theorem prover.

At first glance, this achievement feels like the realization of long-held science fiction tropes regarding autonomous artificial intelligence. The narrative suggests handing a centuries-old puzzle to a machine, stepping away for a few days, and receiving a complete mathematical proof in return. However, a closer examination of how OpenAI reached this milestone reveals a much more nuanced and intricate story involving human collaboration, massive computational scaling, and inevitable academic friction over attribution.

The Story Started With Two Human Mathematicians

Before OpenAI unleashed its swarm of 10,000 agents, human mathematicians Tristan Buckmaster of New York University and Levent Alpöge, who works at Anthropic, had already been making substantial progress on closely related fluid dynamics problems. Notably, their research was also far from traditional pencil-and-paper work.

Buckmaster and Alpöge incorporated advanced computational tools into their workflow, utilizing systems like Claude and OpenAI Codex. Furthermore, they formally verified their results using Lean. Their research successfully constructed a finite-time blowup for the three-dimensional incompressible Euler equations utilizing a smooth forcing mechanism.

While their specific breakthrough did not directly solve the Navier-Stokes equations, it pushed deep into a heavily overlapping mathematical frontier.

The trajectory shifted dramatically at the beginning of September 2026. According to OpenAI, rumors began circulating on September 1 that two separate Millennium Prize problems had been successfully solved. These rumors, paired with exceptionally strong results from a newly trained internal model, prompted company leadership to deploy the system against the remaining open Millennium Prize problems. OpenAI later confirmed that the circulating rumors were tied directly to the ongoing work of Buckmaster and Alpöge.

Consequently, the artificial intelligence did not spontaneously wake up one morning, independently select the Navier-Stokes problem, and derive a solution from a vacuum. Human researchers were actively blazing a parallel trail along the exact same mathematical frontier. However, it was OpenAI’s own internal progress with the closely related Euler equations that ultimately convinced company executives to concentrate their massive computational resources specifically on Navier-Stokes.

Then OpenAI Scaled It to 10,000 Agents

It was at this juncture that OpenAI implemented a strategy that departed significantly from traditional computational research. The organization effectively established a gigantic, hyper-fast virtual research laboratory.

The system was structured as a network of autonomous agents divided into specialized groups. These groups possessed the capability to communicate internally, execute custom code, and access a cached version of the internet. Different sub-groups were assigned to explore entirely distinct mathematical approaches and hypotheses.

In the early stages of the experiment, a cohort of nearly 100 agents spent approximately 50 hours investigating an Euler-related problem, successfully yielding a promising foundational result. At that crucial turning point, OpenAI dynamically reallocated its agent workforce, pulling units away from other Millennium Prize investigations and concentrating them entirely on Navier-Stokes.

As the experiment scaled, the system began facilitating the exchange of useful discoveries between disparate agent groups. OpenAI terms this process cross-pollination. The OpenAI Codex tool consolidated promising intermediate results generated by individual agent clusters, and those synthesized insights were systematically fed back into subsequent system prompts.

Eventually, the final, successful push involved approximately 10,000 concurrent AI agents operating simultaneously. This monumental scaling may ultimately represent the most significant takeaway from the entire experiment.

A single human mathematician can thoroughly explore a handful of distinct hypotheses over a career. A dedicated academic research group might test a few dozen more. By contrast, a network of 10,000 active agents can investigate vast constellations of potential pathways concurrently, instantly discarding dead ends while rapidly stacking proofs on top of promising intermediate results.

Viewed through this lens, the true breakthrough may not be that artificial intelligence suddenly evolved into an intuitive mathematical genius. Instead, the breakthrough lies in the realization that AI can function as an unprecedentedly scalable, high-throughput research workforce.

Then the Controversy Started

The announcement of the proposed solution immediately triggered intense academic debate, spearheaded by Tristan Buckmaster, who raised deeply uncomfortable questions regarding data provenance and academic credit.

Both Buckmaster and Alpöge had utilized OpenAI’s Codex tool extensively while developing their own research. According to Buckmaster, they had routinely uploaded draft materials and working notes from their private project into the platform during their development cycles. Consequently, he approached OpenAI to inquire whether their newly trained model had been exposed to, or trained upon, those specific user sessions.

As documented in contemporary journalistic accounts of the dispute, Buckmaster noted that his initial inquiries regarding whether the model actively looked up user data received evasive or incomplete answers, particularly when pressing for details regarding model training procedures. Buckmaster maintained a cautious stance in public statements, explicitly stating that he did not know what OpenAI’s model actually did or how it processed their inputs, leaving open the question of whether their data had inadvertently influenced the machine.

In response to the mounting concerns, OpenAI conducted an internal review and subsequently updated its public report. The company asserted that Buckmaster’s Codex prompts from the preceding two months could not have influenced the system in any way, including through model training. Furthermore, OpenAI maintained that its internal researchers and autonomous agents had no prior visibility into the unpublished work of Buckmaster and Alpöge before it entered the public domain. Based on currently available public evidence, no definitive proof exists to suggest that OpenAI trained its systems on unpublished proofs or copied private research.

However, a secondary and more intractable dispute quickly surfaced regarding the attribution of credit.

With two separate achievements now on the table—the human-led Euler research by Buckmaster and Alpöge, and the AI-generated Navier-Stokes proof from OpenAI—administrative discussions began regarding joint presentation or publication. According to Buckmaster, OpenAI researcher Sébastien Bubeck presented two primary pathways forward.

The first option suggested that Buckmaster and Alpöge should rush to publish their Euler results independently before OpenAI officially released its Navier-Stokes proof. The second option proved significantly more controversial. Buckmaster stated he was offered the opportunity to serve as an author on a paper presenting OpenAI’s Navier-Stokes proof, provided that the publication explicitly acknowledged that an OpenAI model had generated the underlying mathematics.

Under that proposed arrangement, however, Levent Alpöge would be excluded from the author list entirely due to his current employment at rival artificial intelligence firm Anthropic. Buckmaster ultimately declined the offer.

Bubeck later clarified his position, stating that he had proposed Buckmaster as the lead author for a rewritten presentation showcasing OpenAI’s computational proof, noting that he viewed it as inappropriate for an active Anthropic employee to co-author an OpenAI-generated project. Bubeck also clarified that he had never proposed removing Alpöge from the separate, human-led Euler paper.

This sharp disagreement highlights a systemic structural flaw in modern academic publishing, which was never designed to accommodate hybrid workflows. When human researchers establish foundational concepts, AI tools assist in day-to-day exploration, an automated system generates the final mathematical proof, and human experts must subsequently interpret and package the results for the scientific community, the traditional metrics for assigning credit break down entirely.

So What Did AI Actually Discover?

OpenAI’s ambitious experiment was by no means an instance of 10,000 autonomous agents staring at a blank digital slate and inventing modern fluid dynamics from the ground up. The system stood firmly upon decades of accumulated human mathematics, recent breakthroughs achieved by human researchers in adjacent fields, and theoretical trajectories that were already showing immense promise within the academic community.

Furthermore, human supervisors continually intervened to direct computational power toward specific avenues, while OpenAI engineers deliberately curated and transferred useful intermediate findings between competing agent clusters.

Yet, acknowledging these human-guided parameters does not diminish the sheer significance of the milestone. What OpenAI successfully demonstrated is an entirely novel operational paradigm: thousands of computational agents can explore divergent research paths in parallel, rapidly prune failed hypotheses, synthesize winning concepts, and compress years of arduous manual calculation into a matter of days.

In the wake of the announcement, representatives from the Clay Mathematics Institute noted that the formidable Navier-Stokes problem "has apparently been settled," while simultaneously emphasizing that rigorous evaluation of the proof and the formal determination of credit will require extensive peer review over an extended period.

This development reframes the broader discussion surrounding artificial intelligence. The central question is no longer whether modern systems have achieved general artificial intelligence in a traditional sense. Instead, the experiment demonstrates that humanity can now entrust deeply complex research challenges to thousands of autonomous agents, allowing them to explore thousands of conceptual pathways simultaneously and complete tasks in days that traditionally consume years of human labor.

This represents a profound shift in scientific methodology. Yet, it simultaneously introduces a persistent philosophical and professional dilemma: when an artificial intelligence system builds seamlessly upon decades of prior human research and explores those conceptual frontiers at a massive, automated scale, what specific portion of the ultimate discovery can genuinely be attributed to the machine?

For many observers, the most compelling aspect of the breakthrough is not simply that a machine generated a potential solution to a famous mathematical puzzle within 88 hours. Rather, it is the sophisticated mechanics of how the goal was achieved.

OpenAI did not rely on a single, monolithic model instructed to ponder a problem more deeply. Instead, the organization engineered a vast digital ecosystem, dispatching thousands of agents down disparate intellectual roads, dynamically shifting computational resources toward promising pathways, facilitating cross-talk between isolated sub-groups, and maintaining the cycle until a viable proof materialized.

This operational structure bears less resemblance to a conventional software application or conversational chatbot and more to a high-speed, hyper-scaled corporate research laboratory operating entirely at machine velocity. Indeed, OpenAI has made no secret of its long-term ambition to construct a fully automated artificial intelligence researcher.

At the same time, the Navier-Stokes episode serves as a powerful reminder of the care required when wielding terms like independent discovery. The autonomous agents were invariably operating atop mountains of legacy mathematics, existing theoretical frameworks, human-managed computational budgets, and frontier concepts already taking shape in university offices.

Consequently, the true historical watershed may not be that artificial intelligence has suddenly learned how to autonomously unlock the secrets of the universe on its own. The genuine revolution is that the very nature of scientific research itself has become radically scalable.

If a swarm of 10,000 AI agents can successfully navigate the razor-edge of fluid dynamics today, the most pressing question facing the scientific community is what will happen when this exact computational methodology is unleashed simultaneously upon thousands of other stubbornly unsolved problems across mathematics, physics, biology, and engineering.

Share:

Reynand Wu writes for Tech Maze.

Leave a comment