6 min read · 1,552 words
This article was written with AI. It was drafted from the sources it cites and checked against the full text of those sources before publishing. How we make articles
Roughly 10,000 concurrent AI agents ran for about 88 hours and produced a 166-page mathematical writeup, yet OpenAI says it does not intend to claim the Clay Mathematics Institute’s $1 million Millennium Prize (VentureBeat; Smithsonian Magazine; Science News). Unclaimed prize money is the hook. Behind it sits an AI research credit crisis, in which the company announcing the breakthrough first said that, while unlikely, it could not rule out that de-identified data from an outside mathematician’s use of its own products helped improve its models (Smithsonian Magazine), and later said that, following an investigation, it had confirmed his Codex prompts could not have influenced the system (OpenAI).
A Prize of $1 Million and $6.5 Million in Retail-Rate Output Tokens
OpenAI states that an internal system “significantly more capable than GPT-6 Astra” produced a proof that 3-D Navier-Stokes dynamics can develop a finite-time singularity (OpenAI), formalized in the Lean proof language (OpenAI). Per the announcement, the result resolves the Clay problem’s disproof-oriented formulations, establishing the singularity variants rather than a global-existence proof (Unite.ai). In the race between frontier labs, that headline is worth more than the prize money attached to it.
What it cost is where the trouble starts. Cited inputs: about 130 billion output tokens, which VentureBeat prices at about $6.5 million in retail API charges, an output-only figure it stresses is not OpenAI’s actual infrastructure cost, beside one outside estimate circulating on X, which it calls speculative, of below $10 million to $30 million or $40 million (VentureBeat). Division recovers the rate behind it: $6.5 million ÷ about 130 billion tokens ≈ $50 per million output tokens, the published GPT-6 Astra output price the estimate used. Plug your own assumptions into what follows. Bill ≈ the $6.5 million floor × (1 + r × k), where r is the input-to-output token ratio and k is input pricing over output pricing. OpenAI says follow-up prompts drew on the agents’ own intermediate results (OpenAI), so r above 1 is plausible. Take the published input rate, a fifth of the output rate ($10 against $50 per million tokens) VentureBeat: at r = 1, the floor rises to $7.8 million; at r = 2, $9.1 million. Also unpriced: orchestration for roughly 10,000 concurrent agents across about 88 hours (VentureBeat), the Lean formalization pipeline, and every failed run preceding the success. OpenAI has not disclosed how many input tokens the system processed (VentureBeat), so the true retail-equivalent total is unknown. Unresolved, for the mathematicians involved, is whether their unpublished work flowed into the machine at all.
Mark Chen’s One Careful Sentence
Read the disclosure language like a contract lawyer. Chief Research Officer Mark Chen stated that “no people or AI systems searched through user data to solve this problem,” and expressed disappointment at “the allegations of the huge breach of user trust,” even as the company’s own post said it could not rule out a de-identified-data route (VentureBeat). No direct search, in other words. Then came the qualifier: while unlikely, the company “cannot rule out that de-identified data derived from their usage of our products helped improve our models” (VentureBeat).
NYU mathematician Tristan Buckmaster had been drafting related work inside OpenAI’s Codex and asked the company twice whether the model had been trained on his private sessions. His account: “I was told the model did not look up user data. I asked again, about training, and I did not get an answer” (Smithsonian Magazine). In a statement to the New York Times, reported by Smithsonian Magazine on September 10, 2026, OpenAI said it could say “categorically” that Buckmaster’s prompts could not have influenced the system “in any way, including training” (Smithsonian Magazine).
On this article’s reading, a statement that hardens between one telling and the next is not, by itself, reassurance.
166 Pages, Zero Independent Human Verification
Compounding the credit problem is a verification problem. Lean has checked the formalization, but human mathematicians have not fully verified the proof; as Johns Hopkins mathematical physicist Gregory Eyink put it, “I don’t think anyone has completely verified the proof yet, certainly not on the human side” (Science News). OpenAI researcher Sébastien Bubeck called it “the spectacular culmination” of a year of progress (Science News), an apt description of the excitement and of how much rests on trust.
Eyink also told Science News that the solution won’t have significant practical implications, and it is more about prestige (Science News). Science News adds that the result raises concerns about how credit for a discovery is doled out (Science News). Translation: at retail rates the output tokens alone come to about $6.5 million, which VentureBeat stresses is not OpenAI’s actual infrastructure cost (VentureBeat), for a proof mathematicians are still digesting (Science News).
OpenAI’s Strongest Defense, and Where It Breaks
Steelman the company before ruling on it, because its best defense is stronger than the denial sequence made it look.
Pillar one: if training on de-identified usage data is ordinary practice for frontier labs, then the qualifier VentureBeat reported, that OpenAI “cannot rule out” de-identified data helping, describes that practice rather than confessing anything (VentureBeat). Pillar two: the hardened denial followed an audit, by OpenAI’s own account; its post now says that, following an investigation, it confirmed Buckmaster’s Codex prompts could not have influenced the system in any way, including through training (OpenAI). Pillar three: OpenAI denies that its people or agents directly accessed the mathematicians’ private work to produce the proof (VentureBeat).
On this article’s reading, each of the three has a weak point. Ubiquity is the exposure, not the answer: if every vendor claims the same right to absorb de-identified work, every research contract governing that traffic is mispriced by the same amount, an indictment of the norm rather than a defense of it. An internal investigation is not an independent audit, and a lab that asks anyone to accept “categorically” about its training corpus is asking for trust no outside party has checked. Absence-of-mechanism dodges the live question of whether Buckmaster’s private sessions overlapped the singularity result in substance, not merely in process; the statements quoted here describe the absence of a channel, a much weaker claim than an audit of that overlap. Even at zero technical influence, the sequence survives: Buckmaster says he asked twice and got no answer about training at the time; OpenAI’s categorical statement came later, to the New York Times (Smithsonian Magazine).
The Credit Question Every Research Organization Now Faces
Scale down from Millennium Prizes and the stakes land on every enterprise and university lab. If a vendor’s disclosure concedes it cannot rule out that de-identified usage improves its models (VentureBeat), then every unpublished draft, proprietary dataset, and half-formed proof pasted into a frontier chat is a possible uncompensated contribution to a commercial engine that competes with its author. Not theft, which requires proof. Unpriced credit, which requires only ambiguity.
Any institution that funds research should now treat data-use terms as intellectual-property terms, because the vendor’s hedged language is the disclosure of record until contracts say otherwise. Next to the liability a single uncovered research pipeline represents, a $1 million prize OpenAI says it will not claim is cheap.
Prediction, with a timeline: within twelve months, one of three things happens. Independent mathematicians, who Science News reports are still digesting the 166 pages and have not yet verified the proof (Science News), finish checking it, and verification becomes the story. Or a major lab publishes the first training-data provenance attestation, converting “cannot rule out” into an auditable claim, and wins the compliance market that follows. Or the first AI research credit dispute reaches lawyers while the wording is still soft, and every frontier-model contract gets rewritten in a quarter.
Two Ledgers, Same Opacity
Smithsonian Magazine puts the open question plainly: could OpenAI’s models have used data from two mathematicians working on a similar breakthrough (Smithsonian Magazine)? When even the retail-equivalent cost of the run cannot be pinned down, “cannot rule out” is not a data-governance answer either. A quick check before pasting unpublished work into a frontier chat: search your vendor’s data-use terms for “de-identified” and read what that clause licenses, and date-stamp and privately archive any result before a model touches it.
On this article’s reading, a verdict on the proof will arrive eventually. Until the credit question is settled, anyone pasting unpublished work into a chat window should read their vendor’s data-use terms first.
References
- On the Navier-Stokes Millennium Prize Problem, OpenAI — official announcement of the internal system’s proof and Lean formalization.
- AI may have solved a longstanding math problem, Smithsonian Magazine — Buckmaster’s statement and OpenAI’s statement to the New York Times that his Codex prompts could not have influenced the system.
- OpenAI solves math problem but can’t rule out benefitting from private Codex data, VentureBeat — Mark Chen quotes, the “cannot rule out” disclosure, and the 130 billion output tokens priced at about $6.5 million in retail API charges.
- AI tackles the Navier-Stokes puzzle, Science News, verification status (Eyink: no one has completely verified the proof) and Bubeck on the result’s significance.
- OpenAI says internal system resolved Navier-Stokes, Unite.ai, OpenAI’s account of the result and of the concurrent Buckmaster–Alpöge work.
