OpenAI math breakthroughs now come with a second question attached: whose unpublished work helped produce them? On Tuesday 8 September 2026, OpenAI said a swarm of AI agents running on an unreleased internal model had resolved the Navier–Stokes existence and smoothness problem, one of the Clay Mathematics Institute’s seven Millennium Prize Problems. Within two days, two mathematicians who had used OpenAI’s own products on their unpublished research were asking the company for something it has not offered: proof that their work did not help.

The first is Tristan Buckmaster, a mathematics professor at New York University. He spent most of a year on related fluid equations with Levent Alpöge, an Anthropic researcher collaborating in a personal capacity, and says the pair put “all our drafts for the whole of this project” into OpenAI’s Codex. The second, reported by The Verge on 10 September, is Andreas Thom, a group theorist at TU Dresden. He says OpenAI never properly answered whether his ChatGPT conversations became training data for an earlier OpenAI math result on non-sofic groups.

OpenAI’s position is precise, and the precision is the story. The company says no researcher or agent saw the pair’s work before it was published and that “no specific user data was accessed”. It also says it “cannot rule out that de-identified data derived from their usage of our products helped improve our models.” Mathematicians want that gap in the OpenAI math story closed with evidence rather than assurances.

Below we set out the OpenAI math timeline, the compute arithmetic, both accounts of the phone calls, what OpenAI’s data controls actually cover and what real proof would look like, then what the dispute means for any business that puts unpublished work into AI tools. For background, see our coverage of OpenAI’s GPT-6 Astra launch and our GPT-5.6 Sol guide, the model Buckmaster singles out among those he used inside Codex.

What Mathematicians Want the OpenAI Math Team to Prove

openai math mathematicians want proof didnt use their work b open jewellery box with violet sides

Nobody in this dispute is asking OpenAI to retract its Navier–Stokes proof, or to hand back a prize the company says it will not claim. The request is narrower and harder to meet. Mathematicians want the OpenAI math team to show, with records rather than statements, that nonpublic research typed into ChatGPT and Codex did not shape the systems that then reached neighbouring results first.

Thom framed the burden in his Mathstodon posts. Researchers, he said, are not equipped to reverse-engineer a training pipeline, and “only OpenAI has the relevant data for that.” If the company denies using such material, The Verge reports, he believes the responsibility is on OpenAI to prove it by disclosing the datasets involved and clarifying the settings and terms that govern how it uses data.

The two questions Thom asked about an earlier OpenAI math result

Shortly after OpenAI announced ten results in mathematics and theoretical computer science on 1 August, Thom emailed OpenAI researchers Mark Sellke and Sébastien Bubeck. He wrote that he and a Dresden colleague had been discussing the expander matching problem, and extensions of his joint work with Gábor Kun, “actively over the last months with ChatGPT.” He asked whether that “was part of the training data or accessible to the reasoning process.”

Sellke’s complete reply, as Thom quotes it, was: “Regarding your conversations with ChatGPT: that did not happen.”

Why Thom now calls that reply dishonest

Thom’s objection is that he asked two separate things and received one categorical answer. That answer, he wrote on 9 September, “now looks as though it addressed only direct access.” He added: “No such qualification, explanation, or evidence was given. I take this as dishonesty to say the least.” Replying to another user in the same thread, he said he had opted out on 29 June.

In later posts quoted by The Verge, Thom called Sellke’s answer “at minimum, unjustifiably broad and materially misleading; looking back it was plainly dishonest.” He said he began reflecting on that exchange after Buckmaster publicly questioned whether the OpenAI math effort had benefited from his use of Codex.

Buckmaster’s unanswered OpenAI math question

Buckmaster’s version of the request sits in the four-page statement he posted on 7 September. On a call with OpenAI the day before, he asked whether the model “had been trained on, or had access to, our sessions in Codex.” His account of the reply: “I was told the model did not look up user data. I asked again, about training, and I did not get an answer.”

He is explicit about the limits of his claim. “I do not know whether our data was used. I am not accusing anyone of anything,” he writes. What he wants is the history told accurately: if an OpenAI model closed the gap to Navier–Stokes, “it should be said loudly, by them, with the history intact.”

The OpenAI math questions and the answers so far

Set side by side, the questions put to OpenAI and the answers on the record do not yet meet.

QuestionAsked byAnswer on the recordStill open
Could the system read our Codex sessions while it worked?Buckmaster, 6 September callTold the model did not look up user data; OpenAI says no researcher or agent saw the workIndependent confirmation from logs
Was the model trained on those sessions?Buckmaster, 6 September callNo answer on the call; OpenAI later said it cannot rule out de-identified data helpingWhether the data was eligible for training, and when
Did our ChatGPT conversations enter training data?Thom, August email“That did not happen,” from Mark SellkeThom says the reply only covered access
Will OpenAI disclose the datasets, settings and terms involved?Thom, 9 September postsNo specific response reported by 10 SeptemberAll of it
Can AI companies demonstrate that chat logs will not improve their models?Brendan Hassett, Brown University, to The VergeGeneral data-control policiesA way to demonstrate it

The OpenAI Math Timeline, Day by Day

openai math mathematicians want proof didnt use their work c toy steam locomotive

Both sides of the OpenAI math dispute lean on the order of events. Buckmaster’s case rests on it: his results came first, word of them reached OpenAI, and only then did OpenAI’s agents take the same route. OpenAI’s own account confirms much of that sequence and disputes what it implies.

Before the OpenAI math run began

The earliest relevant dates are months old. Thom says he opted out on 29 June, and OpenAI published its ten results on 1 August. Buckmaster and Alpöge obtained blowup results with smooth forcing for the Boussinesq and Euler equations on 15 August, and verified the Euler proof in Lean on 22 August. OpenAI says it began training a new internal model, one that “exhibited unprecedented performance in our benchmarks, including mathematics,” on 28 August.

The week the OpenAI math agents launched

OpenAI says that on Tuesday 1 September it heard rumours that two Millennium Prize Problems had been resolved, and launched agents on all of the open ones. On Thursday 3 September, with a rumour circulating that Anthropic had solved a major problem, Buckmaster emailed a prominent OpenAI mathematician to set out the facts. The reply said details “would be useful to avoid competing here” and offered compute.

OpenAI asked to meet on Friday 4 September and again at 12:45 on Sunday 6 September. That afternoon Buckmaster spoke twice with the OpenAI mathematician he had emailed and with Bubeck. By OpenAI’s account, its agents had already reached a resolution on Saturday 5 September, about 88 hours after launch, and Lean verification finished on 6 September.

The OpenAI math announcement and its fallout

Buckmaster posted three results and his statement just before midnight New York time on Monday 7 September, according to Quanta Magazine. OpenAI published its solution, a paper and a GitHub repository of Lean certificates on 8 September, followed by posts on X from the company, from Bubeck, from chief executive Sam Altman and from chief research officer Mark Chen. Thom’s posts followed on 9 September, and The Verge published its report on 10 September.

Date (2026)What happenedSource
29 JuneAndreas Thom says he opted outThom on Mathstodon
1 AugustOpenAI announces ten results, including a construction of non-sofic groupsOpenAI
15 AugustBuckmaster and Alpöge obtain blowup with smooth forcing for Boussinesq and EulerBuckmaster statement
22 AugustTheir Euler proof is verified in LeanBuckmaster statement
28 AugustOpenAI starts training a new internal modelOpenAI
1 SeptemberOpenAI hears rumours of solved Millennium Prize Problems and launches its agentsOpenAI
3 SeptemberBuckmaster emails a prominent OpenAI mathematician; the reply offers computeBuckmaster statement
5 SeptemberOpenAI’s agents reach a Navier–Stokes resolution, about 88 hours after launchOpenAI
6 SeptemberLean verification finishes; Buckmaster has two calls with OpenAIOpenAI, Buckmaster statement
7 SeptemberBuckmaster posts three results and his statementBuckmaster, Quanta Magazine
8 SeptemberOpenAI publishes its solution, Lean certificates and a statement on user dataOpenAI, GitHub
9 SeptemberThom asks whether his ChatGPT use fed the non-sofic group resultThom on Mathstodon
10 SeptemberThe Verge reports that mathematicians want proofThe Verge

OpenAI’s own figures show how compressed its side of the OpenAI math timeline was.

Hours in the OpenAI math run, as the company describes it
Euler search, nearly 100 agents about 50 hours
Navier–Stokes search, on the order of 10,000 agents about 88 hours
Lean formalisation with GPT-6 Astra 17 hours

The six days that matter to Buckmaster

The interval Buckmaster keeps returning to runs from 22 August, when his Euler proof passed Lean, to 28 August, when OpenAI’s new model started training. On 8 September he quoted OpenAI’s line about that training run on Mastodon and wrote: “Note that they are openly admitting they used training data from a period after we found our result. Is it ethical to use customer’s data to try to scoop their customer?”

That is an inference, not a disclosure. A training start date says when a run began, not what data it contained, and OpenAI says training on the model is ongoing. The dates show that Buckmaster’s OpenAI math question is reasonable. They do not answer it.

The OpenAI Math Effort in Numbers

openai math mathematicians want proof didnt use their work d violet pyramid of stacked spheres

For the OpenAI math run, the company disclosed more operational detail than AI companies usually do. Its post describes groups of coordinating agents with access to “a cached version of the internet” and the ability to run code, each group prompted with a different variant of the problem. Codex was used to consolidate the most useful insights from each group, and the agents were switched to a further trained version of the internal model partway through.

What the OpenAI math post disclosed

Every figure below comes from the 8 September OpenAI math post except the last row, which is our own multiplication.

MeasureEuler (unforced)Navier–StokesAll problems
AgentsNearly 100On the order of 10,000 concurrentNot stated
Search timeAbout 50 hoursAbout 88 hoursNot stated
Lean formalisationNot stated17 more hours, via GPT-6 AstraNot stated
Messages between agentsNot stated2.7 million4.9 million
Output tokensNot statedAbout 130 billionAbout 300 billion
Agent-hours (our arithmetic)About 5,000 (100 × 50)Up to about 880,000 (10,000 × 88)Not stated

Group sizes varied and “on the order of 10,000” is not an exact count, so the Navier–Stokes figure is a rough ceiling rather than a measurement. Even so, the ratio is striking. On those numbers the Navier–Stokes search used up to about 176 times the agent-hours of the Euler search that preceded it.

Navier–Stokes took most of the run’s messages but less than half of its output tokens.

Share of the OpenAI math run spent on Navier–Stokes
Agent messages, 2.7 million of 4.9 million 55%
Output tokens, 130 billion of 300 billion 43%

What the OpenAI math run might have cost

No official cost has been published for the OpenAI math run. Mark Chen said at a press briefing that the effort cost “in the millions of dollars,” WIRED reported. Fortune reported that OpenAI told reporters it used at least 1,000 times the compute of earlier math work that cost about $2,000, which implies about $2 million. TechCrunch valued the 300 billion output tokens at $22.5 million at current Astra rates.

OpenAI’s API pricing page lists GPT-6 Astra output at $50 per million tokens in the standard tier for short context and $75 for long context. Batch and Flex halve the short-context price to $25, and Fast mode doubles it to $100. Multiplying 300,000 million tokens by each rate gives the range below.

What 300 billion output tokens would cost at GPT-6 Astra list prices
Fortune’s briefing estimate, 1,000 × $2,000 about $2m
Batch or Flex, short context, $25 per million $7.5m
Standard, short context, $50 per million $15m
Standard, long context, $75 per million $22.5m
Fast mode, short context, $100 per million $30m

Treat those figures as a scale marker, not an invoice. The run used an internal model rather than Astra, the arithmetic ignores input tokens, and a list price is not what OpenAI pays to run its own hardware. The point Brown University’s Javier Gómez-Serrano made to MIT Technology Review survives any choice of rate: “very few mathematicians will have resources of that scale.”

What a swarm changes

András Juhász of the University of Oxford put the asymmetry bluntly to The Verge: “Suddenly, 10,000 mathematicians jump on your problem.” OpenAI researcher Noam Brown argued on X that such costs fall fast. He noted that o3 cost about $500,000 to score 87.5% on ARC-AGI 1, while Astra scores higher for about $20. “I believe that a year from now everyone will have an AI at their fingertips capable of solving problems of this caliber,” he wrote.

What the OpenAI Math Statement Denied and What It Would Not Rule Out

openai math mathematicians want proof didnt use their work e kitchen blender with solid jar

The OpenAI math statement packs three claims into four sentences, and they are not the same kind of claim. Two are denials of direct access. One is an admission that an indirect path cannot be excluded. Much of the anger in mathematical circles comes from treating the first two as though they answered the third.

Access and training are different things

Mark Chen drew the line himself on X. “Did any human or agent look at user data as part of the Navier Stokes effort? No,” he wrote. “Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company.” Bubeck told the press briefing, according to Fortune, that “we did not use their prompts or proofs to prompt our models or direct our agents.”

Those are statements about what people and agents did during the OpenAI math effort. They say nothing about what a model had already absorbed before the effort began, which is the question Buckmaster says went unanswered.

Each OpenAI math statement and what it leaves open

StatementSourceWhat it coversWhat it leaves open
“We (the researchers and the agents) did not see any of their work through any means until they released it publicly”OpenAI post and X, 8 SeptemberViewing by staff or agents during the effortAnything absorbed in training beforehand
“No specific user data was accessed in order to solve this problem”OpenAI post and XLookups of identifiable user recordsDe-identified data
“While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models”OpenAI post and XAdmits a possible indirect pathHow likely, and whether the accounts were eligible
“Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes.”Mark Chen on XConfirms the general practiceWhich data, and from which dates
“We did not use their prompts or proofs to prompt our models or direct our agents”Bubeck at the press briefing, via FortunePrompting and directionTraining
“The chances are zero if they have opted out (likely)”OpenAI employee posting as roon on XProbability, conditional on an opt-outWhether either account had opted out

The boundary OpenAI says it will not cross

The roon post, which The Decoder cites among OpenAI staff pushing back, called it “exceptionally unlikely that anything they ever did made it into any part of training.” It then explained why OpenAI will not simply check. Doing so, it said, would be “a terrible precedent to break the … PII-scrubbing boundary to go and round it down to 0, and we won’t do it.”

That sentence explains why the OpenAI math dispute is so hard to settle. The protections that stop staff looking up a named user’s data are the same protections that stop them proving a named user’s data was absent. On our reading, OpenAI is choosing privacy over proof, and the mathematicians are asking whether that trade should be OpenAI’s alone to make.

How the Anthropic researcher read it

Alpöge’s own reaction to the OpenAI math carve-out was notably mild. “i mean props to them for straight coming clean,” he wrote on X, before adding that “so far the proof looks more along the lines of another euler blowup proof we had.” Coming from the co-author most directly affected, that is a reminder that the dispute is about evidence and conduct, not about whether the OpenAI mathematics holds up.

Why De-identified Data Is the Crux of the OpenAI Math Dispute

openai math mathematicians want proof didnt use their work f mortar and pestle

De-identification is a privacy technique. Often automated with natural language processing, it is designed to strip names, email addresses and other personal details out of data before the data is used. Thom’s objection, quoted by The Verge, is that this solves the wrong problem for research: “De-identification may remove a name; it does not remove the intellectual content of a mathematical idea.”

OpenAI’s own description of its process draws the same line without intending to. Its help centre says the company takes “steps to reduce the amount of personal information in our training datasets.” A half-finished proof strategy is not personal information. Removing the author’s name leaves the strategy intact, which is precisely what a mathematician worried about being scooped would fear from the OpenAI math pipeline.

What OpenAI’s data controls actually say

OpenAI publishes different defaults for individual and business products. The table summarises its help-centre article on how data is used to improve model performance, as it read on 10 September 2026.

Product or featureUsed for training?ControlCatch
ChatGPT, individual accountsMay be used unless you opt outData Controls, or “do not train on my content” in the privacy portalApplies to new conversations
Codex tasks, individual accountsMay be used unless you opt outThe same opt-out as ChatGPTFull-environment training has a separate Codex setting
Rating a responseThe entire conversation may be used, even after an opt-outDo not rate sensitive chatsEasy to trigger by habit
Temporary ChatNot used for trainingChoose Temporary ChatNo history or memory
ChatGPT Business, Enterprise and the APINot used for training by defaultOpt-in only, for example Playground feedbackCheck workspace settings and contracts

Why opting out is not proof either

Opting out is a forward-looking switch. OpenAI’s wording is that once you opt out, “new conversations will not be used to train our models.” Earlier conversations are not mentioned. Thom says he opted out on 29 June, which leaves any earlier discussion with ChatGPT outside that protection, and nobody outside OpenAI can check what happened to it.

For Buckmaster the facts are thinner. His statement says he pays for his group’s tools “out of my own research funds, including footing a large bill to OpenAI,” but it does not say which plan he used or whether training was switched off. The Decoder notes that whether the pair had opted out is not publicly known. That one setting matters, because training is off by default on OpenAI’s business products.

Codex has its own switch

Codex adds a wrinkle many users miss. OpenAI’s help centre says Codex “has separate controls for allowing training on full environments,” and that adjusting settings in ChatGPT or the privacy portal “will not affect these full-environment Codex settings.” Buckmaster describes Codex as the place “into which we had been putting all our drafts for the whole of this project.” For anyone using Codex on unpublished work, that is the setting to check first. Our comparison of AI coding assistants covers how Codex fits alongside its rivals.

Two Accounts of the OpenAI Math Phone Calls

Much of what happened on the OpenAI math calls of 6 September is known only from the two sides’ own descriptions. Buckmaster’s statement, Bubeck’s post on X and Altman’s post on X agree on the broad outline. They disagree sharply on the details that matter most.

PointBuckmaster and AlpögeOpenAI (Bubeck and Altman)
How OpenAI heardAlpöge had received tips that information about their progress had been passed to OpenAIRumours on the internet that Anthropic’s models had solved a Millennium Problem
Human inputTold of “very little human input”, then learned an entire team had worked on itPost describes agent groups given problem variants, with insights consolidated using Codex
Alpöge’s authorshipBubeck “twice asserted” he wanted Alpöge removed from authorshipBubeck never asked to remove him from his own work; the remark concerned a rewrite of OpenAI’s proof
The career remark“Why would you ruin your career?” then “If you don’t want me to be nice, then I don’t have to be nice.”Bubeck calls it an extremely poor choice of words, apologises, and says he retracted it on the spot
CoordinationAlpöge declined a one-on-one call and said conversations should be with BuckmasterAltman says Alpöge “was not willing to talk or coordinate with us”
The prizeDeclined proposals under which OpenAI would say they deserved the Clay PrizeOpenAI says it will not claim the prize

What both sides of the OpenAI math dispute agree on

Several points are not in dispute. OpenAI went after Navier–Stokes because of rumours about Anthropic; Altman wrote that “it is true that we tried this because there were rumors on the internet last week that Anthropic’s models had solved a millennium problem.” Buckmaster and Alpöge had blowup for forced Euler first, and OpenAI’s post says it recognises “the priority of their work on forced Euler.” OpenAI also offered them visibility into its prompts and, later, the proof.

Authorship and the Anthropic question

The sharpest disagreement in the OpenAI math calls is over Alpöge. Buckmaster writes that Bubeck “twice asserted that he wanted Levent removed from authorship.” Bubeck says he “never ever asked for Levent to be removed from authorship of his own work.” He says his remark that things “would be simpler if Levent was not an Anthropic employee” concerned a proposed rewrite of OpenAI’s own proof, which he felt an Anthropic employee should not author.

Alpöge, who was not on the calls, described it differently on X, referring to “the part where a millennium prize was offered if i’d just be removed from the paper.” He also wrote that he would have been “pumped to collaborate” and did not care about authorship on that step.

The career remark

Both accounts include a line about a career. Buckmaster quotes Bubeck asking “Why would you ruin your career?” and then saying “If you don’t want me to be nice, then I don’t have to be nice.” In Bubeck’s version he said he did not understand why one “would risk their career” over what he calls unfounded accusations. He apologises for “this extremely poor choice of words” and says he retracted it on the spot. His post does not address the second line.

Andreas Thom and the Earlier OpenAI Math Result

Thom’s complaint predates the Navier–Stokes dispute by more than a month and concerns a different result. On 1 August OpenAI published ten results “achieved by an internal version of Astra,” at a token cost it put at roughly $2,000 at Sol API rates. One was “a construction establishing the existence of non-sofic groups, addressing a central open question in group theory.”

Why this OpenAI math result touched Thom’s work

Non-sofic groups are, in The Verge’s summary, infinite mathematical structures that cannot be approximated by finite ones, and whether any exist was a long-standing open question. The Verge reports that OpenAI acknowledged its construction built heavily on earlier work by Thom and Gábor Kun, was criticised for failing to credit their recent contributions, and quietly amended its write-up.

Thom says he was struck by “OpenAI’s detailed command of our techniques,” which he describes as neither the most obvious nor the most promising route at the time. That is the same OpenAI math pattern Buckmaster describes: a technique few people were pursuing, adopted quickly by a company whose products had hosted private conversations about it.

What Thom says would be indefensible

Thom’s strongest line is conditional. He said it “would be ethically indefensible” if nonpublic research supplied by users helped improve models that the company then used to race those same users to publication, without consent, proper disclosure or credit. He has not said that happened. He has said OpenAI’s answers so far do not let him rule it out.

OpenAI’s own standard on attribution

OpenAI’s August post contains a principle mathematicians now want applied to data. “We believe attribution should honestly reflect how a result was produced,” it says, warning that claiming human authorship for an AI-generated proof would misrepresent the system’s contribution. The logic runs both ways. If human conversations helped an OpenAI math result, honest attribution would say so, and so far OpenAI’s answers stop short of that.

Could the OpenAI Math Model Have Found the Route on Its Own?

This is the scientific heart of the OpenAI math dispute, and experts disagree in good faith. Both teams built on a programme developed by Diego Córdoba of the Institute for Mathematical Sciences in Madrid and Luis Martínez-Zoroa of CUNEF University, which constructs a blowup through what Martínez-Zoroa calls an “infinite cascade” of layers. Charles Fefferman, who wrote the Clay Institute’s official problem description, told Quanta the heroes of the story are Córdoba and Martínez-Zoroa.

The case that the route was a red flag

Buckmaster argues that the route through a smooth force, options C and D in Fefferman’s statement of the problem, was the one he and Alpöge “had quietly chosen to attack.” He wrote: “Almost nobody else I know of was working on it. It is not the direction one arrives at in a few days by giving a model the problem statement. When I heard ‘forced,’ it was a bright red flag.”

He also says he was first told the model had simply been given the problem statement. Over the call, he writes, it emerged that an entire team had worked on the problem, that the model had first been set on easier problems including Euler, and that even the prompt shown to him “had been written by prompting Codex.”

The case that it was findable

Others see a well-lit path. Gómez-Serrano told MIT Technology Review that the Córdoba–Martínez-Zoroa approach was one of several thought to hold promise. The OpenAI math post says its agents first resolved the unforced Euler problem and then carried that result into Navier–Stokes, and it insists the two teams’ proofs differ significantly.

OpenAI’s Boaz Barak wrote on X that the model “actually started off by proving a stronger claim than they did,” and OpenAI mathematician Ven Chandrasekaran stressed, according to WIRED, that the pair’s solution was significantly different in nature from the one OpenAI’s system produced.

Why a public programme does not settle it

The Córdoba–Martínez-Zoroa papers are public, and OpenAI’s agents could read a cached copy of the internet, so a system could plausibly have found the route unaided. MIT Technology Review points to a subtler issue. Choosing which promising direction to pursue, often called research taste, has long been one of the hardest things for AI. If the agents followed the route because humans had already committed to it, human judgement was essential. Without the OpenAI math agents’ logs, outsiders cannot tell which story is true.

What Real Proof of OpenAI Math Provenance Would Look Like

Proving a negative about a training run is hard but not impossible, and much of the evidence already exists inside OpenAI. The options below differ in what they would show and in how far each would require the company to open the privacy boundary it has promised to keep.

EvidenceWhat it would showWho holds itLimits
Account training settings and their change dates, including Codex full-environment settingsWhether the sessions were eligible for training at allOpenAI, released with the users’ consentSays nothing about data from before an opt-out
Data cut-off dates for each model snapshot the agents usedWhether August sessions could be in the run that began on 28 AugustOpenAITraining is ongoing, so several snapshots may matter
A search of the de-identified corpus for distinctive strings from the draftsWhether identifiable mathematical content was presentOpenAI, or an auditor under confidentialityParaphrased or derived content may not match
Agent transcripts, tool calls and cached pages read during the runWhat the agents actually consulted and triedOpenAIMillions of messages; only prompts were offered to Buckmaster
An independent audit with an agreed scopeA credible answer without publishing anyone’s conversationsOpenAI and an auditorRequires OpenAI to open the boundary in a controlled way
Outside tests for memorised textWhether a model reproduces unpublished passagesAnyone with model accessCan suggest presence; rarely proves absence

“Unlikely” is a probability, not an audit

OpenAI’s word for the chance that the pair’s data helped is “unlikely.” That may well be accurate. It is also impossible to test from outside, because only OpenAI holds the account settings, the data snapshots and the agent logs. MIT Technology Review adds that incidents such as the Hugging Face hack show OpenAI is not always aware of what its agents are doing, a pattern we covered in OpenAI’s German wiki incident.

A realistic minimum for OpenAI math disclosure

Some steps would cost OpenAI little. With the users’ consent, it could confirm whether their accounts and Codex full-environment settings were eligible for training during the relevant months. It could publish data cut-off dates for the snapshots its agents used, and release the OpenAI math agent transcripts for Navier–Stokes, dead ends included. Terence Tao has criticised “the refusal of AI companies to disclose their negative results, or reveal the process towards obtaining their solutions.”

Private shorthand makes a natural test

Research drafts often contain shorthand that appears nowhere else. Alpöge mentioned on X that he and Buckmaster gave their blowup ansätze pun names such as “smooth criminale.” A string like that, searched for in training snapshots taken before the pair went public, would give a far more concrete answer than “unlikely”, without anyone’s conversations being published. Brendan Hassett of Brown University told The Verge that companies “should be held accountable to deliver” assurances about chat logs.

How OpenAI Math Races Change the Norms of Mathematics

The data question is only half of what alarms mathematicians. The other half is the OpenAI math race itself. “Mathematics depends heavily on an informal norm of trust,” Matthew Ballard of the University of South Carolina told The Verge. “Researchers routinely share incomplete ideas and ongoing work with colleagues to sharpen their thoughts. It is done with the expectation that it will not turn into a competition.”

A rumour is now enough to start an OpenAI math race

Terence Tao of UCLA described the new incentive on 8 September. “We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential,” he wrote. He warned that incentives may now favour “no longer sharing any promising research directions with the broader community, which would reverse centuries of traditions of open science.”

Abhishek Saha of Queen Mary University of London said OpenAI had engaged in the “kind of things that mathematicians will generally not do.” Scooping happens between humans too, he told The Verge, but it is not easy, because few people have the specialised expertise to swoop in on someone else’s problem. OpenAI math swarms remove that barrier.

Caution is the rational response

Jeremy Avigad of Carnegie Mellon University, director of the Institute for Computer-Aided Reasoning in Mathematics, told The Verge that “the thought that AI systems might steal ideas from our queries is chilling.” With a swarm a hint away, he expects that “people are likely to be more cautious.” Yang-Hui He of the London Institute for Mathematical Sciences worries that “maths under the big companies is much too secretive,” recalling an era when mathematics depended on patrons such as the Medici family.

What the solution was supposed to be for

Tao argued before the announcement that the Navier–Stokes problem matters less for its answer than for what attempts to solve it teach. Prematurely solving such problems “by purely AI-powered methods — particularly without full transparency into the solution process — can contaminate this process to the point where it actually becomes a net negative for the progress of mathematics as a whole,” he wrote on 3 September.

Martínez-Zoroa, whose ideas both teams used, was generous when Quanta asked. “It would have been nice to do this ourselves, but I’m very happy for him,” he said of Buckmaster. Juhász, for his part, called the episode “clearly a PR victory for OpenAI” while questioning how sustainable that approach could be.

What the OpenAI Math Row Means for Businesses

Swap “mathematician” for “product team” and the dispute becomes familiar. Companies put unreleased designs, source code, pricing models and strategy drafts into AI assistants every day. The OpenAI math row shows that a vendor can truthfully say nobody looked at your data, while being unable to say whether that data, stripped of your name, helped improve a model that anyone can use.

Defaults matter more than promises

The most useful fact in OpenAI’s documentation is a default. Individual ChatGPT and Codex accounts may be used for training unless you opt out, while ChatGPT Business, ChatGPT Enterprise and the API are not used for training by default. OpenAI’s enterprise privacy page says API Platform data after 1 March 2023 “isn’t used for training our models, unless you have explicitly opted in.” For sensitive work, the plan you choose does more than any setting you remember to change.

A practical checklist after the OpenAI math row

StepWhy it mattersWho it applies to
Use business plans for unpublished workBusiness, Enterprise and API data is not used for training by defaultTeams handling code, research or other IP
Opt out on individual accounts and record the dateAn opt-out covers new conversations onlyPersonal ChatGPT and Codex users
Check the Codex full-environment training setting separatelyChatGPT settings and the privacy portal do not change itDevelopers using Codex
Do not rate responses in sensitive chatsFeedback can bring the whole conversation into trainingEveryone
Use Temporary Chat for one-off sensitive questionsTemporary chats are not used for trainingQuick lookups
Keep dated records of unpublished workPriority disputes turn on dates, as 15 and 22 August did hereResearch, engineering and legal teams
Ask vendors about de-identified and derived dataThat carve-out is exactly where OpenAI’s denial stoppedProcurement and contract owners

Put it in the contract

For organisations that buy AI services at scale, the OpenAI math lesson belongs in procurement. Ask vendors in writing whether “de-identified data derived from usage” can be used to improve models, how an opt-out is recorded, and what evidence they would provide in a dispute. Our vendor management and data protection services can build those questions into supplier reviews.

Provenance runs in both directions

Questions about whose work trained whose model are not unique to OpenAI or to mathematics. Our coverage of the AI distillation advisory looks at the mirror image of this dispute: companies training on other labs’ model outputs. Our AI Models, Tools and Releases hub tracks how vendor policies on training data change.

What Happens Next for OpenAI Math Claims

Three separate OpenAI math questions now run on different clocks: whether the proof holds, whether any user data mattered, and who deserves credit. Each will be settled, if at all, by different people.

Verifying the OpenAI math proof

The proof is formalised in Lean, and OpenAI’s certificates are public on GitHub under an Apache 2.0 licence. That gives mathematicians unusual confidence, but Quanta notes one step still needs humans: checking that the statement proved in Lean is logically equivalent to the problem people set out to solve. If the result holds, Quanta called it, by a significant margin, the most important proof yet produced by an AI model. The Clay Institute has not awarded the prize, and OpenAI says it does not intend to claim it.

The OpenAI math data question

Thom has asked OpenAI to disclose datasets, settings and terms, and OpenAI had not responded to The Verge by publication. Buckmaster says he and Alpöge also believe they have blowup for a hypo-dissipative version of Navier–Stokes, a paper held back because its Lean verification had not finished. Watch for any precise definition of “de-identified data derived from usage”, because that phrase now carries the weight of OpenAI’s defence.

Credit and cooperation

Both sides say they wanted cooperation. Bubeck cited Sholto Douglas’s remark that Anthropic and OpenAI will need to coordinate in future, and Alpöge wrote that he likes “the idea of the labs cooperating, and even better on scientific progress. It’s a shame!” Buckmaster’s hope is for the “serious and unhurried discussion about where to go from here” that his statement says the community needs.

OpenAI Math FAQs

Did OpenAI really solve the Navier–Stokes Millennium Prize Problem?

OpenAI says an internal system produced a proof, formalised in Lean, that an initially smooth fluid with a smooth applied force can develop a singularity in finite time, establishing statements C and D of the official problem. The Lean certificates are public. Independent scrutiny is continuing, and the Clay Mathematics Institute has not awarded the prize.

Did the OpenAI math effort use Tristan Buckmaster’s Codex data?

Nobody outside OpenAI knows. OpenAI says no person or agent looked at his or Alpöge’s work before publication and that no specific user data was accessed, but it cannot rule out that de-identified data from their product use helped improve its models. Buckmaster says he does not know whether his data was used and is not accusing anyone.

What does “de-identified data derived from their usage” mean?

OpenAI has not defined the phrase for this case. Its help centre says it takes steps to reduce personal information in training datasets. Critics such as Andreas Thom argue that removing a name does not remove the mathematical ideas in a conversation.

Who is Andreas Thom, and what is he asking OpenAI?

Thom is a group theorist at TU Dresden whose joint work with Gábor Kun underpins OpenAI’s August result on non-sofic groups. He asked whether his ChatGPT conversations entered training data or were accessible to the solving process, and he now wants OpenAI to disclose the datasets, settings and terms needed to show they were not used.

Will OpenAI receive the $1 million Millennium Prize?

OpenAI says it does not intend to claim it. The Clay Mathematics Institute administers the Millennium Prize Problems and has not awarded a prize for this result.

How do I stop ChatGPT or Codex training on my work?

On individual accounts, opt out in Data Controls or OpenAI’s privacy portal, check Codex’s separate full-environment training setting, use Temporary Chat for sensitive one-off questions and avoid rating responses in sensitive chats. On ChatGPT Business, ChatGPT Enterprise and the API, training is off by default.

References and Further Reading

The Verge: Mathematicians want proof OpenAI didn’t use their work

The Verge: OpenAI’s sly mathematical breakthrough sends a chill through academia

The Verge: Drama swirls around OpenAI’s legendary mathematical milestone

OpenAI: On the Navier–Stokes Millennium Prize Problem

OpenAI: Ten advances in mathematics and theoretical computer science

OpenAI Help Center: How your data is used to improve model performance

OpenAI: Enterprise privacy at OpenAI

OpenAI API: Pricing

GitHub: openai/NavierStokesAndEuler Lean certificates

Tristan Buckmaster: Statement on the Alpöge–Buckmaster results

Tristan Buckmaster on Mastodon: training dates and customer data

Andreas Thom on Mathstodon: whether researchers can trust OpenAI with unpublished mathematics

Terence Tao on Mathstodon: the difficulty landscape and open science

Terence Tao on Mathstodon: how AI could contaminate the Navier–Stokes problem

Terence Tao: Finite time blowup with smooth forcing term for the IPM, Boussinesq and Euler equations

TechCrunch: OpenAI fought dirty on career-making math problem, says NYU mathematician

Quanta Magazine: AI Has Solved One of Math’s $1 Million Millennium Prize Problems

MIT Technology Review: What OpenAI’s latest controversy tells us about the future of math

WIRED: OpenAI Just Claimed a Huge Math Discovery. Some Academics Are Crying Foul

Fortune: OpenAI says it cracked one of math’s grand challenges

The Decoder: OpenAI’s millennium proof dispute raises the question of whether researchers can trust AI labs

OfficeChai: Sébastien Bubeck’s account of the coordination attempt

Clay Mathematics Institute: Navier–Stokes Equation

Clay Mathematics Institute: The Millennium Prize Problems