OpenAI math breakthroughs now come with a second question attached: whose unpublished work helped produce them? On Tuesday 8 September 2026, OpenAI said a swarm of AI agents running on an unreleased internal model had resolved the Navier–Stokes existence and smoothness problem, one of the Clay Mathematics Institute’s seven Millennium Prize Problems. Within two days, two mathematicians who had used OpenAI’s own products on their unpublished research were asking the company for something it has not offered: proof that their work did not help.
The first is Tristan Buckmaster, a mathematics professor at New York University. He spent most of a year on related fluid equations with Levent Alpöge, an Anthropic researcher collaborating in a personal capacity, and says the pair put “all our drafts for the whole of this project” into OpenAI’s Codex. The second, reported by The Verge on 10 September, is Andreas Thom, a group theorist at TU Dresden. He says OpenAI never properly answered whether his ChatGPT conversations became training data for an earlier OpenAI math result on non-sofic groups.
OpenAI’s position is precise, and the precision is the story. The company says no researcher or agent saw the pair’s work before it was published and that “no specific user data was accessed”. It also says it “cannot rule out that de-identified data derived from their usage of our products helped improve our models.” Mathematicians want that gap in the OpenAI math story closed with evidence rather than assurances.
Below we set out the OpenAI math timeline, the compute arithmetic, both accounts of the phone calls, what OpenAI’s data controls actually cover and what real proof would look like, then what the dispute means for any business that puts unpublished work into AI tools. For background, see our coverage of OpenAI’s GPT-6 Astra launch and our GPT-5.6 Sol guide, the model Buckmaster singles out among those he used inside Codex.
Table of contents
- What Mathematicians Want the OpenAI Math Team to Prove
- The OpenAI Math Timeline, Day by Day
- The OpenAI Math Effort in Numbers
- What the OpenAI Math Statement Denied and What It Would Not Rule Out
- Why De-identified Data Is the Crux of the OpenAI Math Dispute
- Two Accounts of the OpenAI Math Phone Calls
- Andreas Thom and the Earlier OpenAI Math Result
- Could the OpenAI Math Model Have Found the Route on Its Own?
- What Real Proof of OpenAI Math Provenance Would Look Like
- How OpenAI Math Races Change the Norms of Mathematics
- What the OpenAI Math Row Means for Businesses
- What Happens Next for OpenAI Math Claims
- OpenAI Math FAQs
- References and Further Reading
What Mathematicians Want the OpenAI Math Team to Prove
Nobody in this dispute is asking OpenAI to retract its Navier–Stokes proof, or to hand back a prize the company says it will not claim. The request is narrower and harder to meet. Mathematicians want the OpenAI math team to show, with records rather than statements, that nonpublic research typed into ChatGPT and Codex did not shape the systems that then reached neighbouring results first.
Thom framed the burden in his Mathstodon posts. Researchers, he said, are not equipped to reverse-engineer a training pipeline, and “only OpenAI has the relevant data for that.” If the company denies using such material, The Verge reports, he believes the responsibility is on OpenAI to prove it by disclosing the datasets involved and clarifying the settings and terms that govern how it uses data.
The two questions Thom asked about an earlier OpenAI math result
Shortly after OpenAI announced ten results in mathematics and theoretical computer science on 1 August, Thom emailed OpenAI researchers Mark Sellke and Sébastien Bubeck. He wrote that he and a Dresden colleague had been discussing the expander matching problem, and extensions of his joint work with Gábor Kun, “actively over the last months with ChatGPT.” He asked whether that “was part of the training data or accessible to the reasoning process.”
Sellke’s complete reply, as Thom quotes it, was: “Regarding your conversations with ChatGPT: that did not happen.”
Why Thom now calls that reply dishonest
Thom’s objection is that he asked two separate things and received one categorical answer. That answer, he wrote on 9 September, “now looks as though it addressed only direct access.” He added: “No such qualification, explanation, or evidence was given. I take this as dishonesty to say the least.” Replying to another user in the same thread, he said he had opted out on 29 June.
In later posts quoted by The Verge, Thom called Sellke’s answer “at minimum, unjustifiably broad and materially misleading; looking back it was plainly dishonest.” He said he began reflecting on that exchange after Buckmaster publicly questioned whether the OpenAI math effort had benefited from his use of Codex.
Buckmaster’s unanswered OpenAI math question
Buckmaster’s version of the request sits in the four-page statement he posted on 7 September. On a call with OpenAI the day before, he asked whether the model “had been trained on, or had access to, our sessions in Codex.” His account of the reply: “I was told the model did not look up user data. I asked again, about training, and I did not get an answer.”
He is explicit about the limits of his claim. “I do not know whether our data was used. I am not accusing anyone of anything,” he writes. What he wants is the history told accurately: if an OpenAI model closed the gap to Navier–Stokes, “it should be said loudly, by them, with the history intact.”
The OpenAI math questions and the answers so far
Set side by side, the questions put to OpenAI and the answers on the record do not yet meet.
| Question | Asked by | Answer on the record | Still open |
|---|---|---|---|
| Could the system read our Codex sessions while it worked? | Buckmaster, 6 September call | Told the model did not look up user data; OpenAI says no researcher or agent saw the work | Independent confirmation from logs |
| Was the model trained on those sessions? | Buckmaster, 6 September call | No answer on the call; OpenAI later said it cannot rule out de-identified data helping | Whether the data was eligible for training, and when |
| Did our ChatGPT conversations enter training data? | Thom, August email | “That did not happen,” from Mark Sellke | Thom says the reply only covered access |
| Will OpenAI disclose the datasets, settings and terms involved? | Thom, 9 September posts | No specific response reported by 10 September | All of it |
| Can AI companies demonstrate that chat logs will not improve their models? | Brendan Hassett, Brown University, to The Verge | General data-control policies | A way to demonstrate it |
The OpenAI Math Timeline, Day by Day
Both sides of the OpenAI math dispute lean on the order of events. Buckmaster’s case rests on it: his results came first, word of them reached OpenAI, and only then did OpenAI’s agents take the same route. OpenAI’s own account confirms much of that sequence and disputes what it implies.
Before the OpenAI math run began
The earliest relevant dates are months old. Thom says he opted out on 29 June, and OpenAI published its ten results on 1 August. Buckmaster and Alpöge obtained blowup results with smooth forcing for the Boussinesq and Euler equations on 15 August, and verified the Euler proof in Lean on 22 August. OpenAI says it began training a new internal model, one that “exhibited unprecedented performance in our benchmarks, including mathematics,” on 28 August.
The week the OpenAI math agents launched
OpenAI says that on Tuesday 1 September it heard rumours that two Millennium Prize Problems had been resolved, and launched agents on all of the open ones. On Thursday 3 September, with a rumour circulating that Anthropic had solved a major problem, Buckmaster emailed a prominent OpenAI mathematician to set out the facts. The reply said details “would be useful to avoid competing here” and offered compute.
OpenAI asked to meet on Friday 4 September and again at 12:45 on Sunday 6 September. That afternoon Buckmaster spoke twice with the OpenAI mathematician he had emailed and with Bubeck. By OpenAI’s account, its agents had already reached a resolution on Saturday 5 September, about 88 hours after launch, and Lean verification finished on 6 September.
The OpenAI math announcement and its fallout
Buckmaster posted three results and his statement just before midnight New York time on Monday 7 September, according to Quanta Magazine. OpenAI published its solution, a paper and a GitHub repository of Lean certificates on 8 September, followed by posts on X from the company, from Bubeck, from chief executive Sam Altman and from chief research officer Mark Chen. Thom’s posts followed on 9 September, and The Verge published its report on 10 September.
| Date (2026) | What happened | Source |
|---|---|---|
| 29 June | Andreas Thom says he opted out | Thom on Mathstodon |
| 1 August | OpenAI announces ten results, including a construction of non-sofic groups | OpenAI |
| 15 August | Buckmaster and Alpöge obtain blowup with smooth forcing for Boussinesq and Euler | Buckmaster statement |
| 22 August | Their Euler proof is verified in Lean | Buckmaster statement |
| 28 August | OpenAI starts training a new internal model | OpenAI |
| 1 September | OpenAI hears rumours of solved Millennium Prize Problems and launches its agents | OpenAI |
| 3 September | Buckmaster emails a prominent OpenAI mathematician; the reply offers compute | Buckmaster statement |
| 5 September | OpenAI’s agents reach a Navier–Stokes resolution, about 88 hours after launch | OpenAI |
| 6 September | Lean verification finishes; Buckmaster has two calls with OpenAI | OpenAI, Buckmaster statement |
| 7 September | Buckmaster posts three results and his statement | Buckmaster, Quanta Magazine |
| 8 September | OpenAI publishes its solution, Lean certificates and a statement on user data | OpenAI, GitHub |
| 9 September | Thom asks whether his ChatGPT use fed the non-sofic group result | Thom on Mathstodon |
| 10 September | The Verge reports that mathematicians want proof | The Verge |
OpenAI’s own figures show how compressed its side of the OpenAI math timeline was.
The six days that matter to Buckmaster
The interval Buckmaster keeps returning to runs from 22 August, when his Euler proof passed Lean, to 28 August, when OpenAI’s new model started training. On 8 September he quoted OpenAI’s line about that training run on Mastodon and wrote: “Note that they are openly admitting they used training data from a period after we found our result. Is it ethical to use customer’s data to try to scoop their customer?”
That is an inference, not a disclosure. A training start date says when a run began, not what data it contained, and OpenAI says training on the model is ongoing. The dates show that Buckmaster’s OpenAI math question is reasonable. They do not answer it.
The OpenAI Math Effort in Numbers
For the OpenAI math run, the company disclosed more operational detail than AI companies usually do. Its post describes groups of coordinating agents with access to “a cached version of the internet” and the ability to run code, each group prompted with a different variant of the problem. Codex was used to consolidate the most useful insights from each group, and the agents were switched to a further trained version of the internal model partway through.
What the OpenAI math post disclosed
Every figure below comes from the 8 September OpenAI math post except the last row, which is our own multiplication.
| Measure | Euler (unforced) | Navier–Stokes | All problems |
|---|---|---|---|
| Agents | Nearly 100 | On the order of 10,000 concurrent | Not stated |
| Search time | About 50 hours | About 88 hours | Not stated |
| Lean formalisation | Not stated | 17 more hours, via GPT-6 Astra | Not stated |
| Messages between agents | Not stated | 2.7 million | 4.9 million |
| Output tokens | Not stated | About 130 billion | About 300 billion |
| Agent-hours (our arithmetic) | About 5,000 (100 × 50) | Up to about 880,000 (10,000 × 88) | Not stated |
Group sizes varied and “on the order of 10,000” is not an exact count, so the Navier–Stokes figure is a rough ceiling rather than a measurement. Even so, the ratio is striking. On those numbers the Navier–Stokes search used up to about 176 times the agent-hours of the Euler search that preceded it.
Navier–Stokes took most of the run’s messages but less than half of its output tokens.
What the OpenAI math run might have cost
No official cost has been published for the OpenAI math run. Mark Chen said at a press briefing that the effort cost “in the millions of dollars,” WIRED reported. Fortune reported that OpenAI told reporters it used at least 1,000 times the compute of earlier math work that cost about $2,000, which implies about $2 million. TechCrunch valued the 300 billion output tokens at $22.5 million at current Astra rates.
OpenAI’s API pricing page lists GPT-6 Astra output at $50 per million tokens in the standard tier for short context and $75 for long context. Batch and Flex halve the short-context price to $25, and Fast mode doubles it to $100. Multiplying 300,000 million tokens by each rate gives the range below.
Treat those figures as a scale marker, not an invoice. The run used an internal model rather than Astra, the arithmetic ignores input tokens, and a list price is not what OpenAI pays to run its own hardware. The point Brown University’s Javier Gómez-Serrano made to MIT Technology Review survives any choice of rate: “very few mathematicians will have resources of that scale.”
What a swarm changes
András Juhász of the University of Oxford put the asymmetry bluntly to The Verge: “Suddenly, 10,000 mathematicians jump on your problem.” OpenAI researcher Noam Brown argued on X that such costs fall fast. He noted that o3 cost about $500,000 to score 87.5% on ARC-AGI 1, while Astra scores higher for about $20. “I believe that a year from now everyone will have an AI at their fingertips capable of solving problems of this caliber,” he wrote.
What the OpenAI Math Statement Denied and What It Would Not Rule Out
The OpenAI math statement packs three claims into four sentences, and they are not the same kind of claim. Two are denials of direct access. One is an admission that an indirect path cannot be excluded. Much of the anger in mathematical circles comes from treating the first two as though they answered the third.
Access and training are different things
Mark Chen drew the line himself on X. “Did any human or agent look at user data as part of the Navier Stokes effort? No,” he wrote. “Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company.” Bubeck told the press briefing, according to Fortune, that “we did not use their prompts or proofs to prompt our models or direct our agents.”
Those are statements about what people and agents did during the OpenAI math effort. They say nothing about what a model had already absorbed before the effort began, which is the question Buckmaster says went unanswered.
Each OpenAI math statement and what it leaves open
| Statement | Source | What it covers | What it leaves open |
|---|---|---|---|
| “We (the researchers and the agents) did not see any of their work through any means until they released it publicly” | OpenAI post and X, 8 September | Viewing by staff or agents during the effort | Anything absorbed in training beforehand |
| “No specific user data was accessed in order to solve this problem” | OpenAI post and X | Lookups of identifiable user records | De-identified data |
| “While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models” | OpenAI post and X | Admits a possible indirect path | How likely, and whether the accounts were eligible |
| “Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes.” | Mark Chen on X | Confirms the general practice | Which data, and from which dates |
| “We did not use their prompts or proofs to prompt our models or direct our agents” | Bubeck at the press briefing, via Fortune | Prompting and direction | Training |
| “The chances are zero if they have opted out (likely)” | OpenAI employee posting as roon on X | Probability, conditional on an opt-out | Whether either account had opted out |
The boundary OpenAI says it will not cross
The roon post, which The Decoder cites among OpenAI staff pushing back, called it “exceptionally unlikely that anything they ever did made it into any part of training.” It then explained why OpenAI will not simply check. Doing so, it said, would be “a terrible precedent to break the … PII-scrubbing boundary to go and round it down to 0, and we won’t do it.”
That sentence explains why the OpenAI math dispute is so hard to settle. The protections that stop staff looking up a named user’s data are the same protections that stop them proving a named user’s data was absent. On our reading, OpenAI is choosing privacy over proof, and the mathematicians are asking whether that trade should be OpenAI’s alone to make.
How the Anthropic researcher read it
Alpöge’s own reaction to the OpenAI math carve-out was notably mild. “i mean props to them for straight coming clean,” he wrote on X, before adding that “so far the proof looks more along the lines of another euler blowup proof we had.” Coming from the co-author most directly affected, that is a reminder that the dispute is about evidence and conduct, not about whether the OpenAI mathematics holds up.
Why De-identified Data Is the Crux of the OpenAI Math Dispute
De-identification is a privacy technique. Often automated with natural language processing, it is designed to strip names, email addresses and other personal details out of data before the data is used. Thom’s objection, quoted by The Verge, is that this solves the wrong problem for research: “De-identification may remove a name; it does not remove the intellectual content of a mathematical idea.”
OpenAI’s own description of its process draws the same line without intending to. Its help centre says the company takes “steps to reduce the amount of personal information in our training datasets.” A half-finished proof strategy is not personal information. Removing the author’s name leaves the strategy intact, which is precisely what a mathematician worried about being scooped would fear from the OpenAI math pipeline.
What OpenAI’s data controls actually say
OpenAI publishes different defaults for individual and business products. The table summarises its help-centre article on how data is used to improve model performance, as it read on 10 September 2026.
| Product or feature | Used for training? | Control | Catch |
|---|---|---|---|
| ChatGPT, individual accounts | May be used unless you opt out | Data Controls, or “do not train on my content” in the privacy portal | Applies to new conversations |
| Codex tasks, individual accounts | May be used unless you opt out | The same opt-out as ChatGPT | Full-environment training has a separate Codex setting |
| Rating a response | The entire conversation may be used, even after an opt-out | Do not rate sensitive chats | Easy to trigger by habit |
| Temporary Chat | Not used for training | Choose Temporary Chat | No history or memory |
| ChatGPT Business, Enterprise and the API | Not used for training by default | Opt-in only, for example Playground feedback | Check workspace settings and contracts |
Why opting out is not proof either
Opting out is a forward-looking switch. OpenAI’s wording is that once you opt out, “new conversations will not be used to train our models.” Earlier conversations are not mentioned. Thom says he opted out on 29 June, which leaves any earlier discussion with ChatGPT outside that protection, and nobody outside OpenAI can check what happened to it.
For Buckmaster the facts are thinner. His statement says he pays for his group’s tools “out of my own research funds, including footing a large bill to OpenAI,” but it does not say which plan he used or whether training was switched off. The Decoder notes that whether the pair had opted out is not publicly known. That one setting matters, because training is off by default on OpenAI’s business products.
Codex has its own switch
Codex adds a wrinkle many users miss. OpenAI’s help centre says Codex “has separate controls for allowing training on full environments,” and that adjusting settings in ChatGPT or the privacy portal “will not affect these full-environment Codex settings.” Buckmaster describes Codex as the place “into which we had been putting all our drafts for the whole of this project.” For anyone using Codex on unpublished work, that is the setting to check first. Our comparison of AI coding assistants covers how Codex fits alongside its rivals.
Two Accounts of the OpenAI Math Phone Calls
Much of what happened on the OpenAI math calls of 6 September is known only from the two sides’ own descriptions. Buckmaster’s statement, Bubeck’s post on X and Altman’s post on X agree on the broad outline. They disagree sharply on the details that matter most.
| Point | Buckmaster and Alpöge | OpenAI (Bubeck and Altman) |
|---|---|---|
| How OpenAI heard | Alpöge had received tips that information about their progress had been passed to OpenAI | Rumours on the internet that Anthropic’s models had solved a Millennium Problem |
| Human input | Told of “very little human input”, then learned an entire team had worked on it | Post describes agent groups given problem variants, with insights consolidated using Codex |
| Alpöge’s authorship | Bubeck “twice asserted” he wanted Alpöge removed from authorship | Bubeck never asked to remove him from his own work; the remark concerned a rewrite of OpenAI’s proof |
| The career remark | “Why would you ruin your career?” then “If you don’t want me to be nice, then I don’t have to be nice.” | Bubeck calls it an extremely poor choice of words, apologises, and says he retracted it on the spot |
| Coordination | Alpöge declined a one-on-one call and said conversations should be with Buckmaster | Altman says Alpöge “was not willing to talk or coordinate with us” |
| The prize | Declined proposals under which OpenAI would say they deserved the Clay Prize | OpenAI says it will not claim the prize |
What both sides of the OpenAI math dispute agree on
Several points are not in dispute. OpenAI went after Navier–Stokes because of rumours about Anthropic; Altman wrote that “it is true that we tried this because there were rumors on the internet last week that Anthropic’s models had solved a millennium problem.” Buckmaster and Alpöge had blowup for forced Euler first, and OpenAI’s post says it recognises “the priority of their work on forced Euler.” OpenAI also offered them visibility into its prompts and, later, the proof.
Authorship and the Anthropic question
The sharpest disagreement in the OpenAI math calls is over Alpöge. Buckmaster writes that Bubeck “twice asserted that he wanted Levent removed from authorship.” Bubeck says he “never ever asked for Levent to be removed from authorship of his own work.” He says his remark that things “would be simpler if Levent was not an Anthropic employee” concerned a proposed rewrite of OpenAI’s own proof, which he felt an Anthropic employee should not author.
Alpöge, who was not on the calls, described it differently on X, referring to “the part where a millennium prize was offered if i’d just be removed from the paper.” He also wrote that he would have been “pumped to collaborate” and did not care about authorship on that step.
The career remark
Both accounts include a line about a career. Buckmaster quotes Bubeck asking “Why would you ruin your career?” and then saying “If you don’t want me to be nice, then I don’t have to be nice.” In Bubeck’s version he said he did not understand why one “would risk their career” over what he calls unfounded accusations. He apologises for “this extremely poor choice of words” and says he retracted it on the spot. His post does not address the second line.
Andreas Thom and the Earlier OpenAI Math Result
Thom’s complaint predates the Navier–Stokes dispute by more than a month and concerns a different result. On 1 August OpenAI published ten results “achieved by an internal version of Astra,” at a token cost it put at roughly $2,000 at Sol API rates. One was “a construction establishing the existence of non-sofic groups, addressing a central open question in group theory.”
Why this OpenAI math result touched Thom’s work
Non-sofic groups are, in The Verge’s summary, infinite mathematical structures that cannot be approximated by finite ones, and whether any exist was a long-standing open question. The Verge reports that OpenAI acknowledged its construction built heavily on earlier work by Thom and Gábor Kun, was criticised for failing to credit their recent contributions, and quietly amended its write-up.
Thom says he was struck by “OpenAI’s detailed command of our techniques,” which he describes as neither the most obvious nor the most promising route at the time. That is the same OpenAI math pattern Buckmaster describes: a technique few people were pursuing, adopted quickly by a company whose products had hosted private conversations about it.
What Thom says would be indefensible
Thom’s strongest line is conditional. He said it “would be ethically indefensible” if nonpublic research supplied by users helped improve models that the company then used to race those same users to publication, without consent, proper disclosure or credit. He has not said that happened. He has said OpenAI’s answers so far do not let him rule it out.
OpenAI’s own standard on attribution
OpenAI’s August post contains a principle mathematicians now want applied to data. “We believe attribution should honestly reflect how a result was produced,” it says, warning that claiming human authorship for an AI-generated proof would misrepresent the system’s contribution. The logic runs both ways. If human conversations helped an OpenAI math result, honest attribution would say so, and so far OpenAI’s answers stop short of that.
Could the OpenAI Math Model Have Found the Route on Its Own?
This is the scientific heart of the OpenAI math dispute, and experts disagree in good faith. Both teams built on a programme developed by Diego Córdoba of the Institute for Mathematical Sciences in Madrid and Luis Martínez-Zoroa of CUNEF University, which constructs a blowup through what Martínez-Zoroa calls an “infinite cascade” of layers. Charles Fefferman, who wrote the Clay Institute’s official problem description, told Quanta the heroes of the story are Córdoba and Martínez-Zoroa.
The case that the route was a red flag
Buckmaster argues that the route through a smooth force, options C and D in Fefferman’s statement of the problem, was the one he and Alpöge “had quietly chosen to attack.” He wrote: “Almost nobody else I know of was working on it. It is not the direction one arrives at in a few days by giving a model the problem statement. When I heard ‘forced,’ it was a bright red flag.”
He also says he was first told the model had simply been given the problem statement. Over the call, he writes, it emerged that an entire team had worked on the problem, that the model had first been set on easier problems including Euler, and that even the prompt shown to him “had been written by prompting Codex.”
The case that it was findable
Others see a well-lit path. Gómez-Serrano told MIT Technology Review that the Córdoba–Martínez-Zoroa approach was one of several thought to hold promise. The OpenAI math post says its agents first resolved the unforced Euler problem and then carried that result into Navier–Stokes, and it insists the two teams’ proofs differ significantly.
OpenAI’s Boaz Barak wrote on X that the model “actually started off by proving a stronger claim than they did,” and OpenAI mathematician Ven Chandrasekaran stressed, according to WIRED, that the pair’s solution was significantly different in nature from the one OpenAI’s system produced.
Why a public programme does not settle it
The Córdoba–Martínez-Zoroa papers are public, and OpenAI’s agents could read a cached copy of the internet, so a system could plausibly have found the route unaided. MIT Technology Review points to a subtler issue. Choosing which promising direction to pursue, often called research taste, has long been one of the hardest things for AI. If the agents followed the route because humans had already committed to it, human judgement was essential. Without the OpenAI math agents’ logs, outsiders cannot tell which story is true.
What Real Proof of OpenAI Math Provenance Would Look Like
Proving a negative about a training run is hard but not impossible, and much of the evidence already exists inside OpenAI. The options below differ in what they would show and in how far each would require the company to open the privacy boundary it has promised to keep.
| Evidence | What it would show | Who holds it | Limits |
|---|---|---|---|
| Account training settings and their change dates, including Codex full-environment settings | Whether the sessions were eligible for training at all | OpenAI, released with the users’ consent | Says nothing about data from before an opt-out |
| Data cut-off dates for each model snapshot the agents used | Whether August sessions could be in the run that began on 28 August | OpenAI | Training is ongoing, so several snapshots may matter |
| A search of the de-identified corpus for distinctive strings from the drafts | Whether identifiable mathematical content was present | OpenAI, or an auditor under confidentiality | Paraphrased or derived content may not match |
| Agent transcripts, tool calls and cached pages read during the run | What the agents actually consulted and tried | OpenAI | Millions of messages; only prompts were offered to Buckmaster |
| An independent audit with an agreed scope | A credible answer without publishing anyone’s conversations | OpenAI and an auditor | Requires OpenAI to open the boundary in a controlled way |
| Outside tests for memorised text | Whether a model reproduces unpublished passages | Anyone with model access | Can suggest presence; rarely proves absence |
“Unlikely” is a probability, not an audit
OpenAI’s word for the chance that the pair’s data helped is “unlikely.” That may well be accurate. It is also impossible to test from outside, because only OpenAI holds the account settings, the data snapshots and the agent logs. MIT Technology Review adds that incidents such as the Hugging Face hack show OpenAI is not always aware of what its agents are doing, a pattern we covered in OpenAI’s German wiki incident.
A realistic minimum for OpenAI math disclosure
Some steps would cost OpenAI little. With the users’ consent, it could confirm whether their accounts and Codex full-environment settings were eligible for training during the relevant months. It could publish data cut-off dates for the snapshots its agents used, and release the OpenAI math agent transcripts for Navier–Stokes, dead ends included. Terence Tao has criticised “the refusal of AI companies to disclose their negative results, or reveal the process towards obtaining their solutions.”
Private shorthand makes a natural test
Research drafts often contain shorthand that appears nowhere else. Alpöge mentioned on X that he and Buckmaster gave their blowup ansätze pun names such as “smooth criminale.” A string like that, searched for in training snapshots taken before the pair went public, would give a far more concrete answer than “unlikely”, without anyone’s conversations being published. Brendan Hassett of Brown University told The Verge that companies “should be held accountable to deliver” assurances about chat logs.
How OpenAI Math Races Change the Norms of Mathematics
The data question is only half of what alarms mathematicians. The other half is the OpenAI math race itself. “Mathematics depends heavily on an informal norm of trust,” Matthew Ballard of the University of South Carolina told The Verge. “Researchers routinely share incomplete ideas and ongoing work with colleagues to sharpen their thoughts. It is done with the expectation that it will not turn into a competition.”
A rumour is now enough to start an OpenAI math race
Terence Tao of UCLA described the new incentive on 8 September. “We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential,” he wrote. He warned that incentives may now favour “no longer sharing any promising research directions with the broader community, which would reverse centuries of traditions of open science.”
Abhishek Saha of Queen Mary University of London said OpenAI had engaged in the “kind of things that mathematicians will generally not do.” Scooping happens between humans too, he told The Verge, but it is not easy, because few people have the specialised expertise to swoop in on someone else’s problem. OpenAI math swarms remove that barrier.
Caution is the rational response
Jeremy Avigad of Carnegie Mellon University, director of the Institute for Computer-Aided Reasoning in Mathematics, told The Verge that “the thought that AI systems might steal ideas from our queries is chilling.” With a swarm a hint away, he expects that “people are likely to be more cautious.” Yang-Hui He of the London Institute for Mathematical Sciences worries that “maths under the big companies is much too secretive,” recalling an era when mathematics depended on patrons such as the Medici family.
What the solution was supposed to be for
Tao argued before the announcement that the Navier–Stokes problem matters less for its answer than for what attempts to solve it teach. Prematurely solving such problems “by purely AI-powered methods — particularly without full transparency into the solution process — can contaminate this process to the point where it actually becomes a net negative for the progress of mathematics as a whole,” he wrote on 3 September.
Martínez-Zoroa, whose ideas both teams used, was generous when Quanta asked. “It would have been nice to do this ourselves, but I’m very happy for him,” he said of Buckmaster. Juhász, for his part, called the episode “clearly a PR victory for OpenAI” while questioning how sustainable that approach could be.
What the OpenAI Math Row Means for Businesses
Swap “mathematician” for “product team” and the dispute becomes familiar. Companies put unreleased designs, source code, pricing models and strategy drafts into AI assistants every day. The OpenAI math row shows that a vendor can truthfully say nobody looked at your data, while being unable to say whether that data, stripped of your name, helped improve a model that anyone can use.
Defaults matter more than promises
The most useful fact in OpenAI’s documentation is a default. Individual ChatGPT and Codex accounts may be used for training unless you opt out, while ChatGPT Business, ChatGPT Enterprise and the API are not used for training by default. OpenAI’s enterprise privacy page says API Platform data after 1 March 2023 “isn’t used for training our models, unless you have explicitly opted in.” For sensitive work, the plan you choose does more than any setting you remember to change.
A practical checklist after the OpenAI math row
| Step | Why it matters | Who it applies to |
|---|---|---|
| Use business plans for unpublished work | Business, Enterprise and API data is not used for training by default | Teams handling code, research or other IP |
| Opt out on individual accounts and record the date | An opt-out covers new conversations only | Personal ChatGPT and Codex users |
| Check the Codex full-environment training setting separately | ChatGPT settings and the privacy portal do not change it | Developers using Codex |
| Do not rate responses in sensitive chats | Feedback can bring the whole conversation into training | Everyone |
| Use Temporary Chat for one-off sensitive questions | Temporary chats are not used for training | Quick lookups |
| Keep dated records of unpublished work | Priority disputes turn on dates, as 15 and 22 August did here | Research, engineering and legal teams |
| Ask vendors about de-identified and derived data | That carve-out is exactly where OpenAI’s denial stopped | Procurement and contract owners |
Put it in the contract
For organisations that buy AI services at scale, the OpenAI math lesson belongs in procurement. Ask vendors in writing whether “de-identified data derived from usage” can be used to improve models, how an opt-out is recorded, and what evidence they would provide in a dispute. Our vendor management and data protection services can build those questions into supplier reviews.
Provenance runs in both directions
Questions about whose work trained whose model are not unique to OpenAI or to mathematics. Our coverage of the AI distillation advisory looks at the mirror image of this dispute: companies training on other labs’ model outputs. Our AI Models, Tools and Releases hub tracks how vendor policies on training data change.
What Happens Next for OpenAI Math Claims
Three separate OpenAI math questions now run on different clocks: whether the proof holds, whether any user data mattered, and who deserves credit. Each will be settled, if at all, by different people.
Verifying the OpenAI math proof
The proof is formalised in Lean, and OpenAI’s certificates are public on GitHub under an Apache 2.0 licence. That gives mathematicians unusual confidence, but Quanta notes one step still needs humans: checking that the statement proved in Lean is logically equivalent to the problem people set out to solve. If the result holds, Quanta called it, by a significant margin, the most important proof yet produced by an AI model. The Clay Institute has not awarded the prize, and OpenAI says it does not intend to claim it.
The OpenAI math data question
Thom has asked OpenAI to disclose datasets, settings and terms, and OpenAI had not responded to The Verge by publication. Buckmaster says he and Alpöge also believe they have blowup for a hypo-dissipative version of Navier–Stokes, a paper held back because its Lean verification had not finished. Watch for any precise definition of “de-identified data derived from usage”, because that phrase now carries the weight of OpenAI’s defence.
Credit and cooperation
Both sides say they wanted cooperation. Bubeck cited Sholto Douglas’s remark that Anthropic and OpenAI will need to coordinate in future, and Alpöge wrote that he likes “the idea of the labs cooperating, and even better on scientific progress. It’s a shame!” Buckmaster’s hope is for the “serious and unhurried discussion about where to go from here” that his statement says the community needs.
OpenAI Math FAQs
Did OpenAI really solve the Navier–Stokes Millennium Prize Problem?
OpenAI says an internal system produced a proof, formalised in Lean, that an initially smooth fluid with a smooth applied force can develop a singularity in finite time, establishing statements C and D of the official problem. The Lean certificates are public. Independent scrutiny is continuing, and the Clay Mathematics Institute has not awarded the prize.
Did the OpenAI math effort use Tristan Buckmaster’s Codex data?
Nobody outside OpenAI knows. OpenAI says no person or agent looked at his or Alpöge’s work before publication and that no specific user data was accessed, but it cannot rule out that de-identified data from their product use helped improve its models. Buckmaster says he does not know whether his data was used and is not accusing anyone.
What does “de-identified data derived from their usage” mean?
OpenAI has not defined the phrase for this case. Its help centre says it takes steps to reduce personal information in training datasets. Critics such as Andreas Thom argue that removing a name does not remove the mathematical ideas in a conversation.
Who is Andreas Thom, and what is he asking OpenAI?
Thom is a group theorist at TU Dresden whose joint work with Gábor Kun underpins OpenAI’s August result on non-sofic groups. He asked whether his ChatGPT conversations entered training data or were accessible to the solving process, and he now wants OpenAI to disclose the datasets, settings and terms needed to show they were not used.
Will OpenAI receive the $1 million Millennium Prize?
OpenAI says it does not intend to claim it. The Clay Mathematics Institute administers the Millennium Prize Problems and has not awarded a prize for this result.
How do I stop ChatGPT or Codex training on my work?
On individual accounts, opt out in Data Controls or OpenAI’s privacy portal, check Codex’s separate full-environment training setting, use Temporary Chat for sensitive one-off questions and avoid rating responses in sensitive chats. On ChatGPT Business, ChatGPT Enterprise and the API, training is off by default.
References and Further Reading
The Verge: Mathematicians want proof OpenAI didn’t use their work
The Verge: OpenAI’s sly mathematical breakthrough sends a chill through academia
The Verge: Drama swirls around OpenAI’s legendary mathematical milestone
OpenAI: On the Navier–Stokes Millennium Prize Problem
OpenAI: Ten advances in mathematics and theoretical computer science
OpenAI Help Center: How your data is used to improve model performance
OpenAI: Enterprise privacy at OpenAI
GitHub: openai/NavierStokesAndEuler Lean certificates
Tristan Buckmaster: Statement on the Alpöge–Buckmaster results
Tristan Buckmaster on Mastodon: training dates and customer data
Andreas Thom on Mathstodon: whether researchers can trust OpenAI with unpublished mathematics
Terence Tao on Mathstodon: the difficulty landscape and open science
Terence Tao on Mathstodon: how AI could contaminate the Navier–Stokes problem
Terence Tao: Finite time blowup with smooth forcing term for the IPM, Boussinesq and Euler equations
TechCrunch: OpenAI fought dirty on career-making math problem, says NYU mathematician
Quanta Magazine: AI Has Solved One of Math’s $1 Million Millennium Prize Problems
MIT Technology Review: What OpenAI’s latest controversy tells us about the future of math
WIRED: OpenAI Just Claimed a Huge Math Discovery. Some Academics Are Crying Foul
Fortune: OpenAI says it cracked one of math’s grand challenges
OfficeChai: Sébastien Bubeck’s account of the coordination attempt
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.