Math results from an unreleased OpenAI model went public on Tuesday 6 October 2026, and they arrived in bulk. Shortly before 6pm US Eastern time, OpenAI opened a GitHub repository holding 722 manuscripts grouped into 372 “families” of related results. Each family claims to resolve, or make substantial progress on, an open question in mathematics or theoretical computer science. Engadget’s Steve Dent summed it up the next morning: OpenAI just posted hundreds more results on major math problems.
The release follows the Navier-Stokes result in September and OpenAI’s statement last month that its system had resolved “more than 100 long-standing open problems”. It also lands a week after an independent panel on mathematics and artificial intelligence, hosted by the Institute for Advanced Study, published rules for how labs should release machine-generated proofs. OpenAI says it drew on that advice. The panel says the release is “the beginning, not the completion” of the work.
This guide sets out what is actually in the release: the headline math results, how many are checked by the Lean proof assistant, how the release measures up against the advisory group’s rules, what the dates inside the repository reveal, and how mathematicians reacted. It closes with what the episode teaches businesses that are starting to rely on AI output nobody can check by hand.
Table of contents
- The Math Results OpenAI Released on 6 October
- The Headline Math Results OpenAI Is Claiming
- How Many Math Results Are Machine-Checked?
- Did the Math Results Follow the Advisory Group’s Rules?
- What the Dates Behind the Math Results Show
- How Mathematicians Reacted to the Math Results
- One Agent, One Prompt: How the Math Results Were Made
- What These Math Results Mean for Businesses
- What Happens Next for OpenAI’s Math Results
- Math Results FAQs
- References
The Math Results OpenAI Released on 6 October
OpenAI announced the release on X at 22:19 UTC with a short post and a link to a blog entry titled “Sharing AI progress in mathematics”. The blog post is brief. The math results themselves sit in the repository, which carries an Apache 2.0 licence and the plain name openai/math.
The repository, the post and the timing
GitHub’s records show the repository was created at 21:47 UTC and its single “Initial commit” landed at 21:58 UTC, just before 6pm in New York, the release time Scientific American reported. Sam Altman shared the blog post at 00:06 UTC with one line: “We are entering a new era of discovery now.” Greg Brockman, OpenAI’s president, added “towards acceleration of scientific discovery and improving quality of life for everyone.”
By the evening of 7 October the repository had close to 7,700 stars and more than 700 forks, which is a lot of attention for math results in preprint form.
722 manuscripts, 372 families and the 377 puzzle
The README describes “722 manuscripts organized into 372 families”. A family groups related papers: a principal result, companion arguments, consequences or alternative proofs. Some families hold a single paper; one holds fourteen. By our count, 206 of the 372 families contain just one manuscript.
Coverage has quoted two different totals. The Indian Express and Anadolu Agency reported 372. The New York Times headline and Pakistan’s The News reported 377. We checked the catalogue, CONTENTS.md, line by line. The family numbers run from 001 to 377, but five numbers are missing: 045, 061, 070, 123 and 163. So there are 372 families, and the 377 figure appears to come from the highest family number rather than a count of math results.
Three hours of Pro thinking per result
The repository explains how the math results were produced. “The vast majority of results were obtained with the same procedure using an unreleased internal OpenAI model,” it says. “On average, each result used three hours of ChatGPT Pro thinking compute with that model.” Over the evaluation, the model “was posed approximately 4,000 problems”. Filtering for significance left the catalogue. That is a keep rate of roughly 9.3%, which matches heise’s description of “just under a tenth”.
OpenAI also explains why it was running these tests at all: “As part of model development, we evaluate our models on open research problems. We expanded these evaluations after performance on our existing mathematical evaluations saturated.” In other words, the math results are a by-product of benchmarking.
The README names two exceptions to the fixed procedure: work on a zero-free region for the Riemann zeta function, and a proof of the Hodge conjecture for CM abelian varieties. It adds that the write-up for one zero-free region result was “human edited for readability”.
| Measure | Figure | Source |
|---|---|---|
| Manuscripts | 722 | Repository README |
| Result families | 372 | README; our count of CONTENTS.md |
| Highest family number | 377 (five numbers unused) | Our count of CONTENTS.md |
| Problems posed to the model | About 4,000 | README |
| Families kept, as a share of problems | About 9.3% | 372 divided by 4,000 |
| Average compute per result | About 3 hours of ChatGPT Pro thinking | README and blog post |
| Reasoning summaries published | 10 | README |
| Families linking a Lean scope document | 235 | Our count of CONTENTS.md |
| Papers in the formalization catalogue | 162 | Our count of lean/formalization.yaml |
The Headline Math Results OpenAI Is Claiming
Several of the claimed math results would, if they hold up, be among the most significant in their fields for decades. Each family in the catalogue opens with a one-paragraph summary, and those summaries are precise about scope. They are worth reading closely, because the press shorthand sometimes went further than the papers.
Number theory: a zero-free strip and the BSD formula
Family 003 claims “the quasi-Riemann hypothesis”. It proves that every Dirichlet L-function, including the Riemann zeta function, has no zeros where the real part of s exceeds 7/8. A companion paper gives a different proof for the weaker 11/12 bound. This is not the Riemann hypothesis, which concerns the line at 1/2. But nobody had previously proved that any fixed half-plane beyond a bound below 1 is free of zeros.
The README also lists “work on a zero-free region for the Riemann zeta function” among the exceptions to its standard procedure, so this family was not produced quite the way most of the catalogue was.
Family 002 claims the full Birch and Swinnerton-Dyer leading-term formula for every elliptic curve over the rationals whose q-power Selmer group has corank zero or one for some prime q. Combined with family 006, that gives full BSD for a density-one set of quadratic twists of every such curve. It is a large conditional result, not a proof of the whole conjecture.
Analysis and geometry: Kakeya and the digits of π
Family 074 claims the Kakeya maximal conjecture in three dimensions and the Hausdorff-dimension form of the Kakeya conjecture in four. Family 017 claims that the irrationality exponent of π is exactly 2, meaning π cannot be approximated by fractions much better than a typical number can. The summary notes this also proves that the Flint Hills series converges, a question that has circulated for years.
Algebra and operator algebras: Kaplansky and free group factors
Family 196 claims a counterexample to Kaplansky’s zero-divisor conjecture: a finitely presented, torsion-free group whose group algebra over the field with two elements has nonzero zero divisors. Family 197 adds a counterexample to Kaplansky’s direct-finiteness conjecture. Family 287 claims to settle the free group factor isomorphism problem, showing the von Neumann algebras of free groups of different ranks are all isomorphic.
Computer science: faster matrix multiplication
Family 107 claims that the exponent of matrix multiplication over the complex numbers is at most 9/4, so multiplying two n-by-n matrices needs on the order of n to the power 2.25 arithmetic operations. It also claims a bound below 2.371054886006746 over every field. The Lean scope note is careful: the result “concerns arithmetic complexity, rather than bit complexity or practical crossover sizes”. Nobody’s code gets faster tomorrow.
| Family | Claimed result | Lean scope document |
|---|---|---|
| 002 | Full BSD formula when the Selmer corank is zero or one | No |
| 003 | Zero-free half-plane beyond 7/8 for zeta and Dirichlet L-functions | Yes |
| 017 | Irrationality exponent of π equals 2 | Yes |
| 074 | Kakeya maximal conjecture (3D), dimension conjecture (4D) | No |
| 107 | Matrix multiplication exponent at most 9/4 over the complex numbers | Yes |
| 196 | Counterexample to Kaplansky’s zero-divisor conjecture | Yes |
| 287 | All free group factors are isomorphic | Yes |
How Many Math Results Are Machine-Checked?
The strongest argument for taking these math results seriously is formal verification. A proof that compiles in Lean has had every logical step checked by a computer, so a reader does not have to trust the prose. That is why Scientific American wrote that the verified results are “all but certain to be correct”.
What Lean and Comparator add
Lean is a programming language and proof assistant used by a fast-growing community of mathematicians. The repository includes a single large Lean library, a formalization catalogue in lean/formalization.yaml, and “challenge” files for Comparator, a tool that checks whether a formal proof matches an agreed formal statement. The instructions are short: install the tools, fetch the cached library, then run a command such as lake env comparator ComparatorChallenges/QuasiRiemannHypothesis.json.
235 scope documents, 162 catalogued papers
We counted how many math results have formal coverage in two ways. Of the 372 family summaries, 235 link to a Lean scope document, about 63%. The formalization catalogue, however, lists 162 papers “with a formalized main result”, about 22% of the 722 manuscripts. The two lists do not line up neatly. The free group factor paper, for example, has a scope document but does not appear in the catalogue.
Each bar is the stated count divided by its total: 235/372, 137/372 and 162/722, rounded to whole percentages.
What a passing check does and does not mean
The scope documents matter because a formal proof certifies exactly what was formalised, no more. For the quasi-Riemann family, the note says the formal work covers the 7/8 bound but “the paper’s later applications are not included”. A check also says nothing about whether the result is new, whether earlier work should be cited, or whether anyone understands it. The README is candid: “Some of the unformalized results could have issues.”
Did the Math Results Follow the Advisory Group's Rules?
OpenAI formed the Advisory Group on Mathematics and Artificial Intelligence (AGMAI) after the Navier-Stokes dispute. Scientific American dates the announcement to 21 September. The group’s nine members are François Charles, Camillo De Lellis, Timothy Gowers, Martin Hairer, Nikhil Srivastava, Ulrike Tillmann, Ravi Vakil, Edward Witten and Melanie Matchett Wood. We covered its formation in our piece on the OpenAI math advisory group.
What AGMAI asked for on 29 September
The group’s guidelines, “Responsible Release of AI-Generated Mathematics”, were informed by more than 600 replies from mathematicians. They open bluntly: “we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models.”
For math results that nobody yet understands, the guidelines set five technical norms. Labs should search the literature and cite related work, and produce readable write-ups. They should deposit papers in repositories “not controlled by any AI lab”. For each result they should publish the model name, prompts, a summarised chain of thought, time and cost. Proofs should be formalised where possible, and labs should document how many comparable problems failed.
Where OpenAI complied
The math results release meets several of those norms. The Lean library includes the formalization.yaml catalogue and Comparator challenge files that the guidelines name. The README states the number of problems posed and the selection step. It gives an average compute figure, publishes ten reasoning summaries, and promises that “previously released versions” will stay accessible. The blog post commits OpenAI to “funding a series of workshops, conferences, and special programs”.
Where it falls short
On the per-result disclosures, the release does not comply. There are no prompts, no per-result compute times and no model name beyond “an unreleased internal OpenAI model”. The ten reasoning summaries cover under 3% of the 372 families. The papers sit in a repository on OpenAI’s own GitHub account, and OpenAI says it is still “exploring other community-hosted alternatives”. Better citations and exposition are promised “for future releases”. An OpenAI spokesperson told Scientific American the company takes the guidelines seriously and is doing its best to comply, but added that it is not bound by them.
| AGMAI recommendation (29 Sep) | What the 6 October release did | Our reading |
|---|---|---|
| Stop testing on proprietary models | Results come from an unreleased internal model | Not met |
| Cite related literature; readable write-ups | Promised for future releases | Deferred |
| Repository not controlled by an AI lab | OpenAI’s GitHub account; alternatives being explored | Not met yet |
| Model name, prompts, chain of thought, time and cost per result | Average compute only; 10 summaries; no prompts; model unnamed | Largely not met |
| Formalise proofs; comparator files and formalization.yaml | Lean library, catalogue and challenge files included | Met for part of the release |
| Report attempts, failures and selection | About 4,000 problems posed; significance filter described | Partly met |
| Fund understanding through existing nonprofits | Workshops and programmes promised, details to follow | Pending |
The advisory group’s own verdict
AGMAI’s statement on the release, posted on its website, is measured. “This is an important event for mathematics, with consequences both for mathematics and for the mathematical community that extend far beyond the individual results,” it says. But its role “should not be interpreted as a judgment of the impact of these results or an endorsement of the process by which OpenAI obtained them.”
The statement adds that “the future of mathematical research cannot consist only of understanding results produced by AI labs”, and that “equitable access to powerful research tools and adequate computational resources” is essential. Whether its recommendations were followed, it says, is “ultimately up to the mathematical community to assess”.
What the Dates Behind the Math Results Show
Every manuscript folder in the repository ends with a date, such as September-24-2026. OpenAI has not said what the dates mark. They most likely record when each manuscript version was produced. Counting them gives a picture of how the catalogue of math results was assembled.
Two bursts of writing
Of the 721 manuscripts the catalogue links to (the README counts 722), 563 carry dates from 23 to 27 September, a five-day window. A second burst follows on 4 and 5 October, with 136 manuscripts, the day before release. The remaining 22 are scattered, including seven dated before 23 September, and one folder carries no date.
Bars are scaled to the largest day, 24 September (193): for example 176/193 is 91% and 112/193 is 58%.
Most manuscripts predate the guidelines
AGMAI published its guidelines on 29 September. By the folder dates, 570 of the 721 linked manuscripts were dated before that, and 150 were dated between 30 September and 6 October. That fits OpenAI’s account that the results came from evaluations it was already running, with the release format shaped afterwards. It also explains why the per-result disclosures the group asked for are thin: most of the work was done before the rules existed.
How Mathematicians Reacted to the Math Results
Reaction to the math results split along a familiar line. Some mathematicians were awed by the scale and the formal checks. Others asked how anyone could evaluate hundreds of claims from a model they cannot use. Our earlier coverage of mathematicians who feel bulldozed by OpenAI set out why the field was already tense.
“An instant Fields Medal”
Rutgers mathematician Alex Kontorovich reacted to the zero-free strip within an hour of the release. “Quasi-RH?!?!???! Are you kidding me? If a human did this, it would be an instant Fields Medal, no questions asked,” he wrote on X, in a post viewed more than 900,000 times. The next day he worried he would “never actually finish reading/digesting anything” because each release would bury the last.
“We should ask for receipts”
MIT’s Andrew Sutherland was cautious. “Until and unless they release the model and people can replicate their results, I think you should treat any claims about one-shotting problems with a single agent as unverified,” he told Scientific American. “We should ask for receipts.” Daniel Litt of the University of Toronto took the other side: “If we want to know the answers to these math questions, I see no reason why we should ask the company to keep them secret from us.” Terence Tao has criticised the “insane” pace of results from frontier labs.
Too fast for the lab itself
The most striking admission came from OpenAI. Its spokesperson told Scientific American that many of the newly released math results are not yet understood by the company’s own mathematicians. The company does not plan to slow down, arguing that open problems are an indispensable test of whether its systems are really getting smarter.
One Agent, One Prompt: How the Math Results Were Made
The detail that most separates this release from September’s is how the math results were produced. OpenAI told Scientific American that “almost every one” came from a single prompt given to a single AI agent, though some may have taken multiple attempts.
From a 10,000-agent swarm to a single prompt
The Navier-Stokes solution came from a swarm of about 10,000 agents that, by Scientific American’s account, cost millions of dollars in computing. If a single agent with about three hours of Pro-level compute can now produce results of this kind, as Scientific American put it, “unprecedented mathematical power could soon be accessible to anyone”. Our earlier report on whether OpenAI used mathematicians’ work covers the Navier-Stokes dispute in detail.
Why outsiders cannot test it yet
Claims about how the math results were made cannot be tested without the model and the prompts, which is exactly what the guidelines ask for. Reasoning systems are commonly trained with reinforcement learning on tasks whose answers can be checked automatically, which is one reason proof checkers such as Lean have become central to AI mathematics. OpenAI says it is “working to responsibly release the model that produced these results”, but has given no date.
What These Math Results Mean for Businesses
Few companies need math results about free group factors. But the release is a clear preview of a problem many organisations already face: AI systems can now produce far more output than people can review, and some of it is too specialised for in-house staff to judge.
Generation is cheap; verification is the bottleneck
The math results show where the value moves. Producing 722 manuscripts took a model and compute. Making them useful will take months of expert reading, as Scientific American noted. The same pattern applies to AI-written code, contracts, analyses and reports. Budget for review, not just generation, and treat unreviewed AI output as a draft. This is the core argument in our AI strategy work with clients.
Ask your AI vendors for receipts
The advisory group’s list translates well to procurement. Ask which model produced an output, what instructions it was given, how much it cost, and how often it fails. A vendor that will only quote averages is asking you to trust its selection. Writing those questions into contracts and audits is basic IT governance for AI.
Make outputs machine-checkable where you can
Lean is the reason these math results are credible at all. Most business work has no proof assistant, but much of it can be checked automatically: tests for generated code, schema validation for data, reconciliations for financial figures, and second-model review for drafts. Where you can build a checker, require one. Where you cannot, keep a person accountable for sign-off, as the mathematicians insist.
What Happens Next for OpenAI's Math Results
The release starts a process rather than ending one. Three tracks are worth watching.
Peer review and community hosting
Mathematicians will now read, test and, in some cases, try to break the math results. OpenAI has promised version history so corrections remain visible. Whether it moves the collection to an independent archive, as AGMAI asked, will be an early test of how seriously the guidelines are taken.
Workshops, funding and a public model
OpenAI says it will share details of its workshops and programmes “in the near future”. AGMAI wants those decisions made by existing nonprofit institutions, not the lab. The bigger question is model access. Until mathematicians can run the system themselves, every claim about how the math results were produced rests on OpenAI’s word.
More releases are coming
OpenAI says it will “continue to act on feedback from the community and update our standards for future disclosures”. Given that its own mathematicians have not digested this batch, the field should expect the next one before this one is understood. For a wider view of the debate, see our analysis of why AI is changing how mathematics is done.
Math Results FAQs
How many math results did OpenAI release?
The repository holds the math results in 722 manuscripts across 372 result families. OpenAI posed about 4,000 problems and kept those it judged significant.
Which model produced the math results?
OpenAI has not named it. The README calls it “an unreleased internal OpenAI model”, the same one linked to the Navier-Stokes result.
Are the math results proven?
Some of the math results are formally verified. By our count, 235 families link a Lean scope document and 162 papers appear in the formalization catalogue. The rest await human checking, and OpenAI warns that some unformalised results “could have issues”.
Did OpenAI follow AGMAI’s guidelines?
Partly. It published Lean artefacts, attempt statistics and ten reasoning summaries, but not per-result prompts, compute or a model name, and it hosts the papers on its own GitHub account.
Why do some reports say 377 math results?
The catalogue numbers its families up to 377, but five numbers are unused, so there are 372 families.
Where can I read the papers?
All of the math results are in the openai/math repository on GitHub, under an Apache 2.0 licence, with an overview PDF and a manuscript map.
References
openai/math: mathematical manuscripts and Lean formalizations (GitHub)
Sharing AI progress in mathematics (OpenAI)
Responsible Release of AI-Generated Mathematics (AGMAI)
Advisory Group on Mathematics and Artificial Intelligence (AGMAI)
OpenAI just posted hundreds more results on major math problems (Engadget)
OpenAI: Hundreds more AI-generated solutions for math problems (heise online)
OpenAI releases findings on hundreds of math problems (Anadolu Agency)
Sam Altman: We are entering a new era of discovery now (X)
Alex Kontorovich on the quasi-Riemann result (X)
Comparator: checking Lean proofs against agreed statements (GitHub)
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.