- OpenAI says it published 722 manuscripts on GitHub covering what the company describes as 372 major open math problem families in a single day on October 7, 2026.
- The company says the results came from an unnamed, unreleased internal model, with OpenAI disclosing only average compute figures rather than per-problem details, according to AI Weekly.
- An independent benchmark by Epoch AI and the University of Manchester found that under controlled conditions, GPT-6 Astra scored only 3% on open Erdős problems, with all other tested models scoring 0%.
What Folks Are Whispering: The Big Claim
Well, shoot, y'all — on October 7, 2026, OpenAI went and dropped 722 mathematical manuscripts on GitHub like a man dumping a truckload of firewood in the front yard and then driving off without a word. According to Engadget, Scientific American, and The Decoder, the company says those papers cover what it describes as 372 major result families, each aimed at resolving or substantially advancing a significant open problem in mathematics or theoretical computer science. Fortune put the tally at 'more than 370,' and OpenAI itself has used the figure 377 in its own announcement, so the exact headcount is already muddier than a creek after a thunderstorm — but 722 manuscripts is a number every outlet agrees on.
OpenAI also says, according to AI Weekly and Engadget, that this release follows the company's September claim of resolving the Navier-Stokes Millennium Prize problem. The new batch, the company claims, includes reported partial progress on three additional Millennium Prize problems — among them work that some outlets and mathematicians have characterized as touching the Riemann hypothesis. That last part, as we'll get to in a hot minute, is about as settled as a screen door in a hurricane.
What We Actually Know for Sure
Here's the stuff that's nailed down tighter than a good fence post. OpenAI published 722 manuscripts to a GitHub repository on October 7, 2026 — that count is consistent across every outlet that covered it, from Engadget to Scientific American to The Decoder. The company says those manuscripts are organized into 372 result families. Confirmed by multiple independent sources: nearly every paper, according to what OpenAI told Scientific American, was produced in response to a single prompt handed to a single AI agent. OpenAI says the average result required roughly three hours of ChatGPT Pro compute.
Also confirmed: OpenAI's own Advisory Group on Mathematics and AI (AGMAI), which is linked to the Institute for Advanced Study, recommended the release — but, per AI Weekly and Tech Insider, it also explicitly asked OpenAI to publish the underlying model, the exact prompts used, and per-problem compute data. OpenAI only partially met those requests, releasing average figures rather than the full breakdown the AGMAI sought. The model itself remains unnamed and unreleased to the public, which is a detail that matters about as much as knowing whether the fishing hole is stocked before you drive two hours to get there.
The Unverified Swamp: What Nobody Can Confirm Yet
Now here's where we wade into the tall grass. The 722 manuscripts have not passed independent peer review — not a single one, as of the reporting date. Scientific American quoted MIT's Andrew Sutherland characterizing the company's claims about one-shotting problems with a single agent as unverified. Scientific American itself reported that OpenAI's reputation, in the eyes of the mathematics community, for bold claims paired with limited transparency is driving considerable skepticism — and that's the publication's own characterization, not just idle gossip.
The Riemann hypothesis angle is particularly murky. Latent Space and some enthusiastic mathematicians highlighted what they called 'quasi-Riemann' results with genuine excitement, while Tech Insider and other outlets stated plainly that the repository material does not support strong claims of progress on the Riemann hypothesis, and warned that enthusiasm was running ahead of the evidence. That disagreement between informed observers is, itself, unresolved. Per The Decoder and Tech Insider, roughly half the result families reportedly lack Lean formal verification, meaning those proofs rely entirely on manual review — by a math community that is already, according to The Decoder, overwhelmed by the sheer volume of material like a one-pump gas station hit by a tour bus.
The independent FrontierMath Erdős benchmark, produced by Epoch AI and the University of Manchester and published on arXiv, offers a contrasting data point worth keeping in your back pocket: under controlled conditions testing five AI models on 68 open Erdős problems as of August 2026, only GPT-6 Astra managed a 3% score, and every other model tested scored 0%. That benchmark does not directly test the model OpenAI used for this release — which, again, is unnamed — but it gives a sense of how hard these problems are and why independent verification matters.
The Community Catfight: Cheers, Fury, and a Strong Word from Fields Medalists
Lord have mercy, the reaction from mathematicians has been split wider than a log after a good maul swing. On the enthusiastic side, Fortune reported that Harvard's Levent Alpöge called this release the most significant moment in mathematical history — which is a helluva thing to say about a GitHub commit. University of Toronto professor Dan Litt also expressed genuine enthusiasm, according to Fortune.
On the other side of the pasture, 25 Fields Medal winners — the highest individual honor in mathematics, for those keeping score at home — issued a warning, according to Fortune and The Decoder, that mass-producing claimed mathematical truths at this scale and pace could damage the conditions that allow new mathematical ideas to grow, which is the kind of concern that deserves to be taken seriously regardless of how excited the press release sounds. At least one mathematician, per the available reporting, reportedly characterized OpenAI's release approach in terms that invoked organized crime — a colorful description that we will not repeat verbatim but suggests the collegial warmth was not universal.
Our Analysis: The Barn Door Is Open, the Horse May or May Not Have Left
This is analysis, not reporting, so put on your thinking overalls. The scale of what OpenAI says it has done — if even a fraction of the claimed results survive independent scrutiny — would represent a genuine shift in what automated reasoning can accomplish. That much seems worth taking seriously, even with every asterisk attached.
But the credibility deficit OpenAI has created for itself here is the kind of problem that compounds like bad debt. Releasing 722 papers through a GitHub dump rather than through any coordinated peer-review pipeline, using an unnamed and publicly unavailable model, and providing only average compute figures after an advisory body explicitly asked for more — that is a set of choices that makes verification slower and skepticism easier. The Lean formalizations that do exist, as The Decoder noted, confirm logical consistency within a formal system; they do not independently certify that a result is mathematically novel, significant, or correctly framed. The roughly half of result families that lack even that much are sitting out in the open air with no roof over them.
The independent FrontierMath Erdős benchmark result — 3% for the best tested model under controlled conditions, 0% for the rest — is not a direct contradiction of OpenAI's claims, since different models and methods were involved, but it does illustrate why the math community is not inclined to simply take a company's word for it. Whether this release is a watershed or a very loud weather balloon is a question that will take months of expert review to even begin answering, and the community being asked to do that reviewing is the same one that was just handed 722 manuscripts on a Tuesday.
Who is doing the hollering
These links show where the chatter came from. A link is attribution, not our endorsement or independent confirmation.
- OpenAI publishes solutions to more than 370 outstanding math challengesFortune · top tier
- OpenAI unleashes hundreds more math results upon a field already in shockScientific American · top tier
- OpenAI just posted hundreds more results on major math problemsEngadget · specialist
- OpenAI dumps 372 AI-generated math proofs on GitHub, telling the academic world to keep upThe Decoder · specialist
- OpenAI Says 372 Math Results Came Mostly From One PromptAI Weekly · specialist
- Advisory Group on Mathematics and Artificial IntelligenceOpenAI · primary
- FrontierMath ErdősarXiv (Epoch AI / University of Manchester) · specialist
- OpenAI Just Solved Hundreds of Math Problems. The Fight Over How It Did That Is Just Beginning.Vocal Media · specialist
- [AINews] Quasi-Riemann-Hypothesis: OpenAI publishes 722 math papersLatent Space · social signal
- OpenAI 377 Math Problems: 372 Families, One PromptTech Insider · specialist
Last checked Oct 8, 2026, 5:07 AM EDT. Talk Around Town: The 722 manuscripts have not yet passed independent peer review. OpenAI used an unnamed, unreleased internal model and disclosed only average compute figures; per-problem data was not published. Roughly half the result families reportedly lack formal Lean verification. Some claimed results—including progress on the Riemann hypothesis—are disputed or unconfirmed by named outside mathematicians.