Transformation
A weekly dispatch exploring the latest news in math and AI
Konstantin Kakaes
Latest Articles
Cracks appear in OpenAI’s wall of proofs
Mathematicians have started to take stock of the 377 claimed proofs that OpenAI released on Tuesday, October 6 — although the task will likely require months if not years.
The 722 accompanying papers, which span the entirety of modern math, theoretical computer science, and mathematical physics, appear to be of varying quality. Many seem to contain remarkable resolutions to long-standing challenges. However, the technology company has already withdrawn three of the papers, in the area of algebraic geometry, due to a misplaced minus sign. Mathematicians are poring over the remaining 719, some with excitement and others with frustration.
“We’re now expected to either (i) check all this garbage that OAI is producing or (ii) concede that all of these conjectures have been proven,” one algebraic topologist wrote to me. “I believe in politics this is known as ‘flooding the zone.’ We simply don’t have the resources to correct the record.”
OpenAI is claiming formal verification — that is, a guarantee of correctness using code written in a programming language called Lean — for some 42% of the papers (and climbing). But the formalizations do not appear to be ironclad. A handful of the Lean statements had technical issues. Though these could be easily fixed, the errors point to the fact that the mere existence of a Lean proof is not the final word.
Furthermore, to verify a proof in Lean, one needs to translate the mathematical statement being proved into an equivalent statement in Lean code. “That’s not something that can be established automatically, it needs human experts to read the Lean code,” said Johan Commelin, the director of the Mathlib Initiative, which maintains a digital library of proofs that have already been formalized. Commelin added that though he has not confirmed this firsthand, “some of the formalizations [appear to] prove weaker statements than what’s claimed in the paper. Or only formalize parts of the result claimed on paper.”
Other papers, meanwhile, appear to be correct but jumbled. Chaim Goodman-Strauss of the Museum of Mathematics noted in an email that a paper describing a three-dimensional aperiodic tile has some novel mathematical ideas but is “unnecessarily difficult to read… I figure this is pretty representative across the corpus.” It is, he notes, not remotely close to being in publishable form according to traditional mathematical norms. Put another way, “The stuff I’m familiar with is gross to look at,” one topologist wrote on Twitter.
In some cases, bad writing crosses a line from obfuscating substantive ideas to, in the judgment of experts, simply not holding up. Mihalis Dafermos, a prominent expert in the math of black holes at Princeton University, said he couldn’t point to a specific flaw in several papers in his area of expertise, because the papers were too vaguely written to even approach the standard of proof that mathematicians expect. They’re essentially not even wrong. “In their current form at least, the papers are written in a very synoptic style and look far from constituting in of themselves a proof, let alone a proof that a person could verify,” he said.
Many mathematicians are exasperated. But certainly some are excited. “It’s clear that the OpenAI release includes major advances,” wrote Jeremy Avigad, the director of the Institute for Computer-Aided Reasoning in Mathematics, in an email. “We will learn a lot from the processes the mathematical community will develop to address the problem [of translating back and forth from Lean].” But perhaps the best that can be said of the papers at the moment is, as Daniel Litt of the University of Toronto suggests, that they be seen as “something like ‘research findings’ or ‘raw data’ at this point,” rather than as results in and of themselves. Which is, he notes, a “very unusual situation.”
In other news:
The Association for Human Mathematics, a group of over 800 mathematicians including Fields medalist Peter Scholze, issued a call to boycott OpenAI (which was then reposted by UCLA mathematician Terry Tao on his blog): “Releasing over 700 files at once is not a demonstration of scholarship, but a demonstration of power.”
In response to the influence of AI on math, Thomas Bloom announced major changes to the structure of his website erdosproblems.com, which had been a hotbed of early AI-math results. The changes are intended to increase attention to why things are true rather than treating the Erdős problems as a scoreboard.
Deluge of 377 OpenAI proofs bewilders the math world
The technology firm OpenAI has announced the solutions to 377 open questions across nearly every mathematical subdiscipline.
The results are spread over 722 papers, which OpenAI has uploaded to the Github code repository. They were achieved by a new, unnamed internal model that the company intends to release publicly in the near future, once they deem it safe. All of the papers were written by this model; only one, on a zero-free region of the Riemann zeta function, was edited by a person.
The models seem to be growing more capable. Whereas OpenAI’s Navier-Stokes result in September used millions of dollars of computing power spread over a swarm of 10,000 intelligent agents, the computations behind the new results are, according to the firm, “very different in spirit to the Navier-Stokes setup.” With two exceptions — work on the Riemann zeta function and a proof of a particular case of the Hodge conjecture (both advances in pursuit of solutions to Millennium Prize problems) — each proof required only a nominal amount of computing time using the new model. “The average result used the equivalent of roughly 3 hours of ChatGPT Pro,” OpenAI wrote in a press release. That said, the new model failed on the vast majority of the roughly 4,000 problems it attempted to solve.
According to OpenAI, about half of the results have been formally verified in the programming language Lean. They do not expect any major obstacles to formalizing the remaining half, but rushed to share the results on the recommendation of an advisory group of mathematicians they recently convened. In the formal proofs obtained so far, no errors were discovered to have been made by the model, though in a handful, the formalization process revealed mistakes in the existing mathematical literature.
Mathematicians’ reactions to the announcement have been decidedly mixed. Some are expressing measured excitement. “350 is a lot, but time will pass and mathematicians will study the new content obtained here and see what use can be made of it (I am hopeful that the answer will be: a lot!),” Jordan Ellenberg of the University of Wisconsin wrote in an email. “There’s no hurry and it will all get done.”
According to Jeremy Avigad, who runs the Institute for Computer-Aided Reasoning in Mathematics at Carnegie Mellon University, “It’s against the spirit of mathematics to keep secrets, so I think it is generally a good idea for OpenAI to release their results.”
But other mathematicians are concerned about the implications of publishing such a massive trove of proofs that no person fully understands yet. In a series of Mastodon posts written prior to the release, Terence Tao of the University of California, Los Angeles noted that bulk-published proofs short-circuit the ways in which proofs have traditionally catalyzed the creation of knowledge. My colleague, Quanta’s math editor Jordana Cepelewicz, wrote about this phenomenon at length earlier this week. Tao is far from an AI opponent — as he points out, it’s possible to use AI in a way that preserves “all of the valuable followup activities that traditionally flow out of that solution. But at the current time, the opposite is often occurring… solutions to open problems are now being harvested at large scale in an unsustainable fashion, leaving entire fields of mathematics much less fertile than when such problems were solved in the traditional ‘Math 1.0’ fashion.”
For those inclined to anthropomorphize computer software, OpenAI’s anonymous new model today became one of the most prolific mathematicians in history. Seen in another light, as one anonymous online wag put it, “OpenAI Releases the Final Ten Minutes of 500 Previously Unreleased Films, Ushering in a New Era of Movie Watching.”
In other news:
A group of mathematicians announced the creation of Hexagon, a new repository for research in math and theoretical computer science.
Our deep dive on how mathematicians are responding to AI
The practice of mathematics is being remade by AI at a rate that’s often hard to fathom. Young mathematicians are facing those changes with grief, anger, and new ideas. As one doctoral student put it, the constant generation of new AI proofs that humans don’t yet understand “breaks our system. It just takes it to its absolute limit and destroys it.” As researchers work to change this system and save their field from potential extinction, will future generations be attracted to the new discipline that emerges? I reported on the anxious conversations taking place in math departments across the country in Quanta’s latest column: “Is AI the End of Math As We Know It?” —Jordana Cepelewicz
Major lattice proof percolates through math
On August 30, Hugo Duminil-Copin, a 2022 Fields medalist who works in probability theory, posted an essay to the website Proofs and Prompts. He was concerned with the impact AI was having on his field, and on math more generally. He used the example of a famous problem in his area known as the θ(pc) = 0 conjecture. (Read it as “theta of p–c equals zero.”) “This conjecture stands out among percolation problems because it is inspiring and intriguing, not because solving it would obviously unlock a vast new area of mathematics,” he wrote.
Duminil-Copin had worked to prove the conjecture for years. He hadn’t managed to do so, and yet his failed attempts “generated dozens of ideas that I later repurposed in other contexts, leading to discoveries I would never have imagined making.” His essay is a thoughtful exposition of what might be lost if AI skips ahead to the answer.
A few days later, news spread online that Anthropic had apparently already done so. A Lean-verified proof of the conjecture had quietly been posted on the GitHub code repository on August 28. As of this posting, a formal announcement of the result has not yet been made.
Just what is the conjecture? It concerns what happens when links are placed between vertices laid out in an infinite grid:

Imagine that a link between two adjacent vertices might be present with some probability p.

As p varies from 0 (no links at all) to 1 (every link is present), the graph behaves in a shocking way. If p is less than what mathematicians call the critical probability, or pc, any connected cluster has a finite number of edges. But if p is greater than pc, then an infinite cluster is certain to appear. The chance that an infinite cluster will form jumps from 0 to 1. It’s like how (under ordinary circumstances) water below the freezing point will definitely be ice, and water above it will definitely not be.
But what about precisely at pc? Can you find an infinite cluster at this point, or are such clusters essentially impossible, which mathematicians write as θ(pc) = 0?
In two dimensions, θ(pc) is indeed 0; there’s no infinite cluster. Same goes for dimensions greater than 10. But nobody could pin down what happens in lattices from three to 10 dimensions.
That might no longer be the case.
On August 28, a member of technical staff at Anthropic named Justin Leder (who does not appear to have published any prior work on percolation) posted a Claude-generated proof that θ(pc) = 0 in dimensions three to 10, along with a Claude-written summary of the proof. Neither Leder nor Anthropic have said much since. The proof is just there as a Github repository, backing its way into the world. As Leder’s summary recounts, the result relies on a 2024 paper by Gady Kozma of the Weizmann Institute in Israel and Shahaf Nitzan of Georgia Tech.
Ahmed Bou-Rabee, a mathematician at the University of Pennsylvania, said that “for me, the big breakthrough was in 2024, when Kozma and Nitzan released their paper.” The pair showed that if a certain relatively simple inequality holds, then θ(pc) = 0.
It seems that Anthropic has succeeded in proving the inequality, thus establishing the result. Bou-Rabee says the result is clever and technically brilliant, but that it does not contain substantial new ideas. He has been analyzing the Anthropic proof with the help of LLMs, and in a single day, modified it so that it also holds for an alternative type of percolation that involves the presence or absence of vertices instead of links. He suspects that the same methods can be used to attack other problems as well.
Kozma, for his part, is withholding judgment for now. “I don’t have much to say about their claim,” he wrote in an email. “We are still waiting for Anthropic to publish a human readable version of it, and to supply some information about how it was achieved.” (An Anthropic spokesperson has not responded to repeated inquiries from Quanta.)
Bou-Rabee cautions that there is a big discrepancy between whatever model Anthropic is using internally and their external model. “I’ve got zero mathematically out of the consumer model,” Bou-Rabee said, even though it’s quite good at coding and at Lean. Nevertheless, “AI has allowed me to do things I was never able to do before. [There are] projects I’d been working on for eight years with essentially no progress, but with AI I’m very close to the solution.”
In Other News:
A team of mathematicians from Cornell, Princeton, and UCSD announced a formalization of the Poincaré conjecture.
A group of 24 mathematicians convened at Harvard University’s Center for Mathematical Sciences and Applications and issued a report on how to modify math doctoral programs in the age of AI.
An advisory group on math and artificial intelligence based at the Institute for Advanced Study issued recommendations to AI companies about how to announce mathematical results.
Software developer finds first 3D “einstein” tile
More than any other area of modern math, the study of tiling has been hospitable to amateurs. Now a Greek software developer has written the latest chapter in that story, by using AI to find a simple solution to a long-standing three-dimensional mystery.
In tiling, mathematicians are interested in figuring out what patterns are created when shapes fill space. Why is it that triangles and squares can tile an infinite plane — meaning they cover all of it without gaps or overlaps — but regular pentagons cannot? Are there shapes that can tile a plane but only in a way that never repeats?
In the 1960s, a set of thousands of tiles was discovered that did so; in the 1970s this was whittled down to a set of two tiles. But for decades nobody could find a single tile. Eventually, in 2009, an amateur mathematician in Tasmania named Joan Taylor found such a shape (what mathematicians sometimes playfully call an “einstein,” a pun on the German words for “one piece”). But Taylor’s tile had a bunch of disconnected parts. Then in 2022, David Smith, a retired print technician in northern England, found a simple, connected, hat-like shape that can only tile the plane aperiodically. This settled the question in two dimensions. But what about three?
On September 16, Ioannis Tsiokos, a software developer in Athens, shared a paper claiming to have successfully prompted GPT-6 Astra into coming up with a three-dimensional aperiodic monotile, which he has dubbed Chair44. “The 3D Einstein was found by Astra, not me,” Tsiokos wrote in an email, adding, “I have only understood the intuitive geometric argument of the proof, not the mathematics of it.”
Chair44 is a cube with a chunk taken out of it, together with some simple rules about how copies of itself can fit together. These rules can be conceptualized as arrows that must match, or they can be hard-coded as bumps and dimples in the cube.
As Chaim Goodman-Strauss, one of Smith’s co-authors on the 2D monotile result, explained in a paper about the 3D shape, the chair tiles can only fit together in a way that creates a larger “supertile” that has the same shape. And those supertiles in turn can only fit together to form 2nd-level supertiles, which in turn form 3rd-level supertiles and so forth. This phenomenon can be used to show that the tiles can fill 3D space — but only aperiodically.
The arrow diagram above comes from Goodman-Strauss’s paper, which he posted online on September 21. He’d seen Tsiokos’s original 61-page AI-written article, then spent the time and effort it took to express the mathematical core of the argument in a tight eight pages. (Felix Flicker of Bristol University has written another succinct analysis.)
“Hats off to Tsiokos for this discovery,” Goodman-Strauss wrote, before continuing:
However, the paper itself and the formalization that it claims for support are not in a form that is readily usable to anyone who wants to understand or check it.… We must, absolutely, insist on a higher standard for scientific discourse. The aim must be to communicate, clearly, with people in the community — certainly this is an historical requirement for publication.
Tiling has deep connections with mathematical logic, which has inspired other mathematicians in more recent work. As Craig Kaplan, another co-author of the 2D monotile paper, wrote in an email:
In an age of AI-generated proofs, we lose some of that vigorous intellectual process. We receive the final outcome of our question, but in a way that produces no direct understanding, and that fails to spin off useful new ideas, objects, and methods. Some of that will surely come after the fact, but I think it’s not too precious to claim that in the before times, there was value in the fact that human minds had to wrestle mightily with these problems in order to solve them. The path they took was often more important than the final answer.
In other news:
OpenAI announced an independent advisory board made up of eminent mathematicians and the physicist Edward Witten.
In one of the many guest essays Terence Tao has been hosting on his blog over the past week, Grant Sanderson of the YouTube channel 3Blue1Brown called for a Hilbert-style list of open, important problems in mathematical exposition. Count us in.
Additional reporting by Natalie Wolchover.
Fields medalists denounce AI companies
A year ago this week, an unusual collection of mathematicians, philosophers, historians, and social scientists gathered in Leiden, home to the oldest university in the Netherlands, to ponder how artificial intelligence was changing math.
A document written by a group of people who attended the gathering, called the Leiden Declaration, was released on June 2. It sought to mitigate the potential risks that widespread AI use would pose to research mathematics. But it was measured, and written with an eye towards consensus.
Over the course of the summer, the tenor has shifted. Artificial intelligence has gone from success to success in math, but in a way that has turned the potential concerns of the Leiden Declaration into actual ones. On September 8, OpenAI announced that their model had solved one of math’s most famous problems, spawning a dispute over authorship, publishing norms, and other concerns. On September 11, a group of 25 Fields medalists published a polemic that does not mince words:
The push by AI companies to solve mathematical problems as a benchmark is detrimental to the science of mathematics, and to the mathematical community. The goals of the AI companies and the goals of the mathematical community are severely misaligned.
In just a few days, over 7,400 researchers endorsed the new declaration, more than had signed the Leiden Declaration in over three months.
Terence Tao of the University of California, Los Angeles, in many ways embodies the rapid evolution of the mathematical community as a whole. He has worked in the past with Google DeepMind and organized an OpenAI-funded workshop on “Accelerating Math and Theoretical Physics with AI.” Now he has signed both declarations. “I do not regret my past efforts to raise awareness of the potential of AI in mathematics, to engage with industry, and to promote a vision of sustainable incorporation of these tools,” he recently wrote.
But in the last few months, he noted, “the situation has deteriorated markedly.”
I spoke to dozens of mathematicians last winter and spring while reporting a feature about how AI is changing math. I was struck by how guarded they were. More than one mathematician spoke to me on-the-record about the benefits of AI, only to ask not to be quoted when speaking about its drawbacks.
In a sense, the ground was softened for the Fields medalists by a procession of younger, less established mathematicians. Kirwin Hampshire, a graduate student at the University of Victoria, was arguably the first, describing the Leiden Declaration in a late July essay as a “well-muffled scream.” A few days later, another graduate student, Caltech’s Tasmin Chu, called for a boycott of the use of AI to prove new theorems. A few days after that, Max Weinreich, an assistant professor at CUNY Baruch College, went further, arguing for “total opposition to the use of artificial intelligence in mathematics.” In mid-August, Iris Shi, who recently completed her Ph.D. at the University of Florida, published a lament critiquing how AI had changed how math research is done: “Easy incremental progress gets spit out by the LLM while I occupy myself with washing its Godawful prose out of a report in the foreground.”
A few days prior, a group of five young mathematicians had launched a publication, Proofs and Prompts, devoted to hosting a discussion by and for mathematicians about what AI means for math. On August 30, they published a critique of AI by Hugo Duminil-Copin, a Fields medalist and one of the authors of the more recent declaration.
Even as no-holds-barred criticisms of AI have become more acceptable, many, if not most, mathematicians are coming to rely on frontier AI models as a research tool. A total boycott seems unlikely. Nevertheless, after a tumultuous summer, the mathematical community is reasserting itself.
In other news:
University of Toronto mathematician Daniel Litt published an essay laying out “a positive vision of the future of mathematics, and the human practice of mathematics.”
EpochAI announced a benchmark of 68 Lean-formalized Erdős problems, curated by Thomas Bloom of the University of Manchester, the creator of Erdosproblems.com.
The Navier-Stokes proof has blown up the math world
Earlier this week, mathematicians at OpenAI announced that 10,000 autonomous AI agents under their direction had found a “singularity” in the Navier-Stokes equations — resolving one of the six remaining Millennium Prize problems. It potentially marks a major shift in what mathematics research will look like going forward. But it didn’t come without controversy. We wrote about the details of the result, the long research program that made it possible, and who deserves the credit. Read that story here.
Elusive complex structure found for S6, the six-dimensional sphere
A large language model has found an object that has eluded mathematicians since 1947, despite intense efforts to find it or prove it doesn’t exist.
In the mid-20th century, mathematicians sought to learn more about the nature of high-dimensional spheres by asking if it’s possible to create a “complex structure” on them. The idea is to divide a sphere into smaller pieces, and see if each piece can be mapped to a list of complex numbers (numbers of the form a + bi where i is the square root of –1) in a way that allows for consistent, well-behaved transitions between each piece.
Mathematicians have known since the 1850s that every point on the ordinary two-dimensional sphere can be identified with a single complex number. But can each point on the surface of a four-dimensional sphere be identified with a pair of complex numbers? Can each point on a six-dimensional sphere be identified with a triplet? And so on. (Since complex numbers each consist of two real numbers, it only makes sense to ask the question for even dimensions.)
When the German mathematician Heniz Hopf first posed these questions in 1947, he showed that four- and eight-dimensional spheres don’t have a complex structure. By 1951, researchers had shown the same for ten-dimensional spheres and above. But the six-dimensional sphere, S6, remained enigmatic. Was this phenomenon particular to two dimensions? If not, it might suggest something deeper about the relationship between geometry and complex variables. “In some sense, the six-sphere is the simplest manifold for which we did not know the answer,” said Mohammed Abouzaid, a geometer and topologist at Stanford University.
As numerous mathematicians have tried and failed to settle the question one way or the other, the stakes of the problem have grown.
On August 23, Levent Alpöge, a researcher at Anthropic, announced in a tweet that S6 has a complex structure. (A few days earlier, together with Tristan Buckmaster of New York University, he had found a proof of a singularity in the 3D Euler equations, which they announced earlier this week.) The tweet was accompanied by a nearly 100-page paper written using Claude, Anthropic’s AI model, which Abouzaid describes as “so poorly written that it will take a while to check.” (Alpöge agrees with the criticism, writing in a later tweet, “Yea clearly my bad for fucking up the writeup.”) He said he was too excited to take the time to explain the result clearly.
I spoke with several experts in the field who say that although it will take some time for a consensus to solidify about the correctness of Alpöge’s proof, early signs point to it being true. As Philip Engel of the University of Illinois in Chicago, who has written an expository note streamlining and clarifying the arguments Alpöge first publicized, told me, “For myself, I did all the computations necessary to convince myself that it’s true.”
A version of the statement has also been formally proved by Boris Alexeev at OpenAI, but Abouzaid cautions that although the formalization “is making the community confident that the statement is correct, the effort to digest Alpöge’s proof, and its connection to the formalized proof, have just begun.”
Arguably one reason mathematicians struggled to find the object is that an influential 1996 paper asserted that it couldn’t possibly exist. (In the aftermath of the new proof, that paper is now thought to be wrong.) Engel expects the new construction will have some interesting implications, leading to a “little bit of a zoo” as mathematicians generalize it. “Maybe we all would have died not knowing the answer if the AI hadn’t discovered it,” he said.