Sakupljač feed-ova
“Deep theorems were scarce and difficult and so became an effective mechanism to identify deep thought. AI has broken this system.”
[This is a guest post by Bryna Kra. This blog post was initially written in a different file format and converted using AI. — T.]
For generations, mathematicians have treated the production of theorems as a clear measure of success. The stronger the theorem, the deeper the proof, the more surprising the connections, the greater the achievement. Positions and prizes are based on these theorems and the mathematicians making the breakthroughs set the directions for future research.
But theorem production was only part of what we cared about. It was a proxy for something harder to measure: understanding. A major breakthrough meant that years were invested in learning a subject and uncovering hidden aspects, accompanied by work to make the answer apparent to the community. Deep theorems were scarce and difficult and so became an effective mechanism to identify deep thought.
AI has broken this system.
Artificial intelligence has lowered the cost of producing sophisticated proofs. As the models improve, the list of deep conjectures that fall will grow. Someone with little knowledge in a field can now generate a manuscript that reads like a polished article, cites the literature, and combines techniques that would have taken years to master. The machines are producing solutions faster than the mathematical community can read them, much less digest them.
This week the changes arrived in my field. The Nivat conjecture is a beautiful problem at the intersection of combinatorics and dynamical systems. Imagine an infinite grid of square tiles, with each tile colored either red or blue. Look at every rectangular window of a fixed size, say tiles wide and tiles high, and count how many different color patterns show up in that window. The conjecture says that if there are at most such patterns, then the entire coloring repeats under some fixed, nonzero shift of the grid. In other words, some precise limit on local behavior gives rise to simple global behavior, periodicity.
Though this statement can be explained without formulas or higher mathematics, the conjecture resisted our efforts for nearly three decades.
In work that we first circulated in 2012, Van Cyr and I proved a partial result, showing that periodicity holds when the number of patterns is at most half the area of the window. We translated the combinatorial question into one about the geometry of dynamical systems, studying ways in which information in the system does (or does not) propagate. Progress on challenging problems does not happen in a vacuum, and our work began by studying and building on earlier results in the literature [1, 2, 3]. Soon after, Jarkko Kari and Michal Szabados developed a different approach, viewing the combinatorial question as an algebraic object, using certain types of polynomials to find the infinite patterns.
To those of us who worked in the area, this was a tantalizing situation. Two different methods to approach the same problem seemed to shout to us: find the conceptual bridge between them. We tried, without success, and moved on to other problems. Colleagues continued pushing the boundaries of what we knew, combining the varied techniques in increasingly powerful ways [1, 2, 3].
But synthesizing different approaches is one of the ways in which frontier models excel. A machine does not need years to absorb distinct areas of mathematics before it can find the connections. It can translate notation, compare partial results, and try thousands of combinations without being discouraged. The Nivat conjecture has been proposed for inclusion in Google DeepMind’s Formal Conjectures project, turning the problem into a target for automated reasoning and making more people aware of the question.
The manuscripts have started to arrive. This week I received several purported proofs of the full conjecture from researchers making their first foray into the subject. Some acknowledged using AI, but only for polishing the English or checking the proof. When I asked if the authors would meet over Zoom to explain their arguments, none accepted the invite.
Silence does not mean that the authors used AI and a proposed proof is not a theorem. But merely disclosing AI use and checking correctness of the proof, both of which are essential, miss the point. Suppose these results are correct. That is a mathematical contribution. Yet it opens more questions. What did we learn? What did the authors contribute? What impact will this result have?
A proof is more than a certificate that something is true. Instead, it is a story, a picture, an insight, an explanation. A proof highlights novel ideas and opens new directions for what we should ask next. It becomes part of the toolkit of the community. A deep theorem changes how we think, not because of its statement, but because of what it teaches us. As Bill Thurston wrote on MathOverflow in a 2010 response to a question about what mathematicians do: “The product of mathematics is clarity and understanding. Not theorems, by themselves.”
These new texts certainly contain knowledge, just as a data file of bits of 0’s and 1’s contains information. But human authorship is more than the string of words that prove something. It is intellectual responsibility, understanding of a new concept, the ability to respond to questions, the sorting of new ideas from existing scaffolding. And perhaps most importantly, it is incorporating the result into the corpus of our understanding.
I welcome AI-generated proofs and believe that human-only proofs will become rare. This is a profound change of perspective, but mathematicians have always used external tools and machines. AI is already a powerful mathematical instrument and we, as a community, need to incorporate it responsibly into our practice of mathematics. The question is not if machines should participate in discovery. They already do. The question our community needs to address is what we value in our mathematics now that certain kinds of discovery are no longer a scarce resource.
The tradition of journals publishing novel results, hiring committees rewarding those publications, and prize committees recognizing the person who completed the result do not suffice. Peer review is a slow process that relies on the unpaid expertise of a small group of overworked researchers. Any single submission can be handled, but the current volume is drowning the reviewers. The careers of young researchers depend on producing theorems, with incentive to produce as much as possible as quickly as possible. The time for deep reflection, necessary for deep understanding, does not happen when someone types the question into a model and quickly produces a manuscript.
All of these issues existed before AI entered our field. Unfortunately, we no longer have the luxury of addressing them in a leisurely manner.
Mathematics is at a watershed moment. Our incentive structures are misaligned with the prolific output of highly accessible AI models. Careful verification and stellar exposition will not happen when there is no reward for those tasks. Producing a paper is no longer enough: authors must be able to explain the proof’s mechanism and how they arrived at this point. Journals need to distinguish and credit the roles of discovery, proof, formalization and explanation. Exposition must become an important component of any major intellectual achievement. We advance mathematics not by what we write, but by what we learn, check, apply, and teach others.
Finding the right concept, crafting the right definition, and formulating new questions have always been part of our intellectual work. The creation of collective understanding, the ability to communicate ideas in a lecture, and the sharing of ideas informally that spark research have always been admired. Such achievements are harder to quantify than theorem production, and so have been treated as by-products. But these are the key aspects shaping the future of our field and their value is at least as important as knocking down the next conjecture. Our incentive system needs to be realigned to reflect what we truly value, and not just what we can easily measure.
Mathematics is the canary in a coal mine for all parts of intellectual life. Its claims can be carefully checked and so we are witnessing the disruption in real time. But when AI disrupts the tried and tested standards in a field like mathematics, what will it do to fields where evidence and interpretation are more contested? Science, law, public policy, and art will all have to find ways to assess what is valuable amidst the abundant output.
The arrival of machine generated mathematics does not replace the role of human expertise. Instead it reveals why we valued that expertise. The goal of our discipline is not simply to resolve conjectures, but to enlarge our capacity to reason and open our minds to new ways of understanding. It will take our entire community to create new standards, discover new ways of creating mathematics, work with AI models to open new horizons of knowledge, and develop the tools to approach them. It’s truly an exciting time to be a mathematician.
Happy, those able to know the causes of things
[This is a guest post by Nestor Guillen, crossposted from his blog. This blog post was initially written in a different file format and converted using AI. — T.]
Keywords: LLMs, cultural technologies, skateboarding, creative communities, AI industry.
A note on the timing of this essay. I started writing this in the middle of the ICM, and it has taken me this long to finish writing, see also my shorter post from August. In any case the bulk of the essay was done in August and shared with a few close friends for feedback. Earlier this week, Tristan Buckmaster announced a breakthrough finding, done in collaboration with Levent Alpöge, of a finite-time blow up for 3D incompressible Euler with forcing. He also made serious allegations of misconduct by OpenAI (see announcement). The next day, OpenAI announced a finite-time blow up result for 3D Navier-Stokes with forcing, obtained from their internal LLM. There has been lots of press coverage, I recommend Kenneth Chang’s reporting. I have decided not to change the opening of the essay and instead make a note of the news here. Tristan gave considerable credit to the ideas of Córdoba and Martínez-Zoroa in the blow up construction; I found this important detail to be quite validating of the perspective on LLMs I advocate in this essay, and found that to be further reason to leave the content of the essay largely unchanged.
What would be the significance if one day we woke up to the news that output from a Large Language Model contains a definitive answer to the question of blow up for the incompressible Navier-Stokes equations, and to the question of where the zeros of the Riemann zeta function lie? (Or, to mention problems dear to my heart: a proof of a nonlocal version of the ABP maximum principle or a resolution of the Mahler or Komlós conjectures?)
Many of my fellow mathematicians consider this a kind of nightmare situation, and I think this is a very valid position. Here I would like to make a different case, however. Personally, I have come to believe that the news of a Large Model output containing an interesting, insightful, novel mathematical idea can and should be received as a source of communal pride for mathematics and for mathematicians. I also believe there can be a world where LLMs are not harmful but rather supportive of mathematicians as we follow our prime directive: understanding things.
My intention in this essay is to elaborate on what I mean by this. I believe this perspective can assuage fears many of us feel amid the daily hype and doom narratives around AI and the alleged impending obsolescence of mathematicians. I will not be talking about the less existential / philosophical and more practical / material concerns here, such as challenges to the funding of the mathematics occupation, how journals will work, changes to the training mathematicians, and so. Those issues are also important, but I think clarifying the perspective I push here will make the discussion of the more practical matters easier.
Bill Thurston touched on the same idea in his famous 1994 essay, “On proof and progress in mathematics”, which is more relevant than ever. Thurston observes that mathematicians (as a group) prove theorems and solve problems, but that is not all they do: they also absorb these solutions, and find new uses for them once they are proven. This “digestion” [1] in turn allows for the next wave of breakthroughs, and thus advances our understanding of mathematics.
This is not secondary to what mathematicians do, but our very purpose. The value of mathematics to humanity lies not in determining whether a statement or conjecture is true or false, but rather in the understanding gained as we make those determinations (even if the facts themselves are also of great value). To put it briefly, mathematicians are people who try to understand the causes of things.
Felix, qui potuit rerum cognoscere causas
LLMs, the technology AI, the product from large companies
The conversation about LLMs and generative AI within the mathematical community at large must deal with an obstacle before we can get to the substance of the matter. That obstacle is the current and unfortunate identification of generative AI, the technology, with generative AI, the specific product and vision from the large AI companies. It is imperative for us to make that distinction and separate the two. The private AI lab vision is what most people have in mind when discussing AI, and this is no accident, for this vision has been aggressively pushed into every part of our daily lives.
It might seem impossible at the moment, but we can imagine an alternative world with many parallel approaches to the use of LLMs developed by the grassroots, reflecting the variety of circumstances across different communities, industries, and institutions. A seed for this world is the emergence of open-weight models. This is a matter for a different essay.
So in all that follows, I will crucially rely on the above distinction, so that when I say LLMs or “Large Model” I am talking about the technology itself, and very much not any of the major AI companies or their products.
With all this throat clearing behind us, let me get to the matter at hand. Why should the output of an LLM containing new and exciting mathematical ideas be a source of pride for mathematicians, and not of anxiety? My answer owes a lot to this very interesting article by Henry Farrell, Alison Gopnik, Cosma Shalizi, and James Evans. I quote here one of their central assertions:
…Large models should not be viewed primarily as intelligent agents but as a new kind of cultural and social technology, allowing humans to take advantage of information other humans have accumulated.
I will do my best to apply the ideas in their article to mathematics and the work of mathematicians. If, rather than continue reading this, you stop and go read their article and Thurston’s article instead, I would consider this essay a success.
Cultural and social technologies
Humans have always lived in a complex world and their survival has hinged on their ability to process information in amounts beyond what any individual can handle — and it is likely this has been so nearly as long as humans have existed. Herbert Simon’s concept of “bounded rationality” captures this well: in real life, unlike in the simplest control theory models, we have limitations on the time available to decide and on the accuracy of our information. This means that an agent making decisions must give up optimality and settle for “good enough.”
Simon saw that one implication of bounded rationality is that humans have to (and do) develop systems to process information, as a group. We see that humans have built states, stories, bureaucracies, markets, libraries, and more. These are cultural and social technologies. They are means by which humans (not individually but as a group) navigate amounts of information so vast that they cannot be handled by any individual on their own. An important feature of these technologies is that they are to some extent forgetful: you ignore or downplay smaller details in order to make a bigger picture easier to comprehend.
Markets, governments, and bureaucracies mean each human can focus on a few things and still get the aggregate benefits. Thanks to markets, governments, and bureaucracies, I can rely on being able to buy fresh fruits and vegetables around the block from my Manhattan apartment. Thanks to markets, governments, and the field of modern medicine, we now have high confidence that a newborn child will live to be an adult — in stark contrast to what was the norm for most of human history. Cultural technologies are powerful and awe-inspiring things. It is understandable that we tend to anthropomorphize them so often.
This week I started teaching my graduate-level optimal transport course. I wrapped up the first class by showing one neat application of optimal transport: a “one-page proof” of the isoperimetric inequality in . Of course, the proof being just one page depends on what we are assuming as given.
This is standard, not only when teaching advanced courses but also very much in day-to-day research, and it is made possible by the power of cultural technologies. In the proof I don’t have to explain from scratch the notion of Lebesgue measure or sets of finite perimeter. When in the middle of a proof I say “I will do the details only when the domain has a smooth boundary, and I will assume the Lebesgue measure is without loss of generality thanks to scaling,” I am relying on a shared trick that I can trust a first-year PhD student to be familiar with.
All of this is to my benefit and the students’. We can focus on “the one page” proof with the interesting new idea: how Brenier’s theorem provides a map through which the isoperimetric inequality is essentially reduced to the arithmetic-geometric mean inequality.
So, here is basically what I showed on the blackboard: assuming has Lebesgue measure , we want to prove that
where is the -dimensional Euclidean ball with the same volume as , so . Then, I tell my class how Brenier’s theorem says there is always a map from to , that this map is volume-preserving, and that for some convex function . Then, since and all the eigenvalues of are non-negative, the arithmetic-geometric mean inequality says that everywhere in , and so
Then, the divergence theorem and the fact that (where = volume of the unit ball) guarantee that
For a ball of volume , scaling shows that , so we conclude that , as we wanted.
In interacting with this proof, we are engaging in a conversation through space and time with many mathematicians, but especially with Knothe (1957) and Gromov (1986), followed later by separate works of Brenier, McCann, and Trudinger — and even with my younger self in a set of lecture notes with McCann, where we go over the above proof (see concretely Section 1.6). The conversation is mediated by a common mathematical culture, which encompasses the common definitions and methods I can count on the graduate students to know. I don’t have to prove the divergence theorem or the change of variables formula to justify these computations, and I don’t have to worry about the smoothness of , since I can work by approximation.
Without having to go into all that detail, I presented on one blackboard the idea that the Brenier map takes you from the arithmetic-geometric mean inequality to the isoperimetric inequality, and I did this in the last 15 minutes of the class. These are cultural technologies in action! Which cultural technology in particular? In this case it is the technology of a scientific community with standards in what is taught, the mathematical literature, and so on.
Large Language Models as cultural technologies
Farrell et al.’s main point is that LLMs are cultural technologies, just like markets, bureaucracies, and scientific disciplines. Accordingly, the framework of cultural technologies is a very useful one as we seek to understand and navigate the impact of Large Models on our disciplines.
LLMs indeed are a lot like the example from my class, and in a way they are “wrappers” around an older cultural technology: the scientific literature. When you ask an LLM a question about the isoperimetric inequality, the ABP maximum principle, or a PDE, you are effectively interacting with every person who has thought about and worked on those things — provided the data the LLM was trained on contained those contributions in one way or another. Surely the lecture notes I reference above are in the training data of most LLMs in use today, just as a vast portion of the digitized mathematical literature (across all disciplines) was part of the training data. LLMs are lossy, so while a model might not be able to reproduce a given paper exactly as in the training data, it can reproduce large parts of it.
What sets LLMs apart from older cultural and social technologies is the fact that we have digitized all this data, and can now explore it, reproduce it, and recombine it with ease by sending text commands (prompts) from a computer terminal, or even a phone. LLMs are, in essence, top-tier information-processing tools. They are also an excellent tool for generating conversational level-text, and for coherently combining patterns of information even if they come from very different sources.
By a product-design choice at a few private labs, the output of most LLMs is conversational in tone and made easy to anthropomorphize. These properties have amplified an ELIZA-like effect that reaches even people with advanced training in mathematics and computer science. In a way, anthropomorphizing this technology — which, as some have argued, did not have to be the case — does everyone a disservice: the conversational setup of the LLM gives the appearance of a single “thinking” entity, and buries the work by vast groups of people and which are being recombined. This appearance is a feature not of LLMs, but of the LLMs as the AI companies want them. We are not dealing with a very smart, genius-level intelligence; we are interacting with something greater (and yet somehow more human): the aggregate creations of millennia of human ingenuity and creativity, digitized and efficiently compressed — all mediated by state-of-the-art statistical and generative algorithms.
In particular, any mathematical proof found in an LLM output is the collective outcome of lifetimes of work by mathematicians. It must be welcomed as the fruit of our shared mathematical heritage, belonging to all of humanity, carried through millennia by people who at various times went by names other than mathematician (natural philosophers, astronomers, geometers). This applies very much to LLM output containing what we might consider novel proofs or ideas, for in the end, it is all an outgrowth of that gigantic body of knowledge. Now there lies a scaling law that has been going for centuries. The recently arrived software is just a kind of harness or wrapper for it.
Now, let’s reconsider the “nightmare scenario” I opened this essay with. One day we wake up to the news that the output of a Large Model contains a resolution to the Riemann Hypothesis. We are now equipped with the ideas of Farrell et al., and especially the idea of cultural technologies. Reading the news under this framework, we are learning a new mathematical fact, but we are also learning that this fact is the fruit of the collective efforts of researchers in the field up to this moment. The problem was cracked by the ideas of one or two (or maybe many) researchers to the point that a resolution could be reached by the sampling and recombining of ideas an LLM is good at.
This leaves us with a delicate question: what ideas are reachable from the totality of ideas in the mathematical literature at a given time? I am going to borrow from geometry and loosely use the metaphor of “convex combination” of ideas to discuss this.
Convex hulls of ideas
A convex body containing another, from Geometry and the Imagination by Hilbert and Cohn-Vossen
The capabilities of LLMs have progressed quite a bit in the last couple of years, and dramatically so in mathematics this year alone. The rapid improvement makes this discussion difficult. The question “what can LLMs do?” is a rather moving target, at least for the time being. This being said, let’s try to think of a related question that might not be as LLM-dependent, and has a chance of being useful, even if it is still vague.
Ideas sometimes are very clearly compositions of other, more fundamental ones. Sometimes a new idea really stands apart from what came before. What would the finding of a solution the Riemann Hypothesis, or of the 3D Navier-Stokes smoothness problem, within the output of an LLM, tell us? One toy model (which might not reflect empirical reality) is that the result might belong to some sort of “closure” or “convex hull” of the available literature. In broader terms, one could pose the following question: given a collection of ideas, what ideas could be said to stand squarely beyond them, and which are lying closer to their “convex hull”?
I will not give a more concrete definition for this, because I am not able to. However, I can discuss by way of example what this definition ought to cover. I also do not use “convex hull” to mean that the point is derivative or that the finding of a point of something inside a convex hull to be easy — high dimensional convex sets are deeply complex objects!
At the very least, we can agree that the proof of the infinitude of primes with coprime is certainly not in the “convex hull” of number theory that preceded Euler and Dirichlet: using an infinite series to quantify the “amount of primes” is precisely the sort of idea that stood squarely outside of what was considered number theory until it was proposed. However, the question of the infinitude of primes of the form ( coprime) could be posed within the number theory that existed before Euler. In fact, this was known already for special choices of and , say for primes of the form . Likewise, number fields, ideals, and their unique factorization are all ideas lying well beyond all the number theory that preceded them. However, once we have these ideas together, it is slightly less outside the convex hull (or maybe it is even inside the convex hull) to develop a proof of the fact that every ideal class contains infinitely many prime ideals — a combination of Euler’s and Dirichlet’s ideas with the idea of, say, Gaussian integers.
From what I have seen (so far) in research areas that I am closest to, no LLM has produced anything remotely like Euler’s introduction of the methods of analysis in the study of primes. One possibility is that LLMs are not even being used so far for such things. In fact, one could put it like this: Euler was not simply answering a well-known question, but rather saying “here is actually a whole new question that we did not know we could ask, and is actually a very important one that connects to older problems!” The framework of LLMs as cultural technologies, to the extent that it reflects existing LLMs, suggests that this is a type of contribution we cannot expect to find in their output.
According to Farrell et al., cultural and social technologies are preceded by individual capabilities, which are then aggregated and transmitted by the technology. In their essay, they note that “without innovation, there would be no point to imitation.” Suppose Euler had access to an LLM trained on all of the mathematics known up to the moment he started working in number theory. Without an individual first introducing the analytical perspective on prime numbers, it would be very hard for our hypothetical LLM to innovate analytic number theory into existence, and then use this innovation to solve the “conjecture” [2] on the infinitude of primes that are congruent to modulo , for any coprime. All the same, the framework is compatible with the happy possibility that a person working with an LLM might have a eureka moment where they, the person, arrive at this innovation.
Here is another example, of a slightly different character. One of the most important theorems in PDE is the Krylov-Safonov theorem for non-divergence form elliptic and parabolic equations. For those not in PDE, these are versions of the Laplace equation and the heat equation, except the “coefficients” in front of the partial derivatives are not constant but rather change from point to point.
Such equations, linear and nonlinear, appear in the study of stochastic processes as well as in optimal control and geometry. The Krylov-Safonov theorem produces a Hölder regularity estimate for solutions of the PDE that only depends on the uniform ellipticity — what is notable is that this estimate works even if the coefficients are potentially very discontinuous. Consider, for concreteness, a linear equation such as
When the coefficients are smooth, the regularity of solutions to this equation can be understood from the theory for the Laplace equation. Heuristically, the coefficients being smooth means that as you zoom in near a point, the coefficients are closer and closer to being constant, which puts you just an affine change of variables away from Laplace’s equation. Therefore the good regularization properties of Laplace’s equation carry upward to the larger scales and so one proves is smooth. This estimate, however, depends very much on the modulus of continuity of the functions .
Now, I cannot go on at length about this here, but obtaining estimates for that do not depend on the continuity of the coefficients is extremely important. In fact, the importance of this question cannot be overstated — such estimates have profound implications in the theory of stochastic processes, in geometric analysis and differential geometry (say, for the many PDE related to curvature), in stochastic homogenization, and more. For this reason the Krylov-Safonov theorem is one of the most important PDE results we have.
The Krylov-Safonov theorem, however, depends on a mathematical fact that comes from a very different area: convex analysis. There is an estimate by Aleksandrov that says the following: for denoting the unit ball, given any function convex in and with on , we have
where denotes the subdifferential of the convex function. This estimate holds for general convex bodies, not just , and in this generality it is equivalent to the reverse Blaschke-Santaló inequality from convex geometry. The estimate is the key component in what is known as the Aleksandrov-Bakelman-Pucci (ABP) estimate, which in turn is essential for the Krylov-Safonov theorem.
To state it in simple terms: without this fact from convex geometry the Krylov-Safonov theorem cannot be proved, not in any way we know today. This is a significant statement, as many important theorems in analysis have multiple, independent lines of attack that do not necessarily use the same tools. This would be like asking how to prove results in geometric analysis and probability without knowledge of the Poincaré inequality. Without the advances in convex geometry done many decades earlier, the theory of fully nonlinear elliptic PDE, which seems at first sight so remote from convex geometry, could not have advanced the way it did onward from Krylov and Safonov.
Imagine a parallel universe where we came to LLM technology before we knew about the Krylov-Safonov theorem, so that the question of regularity for non-divergence elliptic equations without dependence on coefficients is considered an important open problem. The cultural technologies framework suggests that in this universe an LLM (at least at the current level of power) would not be able to settle this open problem if the mathematical literature in this parallel universe does not have a well-developed field of convex geometry/analysis. By the same token, if in this parallel universe there is a well-developed convex geometry literature, including the Aleksandrov estimate, then it is plausible an LLM could bridge the two things and arrive at the ABP estimate and Krylov-Safonov’s theorem.
Thinking in terms of “convex hull of ideas,” one sees that whether an LLM can solve a given mathematical problem is in large part a function of the status of the mathematical literature at any given time, and is not just a function of how large the LLM is. With this perspective, the event of an LLM output containing the solution of a hard problem ought to be welcome as a happy occurrence: an LLM helped us dig out a gem buried somewhere in our mathematical heritage, or in the convex hull, if you will.
In fact, this has already happened: less than a year ago, we heard news of “Erdős problems” solved by an LLM, which then turned out to be instances of the LLM finding solutions in old (and in some cases, forgotten!) papers. At the time, this realization was seen as a defeat for LLMs. Now, that situation illustrates the perspective I am advocating here, in miniature.
Creative communities during turbulent times
The framework of LLMs as cultural technologies, together with the “convex hull” of ideas it inspired, has addressed my own philosophical concerns with LLMs. It has reinforced my optimism that the mathematics I value will endure. However, this does not change the fact that the mathematical community is facing serious and urgent challenges. The AI companies’ aggressive and narrow-minded push of their vision is hurting the field, exacerbating old fault lines and dysfunctions of our profession, and creating new ones.
This is not the first time a creative community has faced a turbulent period or a big change brought by new technologies or economic realities. There are lessons we can learn from past instances, in mathematics and in other disciplines. One can think of painters facing the advent of photography, musicians facing the arrival of the synthethizer, or mathematicians themselves — think for instance of the creation of the arXiv in the 1990s. However unusual it might seem, chapters from the history of professional skateboarding present uncanny parallels to the circumstances in mathematics, and this is a history with many teachings and warnings for mathematicians.
Rodney Mullen performing “the Impossible” trick in an episode of the Physics Girl (Dianna Cowern)
By the 1980s, skateboarding had been around for several decades. It had professional contests, sponsorship deals, and a specialized press. Like mathematics, skateboarding had an ecosystem that allowed practitioners to develop their craft and advance it, and even make some money along the way. It also had subfields, like freestyle, vert, and street skating. In 1984, Kevin Harris, who came from the subfield of freestyling, observed street skaters with only “six months of experience” were already earning more than Rodney Mullen, who at the time had tens of thousands of hours of experience and a nearly perfect record in professional competitions. Harris has recalled his reaction to this observation as: “Okay, this is the death of freestyle and it’s going to hurt vert skating hugely.”
The recession around 1990 dealt the finishing blow to the professional skateboarding ecosystem as it existed at the time. Within a few years, the discipline lost many of the venues and big public events it was built around, even largely losing its version of “funding” — sponsorships. This crisis drove many pros into early retirement. One could argue that with championships and sponsorships gone, the metric for measuring success was gone, and with it all the incentives for participation in competitive skateboarding. However, this was not the only metric, nor even the main one, for skateboarding continued and evolved in the coming years.
Without public contests with large audiences and related sponsorships, skateboarders had to find some other means to communicate and show their work. The economic and social problems in this case were solved with a technology that already existed. Skating videos became the new medium by which skateboarders recorded and communicated their work. Before, if a skater spent countless days perfecting a delicate trick, they could showcase it to at most a few hundred, maybe a thousand people during a championship. Now, a videotape showcasing the work of several skaters could be seen by tens of thousands of people. Before, a skater who did not win several championships was likely to go unappreciated. Now, a skater needed only one good “part” –a single skater’s segment within a video– to be recognized and contribute new tricks to the skating community. The most successful “parts” would introduce a new trick or a new twist on an old one that would get replicated and adopted by others.
I find it notable that the resolution to the crisis in skateboarding made it a bit less a competitive sport and a bit more a collaborative, creative field, with videotapes playing the role of journals [3], and a “part” playing the role of the articles in the journals. This turbulent transition turned out to be for the best for skateboarding as a discipline, while it also turned out for the worse for many individual skateboarders. This ecosystem remains largely unchanged today (with internet videos replacing videotapes). The motivation for a skater to participate in the community, aside from their love for skating itself, was that their contributions could be showcased in video, and their ideas could be shared with others who could try to repeat them or recombine them in new ways.
Mullen said that an invention only becomes an innovation once the community adopts it. This provides a measure of the value of a work. In mathematics we often talk about “the impact” of a paper or idea, and this is completely analogous to a trick being impactful: has the community taken this idea, and used it, and modified it further? Has this work made it easier for newcomers to the field to understand the state of the art?
Reclaiming our mathematical heritage, and the future of mathematics
One lesson I take from cultural technologies is that the most precious component in a Large Language Model, the part most responsible for whatever value lies in its output, is the human-generated work, the cultural and technical heritage encoded in it. Throughout the summer the largest AI companies have been treating the solution of high-profile mathematical problems as a form of PR for their products. We as mathematicians need to lean into our curiosity and love of math and welcome this new knowledge, without letting its provenance ruin our enjoyment of it. We must counter the damaging misconception that “AI has beaten mathematicians at solving problem X,” which is being promoted by the corporate AI labs. Instead, we can celebrate how our understanding of mathematics, together with our creativity, lies behind the gems of mathematics LLMs are helping us uncover.
Earlier in this essay I touched on the importance of separating LLMs (the technology) from AI (the product and vision of the AI companies). Well, we need to also make this a reality for our mathematical infrastructure. It needs to be a top priority of the community to promote and work towards a world with open LLMs — think of the LLM equivalent of free software. It would be quite bad if our profession became dependent on the whims of the AI industry, whose leaders inhabit a moral universe so distinct from ours in science and mathematics that nothing fruitful or sincere can emerge from associating with them [4]. We must remember that the combustion engine was never the heritage or exclusive creation of car makers in Detroit; and railroads and locomotives were not the exclusive heritage of the robber barons of the Gilded Age; and so it is also true that LLMs do not belong in any fundamental way to large industry labs in the Bay Area. The convex hull perspective in this essay suggests it is worthwhile to look into lighter, mathematics-specific LLMs with well curated data, and asses the scale of the technical (and financial) obstacles for such a thing. It would demostrate that the big labs’ “moat” is not as deep as they’d like it to be.
We also ought to approach the use of LLMs with an open attitude, experimenting and exploring until we find those LLM practices that best serve our values: advancing human understanding of mathematics, as Thurston would put it. This also means working as a community to minimize the disruption to mathematical careers, so that as few people as possible share the fate of those skateboarders who had no choice but to retire early. This must also involve policy advocacy: to dramatically increase public funding for science, and policy advocacy for serious public oversight of the large AI labs. We must also experiment with the creation of new social structures and professional arrangements within mathematics, so that we can allow as many individuals as possible to remain part of the mathematical community. If we are successful in this, the mathematical community will come out stronger from this turbulent time, serving as a model for other communities and industries working to transition to a world of ubiquitous LLMs.
To return to where we started, an LLM’s output “solving” the nonlocal ABP or the Mahler conjecture would be, circumstances aside and in a strict sense, a triumph for the discipline. That does not mean that we cannot have mixed feelings about it. It is not pleasant to be scooped by a living, thinking person, and it is no more pleasant to be scooped by the collective knowledge of the field, processed and recombined via an LLM. Rather than being shocked, we ought to be in awe, not of the technology, but of the immensity and power of the mathematical knowledge we have built to this day. Large Models have made this enormity apparent and impossible to miss. Mathematicians must claim such developments as the rightful fruit of our discipline, and accordingly take pride in them, understand them, adopt them, and continue advancing mathematics. Fortunate are we, able to know the causes of things.
Notes
[1] “Digestion” is a term Terry Tao proposed in his ICM lecture on mathematics in the age of AI, part of a metaphor of LLM generated proofs as bringing us to an era of proof abundance, as in mass produced food.
[2] I say “conjecture” only in the context of this hypothetical scenario of Euler and his 18th century LLM.
[3] In mathematics, meanwhile, we are actually approaching a situation where we will have to switch towards whatever comes after journals, or at least journals as they are presently understood.
[4] Here I am very purposefully borrowing a phrase I have long been fond of. It was written by Bertrand Russell in a response to Oswald Mosley.
Wimbledon, the U.S. Open, and the future of mathematics
[This is a guest post by Steven Strogatz. This blog post was initially written in a different file format and converted using AI. — T.]
WIRED reported today that I began to cry while talking about AI and mathematics. But the article didn’t explain what moved me.
The truth is, I’m not entirely sure myself.
(A disclosure upfront: ChatGPT helped me write this post. I’m tweaking it a little to sound more like myself.)
First, thanks to WIRED reporter Isabella Ward. She had the difficult job of distilling a long, wide-ranging conversation into a timely piece about the amazing events of the past week. And she had to put it together in one day! A lot inevitably got left out.
OK, so why did I get choked up? It wasn’t about AI taking my job. I’m 67 and nearing retirement. I also love playing around with AI and find it exhilarating. I use it in research, brainstorming, and learning—and, as you can see, to put this post together.
So what was it then? I think it was when I tried to convey the magnificent 4,000-year story of our subject to the reporter. Mathematics, I said, is more than its results. It’s a human way of seeing and understanding, a conversation in which one generation explains to the next not only what is true, but why. The pleasure of that understanding, and the new discoveries it makes possible, are part of what mathematics is for.
It’s an honor to be part of that long conversation—all the beautiful and important work, and the benefits that have accrued to humanity.
Then, suddenly, the tidal wave of AI hits math in 2026.
Something about the grandeur of the tradition and the suddenness of the change may have conspired to get me worked up.
But I’m not sure it was that. It might have been when I started riffing about Tristan Buckmaster. I don’t know him, but from what I’ve read over this unsettling past week, he reminds me of many mathematicians I’ve known: pure, clear, and scrupulous, committed to truth, and perhaps unprepared for the corporate world.
During the WIRED interview, I mentioned Wimbledon four times as an analogy. To me, Buckmaster represents the Wimbledon of mathematics: white clothes, tradition, etiquette, and the ideal of being a gentle person. How you play the game matters.
The U.S. Open is different: louder, brasher, full of individual style, crassness, energy, and orientation toward the future. It’s also more open and democratic than Wimbledon. I get what’s appealing about that.
With the men’s final coming tomorrow, I find myself wondering whether AI could make mathematics at the cutting edge more like the U.S. Open than Wimbledon. It could open the gates, allowing millions of people without years of specialized training to pour in and make discoveries.
That’s thrilling. And I don’t want to be the stodgy old guard. So yes, I love the U.S. Open … and I also love Wimbledon. (I went for the first time a few years ago, and got choked up then too!)
For the past two years, Alex Townsend and I have been writing a book that traces the marvelous ideas that brought humanity to this dizzying moment: the power unleashed by supercharging math with computers, the gifts it has given us, and what may be yet to come.
So all of that, I think, may have been what set me off: the grandeur of a 4,000-year human tradition, the inconceivability of what may lie ahead—and the possibility that something precious may disappear along the way.
After Math
[This is a guest post by Silvia De Toffoli and Eamon Duede. This blog post was initially written in a different file format and converted using AI. — T.]
Silvia De Toffoli (University School for Advanced Studies IUSS Pavia)
Eamon Duede (Princeton University and Purdue University)
On September 8th, 2026, OpenAI announced that it had produced an AI-generated solution to the Navier–Stokes existence and smoothness problem, one of the seven Millennium Prize Problems. The announcement kicked off debate over credit allocation and the respective contributions of humans and machines to the result. Moreover, the announcement intensified already circulating comparisons with earlier AI conquests in domains believed to otherwise exemplify human intellectual prowess. In a recent statement, Tristan Buckmaster, one of the mathematicians involved in the Navier–Stokes saga, wrote: “This is a Deep Blue–Kasparov moment.”
Existential questions for mathematics follow naturally: if AI can now provide answers to questions at the very frontier of mathematics, is the discipline on the verge of being “solved” as many have said of chess and Go? Like chess and Go players, should mathematicians just “keep playing” and rearrange their practices?
There is something right about the “keep playing” response. As philosopher C. Thi Nguyen (2019) has been insisting, the purpose of playing a game is not exhausted by its aim (winning). The real point is not only the outcome but the process. This is perhaps clearer with a party game such as Twister than with chess: the aim of playing Twister is certainly not winning. But something similar also applies to deep intellectual games, like chess and Go. For instance, playing a game of Go well can be an achievement in defeat.
But in the context of mathematical practice, this feels like an unnecessary retreat. Instead, we can make a stronger move: reject the characterization of mathematics as a game that makes the retreat seem necessary in the first place.
The question, then, is not simply what comes after math, once AI can answer its hardest questions. It is also what we are after when we do mathematics in the first place.
The narrative that AI has “solved” mathematics rests on two assumptions, both seductive and plausible, but both wrong:
- AI really did solve a problem in mathematics.
- Mathematics is only about solving problems.
The first assumption is wrong because to really solve a mathematical problem, providing a mere answer (even if formally certified) is not sufficient. What is missing is an intelligible proof that human mathematicians can understand and use to advance the aims of mathematics. And, as we will argue below, even if AI were to give us just that, the story would not be over because the second assumption is wrong. Mathematics is clearly a much broader enterprise than just problem solving. Mathematicians strive to develop new concepts and theories, to ask and answer new questions, to unify disparate areas, to educate and sustain scholarly communities, and to produce work that is valued for its beauty and depth.
We can (and should) therefore reject the narrative of AI defeating humans at mathematics and start thinking hard about what mathematics really is and what we want it to be.
Not All Answers Are Solutions
OpenAI produced an answer to the question of whether Navier–Stokes can develop a singularity: yes. In The Hitchhiker’s Guide to the Galaxy, Deep Thought produced an answer to life, the universe, and everything: 42. Neither is exactly what we wanted.
Of course, OpenAI gave us much more than “yes.” Deep Thought offered only a number, whereas OpenAI produced two artifacts that many are willing to call proofs. The first is a Lean formalization certifying validity. This was accompanied by a manuscript that appears to contain the corresponding informal proof. So, why is this still dissatisfying? The reason has to do with the underlying notion of proof itself.
There are, in fact, two notions of proof: a logical notion and an intelligible notion.
Modern logic characterizes proof in terms of deductive validity such that a proof can be checked by a mechanical procedure that does not itself require understanding of the mathematical argument. A Lean formalization meets these standards exactly and a Lean formalization of the Navier–Stokes result is therefore a genuine and important contribution: by meeting the demands of the logical notion of proof, it secures certainty.
But mathematicians also want something else from proof. They want understanding (Thurston 1994). They want to know what makes a proposition true. This kind of knowledge trades in mathematical ideas that they can grasp, communicate to other experts, connect with existing knowledge, and use to make further progress. This is the intelligible notion of proof. As of now, it is not clear that OpenAI’s result has given the mathematical community the kind of value that one expects from the intelligible notion of proof.
Genuine proofs are at the same time logical and intelligible proofs. Historically, the two notions have tended to run together. This is because no mathematician could produce an enormously complicated logical proof without first grasping some of the key shareable ideas that made the theorem true. The logical notion of proof was primarily used to verify the correctness of intelligible proofs (Burgess and De Toffoli 2022).
But with AI, these two notions can now come apart dramatically. We can end up with formal proofs that float free from any intelligible proof.
This is not a criticism of formal proof. The converse problem is at least as serious. An intelligible mathematical argument can convey a grand idea while failing to establish that the result is actually true. Jaffe and Quinn (1993) famously used Thurston’s geometrization theorem for Haken three-manifolds as an example: a major insight accompanied by insufficiently complete proofs could become a “roadblock rather than an inspiration.” And one motivation for Hales’s Flyspeck formalization project was to verify that the intelligible (but hard to check) proof presented for the Kepler conjecture was, indeed, a genuine proof (Hales et al. 2009).
Therefore, falling short of either the logical or intelligible notion creates roadblocks where genuine proofs clear the way for mathematical progress. A real mathematical solution requires both logical correctness and intelligibility.
This is particularly clear in the case of the seven Millennium Prize Problems. They were not selected because mathematicians merely wanted seven answers, but rather because they wanted fruitful solutions. The Clay Mathematics Institute itself explains why proof matters in the case of Navier–Stokes: “Because a proof gives not only certitude, but also understanding.”
What OpenAI has given us is an answer. But it is not clear that they have delivered a fruitful solution. Perhaps, we will find that they have, but at the moment, the situation is far from clear. A genuine solution will provide adequate grounds for believing the result but also an intelligible mathematical argument that allows the result to become part of mathematics as understood and practiced by mathematicians.
Nevertheless, if it turns out that what OpenAI has provided is a mere answer, this is not enough to dispel the existential threat that mathematics is facing. Future AI systems are likely to produce genuine proofs that are at once formally certified and fully intelligible to mathematicians. So, current concerns that mathematics is on the verge of being “solved” by AI are not fully dispelled by simply insisting on genuine solutions rather than mere answers.
You Need More than Solutions to “Solve Math”
If future AI systems will produce genuine proofs, logically correct and intelligible, like those produced by “master” mathematicians, it would still be incorrect to think that mathematics would have been “solved” as some say that chess or Go have been solved.
In chess and Go, we accept radically uneven competition between humans and machines because both are, in the relevant sense, playing the same game.
But mathematics is not (or at least not only) a game. To begin with, there is no winner. Mathematics is not an adversarial game with determinate conditions for victory. It is certainly true that mathematicians compete with one another for fame, prizes, jobs, and credit. Chess players do those things too. But chess players also win chess. There is no corresponding condition for winning mathematics. There is no mathematical checkmate.
In mathematics, it is more natural to treat AI as an assistant rather than as a competitor. As Jeremy Avigad (2026) puts it, “We should keep in mind that AI is nothing more than technology, designed to serve our purposes. It is misguided to think of mathematicians as competing with AI; when we drive a car, we aren’t competing to see who can go faster, and when we use a phone, we aren’t competing to see who can speak louder.”
But there is a deeper, and in many ways prior, problem with the competition framing. It requires accepting the assumption that solving problems is the activity by which mathematical success should be measured.
Genuine problem solving is certainly one of the principal aims of mathematics. It is not, however, its only aim. Mathematics is a body of knowledge engaged with, interpreted, and digested by a scholarly community and not a registry of results in the abstract. This simple point has even motivated an entire movement in the philosophy of mathematics: the philosophy of mathematical practice.
Terence Tao (2026) lists many goals of mathematics beyond problem solving. These include developing new theories and techniques, understanding the world, sustaining a community, training the next generation of mathematicians, contributing to cumulative knowledge, and creating works of aesthetic value. Of course, these have been positively correlated with genuine solutions.
But AI breaks that correlation, for the same reason it separates the two notions of proof. So, even genuine solutions would not satisfy us.
This is not moving the goalposts but recognizing that any specific goalpost is inadequate. If mathematics is a game, it is an infinite one.
This attitude is not reactionary. We reject both the concession that logically establishing a theorem is sufficient for a genuine proof and the reduction of “AI for mathematics” to proving theorems. Accepting this, we may find many opportunities for AI in mathematics to support human mathematical flourishing.
Aftermath
We do not deny that, if OpenAI’s announcement is correct, this is an extraordinary achievement. But we should get clear about what type of achievement it is. At this moment, it is an answer, not a solution. And, even if in time the result reveals itself as a genuine solution, we have argued that, in the practice of mathematics, solutions are not everything.
The urgency of rethinking what we value in mathematics is already being recognized within the mathematical community. In a recent declaration initially signed by 25 Fields Medallists, mathematicians warn of a “severe misalignment” between the goals of AI companies and those of the mathematical community.
Mathematicians need to do more to examine their norms. The priority norm is not the problem here (though it is likely a separate problem). The current failure to distinguish genuine solutions from answers, and the growing focus on problem-solving alone, are. This way of thinking is inspired by the current credit economy in mathematics and, as David Bessis recently discussed in his blog, the credit economy needs rethinking.
AI presents mathematics not with an ending but with a choice about what mathematical practice should become. If mathematical success comes to be identified too closely with the production of certified answers, mathematics risks adapting itself to precisely those features that are easiest to benchmark and automate away.
If, instead, mathematicians treat AI as a technology for advancing its long-standing and centrally human purposes, the technology may come to contribute to an accelerated flourishing and enrichment of the discipline. The important question is, therefore, not whether AI will defeat mathematicians, but which mathematical ends we want AI to serve.
What remains in the aftermath is not merely leftovers for humans to scramble for once machines have devoured all of the real problems. Rather, it is an opportunity to clarify what mathematics is all about. We should ask again what we are after when we do mathematics.
(An extended version of this text will appear elsewhere.)
Crowdsourcing a list of general resources on the purpose, value, and nature of mathematics
Parallel to the other crowdsourced resource drive on this blog, I would like to collect books, writings and other resources that discuss or explain the purpose(s), value(s), and nature of mathematics. As seen in reactions to recent events, one can see many oversimplified or inaccurate perceptions of mathematics, such as
- “Mathematical research is about solving open problems.”
- “The main reason for posing an open problem is to obtain its solution.”
- “We should only work on problems that are of immediate benefit to society.”
- “We should prioritize maximizing the number of problems solved.”
The question of what mathematics actually is is far more complex than these assertions might suggest. Here are some classic books and articles on these topics:
- Vladimir Arnold, “On teaching mathematics“, 1997 (somewhat provocative in tone, but still interesting)
- Philip Davis, Reuben Hersh, “The Mathematical Experience“, 1981
- Timothy Gowers, “The two cultures of mathematics“, 2000
- Timothy Gowers, “Mathematics: a very short introduction“, 2002
- G. H. Hardy, “A Mathematician’s Apology“, 1940 (somewhat dated at this point, but still interesting)
- Rueben Hersh, “What is mathematics, really?“, 1997
- Imre Lakatos, “Proofs and refutations“, 1976
- Bill Thurston, “On proof and progress in mathematics“, 1994
- Cedric Villani, “Birth of a Theorem“, 2015
Here are some more informal links:
- Matt Might, The illustrated guide to a Ph.D., 2010.
- A collaboration between myself and Zach Wienersmith on sphere packing, as a five part comic strip (Apr 9, 2026)
- Illustrating the impact of the mathematical sciences – a series of posters by the National Academy of Sciences. (I was on the committee to create these posters.) (2023)
I can also list a few writings, posts, and interviews I have done on these topics:
- There is more to mathematics than rigour and proofs (2007)
- What is good mathematics? (2007)
- The value in “Dumb” questions (2022)
- What makes for ‘Good’ Mathematics (2024)
- What does it mean to think like a mathematician? (Mar 14, 2026)
- Interview with “Big Think” (relating to my forthcoming book) (Aug 28, 2026)
- On remembering what it is like to be a child: curiosity-driven play, the “largest number” game, and what premature optimization for the nominal goal extinguishes (Sep 9, 2026)
Please contribute more links in the comments below!
The status of the Hodge conjecture
[This is a guest post by Claire Voisin. This blog post was initially written in a different file format and converted using AI. — T.]
Hodge classes can be defined on any compact complex manifold . They are rational Betti cohomology classes on of even degree (eg combinations with -coefficients of classes of oriented codimension closed submanifolds of ) which satisfy a delicate “Hodge condition” necessary for them to be combinations with -coefficients of classes of complex submanifolds (and more generally closed analytic spaces) of . To understand this condition, we need to pass to cohomology with complex coefficients, and represent complex cohomology classes as de Rham cohomology classes of closed forms. The Hodge condition is that the class should be representable by a closed form that in local holomorphic coordinates is written as .
The Hodge conjecture states that a Hodge class on a smooth complex projective variety is “algebraic”, that is, is a combination with rational coefficients of classes of closed analytic (or equivalently, algebraic) subsets. Note that the Hodge conjecture can also be formulated for compact complex manifolds, and in particular compact Kähler manifolds. In this setting, there are easy examples where some nonzero degree 2 Hodge class exists on , while there are no codimension 1 closed analytic subsets in , so one needs in any case to consider not only classes of closed analytic subsets but also Chern classes of holomorphic vector bundles and their singular version (coherent sheaves). However it is proved in [12] that some compact Kähler manifolds have nontrivial Hodge classes, while Chern classes of coherent sheaves are all zero. This does not necessarily say that an analytic approach to the Hodge conjecture is not possible, but this says that any such proof has to use the fact that is algebraic. One early approach, proposed in [6], is to work on an affine Zariski open set , where is a hyperplane section of . One can represent even degree rational cohomology classes of , restricted to , as Chern classes of holomorphic vector bundles on , and then use the Hodge condition (but how?) to algebraize these vector bundles, that is, extend them from to as coherent sheaves. This very interesting construction has alas not been successful.
Another potential approach to the Hodge conjecture is by induction on the codimension. It relies on the following deep fact, which is a consequence of the Deligne theory of Hodge structures and their properties [7]. Namely, in order to prove the Hodge conjecture, it suffices to prove the following statement: for any smooth projective variety and any Hodge class on , there exists a dense Zariski open set such that in .
In this statement, is an algebraic hypersurface in , usually very singular. A class satisfying the above property for some is said to have coniveau . This notion goes back to Grothendieck (see [10]) and the study of the coniveau filtration led to important developments in [3]. Alas, how to prove that a class has coniveau ? As explained by Grothendieck in [10], this is very restrictive on and most cohomology classes (eg, classes of holomorphic forms) do not satisfy this property.
The Hodge conjecture thus motivated beautiful developments in Hodge theory and on the topology of algebraic varieties, but it is fair to say that, beyond the formal definition, there is no good understanding of what is a Hodge class from the viewpoint of algebraic geometry. The reason is that one understands well in the setting of algebraic geometry cohomology with complex coefficients and the Hodge condition, but not Betti cohomology with rational coefficients.
However, there are Hodge classes that we understand very well, which are constructed by (multi)linear algebra tricks. Probably the simplest example is the following: let be a smooth projective complex variety and let . Then is of dimension 1 and it is naturally contained in . It is rather obvious that it is generated by a Hodge class on . This class is not known to satisfy the Hodge conjecture.
Similarly constructed examples need a little more knowledge of topology. The standard conjectures like the Künneth conjecture (the Hodge conjecture for the Künneth components of the diagonal), or Lefschetz standard conjecture, are particular instances of the Hodge conjecture for certain Hodge classes that can be constructed on the square of any smooth projective complex manifold. The algebraicity of these classes is rather crucial in the theory of motives.
A more involved example is that of Weil classes on Weil abelian varieties, on which an important recent progress was made by Markman. An abelian variety over is a complex torus that has holomorphic embeddings in complex projective space, and a Weil abelian variety is one which admits a quadratic endomorphism , , . The variety has to be of even dimension and one assumes that the action of on the tangent space of has both eigenvalues with the same multiplicity . A formal argument then produces a two-dimensional space of Hodge classes of degree on , called Weil classes. A big recent progress on the Hodge conjecture is the proof by Markman [11] that Hodge classes on abelian fourfolds are algebraic. This problem is classically reduced to proving the algebraicity of Weil classes, and to prove this, Markman makes a detour through Weil classes on certain families of Weil abelian 6-folds.
Variational aspects. Smooth projective complex varieties come in families , where and are themselves algebraic varieties which can be chosen defined over a number field, as the algebraic map . The topology of the fibers does not change with ( is a fibration) but the complex structure of changes, and so does the set of Hodge classes of given degree on . Given a Hodge class of degree on a fiber , its Hodge locus is (grosso modo) the set of points such that along a path from 0 to in , the locally constant class remains Hodge. An important result concerning is the fact (proved by Cattani–Deligne–Kaplan [5]) that it behaves as if the Hodge conjecture was true, namely is a closed algebraic subset in . However, one missing information in this result is that this locus is defined over a number field, as predicted by the Hodge conjecture (a more precise statement is that Hodge classes are absolute Hodge). One could thus imagine disproving the Hodge conjecture by exhibiting a Hodge class on a fiber with Hodge locus not defined over a number field (this cannot be done with the Hodge classes explicitly described above). Such counterexample would not damage too much our understanding of the theory of motives and could lead to a corrected version of the Hodge conjecture stating that absolute Hodge classes are algebraic.
Variational Hodge conjecture. Given as above and a Hodge class , assume that and that is algebraic on . Is also algebraic on for any ?
A negative answer to that question would be the worst scenario for the theory of motives. In particular, this would disprove the Lefschetz standard conjecture (see [1]).
Deformation theory can in some cases be used to answer affirmatively the question above. Assuming that is the class of an algebraic subvariety (or a Chern class of an algebraic vector bundle on ), the semi-regularity condition of Bloch, (respectively of Buchweitz–Flenner), is a subtle cohomological property of (resp. of ) guaranteeing that deforms to , (resp. that deforms to if all its Chern classes remain Hodge on ). Buchweitz–Flenner semiregular sheaves have been used successfully by Markman in [11] for Weil classes on 6-dimensional Weil abelian varieties. Unfortunately it seems very unlikely that semi-regularity suffices to solve the variational Hodge conjecture in general. One reason is that the variational Hodge conjecture for integral Hodge classes is not true, which shows that in some cases, cycles , , are not homologous to any semiregular cycle (a multiple is in any case needed).
Another reason is that, associated to a given cycle of , there is a Deligne–Beilinson class which lives in Deligne–Beilinson cohomology of and lifts the Hodge class of . When the class is 0, then the class is the Abel–Jacobi invariant of (see [9]). There are examples of families as above where the Hodge class of a cycle of deforms to a Hodge class on but the Deligne cycle class of does not deform to the Deligne cycle class of any algebraic cycle of (see [8]). These facts suggest that semi-regular objects are very hard to construct, and too restricted to lead to a general solution of the variational Hodge conjecture.
References
[1] Y. André. Déformation et spécialisation de cycles motivés, J. Inst. Math. Jussieu, 5 (2006), 563–603.
[2] S. Bloch. Semi-regularity and de Rham cohomology. Invent. Math. 17 (1972), 51–66.
[3] S. Bloch, A. Ogus. Gersten’s conjecture and the homology of schemes, Ann. Sci. Éc. Norm. Supér., Sér. 4, 7, 181–201 (1974).
[4] R.-O. Buchweitz, H. Flenner. A semiregularity map for modules and applications to deformations. Compositio Math. 137 (2003), no. 2, 135–210.
[5] E. Cattani, P. Deligne, A. Kaplan. On the locus of Hodge classes. J. Amer. Math. Soc. 8 (1995), no. 2, 483–506.
[6] M. Cornalba, Ph. Griffiths. Analytic cycles and vector bundles on non-compact algebraic varieties. Invent. Math. 28 (1975), 1–106.
[7] P. Deligne. Théorie de Hodge II, Inst. Hautes Études Sci. Publ. Math. No. 40 (1971), 5–57.
[8] M. Green. Griffiths’ infinitesimal invariant and the Abel–Jacobi map. J. Differential Geom. 29 (1989), no. 3, 545–555.
[9] Ph. Griffiths. On the periods of certain rational integrals. I, II. Ann. of Math. (2) 90 (1969), 460–495; 90 (1969), 496–541.
[10] A. Grothendieck. Hodge’s general conjecture is false for trivial reasons. Topology 8 (1969), 299–303.
[11] E. Markman. Cycles on abelian 2n-folds of Weil type from secant sheaves on abelian n-folds, arXiv:2502.03415.
[12] C. Voisin. A counterexample to the Hodge conjecture extended to Kähler varieties. Int. Math. Res. Not. 2002, no. 20, 1057–1075.
On the Hodge conjecture
[This is a guest post by Burt Totaro. This blog post was initially written in a different file format and converted using AI. — T.]
These are strange times for mathematicians. It now seems possible that AI companies will burn through vast resources in order to prove some new fact about the Hodge conjecture. I’d like to discuss the current status of the Hodge conjecture, in order to think about the value to the mathematical community of exploring such hard problems. Thanks to Terry Tao for suggesting this guest post.
The Hodge conjecture crystallizes a big mystery: the relation between topology and algebraic geometry, or (what amounts to the same) between real and complex geometry. In short, real submanifolds are flexible, whereas complex submanifolds are quite rigid, and it is a challenge to relate the two. The setting is a “smooth complex projective variety” , a complex manifold defined by algebraic equations. We want to understand the possible shapes of complex algebraic subspaces of . (By a famous result of Wei-Liang Chow (1949), every complex analytic submanifold of is in fact algebraic.) One could ask whether every real submanifold of can be moved continuously to a complex submanifold. That is far too optimistic. Still, to a first approximation, the Hodge conjecture predicts which real submanifolds can be moved continuously to a complex submanifold, in terms of whether certain integrals are zero. (More precisely: every Hodge class in the rational cohomology of should be the class of an algebraic cycle.) The conjecture goes back to William Hodge (1950).
Now suppose in some counterfactual world that the day after the conjecture was made some inhuman oracle told us that the answer was “yes” and the conjecture was considered “solved.” We can see, by examining our real world, the vast body of mathematical knowledge that would have been lost. In our world, what we want from a great conjecture is a challenge that inspires all kinds of other discoveries, even if the original question remains open. The Hodge conjecture has been that kind of conjecture for geometers.
A key piece of evidence for the Hodge conjecture is the “Lefschetz -theorem“, which says that the Hodge conjecture is true for algebraic cycles of codimension 1 (that is, of complex dimension , if has complex dimension ), and also for cycles of dimension 1. Solomon Lefschetz’s work was quite early (around 1924), well before the Hodge conjecture was stated in general. Interestingly, some of the most important proofs by both Lefschetz and Hodge have gaps, from our current perspective. They had tremendous insight, but the full justification of their results required hard analysis by many mathematicians, notably Kunihiko Kodaira (around 1950).
By some measures, one could say that progress on the Hodge conjecture has been limited. In view of Lefschetz’s results, the first open case is for codimension-2 cycles on a variety of complex dimension 4, and that case (for arbitrary varieties of dimension 4) still seems far out of reach. But the Hodge conjecture has been enormously successful for inspiring new mathematical theories, such as Phillip Griffiths’s “Variations of Hodge structure” (starting in the 1960s) and Pierre Deligne’s “absolute Hodge cycles“. Broadly speaking, the Hodge conjecture suggests that the structure of algebraic cycles should be controlled by Hodge theory, in other words by integrals of algebraic functions. This hope has led to vast numbers of proved insights about the structure of algebraic cycles. (For example, Griffiths disproved Grothendieck’s conjecture that algebraic and homological equivalence were the same, which was a surprise; but he used Hodge theory in order to do it.) In recent decades, Claire Voisin has been a leader in using Hodge theory in new ways to get information about algebraic cycles.
An exciting recent development is Eyal Markman’s 2025 proof of the Hodge conjecture for abelian varieties of dimension 4 and 5. (Abelian varieties, the tori with a complex algebraic structure, are very special compared to all algebraic varieties; but they are a particularly important class in many ways, for example in number theory.) This is a monumental piece of work that builds on many earlier developments. André Weil (1977) identified a class of abelian varieties, those of Weil type, for which there are “unexpected” Hodge classes that do not appear on most other abelian varieties. In low dimensions such as 4, Ben Moonen and Yuri Zarhin (1995) showed that the Hodge conjecture for all abelian varieties would follow from the special case of abelian varieties of Weil type. Spencer Bloch (1972), extended by Ragnar-Olaf Buchweitz and Hubert Flenner (2003), defined a property called “semi-regularity” which, when it holds, allows proving the Hodge conjecture for a continuous family of varieties when it holds for one variety in the family.
For a long time, however, semi-regularity seemed far too strong a condition ever to apply to hard cases of the Hodge conjecture. Markman found, with great ingenuity, how to construct enough examples of semi-regular sheaves to prove the Hodge conjecture for abelian varieties in dimensions 4 and 5. His ideas use several big theories that have been developed in recent decades, notably about derived equivalences between algebraic varieties and about hyperkähler manifolds. In the end, Markman actually needed an extension of Buchweitz–Flenner’s semi-regularity theorem to “twisted sheaves”, which was supplied by Jonathan Pridham (2024).
At this point, I am confident that with enough effort, it will be possible to push these ideas further. As we have seen, AI companies are willing to mobilize resources on an incomprehensible scale to attack problems that attract their interest. We don’t know how much insight may come from any given AI-generated proof. But there is reason to worry about the commercial pressure on AI companies to claim advances at high speed and the knock-on effect this has on mathematicians. The power of mathematical ideas comes from a continual negotiation among people. What we care about are not so much isolated facts, but rather the ideas and inspiration that the search leads to.
A Severe Misalignment of AI in Mathematics
I am proud to be among the list of 25 initial signatories — all Fields Medallists — to the declaration below, which grew out of discussions between ourselves over the last week. We have also posted our declaration on this web page, and (similarly to the Leiden declaration) invite further signatures. (It is unfortunate that we did not have the time to have a more consultative process, as with Leiden; but we decided that the urgency of the situation was such that we needed to release a statement sooner rather than later.)
See also this recent article in the Economist regarding our declaration, and a brief interview with James Maynard on this topic. A French version of this declaration was published in Le Monde.
Over the last few months, the mathematical capabilities of LLMs have improved dramatically, to the point that they can solve major outstanding problems in many fields of mathematics. However, the push by AI companies to solve mathematical problems as a benchmark is detrimental to the science of mathematics, and to the mathematical community. The goals of the AI companies and the goals of the mathematical community are severely misaligned. We see these as part of broader alignment issues impacting other scientific and creative professions, as well as the whole of society.
Research mathematics deals with understanding basic structures of shapes, numbers, and natural phenomena. Over the course of generations, it has built a large corpus of sophisticated ideas, methods, abstractions, and other tools to comprehend the mathematical landscape. In turn, modern technologies and sciences are based on mathematical tools.
Famous problems have often served as landmarks and lighthouses against which one can measure an improved understanding of this landscape. Solving one of these problems has been a certain sign of new insights and interesting methods, which would then be studied by a community of mathematicians, through a long and arduous process of talks, discussions, simplifications. At the end of this process, one will ideally find a textbook presentation of the results suitable for any graduate or even undergraduate student to study. Some of the mathematical ideas pursue their journey even further to become, decades or centuries after, tools that are understood and used by the whole population.
The mathematical community functions, in many ways, as a miniature version of humanity. It consists of individuals using a wide variety of different approaches, joined by core values. The most precious resources of our profession are students and ideas, and these we nurture with great care. We feel responsible to let them grow to their full potential, until they can live a life of their own in the mathematical world. For students we often suggest problems with the core intention of developing skills making them well-positioned for advances in research and elsewhere. Our ideas we disseminate in talks, private discussions and careful writeups, connecting them to the previous ideas of others. These processes invariably take time and are based on human interaction.
In recent months, the success of AI in solving major mathematical problems has made headlines even outside mathematical circles. But solving problems is only a tool and proxy for achieving the primary goal of conceptual understanding and insight. Forgetting this in the world of AI may turn the tool against the primary goal. Indeed, the mass production at faster and faster pace of “true/false” statements could destroy fertile ground instead of breathing life into new ideas.
Often these solutions are announced in a rush, leaving no time for a proper writeup, the isolation of new methods and ideas, and citing relevant previous work of others. As in all creative professions, this raises severe attribution and plagiarism questions. Moreover, without the willing mathematicians who must take care of their development and integration into the mathematical canon, AI-conceived ideas would never become fully alive and the crucial human transmission chain between mathematicians would be lost.
We are witnessing a general threat to intellectual work, with misalignment between the outcome of the use of AI and its initial purpose. In many fields and activities, years of training have traditionally served not only to produce a final answer or product, but also to develop understanding and the ability to formulate new questions and ideas. However, building on a vast body of previous human work, AI systems are becoming increasingly capable of producing the results of such work directly, and these goals cease to align. The issues the mathematical community faces now are similar to issues that other scientific and creative professions are facing, and indicate issues that all of humanity might face: how to make sure that, as AI changes the way work is done, we do not lose sight of what that work was meant to achieve in the first place.
AI offers the potential of enhancing and accelerating genuine mathematical study and understanding. Mathematics as a profession will need to adapt to these changes in several ways. However, whether these changes ultimately benefit the field or have a destructive effect will in large part be determined by the decisions of the humans in control of this new technology.
These issues must be addressed urgently, in the mathematical community, by the companies developing these technologies and, more broadly, by a society that will confront similar problems in many other forms of intellectual work.
Artur Avila (Fields Medal 2014)
Manjul Bhargava (Fields Medal 2014)
Caucher Birkar (Fields Medal 2018)
Pierre Deligne (Fields Medal 1978)
Yu Deng (Fields Medal 2026)
Simon Donaldson (Fields Medal 1986)
Hugo Duminil-Copin (Fields Medal 2022)
Alessio Figalli (Fields Medal 2018)
Martin Hairer (Fields Medal 2014)
June Huh (Fields Medal 2022)
Maxim Kontsevich (Fields Medal 1998)
Elon Lindenstrauss (Fields Medal 2010)
Pierre-Louis Lions (Fields Medal 1994)
James Maynard (Fields Medal 2022)
Curt McMullen (Fields Medal 1998)
Shigefumi Mori (Fields Medal 1990)
Ngô Bảo Châu (Fields Medal 2010)
Andrei Okounkov (Fields Medal 2006)
Peter Scholze (Fields Medal 2018)
Stanislav Smirnov (Fields Medal 2010)
Terence Tao (Fields Medal 2006)
Maryna Viazovska (Fields Medal 2022)
Cédric Villani (Fields Medal 2010)
Wendelin Werner (Fields Medal 2006)
Efim Zelmanov (Fields Medal 1994)
On the existence of non-sofic groups
[This is a guest post by Andreas Thom. This blog post was initially written in a different file format and converted using AI. — T.]
When I woke up on August 1st, 2026, I had received a few emails from colleagues asking for my opinion on a remarkable result that had circulated the previous day. The result was a solution to a long-standing open problem in geometric group theory, specifically the existence of a non-sofic group. I was astonished and at the same time, looking at the first draft, also in a way happy to see that Kun’s work on expander decompositions and my joint work with Gábor Kun played a decisive role in the crucial Proposition 2.3 of the OpenAI paper. I had always hoped that the theory of centralizer rigidity would eventually have significant applications, but I had not found the right setting in which it could be used so effectively. In that sense, the solution also came as a relief and I was happy to explain the ideas in a post on MathOverflow a few days later.
From the start, colleagues pointed out that the framing in the public announcement that appeared shortly afterwards was misleading, in that it spoke of “no progress” in the last decade, while relying on our 2019 paper (not to mention subsequent work by many hands that was not directly relevant for the OpenAI paper but would still be considered to be progress by many).
So I wrote to Mark Sellke and Sébastien Bubeck: “[…] I find the framing intellectually dishonest. You (and I am talking about you personally, since I have no one else to address this to) cannot speak in the public announcement of a decade without progress and then use a 2019 paper in a crucial way. It is true that Proposition 2.3 is a really clever use of the centralizer-rigidity theorem, but neither does its short proof require new techniques […].
“It is true that the last stone finishes the building and usually those who can put it get the credit for solving the problem, that is fair enough. I have no problem with that and I personally do not care much about credit. However, I guess you would get enough praise without downplaying the previous contributions.”
Sellke replied to this and basically agreed to the need for a revision; as a result the public announcement was changed to the form it has now. I was glad to have received an early draft from Sellke also on August 1st, otherwise there would have been no way to react to the first public announcement at all, since neither the PDF nor the website of the announcement contained contact information. Anyway, I was happy that this was resolved and the matter closed.
It is fair to say that the approach of Kun and myself had not been viewed as the main line of attack on non-soficity prior to OpenAI’s announcement. In fact there were other more promising approaches along the line of quantum games etc. at the time, that had already led to a negative solution of the famous Connes Embedding Problem and, later, the disproof of the Aldous–Lyons conjecture. Hence, OpenAI’s detailed command of the techniques of Kun and myself made me wonder how the model found this route, especially since I discussed these techniques and their use in extensive sessions with ChatGPT over the last months.
So in the same email I asked Mark Sellke and Sébastien Bubeck: “Another point is that I and a colleague in Dresden were discussing the expander matching problem and various extensions of the work with Gábor Kun actively over the last months with ChatGPT, so that we are of course curious if that was part of the training data or accessible to the reasoning process. There is a certain (frankly unacceptable) lack of transparency here; and I fear it will damage the communal process of math more than the new AI-generated results will benefit the subject.”
Mark Sellke’s complete answer to this part of my email was: “Regarding your conversations with ChatGPT: that did not happen.”
Anyway, I thought, these techniques were public, so their use is not evidence that our conversations influenced the model. But because this was not the main line of attack, and because I had recently discussed precisely these techniques and possible extensions with ChatGPT at length, I thought the question had to be asked. Back at the beginning of August, I then returned to mathematics and wrote a subsequent paper with Gábor Kun on applications of the ideas that were the basis of Proposition 2.3. This was my way to react; after all, the integration of the new result in the math landscape seemed like a natural next step.
However, after reading up on the controversy around the Buckmaster–Alpöge case, the whole story came back to me and I realized that the answer I received from OpenAI was misleading, to say the least. I already wrote about this briefly on Mathstodon.
I had explicitly asked about two different things: (1) whether our conversations entered training data, and (2) whether they were accessible to the solving process. OpenAI said in the Buckmaster–Alpöge case that no specific user data was accessed, but added that it “cannot rule out that de-identified data derived from their usage of our products helped improve our models.” In light of OpenAI’s later wording, I cannot tell whether Sellke’s answer denied both possibilities or only direct access under (2). No qualification, explanation, or evidence was given. Whatever its intent, I regard the answer as materially misleading.
OpenAI was drawing a distinction that its answer to me erased, despite the fact that my question explicitly made that distinction. We are not required to reverse-engineer OpenAI’s internal training pipeline to establish what happened. Only OpenAI has the relevant data for that. For such a categorical denial by OpenAI to be credible, OpenAI should disclose its basis: product and privacy settings, relevant datasets and checkpoints, and what “de-identified data derived from usage” means.
I disabled model training on 29 June. That control is still only a promise whose implementation users cannot audit, and it is prospective: it does not answer what happened to earlier conversations or to derivatives already selected.
If nonpublic research supplied by users improved a model and the provider then used that model to race those users to publication—without informed consent, disclosure, or credit—that would be ethically indefensible. De-identification may remove a name; it does not remove the intellectual content of a mathematical idea. Sellke and Bubeck seem to be blind to this simple moral aspect.
Sellke gave me a categorical assurance without explaining its basis; I regard that response as materially misleading. If he lacked the information needed to rule out training use, he had no basis for giving that assurance. Bubeck’s acknowledged career-related remark in the conversation with Buckmaster and his objection to including Alpöge in a proposed paper presenting OpenAI’s proof deepen my concern about their commitment to academic standards. Taken together, these episodes raise serious questions about their judgment and personal integrity.
I am not claiming that anyone read individual chats or that our conversations were in fact used in training; I do not know that. My criticism concerns the categorical denial. If they did not know what entered the training data (the most likely scenario), they should have said so.
After I finished writing this post, I received a message from Mark Sellke, who acknowledged understanding how I “reasonably arrived at [my] conclusions given the evidence available”. He pointed me to a discussion citing OpenAI’s new statement that Buckmaster’s Codex prompts from the preceding two months could not have influenced its system, including through training. I wish I could trust this more. In any case, it suggests that the math community can successfully put pressure on the industry to take these issues at least somewhat more seriously.
So what does that all mean and how do we as a community proceed? Setting aside these particular cases (which might also be very different in what really happened behind the scenes), we have to see the broader picture and I believe there is no way of going back.
As far as I see it, a mathematical publication used to bring together three things. It announced a result, identified the people who had produced it, and added something to our human understanding of mathematics. It could therefore serve at the same time as a record of knowledge, a basis for assigning credit, and evidence of a mathematician’s ability. AI breaks this connection.
An AI system can produce a correct proof even if no human being discovered or even understood the argument in the usual sense. A typical form that this can take nowadays varies from somewhat verified AI Slop to a Lean certificate, or a combination of both. The proof may be worth publishing, but the publication then records only that the result has been established. It does not necessarily tell us how the result was found or who understood it.
The question of credit and contribution may have no satisfactory answer. A model may draw on published work, feedback, conversations, prompts, and its own search in ways that apparently cannot be reconstructed clearly. The person or the company who ran the model should not simply receive whatever credit cannot be assigned elsewhere. As another consequence, a publication record can no longer serve as a reliable measure of a person’s mathematical quality. If the theorem, proof, and written explanation may all have been generated by AI, a list of papers tells us very little about what the named author contributed or understands. Even if matters of data privacy are resolved, I see only little hope that the current system of publication and credit can be salvaged. The whole idea of personal credit will not work in such a highly connected environment anymore and I actually think that the focus on “who got something first” (not to speak of prizes for a solution to particular problems) was always misleading, even though a powerful driving force.
The more important question is therefore what someone contributes to human mathematical understanding. This includes explaining why an argument works, separating the main idea from technical details, connecting a result with other areas, finding the right questions, teaching new methods, and helping other mathematicians make use of them. Such contributions may happen through papers, but also through lectures, discussions, teaching, and collaboration.
A formal proof certificate is comparable, in a sense, to the detection of a new star. It confirms that something is there, but further work is needed to understand what it is, why it matters, and where it belongs in the larger landscape. Mathematics is only to a lesser extent the production of correct statements. It is more the human process of making sense of them.
The institutional issue goes far beyond mathematics. Advanced AI is becoming a general supply of “intelligence” on which science, education, public administration, industry, and ordinary life may all depend. It should therefore be provided according to standards comparable to those governing water or electricity: reliable and broadly available, with clear public duties, strong privacy rules, independent oversight, and protection against discriminatory or self-serving use.
Intelligence of this kind should not be treated as an ordinary consumer product whose conditions are set entirely by a few companies. A provider should not be able to collect people’s ideas and information, control all evidence about how they were used, and then exploit that advantage against its own users. Once intelligence becomes basic infrastructure for society, it must be highly regulated and governed in the public interest.
SAIR competition: Andrews-Curtis challenge
[This is a guest post by Lucas Fagan. This blog post was initially written in a different file format and converted using AI. — T.]
I am excited to announce the Andrews–Curtis Conjecture Challenge, which opens today. This challenge is a collaboration between the SAIR Foundation and the Math-AI group at Caltech, organized by Sergei Gukov, Terence Tao, and myself.
The Andrews–Curtis conjecture is one of the most prominent open problems in combinatorial group theory and also has deep connections to low-dimensional topology. Its potential counterexamples are relevant to the search for exotic smooth four-spheres and the smooth four-dimensional Poincaré conjecture, as well as the Generalized Property R conjecture about surgery on links. Yet unlike many open problems at its level, Andrews–Curtis can be formulated as a combinatorial search problem with easily checkable solutions, making it ideal for a challenge of this form.
At Caltech, we have been developing reinforcement learning and combinatorial search methods for this problem to resolve potential counterexamples (see What makes math problems hard for reinforcement learning: a case study and The Two-Hump Problem). However, many important cases have resisted all our efforts. We hope that this challenge will lead to resolving these and more (or disproving the conjecture), especially given the recent progress in AI. We give more details about the different tracks of the competition below.
To briefly introduce the problem: the Andrews–Curtis conjecture says that any balanced presentation of the trivial group can be transformed to the trivial presentation using the following moves (which do not change the underlying group):
- (AC1) Invert a relator: .
- (AC2) Multiply a relator by another: for some .
- (AC3) Conjugate a relator by a generator or its inverse: for some .
Two presentations connected by these moves are called AC-equivalent; a presentation that is AC-equivalent to the trivial presentation is called AC-trivial. As a simple example, is AC-trivial:
It is generally suspected that the conjecture is false; there are many simple potential counterexamples in which all computational efforts have failed to find a path to the trivial presentation. The most notable of these is the Akbulut–Kirby family
whose AC-triviality is open for . Indeed, , with total relator length , is the shortest possible counterexample on two generators up to AC-equivalence: all other such candidates with total relator length are AC-trivial or AC-equivalent to (see Miasnikov and Myasnikov and Havas and Ramsay).However, difficulty in finding a path is hardly evidence against a path’s existence: Bridson and Lishak demonstrated families of AC-trivial presentations whose trivialization path lengths grow faster than any fixed-height tower of exponentials in relator length. Bridson also explicitly gives a relatively small four-generator presentation that requires more than moves to trivialize.
The competition will also explore the stable Andrews–Curtis conjecture. The stable version asks the same question with two additional allowed moves:
- (AC4) Add a generator and relator :
- (AC5) Undo (AC4), removing a generator and its matching relator:
Of course, AC-triviality implies stable AC-triviality, but it is unknown whether the converse holds. Even with stabilization moves, trivializations can be extremely long: Bridson’s lower bounds still hold with these extra moves allowed.
The stable version of the conjecture is particularly interesting because of its connections to topology. C. T. C. Wall proved that in dimension 3 and higher, we only need one additional dimension to realize any simple homotopy equivalence by elementary expansions and collapses. The stable AC conjecture is equivalent to the statement that this result also holds in dimension 2 in the case of finite contractible polyhedra. The stable AC conjecture is also equivalent to a restricted form of another easy-to-state hard-to-prove open conjecture in low-dimensional topology: Zeeman’s conjecture, which says that for any such polyhedron , the product collapses to a point.
For the competition, the Discovery Track will cover both AC and stable AC and will open today. Both will use the same pool of 10,115 balanced two-generator presentations of the trivial group. The examples range from easy to open research problems, including many from the series mentioned above. Since difficulty is hard to predict, we leave participants to discover which examples are within reach and do not label presentations by difficulty or origin.
The goal of the Discovery Track is to find short trivialization paths. For AC, paths end at , and for stable AC, paths end at the empty presentation. For the stable AC search problem, we allow up to eight generators. Participants submit move sequences, which are checked automatically, and you can download the verifier to check solutions locally before submitting.
The AC and stable AC problems will each have their own leaderboard. For each problem, only teams with the shortest accepted solution receive points, with reduced credit for ties. Finding a shorter solution therefore takes the points from the previous record holders. However, we will note the first solver of each presentation separately. Submitted paths will be private during the competition, and all valid solutions will be released afterwards to form a public benchmark.
The Proof Track will open later today. It accepts proofs or disproofs of either full conjecture, and counterexamples here do not need to come from the competition pool. Submissions will be public for community review, and organizers may assess selected claims for competition recognition.
The competition closes on November 30, 2026. AI tools are welcome, and participants can enter individually or as teams. Registration is open on the competition page. We also encourage participants to exchange ideas and discuss the challenge on the SAIR Zulip.
Rotating needles in space: the road to the Kakeya conjecture, and why it matters
A recent tradition of the ICM is to have a non-technical popular article written for each of its medallists. (This is separate from the older tradition of having a laudatio for the winner, which is similar but aimed at a more advanced audience.) I was approached to write such an article for Hong Wang on the occasion of her Fields Medal. I have uploaded my initial version of this article, “Rotating needles in space: the road to the Kakeya conjecture, and why it matters“, to the arXiv. Chris Sogge gave the corresponding laudatio; both should eventually appear in the Proceedings of the ICM.
Mathematical Discourse
Mathematical Discourse is a new online, peer-reviewed mathematics journal, publishing videos of mathematics research talks of the highest quality. Our inaugural scientific and editorial boards are listed at the end of this message.
Our goal is to promote and nourish a culture which values communication as an essential part of the research process. Giving a mathematical talk captures the human side of mathematics in a unique way, and supporting these human aspects of our discipline will be ever more important in the coming years.
We are now accepting submissions for our first issue. If you have seen (or given!) a truly excellent, mathematically stimulating research talk which was video recorded, we hope you will encourage the speaker (possibly yourself!) to submit it.
Our current call for submissions is limited to ~1 hour research seminar or colloquia style talks in pure mathematics. A detailed description of the scope of the journal and submission requirements are available at https://www.mathematicaldiscourse.org/submit.
We hope you will spread word of this new initiative to your colleagues and collaborators!
Sincerely yours,
Katie Mann, Akshay Venkatesh, Rachel Webb
Managing Editors of Mathematical Discourse
Scientific Board: Hugo Duminil-Copin, Larry Guth, Bryna Kra, Yair Minsky, Bjorn Poonen, Terence Tao, Ravi Vakil, Geordie Williamson
Editors: Matt Baker, Richard Bamler, Andrej Bauer, Alexei Borodin, Yaiza Canzani, Daniel Cristofaro-Gardner, Jordan Ellenberg, Hélène Eynard-Bontemps, Jessica Fintzen, Sergey Fomin, Zaher Hani, Peter Hintz, Shrawan Kumar, Justin Moore, Samuel Taylor, Richard Thomas, Isabel Vogt, Yilin Wang, Alex Wright, Zhiwei Yun
Quantitative bounds for sets lacking polynomial progressions with shifted prime difference
Ben Krause, Hamed Mousavi, Joni Teräiväinen, and I have just uploaded to the arXiv our paper Quantitative bounds for sets lacking polynomial progressions with shifted prime difference. The purpose of this paper is to obtain quantitative versions of this theorem of Wooley and Ziegler:
Theorem 1 Let be polynomials of one variable with integer coefficients with zero constant term, and let be a set of integers of positive density. Then there exist infinitely many primes such that contains a progression of the form for some integer .
This generalizes the famous theorem of Szemerédi in two ways: firstly, by considering “polynomial progressions” instead of arithmetic progressions, and secondly by requiring the shift parameter to be one less than a prime . The first extension of Szemerédi’s theorem is a theorem of Bergelson and Leibman; and the second extension is also obtainable by combining the arguments of Frantzikinakis, Host and Kra with the results of Green, Ziegler, and myself.
The proof of the above theorem uses ergodic theory, which makes it difficult to extract quantitative bounds from it; and standard methods of “finitizing” ergodic theory results, for instance by replacing Host–Kra seminorms by their Gowers uniformity norm counterparts, run into a number of technical difficulties here due to the need to work with multiple scales due to the presence of polynomials, as well as the fact that many of the conjectural uniformity properties of the prime numbers at small scales remain unproven.
Nevertheless, we are able to get reasonable quantitative results (with density bounds that are roughly single or doubly logarithmic in scale) in the following special cases:
- linear polynomials;
- polynomials of distinct degree; and
- multiples of a fixed polynomial.
Previous quantitative bounds in the linear case were obtained by Leng and by Teräväinen and myself, but our new method improves upon these bounds by roughly one iterated logarithm. On the other hand, in the case of two term progressions of spacing , there is a much stronger quantitative result (with polynomial dependence of constants) due to Green; our method do not recover that result.
Our methods use a variety of old and new methods in the subject. For instance, we use the (now quite standard) “-trick” to restrict the primes to a single congruence class to improve their uniformity properties (which, thanks to the recent work of Leng and Matthiesen–Teräväinen–Wang, are now quite strong quantitatively); and for good configurations of polynomials, one can also use a recent transference theorem of Altman and Sawhney to compare the polynomial averages with simpler linear ones, without having to pass to short scales. In order to get relatively strong bounds unconditionally, a “Siegel approximation” for the primes is used taking into account the potential influence of a Siegel zero. It will not be surprising to the experts that quantitative inverse Gowers theorems and nilsequence equidistribution theorems also play a major role.
A key technical difficulty is the presence of the modulus in the coefficients of the polynomial progressions after changes of variable, which forced us to make several of the existing estimates uniform over such coefficients (assuming they are not unreasonably large).This caused several complications that made a correct argument to more time-consuming to locate than initially planned.
AI usage in this work was fairly light, being restricted to proofreading and literature search only.
Palomar – a registry of Lean verified mathematics
In recent months there has been a proliferation of AI-generated proofs of various old and new results, some of which have been formalized in the proof assistant language Lean. However, checking that a given Lean repository actually proves the claimed statement is somewhat non-trivial, especially for an audience which is not expert in the use of Lean: one has to first check that the claimed formal Lean statements have proofs that typecheck, that the proofs do not contain any “cheats” such as adding additional axioms, and that the formal statements also match (in a semantic sense) the informal description of the claimed results.
To help bring some clarity to this situation, I am happy to announce that Palomar registry of Lean verified mathematics, which is an initiative incubated by the Lean FRO and by ICARM, is now open for submissions. I am serving in several roles on this registry, including on the scientific advisory board, together with Jeremy Avigad, Matthew Ballard, Jaume de Dios, Nestor Guillen, Bryna Kra, Kim Morrison, Ravi Vakil, and Akshay Venkatesh.
A detailed motivation for Palomar can be found here, and further information about Palomar can be found here. A zeroth approximation of what Palomar intends to be is the analogue of a preprint server for Lean proofs. More precisely, Palomar (which is named after the astronomical observatory) is a registry of external Github repositories (or more precisely, “snapshots” of such repositories, as represented by a specific Github commit) containing Lean code adhering to the current best practices for such formalizations, in particular containing
- A “challenge file” containing a short, human readable description in Lean of the results claimed.
- A “solution module” containing an (arbitrarily long) proof of the results claimed in the challenge file.
- A “formalization.yaml” file describing the results in informal language, and also containing a number of other relevant metadata and disclosures.
(There are also some additional technical requirements for the repository which I will omit here.) If a snapshot of a repository is submitted to Palomar, it will check both (a) that the solution module typechecks and proves exactly the results claimed in the challenge file, and that (b) the informal description of the result in the formalization.yaml file appears to match the result claimed in the challenge file, and that the repository meets various minimal standards required for a registry entry. The first check (a) is purely mechanical, using the Lean tool Comparator; the second check (b) is non-deterministic, being performed by a large language model. If a repository passes both checks, it can be registered on Palomar. It is worth stressing that the checks in (a) and (b) fall well short of what a proper human peer review of a submission for novelty, interest, and accuracy would give; in particular, Palomar is not a peer-reviewed journal.
The submission process is thorough, but achievable: as a test, I successfully managed to submit my own recent formalization of the proof of Sendov’s conjecture to Palomar, and also plan to submit some older formalizations to the registry soon.
In any event, the registry is now open for formalizations of both old and new results. Submissions (whether human-generated, AI-generated, or some mixture of both) are welcome; please read the (somewhat detailed) instructions here before starting a submission. (I will however note that modern AI agents are quite helpful in assisting with the mechanical details of the submission, though a human review is still strongly recommended.)
Discussion and feedback on Palomar will occur on this Zulip channel.
A digestion of the proof of Sendov’s conjecture
This post concerns the following conjecture of Sendov, as well as its strengthening by Phelps–Rodriguez:
Conjecture 1 (Sendov’s conjecture) Let , and let be a degree polynomial with all zeroes in the unit disk. Then for every zero of , there exists a critical point of with .
Conjecture 2 (Phelps–Rodriguez conjecture) Let , and let be a degree polynomial with all zeroes in the unit disk. Then for every zero of , there exists a critical point of with , unless is on the unit circle and is a scalar multiple of .
By applying a rotation around the origin, we can normalize to be a real number with .
From the work of Rubinstein, both conjectures were already established in the case, so one can restrict to the case. Both of these conjectures then follow from
Conjecture 3 (Sendov’s conjecture in interior) Let . Let be a degree polynomial with all zeroes in the unit disk. Then if is a zero of , there exists a critical point of with .
All three of these conjectures were established for (in a sequence of papers culminating in this paper of Brown and Xiang) and for sufficiently large (in a paper of myself, which in turn built upon several partial results in this setting). This left the case of intermediate to be settled. My arguments used some qualitative ingredients (most notably analytic continuation) and as such did not easily lend themselves to quantifying the threshold of above which the argument was valid.
Recently, Lech Mazur was able to use an AI tool to resolve Sendov’s conjecture for all , with the proof verified in Lean. However, the AI-generated proof was not human-digested to be in the form of a publication-ready preprint; and it has taken me several days (with heavy AI assistance) to perform such a digestion, to place the proof in proper context with previous literature and to simplify and streamline the argument to highlight the main ideas. (Note: the above chat log only represents a portion of the digestion work: the rest was performed with pen and paper, or using some further AI agents.) The same arguments also give a new proof of Rubinstein’s theorem, which I also give below the fold.
One consequence of this digestion is that the argument in fact demonstrates Conjecture 3, and thus resolves both the Sendov conjecture and the Phelps–Rodriguez conjecture in full generality.
The proof ends up being remarkably elementary. No complex analysis is used other than the fundamental theorem of algebra (and very basic facts about Möbius transformations); and the deepest inequality used as input is the Maclaurin inequality (and we only need a special case of that inequality which can be derived from the arithmetic mean-harmonic mean inequality and an induction argument).
Using an AI agent, I have been able to formalize the entire argument in Lean, extended to by some minor modifications to the proof. This formalization is more streamlined than the original formalization (it has about 15,000 lines of code, compared with around 90,000 for the original proof).
We now prove Conjecture 3. The cases have long been known but need to be treated separately; a short proof using the machinery developed here is provided at the end of the post. Suppose now that we have a counterexample for some , thus one can find a degree polynomial with zeroes
for some and , in the closed unit disk, whose critical points all lie a distance at least from . We use notation here in the non-asymptotic sense, thus means that for some absolute constant (independent of ). We will also use the notation to denote a quantity that is bounded in magnitude by .To capture the fact that the critical points lie at a distance at least from , we write these critical points as
for some (non-zero) in the closed unit disk.Example 4 If and , then are the non-trivial roots of unity, while the are all equal to . Strictly speaking this is not actually a counterexample to Conjecture 3, because is not strictly less than one; nevertheless this is an important motivating near-counterexample for the arguments below.
Example 5 A generalization of the previous example was studied in Section 4 of my paper. Here one took where was an asymptotic parameter going to infinity, was a low-degree polynomial for some , and were constants. This polynomial has a zero at , critical points at , and additional critical points near . If all the critical points were at distance at least one from , one would have and while if all the zeroes were in the unit disk, the calculations in my paper showed that Here denotes a quantity that goes to zero as . If one ignores the errors, one can show that these conditions are only simultaneously feasible if and all the vanish, but the argument was somewhat subtle (I had to proceed by inspecting the second Fourier coefficient of (1)). This illustrates the fact that the regime is particularly delicate.
We now have two sets of points in the closed unit disk: and . They “communicate” with each other through the polynomial and its first derivative , both of which can be expressed in terms of either set of points (as well as and ). Indeed, if we normalize to be monic, then we can factor in terms of the zeroes as
and thus upon differentiating Here and in the sequel we adopt the convention of removing singularities when dealing with expressions that involve multiplication by both and , by cancelling such terms first in the event that .In a similar vein, can be factored
and thus on integrating (and using )It is convenient to rule out the easy case right away. In this case we see from (3), (4) that
which is absurd since the first product has magnitude at most one, and the second product has magnitude at least one. Thus we can assume henceforth that .By inspecting or at various natural locations, we can thus obtain a number of identities relating the to the . We record the ones that we actually need here:
Lemma 6 (Communication identities) Let denote the function
- (i) (Centroid identity) We have That is to say, the centroid of the zeroes equals the centroid of the critical values.
- (ii) (Polar identity) We have
- (iii) (First origin identity) We have
- (iv) (Second origin identity) We have (Again, we are using the convention of removing singularities to deal with the case where some of the vanish.)
Proof: For (i), we inspect the behavior of as . From (2) we have
and thus on differentiating term by term Meanwhile, from (4) we have Comparing coefficients, we obtain the claim.For (ii), we consider the expression . On the one hand, from (2), (3) one has
(Note from hypothesis that cannot be a critical point, so the denominator is non-zero.) On the other hand, from (4), (5) one has Equating the two identities, we obtain (ii) after some algebra.For (iii), we evaluate . From (2) we have
while from (5) we have Equating the two identities, we obtain (iii) after some algebra using (6).For (iv), we similarly evaluate . From (3) we have
while from (4) one has Equating the two identities, we obtain (iv) after some algebra using (6).Remarkably, the polynomial will play no further role in the argument: the identities in (i)-(iv), together with the hypotheses that and lie in the closed unit disk, will be sufficient by themselves to obtain a contradiction.
Example 7 Continuing the example in Example 4, in (i) both sides vanish. In (ii), both sides are equal to one. For (iii) and (iv), we have , with both sides of (iii) equal to one, and both sides of (iv) equal to zero.
Remark 8 The centroid identity is extremely classical, going back to this 1948 paper of Popoviciu. The comparison of the polynomial at a location and at the polar inversion of that location across the closed unit disk is a familiar trick in the literature; see, e.g., Lemma 5 and Theorem 8 of Dégot. The specific form of the polar identity is implicit in the first part of Section 5 of Mazur’s AI-generated proof, while the origin identities are extracted from equation (6.3) of that proof. The first origin identity is also very close to Theorem 6 of Dégot, while the second origin identity is similar to some identities appearing in the proof of Lemma 6 of Dégot, as well as the work of Mir–Nazir–Wani and (in the case) Rubinstein. The work of Meir–Sharma and Mir–Nazir–Wani also contain several further identities relating the to the ; see in particular Lemma 15 below. Variants of (5) also appear in Proposition 10 of Miller.
Remark 9 The first origin identity (9) is already strong enough to handle asymptotically all examples of the form in Example 5, except in the endpoint case where vanish and the are all . Indeed, as the are in the closed unit disk, (9) implies that On the other hand, routine calculations (omitted here) show that leading asymptotically to the constraint But all terms here are non-negative (since ), so this forces a contradiction unless (and hence also ) and the all vanish.
As mentioned in Example 5, the most delicate regime occurs when . It is convenient to introduce the normalized version
of , thus , and the case corresponds to . Informally, measures how close is to (at the scale of ).A key role in the argument will be played by the mean
of the , particularly the real part . As the all lie in the unit disk, the mean does also, so that and On the other hand, in the example in Example 4, is equal to the extremal value of , and . In Example 5, we have (and ).It will be convenient to work with the quadratic polynomial
with a particular emphasis on the value at : One should primarily think of as a measure of how close is to . Clearly we have for all (note that is strictly less than ).The arguments will revolve around the relationship between and . Specifically, we will establish the following two inequalities below the fold. The first inequality, which we call the “polar inequality”, comes in three forms:
Proposition 10 (Polar inequality)
- (i) (Raw polar inequality) We have
- (ii) (Polar inequality in , form) We have
- (iii) (Simplified polar inequality) We have In particular, since , one has
It will be the inequality (18) that we use in practice, but it will be derived from (17), which in turn is a consequence of (16), which will follow from the polar identity (8) together with the fact that the and lie in the unit disk. The bound (18) is only slightly weaker than (17); see the (Gemini-generated) image below.
I was not able to find an exact duplicate of the above polar inequalities in past literature, but the paper of Dégot contains several similar inequalities. The inequality (16) was extracted from (5.1) of Mazur’s AI-generated proof; the subsequent bounds (17), (18) arose from my attempts to simplify the arguments after that point.
The second inequality, which is more difficult, also will come in several forms:
Proposition 11 (Origin inequality) Let .
- (i) (Raw origin inequality) We have
- (ii) ( bound) We have
- (iii) (Origin inequality in , form) We have
Part (i) (which was extracted with some effort from Section 6 of the original AI-generated argument) will be deduced from the first and second origin identities (9), (10), as well as the centroid identity (7). Part (ii) will follow from (i) and the polar inequality (18), while part (iii) is an elementary consequence of (i).
As it turns out, the last three terms in (21) are asymptotically negligible as . Dropping those terms gives a competing feasibility region for and which is disjoint from the one coming from the polar inequality (17) (or (18)):
This already suggests that one can use this approach to recover my previous result on Sendov’s conjecture holding for all sufficiently large . In fact, even with the three error terms in (21) added, there is enough room between the two inequalities (18), (21) to obtain a contradiction for all (using the additional bound to control these errors), although showing this for medium-sized (such as ) requires a certain amount of computer assistance.
For fixed , the right-hand side of (21) is monotone increasing in (or equivalently, monotone decreasing in ). In view of (18), we can thus replace by in this inequality, so that is replaced by
and replaced by . The inequality (21) then becomes an inequality involving only and : We also note that the bounds force the constraint This prevents from getting too close to the upper limit (or getting too close to zero).We can now eliminate all large degrees, e.g., , as follows. The quadratic attains its minimum at . For we have
while for (if this region is non-vacuous) we can bound the quadratic by its value at . Thus Evaluating these expressions, we arrive at Since , we have . Next, we claim that . As is monotone increasing in , it suffices to do this when . Here one can directly compute that since the discriminant of the numerator is negative, we conclude that as desired.Dropping some and terms, we conclude that
Every term on the right-hand side can be seen to be decreasing in for . Thus the right-hand side can be bounded by giving the desired contradiction.The remaining range to handle is when
It turns out that (22) remains infeasible in this range. This can be illustrated numerically without much difficulty: see this applet. For instance, in the most delicate case , the right-hand side of (22) only gets as large as (and in particular stays below ) throughout the range :
I have also verified this bound in Lean.
— 1. The polar inequality —
We begin with a proof of Proposition 10.
As is well known, the Möbius transform maps the closed unit disk to itself. In particular, we have
for all of the zeroes . Inserting this into the polar identity (8) and using the triangle inequality, we conclude the lower bound We now convert this bound to a bound involving the quantity in (12). From the arithmetic mean-geometric mean inequality we have and from (12) we have Since , we thus have giving the raw polar inequality (16).Bounding by and using the quantities from (11), (15), we observe that
Using the basic inequality , we thus have with strict inequality for . From (16) we conclude (17). This also implies , since otherwise the integrand is always bounded by , which is absurd.On evaluating the integral in (17), we obtain
and thus so on taking logarithms we obtain It remains to establish the bound Here we use an AI-generated argument. One can directly calculate where and . If we can show that for all , then taking logarithms in (17) yields from which (26) will follow by routine algebra.Both sides of (27) vanish at . Taking derivatives, it suffices to show that
which rearranges to To expand the left-hand side, we use the double angle formulae and to rewrite it as Collecting the coefficient of for and extracting a common factor of , one is left with where . (The remaining coefficients, which also receive contributions from the polynomial terms, all vanish.) Thus the left-hand side has the Taylor expansion in which every coefficient is non-negative, giving the claim.Remark 12 As the image in the introduction suggests, the bound (18) is only slightly weaker than (17). For small , one can perform Taylor approximation on the latter bound to obtain while the former bound is Note that is slightly smaller than .
Relating to this, the constant in (27) cannot be improved.
— 2. The origin inequality —
Now we turn to the proof of Proposition 11, which is more difficult and revolves around an analysis of the function defined in (6). We begin with a heuristic analysis. Inserting the approximation for small into (6) and using (12), we are led to the approximation
at least when is small (which turns out to be the dominant regime in applications). This suggests a relation between the two expressions involving in the origin identities in Lemma 6. Substituting in this approximation, we obtain some (slightly complicated) approximation for the sum in terms of , , , , and the product .As lie in the closed unit disk, the product does also. However, past experience with the Sendov conjecture has taught us that the worst cases tend to be when lie very close to the boundary of the disk, so that is close to one. For instance, in Example 4 all the and lie on the unit circle, and . See Remark 3 of Dégot or Theorem 1.10(ii) of my own paper for other places where this heuristic is noted. To simplify the discussion, let us assume for now that is exactly one, so that all lie on the unit circle. This leads in particular to the inversion identities
The centroid identity in Lemma 6(i) relates the sum of the with the sum of the . Using (31), this gives a similar identity relating the sum of the with the sum of the . The latter sum is of course just . This combines well with the previous approximation, thus giving an approximate identity relating , to , , and . As it turns out, the roles of and are minor and can be quickly eliminated for the purposes of obtaining useful bounds, leading eventually to the relation in Proposition 11.We turn to the details. To make the approximation (30) more precise, we note that , and hence by the fundamental theorem of calculus
The heuristic (29) predicts that , which would give (30). If we actually differentiate (6) carefully, we obtain the exact identity Bounding , we write thisWhen faced with a similar expression in (24), we used the arithmetic mean-geometric mean inequality. Here, the analogous tool is Maclaurin’s inequality, which gives
and hence by Cauchy–Schwarz Repeating the calculations used to show (25), we have and so we obtain the bound Integrating this, we obtain a rigorous analogue of (30), and thus by the triangle inequality From the first and second origin identities (9), (10) we haveThe next step is thus to estimate . When , then all the were on the unit circle and we could use (31) (and the centroid identity) to proceed. Now, we are no longer assuming to equal , but we can still adapt the previous arguments with a loss proportional to . The key lemma is
Lemma 13 (Defect lemma) Let be some points in the closed unit disk. Then
Proof: By a limiting argument we may assume that none of the vanish. If we write for some , then we can calculate that
and Thus the desired inequality reduces to the superadditivity property But from the sinh addition formula we have for all non-negative (this also follows from the convex nature of together with ), and the claim follows by induction.We remark that the lemma can also be proven by direct induction, without an appeal to hyperbolic trigonometry.
From taking complex conjugates of the centroid identity (7) and performing some algebra, we have
Using the defect lemma (applied to the points ) and the triangle inequality we conclude that where as before we are removing singularities when some of the vanish. Applying (12) and some algebraic manipulation, we arrive at Substituting this back into (33), we conclude that and hence after some algebra and the triangle inequality Inserting this into (32), we obtain We can simplify (34) by reducing to the case. Indeed, we shall show that which implies that the right-hand side of (34) is non-decreasing in in the range . Thus we may replace by in (34) to conclude that Let us now verify (35). Using and the triangle inequality, we can lower bound Inserting this into (35) and clearing denominators, we reduce after some algebra to But as a quadratic polynomial in , the left-hand side has discriminant , which one can check to be negative for sufficiently large (in fact suffices), giving the claim (35).Next we eliminate the role of the imaginary term . Observe for any complex number with positive real part that
as can be seen by squaring both sides. The expression has real part which lies between and (in particular, it is positive), and imaginary part of magnitude at most by (13). We conclude that The right-hand side can be rearranged using the quantity from (11) as so the bound (36) gives (19).
— 2.1. Upper bound on —
Now we can prove (20). Suppose for contradiction that ; since , this implies that . Crudely discarding the term in (19) and bounding by , we have
The quadratic polynomial equals at and attains its minimum at with value . By convexity, we thus have for and for (this latter statement is vacuous if ). Since , we can therefore crudely bound and hence From (15) we have , thus by (18) one has From another application of (18) one has We conclude that It is now convenient to introduce the quantity , thus with and Inserting these bounds and dividing by , we conclude Since , we obtain Since and , we conclude that Routine calculus shows that has a maximum of at most , and that the right-hand side here is at most , giving the required contradiction. This proves (20).
— 2.2. A simplified estimate —
Now we show (21). Note from (11) that
while from (15) we have and hence also From (15) we have By the mean value theorem (noting that is non-negative) we thus have From the standard beta function identity (and the fact that ) we can thus replace (19) by From (15) we have Thus by (37), (38), (39) Dividing by the positive quantity gives the claim.
— 3. Rubinstein’s theorem —
We now adapt the arguments to give a proof of Rubinstein’s theorem that the Phelps–Rodriguez conjecture holds in the case, i.e.,
Theorem 14 (Rubinstein’s theorem) Let , and let be a degree polynomial with all zeroes in the unit disk. If , then there exists a critical point of with , unless is a scalar multiple of .
The argument here is essentially in Remark 5.1 of this paper of Tang and Zhang.
Taking contrapositives, we may assume that the critical points are of the form for some in the closed unit disk, and normalize to be monic; our task is to show that .
The polar identity (8), based on calculating degenerates to a triviality when , but we have the following usable substitute, valid for any choice of , first observed in equation (3.2) of Meir–Sharma:
Lemma 15 (Meir–Sharma identity) If and the critical points are of the form then all the zeroes are not equal to , and
Proof: By hypothesis, is not a critical point of , so and for all . Instead of computing , we instead consider the expression . On the one hand, from (4) we have
while from differentiating (4) we have Meanwhile, from (3) we have and from differentiating (3) we have Using these identities to compute in two different ways gives the claim.Now take . Since lie in the closed unit disk, has real part at least , while is at most . Thus, the only way that the above identity can hold is if for all , hence for all . Thus all critical points are at the origin, which forces for some . Since , we conclude that , giving the claim.
— 4. The cases —
We now prove the cases of Conjecture 3. The starting point is (23). Using the triangle inequality and , this implies that
(This also follows from (16) and .) From Hölder’s inequality and we conclude that The right-hand side can be computed to equal which is obviously less than for , giving the contradiction.Remark 16 The same argument also works for , but breaks down for higher .
— 5. Further directions —
The Sendov and Phelps–Rodriguez conjectures are now resolved, but several related conjectures remain open. The following strengthening of Sendov’s conjecture, by Borcea, is open for any :
Conjecture 17 (Borcea conjecture) Let and , and let be a degree polynomial with zeroes satisfying . Then for every zero of , there exists a critical point of with .
Sendov’s conjecture is the limiting case of this conjecture. There has been relatively little progress on this conjecture: the cases were established by Khavinson, Pereira, Putinar, Saff, and Shimorin, and in this previous paper we reported the negative result that AlphaEvolve failed to find a counterexample to the conjecture. The proof methods here do not seem to extend easily; all the identities relating zeroes and critical points continue to hold, but now that the are only constrained to the unit disk in an averaged moment sense, all of the inequalities developed above now fail.
Another strengthening of Sendov’s conjecture that remains open is Schmeisser’s conjecture:
Conjecture 18 (Schmeisser’s conjecture) Let , and let be a degree polynomial with all zeroes in the closed unit disk. Then for any in the convex hull of the zeroes of , there exists a critical point of with .
Schmeisser proved several special cases of this conjecture, and AlphaEvolve again failed to find a counterexample, but there has not been much further progress. Here, the are now back in the closed unit disk, but we no longer have , again rendering most of the previous identities invalid. But perhaps some modification of the arguments here can make some progress on this conjecture.
A common generalization of the Borcea and Schmeisser conjectures was proposed in Conjecture 2.4 of this paper of Zhang. A slightly different strengthening was also proposed in Conjecture 1.10 of Tang and Zhang:
Conjecture 19 (Tang–Zhang conjecture) Let , and let be a degree polynomial with all zeroes in the closed unit disk and critical points . Then for any , one has .
Sendov’s conjecture is the limiting case . By Hölder’s inequality, the case is the strongest form of the conjecture.
Another well known variant of Sendov’s conjecture is Smale’s problem:
Conjecture 20 (Smale’s problem) Let , and let be a degree polynomial. Then for any zero of , there exists a critical point of with .
The constant is best possible, as can be seen by the example and . Using the Koebe one-quarter theorem, Smale proved this conjecture with replaced by . Some slight improvements of this bound have been obtained over the years; for instance for , the improved bound of was obtained by Crane. Again, AlphaEvolve failed to find a counterexample to this conjecture. This problem does not seem to have a direct relationship with Sendov’s conjecture, and there is no useful normalization of the zeroes and critical points that is confined to the unit disk. Nevertheless there may be some hope of making progress on this conjecture, perhaps working first in the asymptotic regime .
Needless to say, I did try some desultory attempts to use AI tools to attack these questions, but without much notable success.
One potential way forward is to find further proofs of Sendov’s conjecture that utilize other techniques that might be more broadly applicable to this larger family of problems. The proof here is remarkable in that the zeroes and critical points are treated almost as independent mathematical objects, communicating with each other only very narrowly through four identities in which one only inspects the underlying polynomial (and its derivative) at a small number of points. It could be that an approach focusing on more global features of the polynomial may lead to new proofs of Sendov’s conjecture, and perhaps also of its generalizations.
What sort of maths are LLMs good at?
For the sake of anyone who might read this blog post in the distant future (a month from now, say), let me mention that I am writing it a few days after OpenAI announced that it had solved ten major problems in mathematics and theoretical computer science, including the first construction of a non-sofic group, and a proof that the multicolour Ramsey number (where there are 3’s) grows superexponentially in . The first was, to judge from various talks I have been to, one of the most important unsolved problems in group theory, and the second was a major open problem in Ramsey theory that I didn’t necessarily expect to see solved in my lifetime, though of course such expectations now have to be revised. The reason I want to be clear about the timing is that I shall be discussing the current capabilities of LLMs in the full expectation that those will continue to change rapidly. So it is likely that in not too long from now, if there is anything interesting in what I write, it will be interesting mainly as a record of what the situation looked like in early August 2026.
These results, and the other eight on the list, are extraordinarily impressive, but it still doesn’t seem to be the case that LLMs are better than all humans at all aspects of mathematics. If they were, then their big speed advantage over us would mean that there would be much more of a flood of results. So it is natural to wonder about what kinds of problems LLMs are good at, and about where there is still room for improvement. I don’t pretend to have a good answer to this question, where a good answer would be a crisp classification that would fit the current examples well, but it is an interesting exercise to try to rule out some bad answers, and to try to identify potential answers that aren’t obviously contradicted by the evidence.
Are LLMs particularly good at finding counterexamples?A first remark here is that LLMs are not just good at finding counterexamples: they can find proofs of difficult statements as well. However, it is notable that the most famous problems they have solved have almost all been with counterexamples rather than proofs. That is true of the two problems mentioned above, and also of the Jacobian conjecture and the unit distance conjecture.
If one wants to theorize that LLMs are particularly good at finding counterexamples, then there are two things it would be good to do to make the theory more convincing. The first may sound unproblematic: it is to decide when solving a problem counts as finding a counterexample. Once that is sorted out, the second is to come up with a potential explanation of why LLMs would be particularly well suited to solving problems of that particular kind.
What does it mean to find a counterexample?Why am I suggesting that it is not completely obvious what it means to find a counterexample? Surely, one might suggest, all it means is that you have a statement of the form “Every object of such and such a type has such and such a property,” and you exhibit an object of the given type that does not have the given property.
However, this doesn’t always work. Consider a famous result of Vinogradov, which states that every sufficiently large positive integer is a sum of three primes. The negation of this statement is (or is equivalent to) the statement that for every positive integer there exists an integer such that is not a sum of three primes. In other words, it states that every positive integer has a certain property. Seen in this light, Vinogradov found an example of a positive integer that does not have the given property. Do we want to say that Vinogradov found a counterexample? Clearly not — the result should obviously be classified as a theorem and not a counterexample.
Thus, we cannot just naively say that LLMs are particularly good at negating universally quantified statements: there has to be something about the nature of the universal quantification. With the three-primes example, it is clear that Vinogradov did not think, “How am I going to find with this property?” Rather, what he thought would have been more like, “I’ve got an integer that is very large. How am I going to show that it is a sum of three primes?” In other words, all his focus would have been on the universally quantified , with the existentially quantified being a sort of afterthought once the details of the proof have been worked out.
In general, many interesting results, when they are stated formally, begin with an alternation of two or three (or more) quantifiers. The question then becomes to determine which is the first “interesting” quantified variable in some sense. Here’s another example to illustrate the point, from the theory of finite-dimensional normed spaces. I’ll give a few mathematical details for those curious, but if you don’t care about those, then you can skip the next three paragraphs and should get the gist of what I am saying about this example.
Let and be two -dimensional normed spaces and let be a linear map from to . We say that is a –isomorphism if there exists such that for every . By rescaling we can always take to be 1, in which case we have that for every . If , then this tells us that is an isometry. In general, the Banach-Mazur distance between and is defined to be the smallest such that there exists a -isomorphism from to . It is easy to see that the logarithm of the Banach-Mazur distance is a metric on the set of isometry classes of -dimensional normed spaces. A less easy fact, but still not too hard, is that the resulting metric space is compact: in fact, it is known as the Banach-Mazur compactum.
It is natural to wonder what the diameter of the Banach-Mazur compactum is, and here things get interesting. A result of Fritz John states that every -dimensional space has distance at most from . (The idea of the proof is as follows: pick inside the unit ball of an -dimensional ellipsoid of maximal volume; that is the unit ball of a normed space that is isometric to ; it can be shown that the identity map is a -isomorphism between and .) From Fritz John’s theorem and the (multiplicative) triangle inequality, it follows that for any two -dimensional normed spaces. That is, the diameter of the Banach-Mazur compactum is at most . But might it be substantially less than that?
An indication that the answer is not obvious comes from looking at the spaces and . The identity map between these two spaces is an -isomorphism, but one can do much better by mapping the standard basis vectors not to themselves but to vertices of the unit cube, with the vertices chosen to be as orthogonal as possible. In particular, if there exists an Hadamard matrix, then the corresponding linear map is a -isomorphism. One can push this observation and deduce that for any the Banach-Mazur distance between and is . It is also easy to show that , so -spaces hardly improve on the easy lower bound, and do not improve on it at all in dimensions for which an Hadamard matrix exists.
In 1981, Gluskin famously solved the problem by determining the correct asymptotics for the diameter of the Banach-Mazur compactum. Informally, what he showed was that the diameter is within a constant of the upper bound that follows immediately from Fritz John’s theorem. If we make the quantification explicit, then the statement we end up with is
,
where I have written for the set of all -dimensional normed spaces. (If you want to argue that it is not a set, then let me specify in addition that the underlying vector space is .) In words, there is a positive constant such that for every positive integer there are -dimensional normed spaces and such that the Banach-Mazur distance between and is at least .
I can’t continue without very briefly describing the beautiful and highly influential idea Gluskin had for solving this problem. He took and to be normed spaces whose unit balls were random symmetric convex sets defined as follows: take the standard basis vectors and a handful of other random unit vectors, as well as the negatives of all these vectors, and take the convex hull. Gluskin then showed that if two normed spaces are chosen from this distribution, then with high probability their Banach-Mazur distance is at least .
But back to the main point, which is that the logical form of the above statement is very similar to the logical form of Vinogradov’s theorem, which is
where I have written for the set of primes. And yet, Vinogradov’s result is unquestionably a theorem, while Gluskin’s result is unquestionably a counterexample, or at least an example.
What is the important difference between the two statements? It seems to be that in Vinogradov’s three-primes theorem the number plays a more essential role in the statement that is to be proved about the various quantified variables. In Vinogradov’s theorem, that statement is , whereas for Gluskin’s theorem the statement to be proved is
and ,
which we can write equivalently as
and .
In the case of Vinogradov’s theorem, the whole challenge is to get those three primes to add up to , whereas for Gluskin it is not remotely challenging to get the dimensions of and to equal : the challenge is to get and to be very far from each other, relative to their common dimension.
There is a further complication to bear in mind here, which is that via the process known as Skolemization, a universally quantified statement of the form can be converted into an existentially quantifed statement . (For this to be an equivalence one needs the axiom of choice, but it is certainly a sufficient condition.) This is not just a piece of logical trickery, but it often reflects quite accurately how we think about some problems. For instance, it is more natural to think of Gluskin’s example as a recipe for constructing (or at least proving the existence of) a pair of suitable normed spaces for any given dimension , or in other words to construct a suitable function from to pairs of normed spaces by giving its value at each , than it is to think of it as a statement that says that every positive integer has a certain complicated property.
Yet another complication is that some universally quantified statements follow naturally from existentially quantified statements, or may even be equivalent to them. For example, the theorem that a 2-dimensional torus is not homeomorphic to a 2-dimensional sphere is a universally quantified statement (every map from the torus to the sphere fails to be a homeomorphism), but the natural way to prove it is to prove the existential statement that there is an invariant that distinguishes the two spaces. For an example of where a universal statement is equivalent to an existential statement, consider a statement of the form that a vector does not belong to the convex hull of a certain compact set . The statement that no convex combination of elements of is equal to is equivalent to the existence of a linear functional and a such that and for every . In both these cases it feels natural to regard the result as a theorem that is proved via an existential statement, perhaps because it is the theorem that is ultimately what interests us. But using “what interests us” as a criterion to determine what counts as a counterexample seems a little vague, and is a difficult criterion to use if we want to explain convincingly why AI should be good at finding counterexamples.
A more general argument against the notion that there is something about existential statements that is particularly suited to AI is that the need to establish existential statements pervades almost all of mathematical research, regardless of the nature of the headline result being aimed for. For example, if I want to prove a statement by induction, I may well look for a strengthening of the statement that serves better as an inductive hypothesis. Or if I want to prove that every object of type with property also has property , then I may well look for a property that follows from and can be used to prove . These are more metamathematical existence problems, but the distinction can be somewhat blurred, and more importantly, when trying to prove a statement , it is often the case that the main question in our minds is less, “Why is true?” and more, “What could a proof of be like?” To give an example, I feel I understand pretty well why Goldbach’s conjecture is true — a highly plausible probabilistic model of the primes implies it and agrees closely with computational data — but if I were making a serious attempt to prove it, that understanding, which many mathematicians have had for a century or so, would be of limited help. Rather, my main task would be to try to find proof techniques that were powerful enough to make those heuristic ideas rigorous.
What is the difference between an example and a counterexample?Logically, every statement of the form is a counterexample to the universally quantified statement . However, we do not describe all existential statements as counterexamples. For example, if I were to say, “The -spaces with are all separable, as is , but is not separable,” I would not describe the second part of that assertion as a counterexample to the claim that all Banach spaces are separable. Rather, I would present it as probably the most basic example of a non-separable space. The important point seems to be that there was no particular reason to think that all Banach spaces would be separable, and finding an example of a non-separable space is not very difficult.
I think the first point is more important here: we are more inclined to call an object a counterexample if the existence of that object disproves a statement that we had quite good reason to believe. It often happens that after repeated unsuccessful attempts to prove a statement, mathematicians begin to feel that it has no particular reason to be true, even if it seems to be hard to come up with a counterexample to it. In such a situation, if a counterexample is eventually found, it may have lost something of its “counter” feel. My impression is that the construction of a non-sofic group comes into this category. There have been several proposals in the literature for how one might construct such a group, and I don’t think there were many (or even any?) experts who strongly believed that all groups were sofic. So it feels more natural to say, “OpenAI came up with the first example of a non-sofic group” than to say, “OpenAI found a counterexample to the soficity conjecture” (despite the fact that that section of their paper is entitled “A counterexample to the soficity conjecture”).
Likewise, it seems to me that the new lower bound for multicolour Ramsey numbers is more of an example than a counterexample. I think quite a lot of people believed that the bound should be exponential, so for them it was a counterexample, but others, myself included, were more neutral about it. As a matter of fact, I have worked on the problem in the past (a long time ago) in an equivalent formulation, which asks how many triangle-free graphs on vertices you need if you want their union to be the complete graph . If you take bipartite graphs, then it’s easy to see that you need of them, but that bound can be improved if instead you observe that a complete 5-partite graph can be written as a union of two triangle-free subgraphs, and therefore it is possible to write the complete graph as a union of triangle-free graphs. It is then tempting to try to do better, with triangle-free graphs that are less dense but that make up for it with unbounded chromatic number — a necessary condition if one wishes to use a sublogarithmic number of graphs, which is equivalent to showing a superexponential lower bound for . All this is to say that when I worked on the problem, my efforts were concentrated on what turned out to be the right direction, so for me OpenAI found an example of what I (weakly) expected, rather than a counterexample.
Where does this leave us?I would like to find a coherent explanation of the conjunction of the following facts.
- The most notable mathematical results proved by LLMs have tended to be ones that we would classify as examples or counterexamples, where counterexamples are, broadly speaking, existence statements that disprove statements that we expected to be true.
- Many statements can be formulated as existence statements when we would usually think of them as universal statements, and vice versa, so what we consider to be an example depends on the mathematical context of a statement as well as its logical form.
- LLMs are pretty good at proving universal statements as well: it’s just that the strongest statements they have proved that we would think of as theorems have mainly not been at the level of the strongest statements that we would think of as counterexamples.
Given these facts, it seems likely that what LLMs are good at is something else, which happens to have as a consequence that they are good at the kind of existence problem that we would normally classify as asking to find a non-trivial example.
Let us consider two things that we can be confident that LLMs are good at. One of them is knowing a lot of mathematics: if a problem can be solved by means of a relatively standard argument, it is highly likely that an LLM will be able to find and use that argument. The other is the ability that an LLM has simply by virtue of being a computer: it can work at huge speed (compared with humans at least) and can therefore afford to make a large number of unsuccessful attempts at a problem before it finds a solution.
Without even looking at what LLMs have actually managed to solve, one might guess that these two features would lead to their having a somewhat different style from human mathematicians. Very roughly, LLMs would have the edge when there is more of a probabilistic element to the proof-finding process: they would be good at problems for which the best method is to try a lot of ideas, not necessarily particularly novel, until at some point you get lucky. Humans on the other hand would be better (for the moment) at finding more “surprising” and “conceptual” arguments, where the appropriate method is to dig deeper and deeper into a problem until the solution reveals itself. (It is hard to say exactly what this means, but I hope that any experienced researcher reading this will know what I am talking about.)
This raises two questions: does the guess above correspond at all to the reality that we are observing, and is there any reason to suppose that what I have tentatively described as the “LLM style” of doing mathematics would lead naturally to LLMs discovering several counterexamples (or just examples) to long-standing conjectures, even if that was by no means all they could do?
I don’t pretend to have a scientific answer to either question, but the reactions of experts to several of the remarkable solutions that ChatGPT has found do lend some support to the idea that LLMs work in more of a try-lots-of-things-till-you-get-lucky way. People often seem to react by saying something like, “Initially I was amazed that the problem had been solved, but on closer inspection I realized that the approach was actually not all that novel, and one that with the right small hint a suitably expert human could have found quite easily.”
For the second question — whether the LLM style is well suited to finding (counter)examples — I think matters are less clear, because there are many ways of searching for a counterexample, and some of them fit better than others the style I have described. Here are a few general methods. (I don’t claim that the list is exhaustive.)
- Look for an off-the-shelf example. Here one has a stock of fairly standard examples and one simply tries them out one after another to see whether any of them fails to satisfy the given statement. For example, Ryan O’Donnell ends his wonderful book on the analysis of Boolean functions with some tips, one of which is, “If you have a conjecture about Boolean functions, test it on dictators, majority, parity, tribes (and maybe recursive majority of 3). If it’s true for these functions, it’s probably true.”
- Build an example from basic examples and standard construction methods. For an algebraic problem, for instance, one might start with some standard examples, but then take products or quotients or limits.
- Make heavy use of metavariables. The word “metavariable” comes from computer science, and in particular from automatic theorem proving, and refers to the practice that in mathematics would correspond to writing, “where is to be chosen later,” (in which case is the metavariable). In a paper we usually do this only in fairly simple situations such as when we need to choose a number that is small enough for later arguments to work. But when we search for an example of an object that satisfies some property (which may well be a conjunction of simpler properties ), it is often not a good strategy to specify completely and only then to check whether it satisfies . Instead, it can be more fruitful to do almost the opposite: we start by saying virtually nothing about and simply launch into proving that it satisfies . In the course of doing so, we find that we need to satisfy a property . If we are lucky we can describe in a nice way a very general class of objects that satisfy . For instance, we may be able to find a parametrized class: we identify some function and show that satisfies for every of a certain type. The problem is then reduced to finding such that $Q(f(y))$ holds, which is a more specific version of the original problem. There may be many iterations of this process, or a mixture of this process and other processes, before an example is eventually found.
- Try to prove the opposite. If one wishes to find such that , it can be surprisingly helpful to start by attempting to prove the statement . The reason this can be helpful is that using our standard methods of attempting to prove something, we may end up identifying a key lemma that would suffice: that is, we may find an intermediate property that implies in a non-trivial way and thus reduce the problem to . Turning things round again, it may well then be that finding a counterexample to is easier than finding a counterexample to (that is, an example that satisfies ). Of course, there is no guarantee that a counterexample to will be an example of , but sometimes we are lucky and it is. More often, we can use the idea of the previous method, noting that it is at least a necessary condition of an example of that it should not be an example of , so one can try to describe a general class of objects that fail and in that way reduce the problem.
- Successive approximation. Sometimes, when we are searching for an example of such that , we write down a moderately plausible guess not because we think it has a chance of working (if we did, then we would be using the first strategy), but because we hope that if does not satisfy , then we will be able to diagnose what went wrong and specify a new guess that does not have that defect. Again, this strategy can either be iterated or combined with one or more of the other strategies.
- Just-do-it proofs. Sometimes we need to satisfy infinitely many properties , each of which is, individually, quite easy to satisfy. In such situations, we often “build” inductively bit by bit, ensuring at the th stage of the process that however the building process continues, will satisfy .
- Pick a random example. Often it is very hard to give an explicit example of an that satisfies , but there is a natural probability distribution for which one can show that if one chooses randomly from that distribution, then with high probability (or at least non-zero probability) it will satisfy .
- Pick a generic example. In more infinite contexts, it may again be quite hard to give an explicit example of an that satisfies , but one may be able to show that the set of that fail is or measure zero, or is a meagre set, or is small in some other way.
There is no particular reason to suppose that LLMs would be equally good at each of the methods above. So perhaps what we are observing is not quite that LLMs have a particular ability to find examples, but more that they are particularly good at finding examples (and proofs) in a certain way. Looking at the above techniques, one might imagine that they would be very well suited to checking off-the-shelf examples, finding just-do-it proofs (since that is a rather standard method with lots of instances in their training data), using the probabilistic method (unless, as often happens, significant new ideas are needed to show that the probabilities work out), and picking generic examples. The other three methods described above — use of metavariables, trying to prove the opposite, and using successive approximation — require more of an ability to judge whether the approach one is taking is likely to be fruitful. Here it seems at least possible that humans will sometimes have an advantage, but the conditions that a problem would need to satisfy are quite stringent. One would need an example to be one that lies at a leaf of a very large search tree — too large to be searched for by a combination of moderate mathematical ability and brute force — but that can be found by a mathematician with a sufficiently good nose for when they are making progress that they can prune the search tree very substantially.
Why wouldn’t LLMs also have that “nose”? I don’t rule out that “nose” is an emergent property of the way LLMs are trained, and that within a year or two they will have it to the same extent that we have it. But for now, in my interactions with ChatGPT, I do have a distinct impression that they haven’t got there quite yet. When I discuss an open problem with 5.6 Pro, I am often presented with approaches that sound promising until I think about them carefully, and then seem quite a lot less promising. And they will also often end a response by saying, “I have not managed to answer the question you asked, but have managed to reduce it to the following much narrower and more precise question,” which sounds very promising until it has happened five times without any obvious progress having been made. It isn’t completely obvious how they will get better at this, since their training data will not be full of examples of fruitful and less fruitful directions to pursue when trying to solve problems: all they will typically see is tidied up proofs that hide the thought processes of their discoverers. Of course, human mathematicians also don’t get to learn much about how to do research from the experience of other mathematicians, and yet we somehow manage to pick it up. But the situation is a little different for us, in that a lot of what we learn is by doing rather than emulating.
Another reason it is not obvious that “nose” is a property that emerges naturally when LLMs are scaled up is that if LLMs make heavy use of their broad knowledge and can afford to do a lot more brute-force search than humans can, then they will lack the incentive that humans have to prune the search tree ruthlessly. It could conceivably be that their successes so far are achieved using methods that for a human would be considered extremely inefficient, but that because of their superior speed and knowledge, the combinatorial explosion these methods will lead to has not yet become apparent.
It would be very interesting to try to test this experimentally, but it is also difficult, because if an LLM has what looks like the kind of idea that could only be the result of “deep thought” about a problem, we can never be sure that it has actually carried out that deep thought, as opposed to finding a model argument already in the literature, or in other words exploiting the deep thought of a human mathematician. It would probably be easier (but still not easy) to test it by using models that are less powerful than the latest ones and that have been to some extent shielded from the mathematical literature: one could give them a carefully designed suite of problems and see whether the ones that the LLMs solve have particular characteristics.
It may seem as though I am desperately clinging to the hope that humans will continue to be able to make meaningful contributions to mathematical discovery for a while yet, but while I do indeed hope that, I am not making any assertions of the form “LLMs will never be able to do X”. I think it is likely that they will, and given the pace of progress over the last three years it will probably happen quite soon. But I do think that there may be a hurdle for LLMs to clear and it seems at least possible that it won’t be cleared as straightforwardly as some of the previous hurdles.
In that connection, it would also be interesting to see whether a different reward structure leads to LLMs being able to solve different kinds of problems. For example, if during training an LLM (or machine-learning system of some other kind) is not just rewarded if it ends up with a solution, but also penalized if it explores too many dead ends or if it “cheats” by getting the answer from the literature, perhaps it would be incentivized to go about the research process in a more human way and thereby achieve better results for classes of problems where it is yet to make a big impact.
If the hurdle is cleared, either by pure scaling up or by some more thoughtful method, it will be quite difficult to know when that has happened, since, as just mentioned, an idea that seems very original and surprising may just be lurking somewhere in an LLM’s training data. But I would be confident that it had been cleared if an LLM were to come up with a proof that was as surprising to me as the solution of the cap-set problem was in 2016: the previous best known bounds were completely eclipsed, the method was utterly different from anything I had thought about trying, and afterwards there was a flurry of activity as people came to understand what this wonderful new technique was capable of.
ConclusionI wasn’t quite sure where I would end up when I started this post, and now that I’ve got to the end, I feel that my main conclusions are not particularly new or surprising, but I hope that the route to them is of some interest. The main points I have made are the following.
- “Finding an example” is in practice not the same thing as proving a statement that begins with an existential quantifier.
- If it is true that current models are particularly good at finding examples, that is probably not because they have a particular affinity for existential statements, but more because the proof-discovery methods that are appropriate for finding certain kinds of examples play to the obvious strengths of LLMs: wide knowledge and the ability to explore many paths of the search tree that humans would judge to have a low probability of success.
- It seems likely that LLMs will carry on improving very quickly. However, if, contrary to expectations (mine at least), there turns out to be some residual class of problems (or other mathematical activities) for which humans continue to have the edge for a while, it is likely that those will be problems for which the mysterious human ability to prune the proof-discovery search tree is particularly advantageous: that is to say, problems where the search tree is deep and has a large amount of branching, so that without rigorous pruning a search is not feasible even for a computer.
- A good sign that LLMs have reached human level for a much wider class of problems will be if they start proving theorems using methods that, like much of the very best human mathematics, are new and surprising but that with hindsight come to seem beautiful and natural. They should also be methods that are difficult to stumble on by accident. It is hard to say precisely what would count as such a proof, but I think we’ll recognise it when we see it.
A partial digestion of the HRT counterexample
A function of one variable can be translated in space by a spatial shift to obtain a new function
and also modulated in frequency by a frequency shift to obtain a new function One can compose these two operations to obtain a time-frequency shift: As per the time-frequency uncertainty principle, the two shifts do not quite commute with each other. For instance, we have As such, is not a representation of the abelian group , but rather a portion of the Weyl representation of the Heisenberg group, but we will not adopt a representation-theoretic perspective here.Some functions obey finite linear relations between their time-frequency shifts. For instance, a sinusoid obeys the relation
However, the Heil-Ramanathan-Topiwala (HRT) conjecture states that once one imposes some reasonable decay condition on , no such relations exist:Conjecture 1 (HRT conjecture) If is non-zero, then there is no relation of the form for some distinct time-frequency shifts and some coefficients , not all zero.
A special case of the HRT conjecture, which was also open, makes the additional assumption that was Schwartz.
Many positive results towards this conjecture were known. I will mention only a few here. A simple case is when we only have frequency shifts rather than spatial shifts:
In this case, the operator is simply a physical space multiplier with symbol so that the relation (2) now takes the simple pointwise form If the coefficients are not all zero, then is a non-zero analytic function and thus has isolated zeroes. It is thus not possible to solve this equation for any . Thus the HRT conjecture is true when all the time-frequency shifts lie on the vertical axis. Using the metaplectic representation, one can then handle the case when all the are collinear.What about the non-collinear case? Suppose first that all the lie in the lattice , thus for some integers . Here, the phase shift in (1) disappears, and all the time-frequency shifts commute with each other. This suggests that it should be possible to diagonalize the situation with a suitable transform to convert (2) to a pointwise equation similar to (3). To find this diagonalization, observe that if one restricts the function to a coset of the integers, then just multiplies the function by the scalar , while shifts the function on this coset by . The latter translation operation can also be converted to pointwise multiplication by performing the Fourier transform on the integers. Thus, if one introduces the Zak transform
of then the equation can be transformed after a brief calculation to the equation where the symbol is now given by the formula As before, if the coefficients are not all zero, then is a non-zero analytic function and thus non-zero almost everywhere. Thus has to vanish almost everywhere, which for can be used to show that also vanishes.More generally, there is a result of Linnell that the conjecture is true if lie in a translate of a discrete subgroup of ; this (together with the argument handling the collinear case) establishes all cases where , and several partial results involving the cases are also known. The conjecture is also known if is decays at a suitably super-exponential rate, by work of Bownik and Speegle.
I was aware of this conjecture through various talks and conversations with colleagues, and even briefly tried my hand at it for a while, though not with particularly serious effort (or progress). It was thus a nice surprise to see that it has just been resolved by Faulhuber, Petersen, van Velthoven, and Voigtlaender, even in the Schwartz case:
Theorem 2 There exist complex numbers , not all zero, distinct points , and a non-zero Schwartz function such that
It is perhaps unsurprising that this result is AI-assisted. However, I think the authors have disclosed their AI use responsibly, with the final arguments written by hand with a readable overview of the argument, as well as proper discussion of methods, relation to past literature, and other independent numerical checks on the result.
The negative result lies only a little beyond the positive results: is now increased to , and all but one of the points lie in (a translate of) a discrete subgroup of (in fact the explicit subgroup is used). The functions constructed are smooth and rapidly decaying, but not analytic or super-exponentially decaying, which would start being in conflict with the known positive results.
In addition to AI being used to come up with the initial proof strategy, a more traditional numerical computation was used to verify one step of the argument.
I have not had the time to do a full digestion of the result, but (after reading the introduction, and using a little AI assistance of my own) I was able to understand the main ideas at a high level. The first few reductions are relatively standard. Setting and , one can view the problem as one of solving an eigenvalue problem
The time-frequency shifts are chosen to lie in a translate of the discrete subgroup by a certain irrational shift . As mentioned previously, if in the shifts of a standard lattice , it would be natural to work with the Zak transform of , but it turns out that the approach does not quite work when doing this for topological reasons (relating to the fact that scalar quasiperiodic functions of mean zero are forced to have zeroes), and so the authors used the slightly denser lattice instead , which relates to a vector-valued version of the Zak transform taking values in rather than . Here, the phase shift in (1) does not completely disappear, but becomes a sign change. This slight loss of abelianness means that we cannot hope to diagonalize the problem all the way to a scalar problem, but we can still hope to reduce it to a two-dimensional vector-valued problem. Indeed, by applying a suitable vector-valued version of the Zak transform, the eigenvalue problem can be transformed to a a “vector cocycle problem” where is a non-zero smooth quasiperiodic vector-valued function, is an irrational shift , and is a certain explicit matrix-valued function depending on the choices of , , and .How to solve this equation? The motivating scenario here is if the matrix function was replaced by a rank one function
for some smooth vector-valued function of unit magnitude. Then one could solve the equation by taking and . It is not possible to make the function exactly of this form, but through some numerical computation and clever AI-assisted guesswork, the authors were able to find a choice of and , and that made approximately equal to a rank one function of this form, in fact getting a uniform estimate As it turns out, such an approximation is sufficient to run a contraction mapping argument to find a solution to a variant of (4), namely for some smooth and . (Here it was important to get the operator norm bound below ; they are barely able to do this, with a numerically obtained bound of , though this bound might not be optimal.)The main remaining obstacle is that the “eigenvalue function” is varying in the parameter rather than constant. (This issue was, by the way, anticipated to some extent in previous work of Demeter, who observed that eigenfunctions of the almost Matthieu discrete Schrödinger operator gave a near-miss counterexample to the HRT conjecture, but with an eigenvalue that depended on an auxiliary phase shift parameter rather than constant.) However, if one was able to solve the scalar cocycle equation for some smooth , then one could solve the equation (4) by setting . The approach to solve (5) is standard: take logarithms, apply a Fourier transform, and then divide out by the multiplier associated to the shift. This can cause a well-known “small divisor” problem (which arises in various dynamical contexts, such as in the KAM theorem) if behaves too much like a rational vector, but the standard resolution to this is to select a shift that obeys good Diophantine approximation properties. For the purposes of numerics the authors selected an extremely concrete shift, namely
but I get the impression that the exact choice here was not crucial for the argument, and that many other irrational algebraic numbers could have worked here.
Thoughts about the Leiden Declaration
Last September I went to a workshop at the Lorentz Centre in Leiden to discuss mathematics and AI with historians, philosophers, computer scientists, AI researchers, and mathematicians of several different flavours (though there was a surprising preponderance of algebraic geometers). The whole event was extremely stimulating, with some talks but also a lot of time set aside for discussion. One of the concrete outcomes of the workshop was the Leiden Declaration, which has now been signed by over 3000 people. Given that I was part of the workshop, it might seem a bit strange that I am not one of the signatories of the resulting declaration. The reason is not so much that I disagree with it in any concrete way, but more that in several places it makes confident assertions and recommendations that I feel somewhat uncertain about. So instead I prefer to try to articulate my views about the issues raised by the declaration and put them in this blog post. Before I do that, I would like to make clear that I am very glad that the Leiden Declaration exists and I think that it has done a lot of good in focusing people’s minds on the issues that AI is forcing the mathematical community to grapple with, which are more acute now than they were last September.
Let me begin by quoting a passage from the declaration that sets out “what we take to be characteristic values of mathematical research that we have a joint interest in preserving”.
- There are many reasons to pursue mathematical research, ranging from intellectual curiosity to a desire to solve practical and societal problems. Underlying much of mathematics is the activity of proof. Mathematical proofs are regarded as conferring the highest degree of certainty to their conclusions, as well as imparting understanding of why their conclusions are true. These characteristics of proof support the scientific integrity of mathematics.
- Results are attributable to specific authors who take credit for their discovery and assume responsibility for their correctness. These principles ground the merit-based standards to which we aspire in mathematical research.
- Mathematical arguments are regarded as transparent and subject to independent verification. They may be extremely long or difficult, but in principle no proprietary knowledge or equipment should be required to understand them.
- Mathematicians share a concern for proper evaluation of mathematical work relative to shared standards of depth, difficulty, and significance.
- Mathematics produces not only a body of results, but also understanding, clarity, and judgment among the communities of mathematicians who have shaped them, often in the context of their own autonomously guided research. This expert knowledge is essential, both to effectively use mathematics, and to continue to articulate new and significant research questions. A key source of strength of the discipline has long been the autonomous shaping of the direction of research and the methods used to pursue it.
The first thing I would say about these values is that they are undoubtedly values that are widely held by mathematicians, including, with some qualifications, me. The main qualification I have concerns point 4: I find the notion of “proper evaluation” somewhat problematic, given that different mathematicians can have very different judgments without either of them being clearly wrong, especially when it comes to the significance of a piece of mathematics. Also, these judgments are used for purposes such as the acceptance of papers in journals, hiring and promotion decisions, the awarding of prizes, and so on, that are part of a system that copiously rewards a few people — I myself have hugely benefited from it — but doesn’t necessarily adequately reward a lot of people who are doing less visible work that is essential to keeping the whole enterprise going.
But the more important point is whether these values are ones that we should fight for in the future, as the Leiden Declaration suggests. I find that clearer for some of them than others. For example, it seems to me that the importance of rigorous proof will be even greater in an AI age than it was before — if the output of AI is not underpinned by rigorous proof, then the kinds of difficulties one already hears about with certain areas of human mathematics (see for example many talks by Kevin Buzzard arguing for the value of formalization) would be hugely magnified. But what about the attribution of results to specific authors, who take both credit and responsibility for them? Suppose that at some point in the future AI becomes more autonomous, reading the literature and solving many problems that it finds. Suppose also that its solutions are autoformalized, so there is no serious doubt about their correctness. In such a situation, there would be nothing for a human to take credit for or responsibility for. Does that mean that we should declare such results undesirable and threatening to mathematical values?
Of course, something could well be missing in such a situation: perhaps the proofs would be badly written and hard to follow, which would mean that they lacked something we all very much value. So let me extend the thought experiment slightly. What if by that stage one could take one of these outputs and ask an LLM to explain the ideas, and what if LLMs did a very good job at that? That is not particularly hypothetical, since they are often pretty good at this job already, but I am imagining a world in which they are much better than they are now, as they will presumably become.
So now we would have a world in which a lot of problems had been solved, we were sure that the solutions were correct, and we had an LLM ready to explain those solutions in as much or as little detail as we wanted. Is that a future we should resist, and if so, why?
One obvious reason is that it would take a huge part of the fun out of the subject. It is extremely satisfying to struggle with a mathematical problem for months or even years and eventually solve it. But I worry about that argument, because it seems to be saying that we should resist doing mathematics the easy way because a tiny fraction of the world’s population gets huge pleasure from taking orders of magnitude longer to do it. That is not to say that I wouldn’t be sad that a way of life that has sustained me for the last forty years was not available any more — of course I would. I just find it hard to use it as a reason to argue that we should try to preserve the “ownership structure” of mathematical results. If we arrive at a world where mathematical theorems are no longer associated with mathematicians, maybe that won’t be any more problematic than the fact that stars aren’t named after astronomers and most aren’t named at all. I’m not necessarily in a hurry for that world to exist, but maybe once the transition had happened, people would be OK with it.
The third value I share in an uncomplicated way, and I have already discussed the fourth. The fifth value is one that I hold very strongly, though I’m not so keen on the idea of experts consciously “shaping the direction of research”, something that I see as happening more organically. Obviously there are some notable examples of mathematicians who have created wonderful programmes of research, but even there I would like to credit other mathematicians with understanding what is wonderful about those programmes and contributing to them enthusiastically as a result, rather than being told what direction to pursue and meekly doing so (which is probably not what the declaration is actually trying to suggest, but it has a slight flavour of that for me).
But that’s a minor quibble when set against my main worry about the effect of AI on mathematics, which is the possible destruction of mathematical culture. There is at the moment an extraordinary body of knowledge and expertise that exists not just in the mathematical literature but in the heads of mathematicians all round the world. Imagine if AI didn’t exist and a pandemic broke out that for some reason wiped out all mathematicians and nobody else. All the literature would still be there, but nobody would have the faintest idea what to do with it. To revive a mathematical tradition under those circumstances would be extremely difficult and take decades. Now imagine a slight variant of that, where AI does exist and because of it people are no longer motivated to put in the years of effort it takes to reach the level of expertise that a typical research mathematician has now. After a decade or two, we might arrive at a situation where the mathematical literature has, in some form, been vastly expanded, but there is no corresponding community of human experts who have a shared understanding of parts of it. Almost all of mathematics would be like the areas that we have more or less forgotten about today, areas that exist in papers written many decades ago that nobody reads any more. (I won’t name any such area because I don’t want accidentally to suggest an area that many people still love and work on.)
This, it seems to me, is a possibility that we should try very hard to resist, but I agree with many other commentators who say that in order to resist it, we will need to give less priority to some of our current values — and I would include ownership of mathematical results in that list — and more to others. For example, if Person A gets an LLM to one-shot a solution of an important open problem (which is formalized, possibly automatically, so there is no doubt about its correctness) but Person B makes the effort to digest the solution and explain it in a way that other mathematicians can understand and learn from, then I think we will want Person B to get the lion’s share of the credit. The credit would be of a slightly different from what it is now, which could be described as admiration for somebody’s talent, insight, speed (I mean here the purely factual statement that speed is often admired — I would prefer that to be less the case) and hard work. It would be more like the gratitude that one feels already for somebody who writes a beautiful textbook that makes a whole area of mathematics coherent and accessible.
Maybe that is what the “research mathematicians” of the future should do: make a selection from a vast sea of AI-generated mathematics and write a book about it in such a way that other mathematicians can read the book and feel the kind of enrichment that we feel when we get to grips with an area of mathematics.
At this point I have to admit that there’s a pessimistic side of me that asks the following general question whenever anyone says anything about what the role for humans might be in the future: why do you think that AI wouldn’t be able to do it? For example, with the suggestion I’ve just made, what reason is there to suppose that ChatGPT 8.2 wouldn’t be able to have a short interaction with you about your mathematical tastes and background and then write the ideal textbook just for you? Humans are likely to be better at this kind of curating for a little while yet, but is it a fundamentally human ability that AI could never hope to emulate?
In a world where AI wrote bespoke textbooks (or more likely, just taught people in some more direct way), something would be lost that feels important: mathematics as a collective endeavour. If we all just learnt cool bits of maths for our own private satisfaction, we would miss the considerable pleasure that comes from discussing mathematics with others, though even that could in principle be restored by a benign LLM that deliberately taught many people the same cool bits of the subject, though an LLM that could do that sort of social engineering would raise all sorts of safety issues.
Let me now turn to the section of the declaration about potential threats. I’ll put my comments on each one in square brackets.
- Current automated techniques can produce plausible but unreliable (or even incorrect) arguments which are difficult to distinguish from correct mathematical proofs. This applies not only to informal arguments, but also to formalizations, where the difficulty lies in the translation between computer-encoded and human presentations of concepts. These fast-moving developments put our present system of review under increasing pressure, jeopardizing our ability to implement traditional standards for the correctness, transparency, and independent verifiability of proof. [This feels like less of a problem now than it did last September, partly because the best LLMs hallucinate a lot less than before, and partly because autoformalization is improving all the time — I have just used harmonic.fun’s Aristotle system to formalize a complicated paper in Lean and I didn’t need to know any Lean to do it.]
- Technologies that draw extensively on the published mathematical commons undermine the traditional system of attribution. Models trained on published works frequently return outputs that do not properly cite the human works they synthesize. Many current models are also built on data obtained by systematically exploiting licenses and access arrangements that were not made with artificial intelligence in mind, or indeed by simply violating copyright protections. [This is a problem at the moment, when ownership of results is important, and I am very much in favour of people making an effort to give appropriate credit for mathematical ideas that AI may have used. However, in the longer term, as I have already discussed, I think this ownership structure will break down and the issue will become less important. It also seems possible that LLMs will become better at revealing their sources.]
- Technologies which affect the way in which mathematics is practiced may disturb the current system of incentives. The use of artificial intelligence — and thus also the sort of problems which it can address — may become incentivized for its own sake, disrupting our mechanisms for hiring, funding, and recognition. This disadvantages researchers who do not have access to the technologies or decision-making related to them, or who are unwilling to use technologies controlled by organizations whose values they do not share. [These seem to me to be genuine problems. I think there is simply no point in hoping that our current system of incentives will not be disturbed — it obviously will. I am not necessarily too worried if our mechanisms for hiring, funding and recognition are disrupted, as I don’t find those mechanisms unproblematic as they are, but disadvantaging researchers who do not have access to good LLMs is something I certainly think we should worry about.]
- Proper evaluation is endangered if results are communicated through informal channels such as press releases or blog posts, often without any research paper or other disclosure of information necessary for scientific evaluation. This practice seeks publicity for new results on market timelines before the accepted processes of community evaluation in mathematics can take place. In many cases this leads to simplifications in reporting, such as overemphasizing the significance of automated tools and undervaluing the prior human contributions which have made those tools possible. Such oversimplification risks influencing public opinion in a way that not only damages perceptions of mathematics, but also misleadingly uses specific mathematical tasks as metrics for the general reasoning capacities of commercial products. [I think this can be a problem, but I think it is not as serious a problem as some of the others, since when results get overhyped, there seems to be no shortage of people publicly (and rightly) pointing that out.]
- These developments put the autonomy of mathematics under threat. The increasing involvement of technology companies in mathematical research raises the risk that research questions may come to be prioritized because of their amenability to automated mathematics, rather than expert judgment of their deeper significance. Indeed, broader understanding of the field may be permanently lost in the process of automation. With university budgets under pressure, this reshaping also changes professional incentives in a manner which encourages the collaboration of researchers with technology companies on asymmetric terms. If left unchecked, these trends go beyond threatening researchers’ autonomy, affecting the scope and depth of mathematical research itself. [I think this could be a problem, but it also seems to me that mathematicians have a lot of power here. For instance, if a technology company were to produce a lot of research that mathematicians did not find all that interesting or important, I don’t think they would be able to use their financial and other resources to persuade us to change our minds. Rather, what seems to happen is that mathematicians say, “Yes that does X but it doesn’t do Y,” and the tech companies then feel challenged to do Y.]
There follow eleven recommendations for individual mathematicians. I agree with almost all of them. The one that I’m not so sure about, for reasons I’ve basically already gone into, is this.
Affirm the humanity of authorship. Credit and responsibility continue to belong to humans within the mathematical community and should not be given to automated systems. Artificial intelligence may obscure, but does not replace, the collective human labor behind a result.
I’m not sure what that really means. For example, should we affirm the humanity of authorship in the case of the solution to the unit-distance problem? Some humans did a wonderful job of explaining the proof that OpenAI’s model came up with, and the model made use of some highly non-trivial mathematics produced by humans, but the solution itself has not been credited to any human, and nor should it be in my view.
Under recommendations for mathematical organizations and not-for-profit research funders I again agree with several of them but have my doubts about some. An interesting case is the following.
Protect the rights of authors. Automated mathematics presents new challenges to the rights of authors, and societies should be proactive in the development of sample licensing agreements to protect these rights. In particular, material should not be used as training data without consent, and publishing agreements should allow authors to opt-out [sic] of the use of their work in this way.
This recommendation seems to belong to a world in which journal articles are the main means of dissemination of mathematics. But that has long since ceased to be the case: almost all dissemination now takes place via arXiv preprints, with journals limited to providing a little extra mark of prestige. Once an article is on arXiv, it is on the internet and one can hardly ask for it not to be used as training data. So this recommendation, if it applies at all, will apply to a tiny fraction of articles that are published without first appearing on arXiv. More generally, what right of an author is being compromised when an article is used as training data? We don’t object if human mathematicians use our articles to help train themselves to become better mathematicians — indeed, we will typically be delighted that somebody else thought our articles worthy of their attention. So the objection to a machine doing the same would have to be that for some reason one did not want machines to get better at mathematics in a similar way. I can imagine grounds for such a wish: perhaps somebody is worried about the threat that LLMs pose to traditional mathematical practice, or perhaps they worry that mathematical ability of LLMs will transfer to much more dangerous reasoning ability. But there’s a more complicated discussion to be had here than one might think from reading the recommendation.
The next recommendation is this.
Insist on appropriate publication outlets. Demand that mathematical results continue to be published in peer-reviewed venues such as journals, proceedings, and books. Informal mechanisms such as press releases or blog posts can provide a valuable supporting role, but they cannot replace peer-review or community scrutiny.
For reasons that I’ve gone into many times, I am not too fond of the current publication system, so I can’t get behind this recommendation. Indeed, if the current system becomes unsustainable because of a flood of AI-generated and AI-aided content, I would regard that as a beneficial consequence of AI. However, that doesn’t mean that I would advocate a total free-for-all. I’ve already said that one of my worries is that if mathematical content is not sufficiently organized, then the traditions that we all value could die. I just think that what we will want to do to preserve those traditions is likely to be a lot more innovative than clinging on to the peer-reviewed journal system.
I have highlighted in this post the parts of the declaration that I have doubts about, either because I disagree with them or, more typically, because I sort of half agree with them but want to add many qualifications. That may make the post come across as rather negative, but that is not my intention. The parts I disagree with are in the minority, and I think it is important that a declaration such as this should be made. I should also make clear that my views are evolving all the time, largely because the speed of progress of LLMs has taken me by surprise, but also as a result of conversations I have had or opinions that other mathematicians have expressed online.
I’ll end with two further clarifications. The first is that it may seem as though I am taking it for granted that LLMs will soon be better than humans at all aspects of mathematical problem solving, and maybe also problem posing, theory building, formulation of definitions, etc. I do think all that will happen at some point, but whereas some people say that it will obviously happen within the next two to three years, I would say that it might happen as soon as that, but I don’t rule out that we’ll get lucky and find that we can do interesting AI-assisted maths for quite a bit longer than that before AI doesn’t need us any more.
The second is that I think I have acquired a reputation as somebody who celebrates what is going on. But if, for example, I post on Twitter saying that such-and-such an AI solution is a remarkable development, the word “remarkable” is meant to indicate no more nor less than that I found it very surprising. My feelings about the possibility of AI solving all sorts of problems that interest me are much more mixed. I’ve had the experience twice now of seeing GPT 5.6 Pro one-shot a solution to a problem that I very much liked and had thought about hard (in both cases with much younger collaborators, who, with my approval, were the ones who prompted the LLM). It felt very strange and not particularly pleasant to have the rug pulled out from under my feet like that. On the other hand, I was quite pleased to see the problems solved. It’s actually a similar feeling to the one I have had many times when a problem I am fond of and have thought about gets solved by another human mathematician.
Another factor for me is that I have invested a lot of thought into automatic theorem proving of a more traditional kind. One of my main motivations for that was the hope that the work I put into it would extend the state of the art, measured by which problems a computer can solve. That ship has sailed now, and that saddens me. I still think that there is value in the work that I and my group are doing, but it has become a tougher sell.
So I personally have already found AI quite disruptive, and this is just the beginning. I would have preferred the developments to happen at a slower pace. But I don’t see any practical way to slow them down, so the best we can do is probably to face up to the changes that are being thrust upon us and do what we can to maximize the benefits and minimize the damage. The Leiden Declaration may not be perfect, but it makes an important and positive contribution to that effort.
A digestion of the Jacobian conjecture counterexample
The notorious Jacobian conjecture can be formulated concretely over the complex numbers as follows.
Conjecture 1 (Jacobian Conjecture) Let be a polynomial map in complex variables, whose Jacobian is a non-zero constant. Then is invertible (with polynomial inverse).
The condition that the Jacobian is non-zero is equivalent to being locally invertible. (The implication of local invertibility from non-vanishing Jacobian follows from the inverse function theorem; the converse implication can be derived from the Weierstrass preparation theorem, but is omitted here; see also Lemma 5 of this previous blog post.) Also, from the fundamental theorem of algebra, once the Jacobian polynomial is non-zero, it must be constant. So the hypothesis “Jacobian is a non-zero constant” can be replaced with “ is locally invertible”. So the Jacobian conjecture can be viewed as an assertion that local invertibility implies global invertibility. The complex numbers can be easily replaced with other fields of characteristic zero by the Lefschetz principle, but I prefer to work in the concrete setting of the complex numbers.
It was recently shown (using the Fable AI) that the conjecture is false in three dimensions (and thus in higher dimensions as well):
Theorem 2 (Counterexample to conjecture) There exists a polynomial which has non-zero constant Jacobian, but is not invertible.
The conjecture remains open in two dimensions, and is easy to establish in one dimension.
The example can be stated completely explicitly: one can take
and one can verify by a brief calculation that and While this is an extremely quick verification, the construction presented in this fashion appears like a massive miracle. The polynomial has degree seven, so a priori the Jacobian ought to be a polynomial in three variables of degree as large as , so the fact that all non-constant coefficients of this polynomial vanish looks like a massive cancellation involving equations, which is much larger than the degrees of freedom for a generic degree seven polynomial map of three variables. So finding such a polynomial looks highly unlikely to be located by brute force.The example has since been retroactively explained in more geometric terms. As a “digestion” exercise to myself, I sought to write this explanation with relatively little use of algebraic geometry, in a manner that minimizes the amount of “miracles” required, although there are still a few places where some remarkable phenomena occur.
It is convenient to use the local injectivity formulation, and to generalize the domain to an equivalent affine variety. Namely, we will show
Theorem 3 (Counterexample, reformulated) There exists an affine variety that is isomorphic to by polynomial changes of variable, and a polynomial map which is locally injective, but not globally injective.
Clearly one can get from Theorem 3 to Theorem 2 by composing with the isomorphism and using the previously mentioned fact that local injectivity implies non-zero constant Jacobian. Our objective is now to find data , that obeys three separate properties:
- (a) is locally injective on .
- (b) is not globally injective on .
- (c) is isomorphic to by polynomial changes of variable.
(A pedantic remark: strictly speaking, in the arguments below, we not only replace the domain of by an equivalent variety , but also replace the range of by an equivalent variety . But the equivalence between and is a boring linear isomorphism ( will just be a hyperplane in a four-dimensional vector space ), so we do not highlight this aspect of the construction.)
It turns out that and can be built out of the operation of multiplication of low degree polynomials. Namely, consider the following three simple affine spaces:
- The space of linear homogeneous polynomials of two complex variables .
- The space of quadratic homogeneous polynomials of two complex variables .
- The space of cubic homogeneous polynomials of two complex variables .
The map , essentially a map from to , is clearly polynomial; it is given explicitly in coordinates as
The map also enjoys two basic (and commuting) symmetries:- If one applies a scaling for some non-zero complex numbers , then the product is scaled by : .
- If one applies a change of variables for some invertible linear transformation , then the product is transformed by : .
The five-dimensional domain is of course larger than the four-dimensional range , so the map clearly cannot be injective. This can already be seen from the scaling symmetry, as the specific scalings
for modify the linear and quadratic polynomials but not their product . But even if one quotients out by this symmetry (3) to cut the dimension of the domain down to four, the map is still not injective for the following basic reason. A generically chosen cubic polynomial will split into the product of three independent linear polynomials. Then there are three pairs which all map to the same cubic polynomial under the multiplication map , but are not related to each other by scaling symmetry (3). Thus, we see that even after quotienting out by the scaling symmetry (3), the multiplication map is generically non-injective in a three-to-one fashion. Thus we already have achieved something resembling goal (b)!It will be convenient to “spend” the scaling symmetry to obtain a useful normalization. If is a linear polynomial and is a quadratic polynomial, the (homogeneous) resultant can be defined by the determinant
If we have a factorization then the resultant can also be described as Thus the resultant measures whether the linear polynomial and the quadratic polynomial share a common root. A fundamental fact about resultants is that they are -invariant: for any , we have One way to see this is to check it first for translations (which translate the roots by while leaving unchanged) and for inversions (which map to while mapping to and respectively), and then noting that these transformations generate all of . They also interact very nicely with scaling: In particular, the scaling symmetry (3) multiplies by : Thus, we can (generically) normalize away this scaling symmetry by imposing the conditionWe now have a restricted multiplication map (which by abuse of notation we will continue to call ) from the four-dimensional variety
to the four-dimensional space . This map is still not globally injective, as we can take the three pairs in (4) from before and apply the scaling (3) separately to each of the three pairs to obtain the normalization (7). So we have kept property (b). Furthermore, this map retains the -equivariance (and also one remaining scaling symmetry, though we will not make much further use of that symmetry).But we now also have property (a)! Suppose we want to show the local injectivity of in the neighborhood of a pair with . As the resultant is non-vanishing, the root of (which exists in the Riemann sphere, or projective line if you prefer) is distinct from the two roots of (though the latter two roots could be equal to each other). Applying the action (which performs Möbius transforms on the roots), one can assume without loss of generality that is the point at infinity (or equivalently ), thus for some complex number and for some complex numbers , with the resultant condition (7) simplifies to (so in particular are also non-zero). It is then clear that if one perturbs and by a small amount (say, modifying each coefficient by ), then the root of will perturb to something large (), while the roots of stay bounded. Thus, just from knowledge of the product , one can reconstruct which of the three roots of this cubic polynomial will be the perturbed root of , and which two will be the perturbed roots of ; from this and (6), (7) we can also reconstruct the leading coefficient of , and this completely determines both and . This establishes the local injectivity property (a). (In fact it is étale, but we will not need the machinery of étale maps here.)
Unfortunately, (the four-dimensional analogue of) condition (c) fails: the quadric hypersurface (8) is not isomorphic to the affine space . But we can try to get around this by passing to a three-dimensional slice. Let be some three-dimensional affine plane of (which we will take to avoid the origin for technical reasons), then we can restrict as a map from the set
to . is clearly identifiable (by linear changes of coordinate) to . As was already locally invertible, it remains locally invertible under restriction; and because generic cubic polynomials had three preimages under in (8), this continues to be the case after restricting to (9) (unless was somehow so degenerate that it had no generic elements, but this turns out to be impossible). So we have retained properties (a) and (b). The miracle is that, with a good choice of , we can also obtain (c) and obtain the desired counterexample to the Jacobian conjecture: despite appearances, the variety (9) is in fact equivalent to the affine space by polynomial changes of variable!Let’s see how. The affine hyperplanes in avoiding the origin are parameterized by the dual space of avoiding the origin, which one can think of as the non-zero third order homogeneous differential operators in two variables. Indeed, every such operator generates an affine hyperplane that avoids the origin, and conversely by duality every affine hyperplane avoiding the origin arises in this form uniquely. Just as the cubic polynomials in can be factored into three linear polynomials, the differential operators in the dual space can also be factored into three linear differential operators, e.g.,
in the case that is non-zero. The action moves the roots around the Riemann sphere by Möbius transformations. As these transformations are -transitive, the actual selection of such roots is not too important (and the scaling symmetry similarly makes the choice of leading coefficient unimportant); the only thing to keep track of is whether the roots repeat. Up to the symmetries, there are in fact just three different equivalence classes of differential operator (and thus of affine hyperplane ) to consider:- Operators where the three roots are all distinct, thus for independent first-order operators .
- Operators where two roots coincide and one is distinct, thus for independent first-order operators .
- Operators where all three roots coincide, thus for some first-order operator .
It turns out that the affine miracle for (9) occurs precisely in the second case, when has two identical roots. I do not have a completely satisfactory geometric explanation for this miracle, but one can verify it by the following coordinate computation.
By applying the action, we can normalize so that , thus is now the affine hyperplane of cubic polynomials with . Using (2) and (5), the variety (9) can now be described explicitly in coordinates as
At first glance this seems to be a generic-looking variety cut out by a cubic equation and a quadratic equation – hardly a candidate to be affine! But observe that if is non-zero, then the second equation can be solved for , and the first equation can be solved for , Putting these two equations together, we see that as long as one removes the case , the quintuple is uniquely determined by by a change of variables which is Laurent in and polynomial in . Thus we have a nice birational equivalence Thus we have already almost established property (c): the variety (9) becomes birationally equivalent to after cutting out the subvariety. In particular, for each fixed non-zero value of , the corresponding fiber of (10) is equivalent to by polynomial changes of variable, since we can reconstruct from the coordinates by the polynomial formulaeSo we just need to glue back in the fiber. Indeed, from (10) we see that the fiber at is just
Now we observe a key miracle: the cubic equation and quadratic equation have a unique affine solution (as opposed to the six possible solutions that Bezout’s theorem might suggest – the other five solutions live on the line at infinity). So the fiber here is also affine: This is extremely encouraging for the purposes of establishing property (c), as it strongly suggests that the variety (10) has the structure of an -bundle over , which is already extremely close to being isomorphic to the affine space . The main remaining task is to make sure that nothing singular happens in the limit , and that a global polynomial coordinate chart for (10) that covers both the and fibers can be constructed.The standard way to proceed here is to manipulate various tangent spaces using the modern machinery of algebraic geometry and commutative algebra, but given my own background, I prefer to adopt the language of analysis, and in particular big-O notation (in place of the ideals used in algebraic geometry), in order to investigate the limit by hand. On the variety (10), let us use to denote any multiple of by a polynomial expression in . Thus, for instance, the equation implies that
while the equation implies that as well as the more refined estimate In the case we could conclude that . Now we perturb this observation. Multiplying (13) by we have , which on substitution into (14) gives ; substituting this back into either (13) or (14) also gives .We can get some more precise asymptotics by also taking advantage of (15). Substituting into (15), we obtain after some algebra
So if we write more explicitly as , then we have and thus Substituting this back into (11) gives an asymptotic for : Finally, one can insert these estimates into (12), although one only gets a trivial bound in this case:Expanding the error term in (16) as , and doing a little more algebra, we thus have a polynomial change of variables
which completely parameterizes the variety (10) by polynomial combinations of three coordinates . This already gives (c) and thus completes the proof of Theorem 3.The previous computations, when expanded out, also gives polynomial inverse maps:
The map from to the coefficients of (dropping the coefficient which is constrained to equal ), we obtain a polynomial map with which theory predicts to have a constant Jacobian, and indeed one can calculate that the Jacobian is . This is essentially the original example up to trivial changes of variable; indeed, one can check that the map is exactly the map given in (1).AI disclosure: I used an AI chatbot to discuss various aspects of this problem and to confirm several of the calculations made here.
Two more apps: visualizing the zeta process and the motions of the heavens
I believe that the creation of visualization apps to illustrate mathematical or scientific concepts is a particularly favorable use case for modern coding agents, as many of the downside risks attached to other LLM use cases are limited:
- Not mission-critical. As such apps are not authorative sources of truth and only used for secondary purposes, a small positive error rate in the output can be acceptable.
- Stand-alone. As the applets are not destined to be incorporated into a larger codebase or literature, the technical debt incurred by delegating all the coding to an LLM agent is bounded.
- End product is deterministic (and sandboxed). As the applets run on a deterministic language (Javascript), are sandboxed against file or internet access, and do not make any LLM calls at run-time, security and privacy concerns are minimal, and the applet can be maintained without continued premium LLM access or resource-intensive compute.
- Not replacing primary skills. While deskilling is the tradeoff one accepts when relying on these tools to accelerate output, I am perfectly willing to forego the opportunity to keep my Javascript skills at a high level, as this is a tertiary skill for me at best in my chosen profession. (I continue to manually program in Lean and in Python to keep in practice with programming in general.)
- Not competing with humans. To my knowledge, there is no existing human effort that is being duplicated by these applets (the activity in this direction appears to have peaked two decades ago).
I would however caution against unrestricted LLM use when one or more of the above five favorable situations is not in effect.
With these points in mind, I have used such an agent to create two further apps. The first app illustrates the “zeta process” that was introduced in my recent paper with Alexeev, Barreto, Li, Lichtman, Price, Shah, and Tang, though it was first discovered by an AI. For each , the zeta distribution is a random natural number with distribution
It has long been known that this distribution has good number-theoretic properties: for instance, the number of times a given prime divides has a geometric distribution of mean . However, the new observation is that these random variables can be chained together into a single stochastic process, which we call the “zeta process”, which is an infinite divisibility chain. I used an agent to create an app to visualize this process:
The underlying process is generated by several exponential random variables at each prime: in the above instantiation of the process, two such variables are visible at the prime , and one variable at the primes . At a given choice of , is formed by collecting all the variables below this threshold (and for which all predecessors also lie below the threshold); in the above illustration, this amounts to one variable at each of the primes , leading to in this case. Additional visualizations in the app display the distribution of each , as well as the distribution of the hitting probability , which among other things can be used to give a quick solution to Erdős problem #1196.
The second app is rather different in nature, and is a somewhat whimsical attempt to display the motion of the heavens, both at “human” scales of space and time, and at more “astronomical” scales (in which the motion of the planets in particular are more apparent). It is very loosely inspired by the game “Katamari Damacy“, in which one absorbs both terrestrial and celestial objects of many different scales. Here is how the app typically looks at a human scale:
And here is how it looks when one’s perspective leaves the Earth’s atmosphere:
(As I did not want to render an entire explorable world in this app, the observer in the app is only limited to changing his or her size, from a human to a creature of comparable size to the Earth itself; they cannot move horizontally on the planet.) At the largest scales of space and time, the classic orrery diagram appears:
After lengthy conversations with the agent, I was able to implement many astronomical phenomena, including phases of the Moon, the effect of Earth’s rotation against the fixed stars (though one can also stabilize one’s view against those stars to see the Earth’s rotation more directly), and so forth.
