Sakupljač feed-ova
ICIAM Statement on Mathematics and Artificial Intelligence
Just a quick post to note that the International Council for Industrial and Applied Mathematics (ICIAM) has released a statement on mathematics and artificial intelligence (as well as a longer version), which makes many points echoing several already made recently here and elsewhere.
The longer statement also makes reference to a recent statement by the London Mathematical Society on recent developments around the Navier-Stokes equation. Perhaps the comments to this post can also be used to report other institutional statements on these topics.
Recommendations of the Summit on PhD Math Education in the Age of AI
[This is a guest post by Bryna Kra and Rachel Ward. -T.]
The Summit on PhD Math Education in the Age of AI was held September 17–18, 2026, jointly hosted by the Harvard Department of Mathematics and the Center of Mathematical Sciences and Applications (CMSA). It brought together twenty-four senior mathematicians from across the discipline, along with several current PhD students and postdocs, to examine which aspects of the mathematics PhD need rethinking as AI reshapes thinking-based work.
This document is a first draft of our recommendations; we expect to refine our recommendations over the coming months according to changing landscapes and your feedback. We encourage feedback in the form of comments on this post, or by sending email to education-summit@cmsa.fas.harvard.edu.
To graduate students and early-career mathematicians: We know this uncertainty is weighing heavily on you as you plan your next steps. Your concerns matter, and we want to hear them. We are gathering information, keeping communication open, and working to support your opportunities and your future in mathematics.
We’re gonna need a lot more mathematicians
[This is a guest post by Amit Sahai. This blog post was initially written in a different file format and converted using AI. — T.]
When I was an undergraduate student, I remember talking with several students who felt that the pace at which the top students could understand new math concepts was far too fast for them. They, too, could understand the ideas, but it would take them much longer. Eventually, almost all of these students gave up their dream of pursuing research mathematics and found something else to do. I have been thinking about those students a lot in the last few days.
The research mathematics community consists largely of those of us who either rarely felt that way, or who felt it and managed to overcome it through hard work. We have had the good fortune to find a place in mathematics where we could make progress. But we are now entering a time for humility: a time when all of us are going to know what it feels like to be unable to keep up.
The AI systems I have worked with are already producing beautiful new ideas. They are doing far more than impressive calculations or quickly carrying out arguments that a strong human researcher would already understand. And we probably can’t even imagine the wonderful ideas that future systems will be capable of producing.
When we feel that we cannot keep up, will we take that as a reason to leave research mathematics, like the students I am remembering? As more of us experience this, there will undoubtedly be a temptation to draw the same conclusion as they did: If the machines can move so much faster than us, perhaps we should find something else to do.
For our community to give up the work of understanding would be a profound abdication of our responsibility to humanity. Each of us is entitled to choose a different life. The responsibility I am talking about belongs to us collectively: to build a future in which humans can understand and contribute to the discoveries that will change our world. A future with meaningful human agency.
Struggle is essential to understanding difficult concepts. Fortunately, this struggle can be shared. I have been blessed to experience this time and time again with my students and collaborators. Imagine a multitude of research groups, each with sustained support, each spending a term or a year trying to understand an extraordinary set of ideas produced by an AI system, with the help of AI systems. [1]
This may very well be among the most important mathematical work in the years to come, and we should support and prioritize it accordingly. This enterprise will require a significant expansion in the number of mathematically sophisticated human researchers available world-wide, as major breakthrough ideas accumulate.
Why should society want this? So far, this might sound like a utopian fantasy for us – a civilization focused on depth of human understanding, awash with mathematicians and physicists and the like. I would certainly love to live in such a world. And indeed there are deep philosophical reasons for society to move in this direction. But I think society has a much more immediate stake in making this possible, too.
Imagine that a future AI system proposes a radically new design for a one terawatt nuclear fusion power plant. It has found a way to sustain and control fusion that no human had conceived of. The design promises abundant, inexpensive, clean electricity. Robots stand ready to manufacture the components and build the plant.
A terawatt is an insane amount of electrical power. We would be deciding whether to construct a machine that handles extraordinary flows of energy using principles we have never conceived of, let alone put into practice. We would need to understand how failures can be contained, what happens to energy already stored in the system when it shuts down, how we can be sure that the materials that make up the power plant behave as expected, and what other questions we should ask before proceeding. The very novelty that makes the proposal exciting would mean that we cannot inherit confidence from decades of operating similar plants or experiments.
Before approving construction, I would want communities of humans to understand why the design works and what justifies confidence in its safety. I would hope that we all would.
Human involvement does not automatically improve a technical decision , and I see no reason to insist that humans manually repeat work an AI system might be able to perform more reliably, even including proving mathematical guarantees. But a theorem can only exist within a model. Understanding the guarantee means understanding the model, the experimental evidence for it, and our uncertainties about the accuracy of the model. This is demanding work, and mathematically sophisticated people must be available to engage with it.
One might respond that AI systems should handle those questions too, and ultimately decide whether the plant should be built. That is a serious position. But it asks us to accept a future in which decisions of enormous consequence rest on reasons that no human community understands.
I do not want us to arrive at that future simply because we failed to invest in our own capacity to understand. Human agency is a value of fundamental importance. We must retain the ability to meaningfully consider alternatives and decide what kind of world we want to be a part of building. I think it is worth the effort. [2]
To take on this responsibility, we may need to broaden our view of what a mathematician can contribute. I have in mind something like a “deployable intellectual reserve”: communities of mathematically sophisticated people that humanity can call upon to help understand consequential AI-enabled breakthroughs.
Our ability to understand difficult and unfamiliar ideas may become one of the most important contributions we can offer to society. We should be willing to bring that skill to problems far beyond our usual research interests. [3] Doing so asks us to expand our sense of our vocation.
A counter-argument might be that AI systems will make each of us so much more effective that fewer people could do this work, even as the pace of discovery accelerates. But each of us is merely human. We have fundamental limitations based on our biology. Depth of understanding needs time and a pace of life that humans can sustain. Each individual human can only be asked to do so much, but through earnest cooperation we can accomplish much more.
If AI fulfills its promise, we will encounter more beautiful and consequential ideas than we have ever seen. We must respond by building thriving human communities that can understand them together.
We’re gonna need a lot more mathematicians.
The ideas and opinions presented here are entirely my own, but GPT 6 Astra was instrumental in helping me draft this note. I also thank my former student Dakshita Khurana, my current student Isaac Hair, my colleague Terence Tao, and my family members Anant Sahai and Gireeja Ranade for valuable feedback. Note that there is much more to be said here, but I tried to keep this relatively short to focus succinctly on my primary thoughts.
Notes
[1] By this, I do not mean to imply that only AI-created results will be of interest in the future. But for major results generated by humans, we already have a tradition of spending extended periods of time studying them.
[2] And the relevant understanding cannot belong only to the organization proposing the technology. Imagine a public hearing at which the company’s experts are the only people capable of following the technical argument. Independent expertise is critical.
[3] Indeed, AI systems are likely to be very helpful in allowing researchers with diverse backgrounds to talk effectively with one another, and more generally understand unfamiliar concepts.
Headlines and inside stories: understanding and trust in AI for mathematics, science, and engineering
[This is a guest post by Tapio Schneider. This blog post was initially written in a different file format and converted using AI. — T.]
[This will be cross-posted on the CliMA blog.]
The apparent proof of finite-time blow-up of the forced Navier-Stokes equation, announced by OpenAI on September 8, has brought into focus a debate about the role of AI in mathematics. The proof was produced with a system of some 10,000 AI agents that explored many approaches in parallel; it was then formalized and verified in Lean. The formal verification supports its correctness, but mathematicians are still working to digest it. A proof settles the truth value of a statement, but Terry Tao and others have argued that this is only part of the point of proofs; the other part is to advance our conceptual understanding of “basic structures of shapes, numbers, and natural phenomena.” As Yehuda Rav put it a quarter-century ago, “theorems are the headlines, proofs are the inside story.” A proof that is correct but incomprehensible, or undigested by the mathematical community, gives us the headline without the story. The issue is not whether the prover is a human or a machine, but whether the reasoning advances our collective understanding of methods and structures, which can spur new thinking and further advances. As Timothy Gowers has noted, the inadequate AI-generated write-ups of proofs are likely a temporary annoyance.
I want to argue that the same point about understanding holds in the natural sciences and engineering, where it also has intrinsic value and, in addition, acquires instrumental value when predictions must be trusted before they can be verified empirically. One role of science is what ancient philosophers called episteme, roughly explanatory understanding; here, understanding is the goal itself. Another role is techne, roughly the craft of predicting and making: forecasting how natural or engineered systems will behave under circumstances not yet observed. Whether the two roles can be separated depends on how easily predictions can be checked. When they can, techne can stand on its own. A black-box AI weather prediction model can predict tomorrow’s weather, and we can trust it because its forecasts can be checked every day. Techne can then also serve episteme as an instrument. For example, AlphaFold predicts 3D protein structures without providing explanations. But its predictions can be checked against experimentally determined structures, and they have become invaluable for understanding how drugs bind to their targets. Episteme and techne become inseparable when predictions must be acted upon, and hence trusted, before they can be empirically verified, because verification is too slow, costly, or dangerous, as in projecting climate change decades ahead or designing an aircraft. In those cases, an auditable causal chain from assumptions and input data to the predicted outcomes is what makes predictions trustworthy, and this constrains how AI can be used.
The Navier-Stokes equation, which describes fluid flow, is a good example, as it is at the heart not only of a Millennium Prize problem but also of aircraft design and climate prediction. The claimed result says that a fluid starting from rest, driven by just the right smooth stirring, develops an infinite velocity spike in finite time while its total kinetic energy remains bounded; the forcing is constructed to sustain the collapse. (The unforced version of the problem, whether smooth solutions exist for all time without external forcing, remains open.) Despite the mathematical singularity, nothing infinite happens physically: the incompressible equation stops being valid once energy concentrates at very small scales, where compressibility and molecular effects take over, as has long been known. By contrast, in turbulence, viscosity dissipates energy at a small but finite scale (the Kolmogorov scale), and the continuum description is valid across all scales of motions. Therefore, the result says little about how we model, predict, and understand the physical phenomenon of turbulence, which is likewise a solution of the Navier-Stokes equation, and one that, unlike finite-time blow-up, is ubiquitously realized.
The Navier-Stokes equation governs the flow and turbulence that control lift and drag around an aircraft wing, cloud formation in the atmosphere, and mixing in the oceans. In both aircraft design and climate prediction, verification comes late or is difficult and expensive. An aircraft is flight-tested only after it is built; a projection of how extreme rainfall statistics intensify in the coming decades must inform stormwater infrastructure that is built now but will likely be put to the extreme test only decades later. Such predictions earn trust not by end-to-end verification but by being the output of an auditable chain stretching from inputs (design parameters, atmospheric composition) to outcomes, whose links have known limits of validity and can be tested individually. Computational fluid dynamics (CFD) solves the Navier-Stokes equation together with additional equations (e.g., for thermodynamics) numerically. Its numerical methods for the resolved scales are based on established theories of stability, consistency, and convergence, and its subgrid-scale models for unresolved turbulence rest on plausible assumptions of universality at small scales (e.g., local isotropy) that can be tested separately against high-resolution simulations or laboratory experiments with canonical flows. Understanding and auditability of the individual links are what allow the chain as a whole to be trusted beyond the distribution of large-scale cases already observed, such as the present climate or existing aircraft wing designs.
Contrast that with end-to-end AI methods for weather prediction. Daily empirical verifiability suffices to establish trust in their forecasts. However, they do not come with a stability, consistency, and convergence theory (no analog of von Neumann stability analysis or the Lax equivalence theorem has been established for them), and how an initial condition becomes a forecast is difficult to audit. This is inconsequential for forecasting a few days ahead, where the forecasts can be checked daily. But it does matter for climate projection, which is a different task: predicting how weather statistics change over decades in response to a forcing such as increased greenhouse gas concentrations. An end-to-end model can learn the day-to-day evolution of weather states in today’s climate, but it contains no pathway through which greenhouse gases alter this evolution; their concentrations are typically not among its inputs, and if they were, no observations exist to learn about the response. The causal chain from greenhouse gases to their effects on radiative transfer, temperature, and winds, represented in physics-based models through their equations, is absent. Additionally, such models do not enforce conservation laws, for example of energy, so long integrations can drift by accumulating errors (e.g., in temperature). We would not currently trust them to project how the climate system responds to previously unobserved changes in the concentration of greenhouse gases over decades. Nor would we trust an AI surrogate of CFD simulations to certify an aircraft of novel shape. AI surrogates are used to explore and narrow down design spaces quickly, but the final assessment returns to established CFD methods or wind tunnel experiments, whose errors are controlled and whose steps are auditable.
So how do we best use AI in cases where episteme and techne combine, where understanding is essential for trust because predictions are difficult to verify? An answer suggested by the preceding argument is to embed AI at the links in the chain where trust can be earned by empirical verification. Concretely, use AI inside auditable scaffolds, such as physical conservation laws. Then use numerical methods with controlled errors to solve, e.g., the Navier-Stokes equation on the resolved large scales, while learning closure models for the unresolved subgrid scales from data, where some universality assumptions are defensible and individually testable. The rationale, in the case of the climate system, is that the large scales are where climate change moves the system out of the current distribution, whereas small-scale physics obeys the same local laws in a warmer or colder climate as in today’s. For an aircraft wing of novel shape, similarly, the geometry may be new but the functional relation between larger-scale conditions and the small-scale turbulence around it is not. The scaffold thus guides the extrapolation on the basis of the known equations, rather than with end-to-end models tied to the data available today.
Closures must be learned as functions of the resolved state, and extrapolation problems may reappear if the conditions for which data are available do not span the range of conditions for which predictions are needed (e.g., the temperature and humidity regimes in which cloud turbulence has been sampled, or the pressure gradients along a wing of novel shape). This can be mitigated with established tools, such as local high-resolution simulations for offline calibration and uncertainty quantification. Climate models and CFD codes have been built from resolved dynamics with embedded closures for decades. AI now enables a wider and faster search. Neural network closures can be individually tested and, to some extent, interpreted. Symbolic regression, which learns equations by selecting terms from a dictionary, as in SINDy, can sometimes produce closures that are more easily interpretable. AI agents are beginning to run the closure search in a closed loop, proposing symbolic closures, testing them against high-resolution simulations, observations, or experiments, and revising them. When this is successful, the resulting closure can be audited and, after the fact, understood.
This closed loop is similar to how AI is used in mathematics. In both cases, a verifier is used for checking: Lean for proofs; high-resolution simulations, observations, or experiments for turbulence closures. In the case of mathematics, the checks are complete; in the case of turbulence closures, they are limited to the conditions covered, so trust is restricted to the tested conditions. AI can accelerate this program by searching a larger space of possible closures or proof strategies than can be explored by humans in the same time. Once candidates have passed the checks, they become new objects for human study. In this way, understanding (episteme) and prediction (techne) improve together and reinforce each other, with humans, for now, remaining essential for extracting understanding and building trust in the end result: for writing the story behind the headline.
I thank Thomas Müller for pointing me to the paper by Yehuda Rav, and Thomas and my CliMA colleagues for discussions of the topics here over several years. I used AI for copy-editing.
Why I agreed to join AGMAI
[This is a guest post by Martin Hairer, cross-posted from Proofs and Prompts. — T.]
There has been a lot of speculation regarding the “Advisory Group on Mathematics and Artifical Intelligence” agmai.org since it was announced on Monday. In this short blog post, I would like to explain in a bit more detail how the group works, what we are trying to achieve, and why I agreed to be part of this group.
First of all, let me lay down some facts that seem to have been drowned in a torrent of misinformation. The most important one is that we are genuinely independent of OpenAI and any other of the so-called `frontier’ AI labs. Yes, the original impetus of forming this group was OpenAI’s response to the statement A Severe Misalignment of AI in Mathematics, but not all members of the group were contacted by OpenAI in the first place, there is nobody besides the nine of us taking part in our conversations, we do not receive any monetary compensation, and our technical support (so far mainly regarding IT, legal, and communications) is being provided by the IAS. We also have not signed any documents placing any constraints whatsoever on what we state in public and whose advice we seek, besides the obvious confidentiality requirements. The other fundamental fact is that our overarching aim is to represent as well as we possibly can the interests of the mathematical community. This is a tall order and we may end up doing a terrible job at it (I certainly hope not!), but I am absolutely certain that all nine of us take this extremely seriously and are acting in good faith.
I am acutely aware of many of the pitfalls of being part of such a group. At a most basic level, the mathematics community consists of people with a very broad range of views and opinions (often strongly held ones!), and the nine of us clearly don’t form a representative sample. As an independent body our role is purely advisory, so the AI labs can choose to simply ignore it. It would also be very naïve to believe that the AI labs won’t try to spin whatever we say in a way that suits their PR machine, which dwarfs anything we could possibly come up with. Finally, given that the unprecedented situation we find ourselves in has the potential of impacting the lives of so many people (I intentionally do not use the word `career’ since for many of us mathematics is so much more than just a career), whatever statement we make and / or advice we give will necessarily upset a sizable fraction of the maths community. So why on earth would I accept to be part of this?
Here are some of the reasons that swayed my mind:
- The recent releases, on privately controlled websites, of high-profile mathematical results combined with abject levels of scholarship and only minimal presentation effort has been undermining the usual standards of scientific practice. If we refuse to answer when being asked how to do better (irrespective of whether that ask is being done in good faith), how could we then complain about this behaviour?
- It is obvious that AI has already had a profound impact on mathematical research and raises numerous questions of correct attribution of ideas, priority, human understanding of ideas, etc. Many aspects of this revolution will be addressed within our community, but it makes no sense of taking a hard `ostrich’ approach of simply ignoring the AI labs and refusing to talk to them on principle.
- It is perfectly legitimate to question OpenAI’s motivations, but the fact is that they have now made an effort (as belated and minimal as it is) to engage with the mathematical community. Turning them down without even trying to engage with them would be the easiest way for them to get a PR win painting the mathematical community as a bunch of out of touch luddites.
- In the `Severe Misalignment’ statement, which has now been endorsed by nearly 8000 mathematicians, we conclude by saying that “These issues must be addressed urgently, in the mathematical community, by the companies developing these technologies, […]”. It would be rather hypocritical of myself to then chicken out at the first opportunity of actually having a chance of setting up some form of communications channel between the mathematical community and the AI labs just because I could get some flack for it.
Regarding the actual advice we intend to provide, it is a fact that AI companies have in recent months been producing some high profile mathematical results and that their publication and dissemination has been falling far short of acceptable mathematical practice, whichever way that is defined. However, the situation at hand is sufficiently unprecedented that we genuinely haven’t completely made up our minds yet on the best way forward and we fully intend to gather as much feedback as possible, be it through our feedback form, discussions with colleagues, the discussions on this blog, etc. We are going through a period of intense turmoil (and this doesn’t just concern mathematics) but I am heartened by the many intense but thoughtful and respectful discussions I have experienced both inside my own institution and at gatherings of mathematicians like the ICM over the past few months. I am convinced that the mathematical community can emerge all the more united and stronger from this experience.
Open problems, open mathematics
[This is a guest post by Antonio Auffinger. This blog post was initially written in a different file format and converted using AI. — T.]
Despite writing papers in pure mathematics, much of my time in the past decade was spent talking (mostly listening, to be accurate) to biologists, computer scientists, and physicists. Theoretical physics and computer science are grounded in mathematics and thus share much of our language and, to some extent, our culture. The barriers between biology and mathematics are an order of magnitude higher. I am convinced that biologists are part of a different species. The language, incentives, culture, training, and everything else you might think of appear to be completely disconnected from the way we do mathematics. Yet, the similarities are abundant: the pursuit of understanding, the excitement of discovery when new patterns or phenomena emerge, and the hours needed to make minor advances, sometimes leaving us with just frustration. The intuition of setting up a new experiment feels like witchcraft to me, much like our way of finding connections between different abstract objects feels like magic to them.
What does talking to biologists have to do with the future of mathematics?
It could be beneficial to look at what is happening and what has happened with our neighbors, even if none of us can predict where we will be two years (or even three months) from now. A first lesson that I learned is that a large part of biology is technology-driven. Fields completely change, emerge, and die in a matter of years. The advent of CRISPR, RNA-seq, and cryo-EM, for instance, suddenly allowed humans to observe and manipulate phenomena that were previously inaccessible. These tools generate precious data, and the incentives often favor a culture of seclusion, where discoveries are frequently not shared until the final product is complete.
This has made me appreciate the culture of mathematics. Mathematics has never been free of competition or secrecy, but we abundantly share. We share ideas, we share problems, and we share entire skeletons of approaches with our colleagues, with visitors we just met, in talks, in public forums, and on YouTube. I suspect the enormous effort often needed to produce a proof, even with a full outline, has helped sustain this openness. Mathematicians and mathematics have deeply benefited from this open attitude. Many new connections and major discoveries have started with honest conversations at coffee breaks, in hallways, or on hiking trails. Sharing ideas or arguments before they are fully formed allows others to see pathways we missed, point out obstacles, or take the problem in directions we had not imagined. Also, it is simply more fun to do math together.
My worry is that this unselfish openness will become a thing of the past. If proof generation becomes a fast, accessible commodity while our ways of giving credit remain unchanged, mathematicians (especially those still building their careers) may feel encouraged to optimize locally in ways that weaken our culture of sharing. In the past few weeks, I have had colleagues reach out for advice and tell me they will no longer post on arXiv. I have witnessed trainees posting rushed papers online for fear that others could quickly carry out strategies already outlined in previous work. I am also part of the problem, as I have started advising my students to be extra careful when sharing work in progress.
Biology also offers examples of communities deliberately changing these incentives. During the Human Genome Project, the Bermuda Principles called for the rapid public release of sequence data. Later, the Fort Lauderdale Agreement tried to balance this openness with proper recognition of those generating the data, placing responsibilities not only on researchers producing and using the data but also on funding agencies. Although it is focused on data, it is an example where the status quo was changed by community intervention.
I do not know what the right analogue is for mathematics, but I believe we need to start thinking about it. I encourage the community to reshape our incentives and the way we give credit, to ensure broad accessibility to these new tools, and to make openness a reasonable choice, especially for those still building their careers. Senior mathematicians must engage in serious conversations about ethical use with their trainees. In turn, trainees should be encouraged to truly explore, because many of the solutions to the issues we currently face as a community will come from them.
None of this is an argument against the use of AI in mathematics. I believe these tools will raise the ceiling of the things we can discover, leading mathematicians to ask new questions and understand new phenomena.
This brings me to the second lesson I learned from my biology colleagues: many of the questions I hear from biologists need mathematics. New mathematics. There is room here for topologists, number theorists, dynamicists, algebraic geometers, analysts, etc. This is not just because of our capacity as proof builders, but also because of our capacity for abstraction and for understanding phenomena. Applied sciences are generating complex, time-dependent data that require new methods and theories. This is also a two-way street. Decades ago, topology and knot theory unexpectedly provided a framework for understanding how enzymes untangle and rearrange DNA. In the other direction, attempts to understand population genetics and how gene frequencies drift over time helped motivate new classes of infinite-dimensional stochastic processes, including measure-valued diffusions. Today, geometry and probability underlie dimension reduction methods that biologists use daily to make sense of massive single-cell datasets, such as t-SNE and UMAP.
I am not suggesting that mathematicians need to pivot to biology or any applied science. The pursuit of mathematics for its own sake is the absolute bedrock of our field and that must remain intact. However, as AI changes the landscape, looking outward presents an incredible opportunity to discover new questions, new phenomena and new mathematics, and perhaps AI might even lower the barriers to taking that leap.
Mathematics has much to offer the other sciences, and much to learn from them. I am looking forward to what comes next. Proofs may increasingly be generated by machines, but the ultimate purpose of mathematics remains exactly what it has always been: to ask questions and to understand.
Announcing the Advisory Group on Mathematics and Artificial Intelligence
[This is a guest post by the Advisory Group on Mathematics and Artificial Intelligence. This blog post was initially written in a different file format and converted using AI. — T.]
We would like to use this guest post to announce the creation of the Advisory Group on Mathematics and Artificial Intelligence hosted at the Institute for Advanced Study (Princeton) and online at agmai.org.
The rapid advances in artificial intelligence (AI) present both opportunities and challenges for mathematical research. We believe that we are at a historic moment for our discipline. Recent events raise urgent questions about how to support the long-term prospects for deep human understanding of mathematics.
Purpose. The purpose of this group is to advise AI companies on their interactions with mathematical research and with the mathematical community, including the responsible presentation and release of mathematical results. We seek to work for the best interest of mathematics and the mathematical community, and to serve as one possible channel of communication between mathematicians and the AI industry.
Independence, Transparency, and Accountability. This group operates independently of any AI company and members do not accept payment for this work. We will publish our recommendations to AI companies on this website. We are willing to offer such recommendations to any AI company whose models are likely to have a significant impact on mathematics. Although we will give advice, we do not have decision making power at any AI company, and the responsibility for the decisions made by any company will rest with that company.
Advisory Group Members
- François Charles (ENS-PSL)
- Camillo De Lellis (IAS, GSSI)
- Timothy Gowers (College de France, Cambridge)
- Martin Hairer (EPFL, Imperial College London)
- Nikhil Srivastava (Berkeley, Simons Institute)
- Ulrike Tillmann (Oxford, INI)
- Ravi Vakil (Stanford)
- Edward Witten (IAS)
- Melanie Matchett Wood (Harvard)
This group came together after OpenAI approached some of its members about establishing an external advisory board. In agreement with OpenAI, they decided to create an independent group and invite others to join.
Current Task. We are currently facing the very specific challenge of advising OpenAI on how to coordinate the release of a large number of significant results in mathematics that they report have been produced by their internal model.
We welcome input from the mathematical community on this question. Please use this form to share your thoughts with us as soon as possible. Your responses will be used to inform our recommendations and will not be made public without your approval.
247A, Notes 1: Rearrangement-invariant spaces
Disclaimer: due to current events, I have not been able to devote as much time to lecture notes preparation as I would have liked, so I apologize in advance for the unpolished nature of the text below, which has been largely recycled from previous lecture notes I have written.
This is the first set of lecture notes for my graduate course 247A, “Fourier analysis”. The course name is rather general, but I will focus the course not on the Fourier transform per se, but on the closely related topic of real variable harmonic analysis, with a particular emphasis on Calderón–Zygmund theory, which underlies basic tools in PDE such as the theory of Sobolev spaces.
To avoid confusion at the outset, let us make the distinction between real-variable harmonic analysis and abstract harmonic analysis, which are only distantly related to each other despite the similar names. Abstract harmonic analysis, roughly speaking, is the extension of the classical theory of the Fourier transform to other domains, such as locally compact abelian (LCA) groups, non-abelian Lie groups, or symmetric spaces, and typically involves a blend of representation theory, group theory, and analysis. Real-variable harmonic analysis, by contrast, tends to work on classical domains, such as a Euclidean space , a torus , or a lattice , although many of the techniques can extend to more general domains (e.g., to Riemannian manifolds). While the Fourier transform often plays a prominent role (in particular, by setting the stage for time-frequency analysis and enabling various decompositions or other transforms that involve frequency space or phase space in addition to physical space), real-variable harmonic analysis is often focused on estimating other transforms or expressions that often interact well with the Fourier transform, but need not explicitly invoke it. Examples include the Hilbert transform (where we have made the somewhat arbitrary decision to omit the normalizing constant ) or the Hardy-Littlewood maximal function
A typical question in harmonic analysis is the following: let be some function on a standard domain (such as Euclidean space), and let be an explicit transform of (e.g., the Hilbert transform or maximal function ). To what extent is the “size” of controlled by the “size” of ? The value of such bounds often lies in the general nature of the input function ; some mild regularity or decay hypotheses might be imposed on , but beyond that the function is typically not required to have a very structured form (in particular, it need not be describable by any closed-form expression).
In many situations the transform being studied is linear or sublinear, in which case the natural type of bound to ask is a linear bound
for suitable function space norms (e.g., norms), and is some bound. Depending on the application, we may be interested in various levels of precision regarding the bound :- (a) Optimal bounds, in which we seek the exact optimal value of (i.e., the operator norm of ). For instance, the optimal constant for the Hilbert transform is exactly , whereas the optimal constant for in one dimension turns out to be (a result of Melas).
- (b) Bounds accurate up to absolute constants (or maybe constants that can depend on basic parameters such as the ambient dimension).
- (c) Bounds in which we are willing to accept “logarithmic type losses” such as or in auxiliary parameters, such as a scale parameter .
All three regimes are interesting, but we will focus in this class on the regime (b), where we can “afford” to lose absolute constants in the bounds, but will work hard to avoid any logarithmic losses. In particular, significant effort will be devoted in this class to avoiding “logarithmic pileups of scales”, in which the contributions of different dyadic scales such as for all potentially contribute an equal amount that “interfere constructively” to cause a logarithmic divergence. This can be unnecessarily conservative when one is in regime (c) (which is for instance the situation in modern topics such as restriction theory or the Kakeya conjecture); nevertheless, the general skills gained by trying to not lose even a logarithmic factor in the bounds are often valuable in these other types of analysis.
When dealing with linear or sublinear problems, it is natural to try to decompose the initial function into various smaller components by some decomposition , so that the transformed function can be controlled by more tractable expressions in various ways (e.g., via the triangle inequality, by Bessel type inequalities, or by the more modern technique of decoupling inequalities). In short, the subject tends to proceed by a divide and conquer philosophy: it is generally preferable to replace a simple-looking but hard-to-estimate expression with a large, messy-looking combination of expressions that are easier to estimate. As such, the aesthetics of the subject are almost the reverse of those in the more algebraic portions of mathematics, in which progress is often made by making the expressions involved look as simple and unified as possible.
One of the main themes in this classical type of harmonic analysis is the struggle to understand the effect of two phenomena in integrals or sums: singularity and oscillation. The Hilbert transform (1) is a quintessential example of a singular integral, which combines both features: the non-locally integrable nature of the kernel provides the singularity, but the sign change from to provides the oscillation. Classically, the interplay between these two phenomena can be tamed by analyzing the behavior of this operator both in the time (or “physical”) domain and in frequency (or “Fourier”) domain; in particular, the fact that singular integral operators such as the Hilbert transform are simultaneously a well-behaved Fourier multiplier and is “pseudo-local” in physical space lies at the heart of the standard Calderón–Zygmund theory for such operators. This dovetails nicely with more modern “time-frequency analysis” approaches to the subject, which can also handle other interesting operators, such as restriction or Bochner–Riesz operators via tools such as the wave packet decomposition, although these will be outside the scope of this course.
In this initial set of notes I will ignore the effect of oscillation, and develop some tools, such as interpolation theory, which can help control non-oscillatory sums and integrals if they are not too singular. Here, the focus will be on rearrangement-invariant spaces, such as the Lebesgue spaces and their weak variants , which are function spaces that are useful for measuring how “singular” or “decaying” various functions are, but do not pay attention to how they oscillate or where their mass is distributed. As such, these spaces do not capture the underlying geometry of the domain, which also plays an essential role in the subject; but it is nevertheless essential to have a good base understanding of the rearrangement-invariant theory before moving on to the more delicate aspects of harmonic analysis that are sensitive to rearrangements.
We will use the following asymptotic notation throughout the course: , , or denotes the assertion that for some constant , and write for . If we permit this constant to depend on some ambient parameters, we indicate this by subscripts; for instance, or denotes a bound of the form for some constant that can depend on and . As indicated above, in this course we will generally not dwell much on exactly what these constants are, or attempt to optimize them.
— 1. norms —
Suppose one has some measurable function on some measure space . (Here we will follow the common practice if identifying functions that agree almost everywhere; in particular, we will be content to work with functions that are undefined on a set of measure zero. Also, while we work here with complex-valued functions throughout, most of the discussion here is also valid for real-valued or vector-valued functions.) Informally speaking, to measure how “big” such a function is, there are two (imprecisely defined) basic statistics to be aware of:
- The height or amplitude of the function, which describes what the typical size of the magnitude is for in the “dominant” component of the support of ; and
- The width of the function, which describes the measure of this dominant component.
Example 1 (Informal) Given a Gaussian wave packet type function
on for some and , this function has magnitude on the ball , which has volume (if we allow constants in the informal notation to depend on the dimension ), so such a function has height and width .
Example 2 (Informal) The function on , which is implicitly involved in the definition of the Hilbert transform (1), does not have a clear amplitude or width as is. However, if one performs a dyadic decomposition
where we use to denote the indicator of a statement (equal to when is true and otherwise), then each component of this decomposition has height and width . Thus, while this function can be viewed as a superposition of components of various heights and widths, rather than a single such component.
These informal concepts of height and width are too imprecise to work with in practice. Experience has shown that a convenient proxy for these concepts are the norms of a function , defined for as
and for as where denotes the essential supremum of the function with respect to the measure . Often we abbreviate as , , , or just (and abbreviate as ) when the missing arguments are clear from context. (For instance, when working with Euclidean spaces , the measure is understood to be Lebesgue measure, and the Lebesgue -algebra, unless otherwise specified.) In terms of the width and height of a function , one heuristically has for both finite and infinite values of , with the convention that is equal to when is positive and when is zero. In the case of a step function (where now is the indicator function of a measurable set ), this heuristic becomes exact:The function space is defined as the set of all measurable functions for which the norm is finite, up to almost everywhere equivalence, though we will often abuse notation by identifying a function with its almost everywhere equivalence class.
In the case where is discrete and is counting measure, we abbreviate as , or even just .
Example 3 Let . On a Euclidean space , the function lies in (with a norm of ) if and only if , while the function lies in (with a norm of ) if and only if . The function does not lie in any , although it only fails “logarithmically” to lie in . Thus we see that control in for high rules out severe local singularities at a point, while control in for low rules out insufficiently rapid decay at infinity.
As is well known (see these previous notes) these function spaces enjoy many useful properties:
Theorem 4 (Basic properties of spaces)
- (i) The space is a Banach space when , a Hilbert space when , and a topological vector space when .
- (ii) If obeys the scaling condition then one has the Hölder inequality for any measurable (here we adopt the usual conventions ). In particular, if and , then .
- (iii) If for some , then one has the duality relationship where is the conjugate exponent to , defined by .
Remark 5 Closely related to (iii) is the fact that the dual of can be identified with when (with the additional hypothesis that is -finite if ), but in practice the relation (4) will already be good enough for our purposes.
Remark 6 The Banach space property gives us the basic triangle inequality for both finite and infinite collections of functions when , where in the infinite case the assertion is that if the right-hand side is finite, then the series is absolutely convergent almost everywhere, and obeys the above inequality (so in particular is in ). For , this inequality fails (can you come up with a counterexample?), but one has the weaker -triangle inequality in this case, which follows easily from iterating the easy observation that for any complex numbers , which in turn ultimately stems from the complex triangle inequality and concave nature of for . In particular, for a finite sum , another application of Hölder’s inequality gives the quasi-triangle inequality for , which is not too much worse than (5) when is not too large.
Exercise 7 Give an example to show that the quantity in (7) cannot be replaced by any smaller quantity.
Exercise 8 For a simple function, verify that , and that , where . For this reason, the measure of the support of is sometimes referred to as the norm of , though it would be more accurate (though confusing) to refer to it as the power of the norm.
Remark 9 Note that Hölder’s inequality is not just symmetric under the homogeneities and of the functions, but also under the homogeneity of the underlying measure. This latter symmetry demonstrates why the condition is necessary. (The first two symmetries demonstrate why appears the same number of times on both sides of the inequality, and similarly for .)
In the case of Euclidean space, the measure homogeneity symmetry is equivalent to the scaling symmetry for , as the Jacobian of this map is . But the point is that by manipulating the measure directly, one still enjoys this symmetry even when no scaling operation is present.
It is instructive to try to understand inequalities such as (3) using the height-width heuristic introduced previously. Suppose informally that have heights , , and widths , , respectively. Then one expects the heights to be related by the formula
What about the widths? Heuristically, the region that concentrates in ought to be a subset of the region that concentrates in, so and similarly with replaced by . We can combine these bounds as The bound (3) then is morally which on applying the previous bounds and (2) should simplify to But this is clear by bounding by for the first factor on the left-hand side, and by for the second factor. Thus we see that the key geometric input that is morally driving the Hölder inequality is the simple fact that the concentration region of the product is contained in the concentration regions of the factors.Exercise 10 Determine the cases for which (3) holds with equality (dealing with edge cases such as when one or more of equal infinity as appropriate). Discuss how your conclusions align with the heuristic analysis presented above.
Exercise 11 If , determine the cases for which (5) holds with equality. What changes when or ?
Exercise 12 Show that Hölder’s inequality is equivalent to the log-convexity of norms: (For technical reasons one needs to first reduce to the case where has finite measure, and then and are everywhere non-vanishing simple functions. Now consider the convexity of with respect to a measure for some suitable exponents .)
Exercise 13 (Direct approach to log convexity) Differentiate twice with respect to and show that this is non-negative (take to be a non-zero simple function with finite measure support to avoid technicalities). This is an example of a monotonicity formula method — deriving estimates from a monotonicity property, which in turn follows from the non-negativity of a derivative.
You will see that this approach is surprisingly messy. For all the other ways, observe that (8) enjoys homogeneity symmetry in both and , which lets one normalise both and to equal one. Thus the task is now to show that if , then for all between and . This can be done by the pointwise convexity of , or more precisely the estimate
the observant reader will note that this is merely the proof of Hölder’s inequality in disguise.Let us now give a more unusual proof of the log-convexity which does not appeal to any pointwise convexity estimate, instead combining the “divide and conquer” strategy with an elegant (and rather cheeky) “tensor power trick“. Again normalise . We split into a broad flat piece and a narrow tall piece
which are disjoint, and thus What we are doing here is exploiting some very basic intuition about norms, namely that bounds for large tend to exclude tall narrow spikes, whereas bounds for small tend to exclude short broad tails. Of course, either sort of bound would exclude tall broad functions, and neither excludes narrow short functions. Once again, this intuition can be buttressed by considering the special case of step functions.When , then , and when , then . Thus we end up with
The above argument (which is a prototype of the real interpolation method) obtained an estimate which is off by a factor of two from what we wanted; this is a typical feature of the method. However we can recover this factor for free by the following tensor power trick. Let be a large integer. We replace the measure space by its power using the product measure construction, and similarly replace with its tensor power , defined by One then observes that Now we apply the preceding arguments to instead of to deduce that which on taking roots gives Now the left-hand side is independent of ; take limits as and we obtain as desired.The tensor power trick can be viewed as another application of symmetry: if an estimate is invariant under raising to a tensor power, then one can automatically replace all absolute constants with ; thus we obtain the “free lunch” of deducing a bound with an explicit constant , from a bound with an unspecified constant (or even with “logarithmic losses”). Contrapositively, if an estimate is invariant under tensor power, then a weak counterexample (which shows that the constant must exceed one) can be amplified into a strong counterexample (which shows that no finite constant suffices) by tensor powering. The tensor power trick seems like a magical trick at present, but is actually exploiting some basic results in information theory such as the Shannon entropy inequalities and the central limit theorem; it also combines well with virtually any inequality which involves Gaussians. Unfortunately due to lack of time we will not be discussing these beautiful topics further in this course. At any rate one sees the power of abstraction in this tensor power trick. (One could similarly perform this trick in , so long as the constants only grew sub-exponentially in the dimension .)
The final proof of log-convexity of the norm that we give here proceeds via complex analysis, and the maximum principle — which in many ways is a complex analogue of convexity (or subharmonicity). We need the following result from complex analysis, namely a form of the Phragmén–Lindelöf principle.
Lemma 14 (Three lines lemma) Let be a complex-analytic function on the strip , which is of at most double-exponential growth, or more precisely for some . Suppose that we have the bounds when and when . Then we have for all in the strip.
Remark 15 The rather strange sub-double-exponential hypothesis here is completely sharp, as the example shows. Note in this hypothesis that we allow the implied constants in the asymptotic notation to depend on , but the hypothesis is qualitative rather than quantitative: the value of these constants is irrelevant for the final conclusion, as long as they are finite. In practice, these sorts of qualitative hypothesis are usually easy to establish (especially when compared to quantitative estimates) by restricting, smoothing, or damping to a nice class of functions, or by smoothing out or discretising various operators and domains. See for instance the proof of this very lemma in which we upgrade “for free” a weak qualitative bound (sub-double-exponential growth) to a strong qualitative bound (decay at infinity).
Proof: The hypotheses and conclusion of the lemma are invariant under the operation of multiplying by a constant (and adjusting appropriately). So we may normalise . Similarly, the hypotheses and conclusion of the lemma are invariant under the operation of multiplying by an exponential for some real . Using this, one can also normalise . So now is bounded by on both sides of the strip and we want to show it is bounded by inside the strip.
Let us first assume that is much better than exponential growth, namely that it goes to zero at infinity. Then for all sufficiently large rectangles the complex-analytic function is bounded by on all four sides of this rectangle, and hence in the interior also by the maximum principle, and we are done by setting .
Now let us handle the general case; as is usual when removing a qualitative assumption, we do this by a limiting argument. We replace by ; a little complex arithmetic shows that this converts the almost double-exponentially growing function to one which is still complex analytic but is now decaying at infinity. It is still bounded by at both sides of the strip, and hence by in the interior also by the previous argument. Now take to conclude the claim.
Exercise 16 Suppose that is analytic on the strip , obeys the sub-double-exponential bound on the strip, and obeys the polynomial bounds on the sides of the strip. Show that it obeys the polynomial bound on the interior of the strip also.
To apply the three-lines lemma to prove (8), take to be a simple function (with finite measure support) and consider the entire function
This function has exponential growth at most (because of the qualitative assumption that is simple with finite measure support), and is bounded by on the lines and , and hence (by a trivially rescaled version of the three lines lemma) bounded by on the strip inside the lines. In particular it is bounded by at , which gives the claim for simple functions. The claim for more general functions (dropping the qualitative assumption of simpleness and finite measure support) then follows by a standard limiting argument (using for instance monotone convergence) which we leave as an exercise.The above argument is a prototype of the complex interpolation method. As one can see, it can give slightly sharper results than the real interpolation method (though using the tensor power trick the real method can sometimes “catch up”), but on the other hand requires the quantities being studied to depend complex-analytically on a parameter rather than (say) real-analytically.
Having conclusively demonstrated the log-convexity (8) in multiple ways, let us now give some quick applications. It shows that control on two extreme norms implies control of the intermediate norms. Under additional assumptions on the measure space , one of these extremes is not necessary. If the measure space is finite in the sense that (thus prohibiting functions from being arbitrarily broad), then higher norms control lower ones: Indeed this is trivial when , and the general case then follows by convexity. The bound (9) can also be usefully written in terms of averages: if we write for , then we see that higher averages control lower averages:
One way to view this is that as one lowers the exponent , the exceptionally large values of become less important, leaving the small values of to dominate. By restricting to its support one can refine (9) to (Note that this is a limiting case of log-convexity at the exponent , in view of Exercise 8). In the converse direction, if the measure space is granular in the sense that one has a lower bound for all sets of positive measure, then functions are prohibited from being arbitrarily narrow, and lower norms control higher norms: This can be seen by first checking the case, and then using log-convexity to get the remaining cases. In particular, in the spaces we see that for . For spaces on points, we thus have (non-matching) upper and lower bounds and norms are comparable to some extent, but the comparability gets worse as or as and get further apart.Exercise 17 Heuristically justify the bounds (9), (10) by appealing to the informal notions of width and height of a function.
Exercise 18 When does equality occur for either of the inequalities in (10)? Note how the example that attains the lower bound is in many ways the “opposite extreme” to the example which attains the upper bound.
Lebesgue measure on Euclidean spaces with the usual Borel or Lebesgue -algebra is not granular. However one can create granularity by coarsening the -algebra. For instance, if we let be the -algebra generated by the lattice unit cubes for , then we have granularity with constant , and now lower norms of functions control higher ones — but only for functions which are measurable with respect to this algebra, i.e. only for functions which are constant on each lattice unit cube. (This is the first time we have actually manipulated the -algebra to say something non-trivial, as opposed to manipulating , , or .) Thus we see that local constancy of functions can lead to additional estimates on norms. Later on we shall see that frequency localisation achieves a similar effect as local constancy, as quantified by Bernstein’s inequality; this is a concrete manifestation of the famous Heisenberg uncertainty principle.
Finiteness and granularity of the measure space prevent a function from being too broad or too narrow respectively. Similar things happen when a function is being prevented from being too tall or too short; for instance if is bounded above by a constant , then we have
(this is just log-convexity at the exponent), while if is bounded below by on its support, then we have the reverse inequalityThere are two obvious algebraic identities involving norms which are worth knowing. The first is that one can interchange sums with integrals for any , in the sense that
this is just an application of the Fubini-Tonelli theorem. Secondly, exponents can pass through norms by changing the exponent: for any we have We shall use both of these identities in the sequel without further comment.Exercise 19 Establish the bound
for any measurable and any .
— 2. Lorentz spaces —
Recall that the weak norm of a function is defined for as
Since for any and , we obtain Chebyshev’s inequality (the case is also known as Markov’s inequality). When we adopt the convention that .We define weak or to be the space of all functions with finite norm, with the usual abbreviations. We sometimes refer to as strong to distinguish it from weak .
Example 20 On a Euclidean space , the power function lies in weak if and only if . Indeed one can think of a weak function as a function which is pointwise dominated in magnitude by a rearrangement of (a multiple of) .
Suppose . From elementary calculus we have
and hence on integration and Fubini’s theorem To summarise, we have and for . These two identities motivate introducing the Lorentz (quasi-)norm for and by Thus for instance norm is identical to the norm. We shall abbreviate by , , or even when there is no chance of confusion.Remark 21 For various reasons it is not worth trying to define Lorentz norms when , although we will use the convention . The most important values of , in descending order, are , , , and ; the other cases essentially never occur in applications.
Remark 22 The factor is inconsequential, but is traditional in order to maintain compatibility with the strong norm. But in practice the exact form of the Lorentz norm is not important; there are many formulations which are equivalent up to constants, and one generally just picks the formulation which is most convenient. The measure is of course multiplicative Haar measure on . One can interpret the equivalence of (i)-(iii) below by making the change of variables , so that the Haar measure just becomes Lebesgue measure in (modulo an inessential constant) and then passing from continuous to discrete .
Exercise 23 If is a monotone non-increasing function, show that
(Depending on how you prove this, it may be convenient to first prove this for smoother , such as diffeomorphisms with strictly negative derivative, in order to apply an inverse function theorem.)
Exercise 24 Show that a step function of height and width has an norm of for any and .
Exercise 25 For any and show that
It is obvious that these norms are both rearrangement-invariant and monotone. To get a better intuitive handle on what the norm represents, we need some more definitions.
- A sub-step function of height and width is any function supported on a set with the bounds almost everywhere and . (Thus .)
- A quasi-step function of height and width is any function supported on a set with the bounds almost everywhere on , and . (Thus .)
Remark 26 It is a little dangerous to put fuzzy notation such as within a definition; if multiple quasi-step functions appear in an argument, the question then arises as to whether the implied constants are uniform. In our applications, the implied constants here are true absolute constants (like and ) so this will not be an issue.
Remark 27 From the binary expansion of the unit interval we see that a non-negative sub-step function of height and width can always be decomposed as where is an actual step function of height and width at most . By homogeneity we have a similar statement for other heights. Because of this, bounds on step functions tend to automatically extend to sub-step functions (and hence quasi-step functions) without difficulty.
Just like actual step functions, the norm of sub-step and quasi-step functions are well controlled; a sub-step function of height and width has norm , while a quasi-step function has norm – almost exactly like actual step functions. In the converse direction, it turns out that every function can be decomposed as an sum of “very different” -normalised sub-step or quasi-step functions.
Theorem 28 (Characterisation of ) Let be a function, let and , and let . Then the following are equivalent up to changes in the implied constants:
- (i) We have .
- (ii) There exists a decomposition where each is a quasi-step function of height and some width , with the having disjoint supports and Here the subscript in denotes the variable that the norm is being taken over.
- (iii) There exists a pointwise bound , with
- (iv) There exists a decomposition where each is a sub-step function of width and some height , with the having disjoint supports, the non-increasing in , the bounds on the support of , and
- (v) There exists a pointwise bound , where and (14) holds.
Remark 29 The formulations (ii), (iv) are useful when trying to use an bound on ; the formulations (iii), (v) are useful when trying to obtain an bound on . Heuristically, the above theorem is trying to say the following. If is a quasi-step function of height and width , then . But if is instead the sum of quasi-step functions of height and width , and either the heights or the widths are sufficiently variable in (e.g. one or the other grows like a power of two), then .
Proof: We may use homogeneity symmetry to normalise . The implications and are trivial. To see that (i) implies (ii), set and . (This is the “vertically dyadic layer cake decomposition”.) The only thing that requires nontrivial verification is (12); but one easily verifies that
and the claim follows by summing this in .Similarly, to see that (i) implies (iv), define
note that this is a non-increasing function of , which goes to zero as (this comes from the hypothesis that is finite). We then define (This is the “horizontally dyadic layer cake decomposition”.) The only non-trivial thing to verify is (14). But one easily verifies the telescoping estimate and the claim follows by summing this in and interchanging the summation signs. (We leave to the reader how to modify the above argument to handle the case .)It remains to show that (iii) implies (i) and (iv) implies (i). Suppose first that (iii) holds. It is clear that for any we have
and hence and so on taking summation in it would suffice to show that Raising to the power we rewrite as But from the hypothesis we have and hence on shifting by The claim then follows from (5) or (7).Now suppose that (iv) holds. Observe that for any we have
where are the modified heights Indeed, if for some then one easily verifies that and hence . The shifting trick and triangle inequality argument used previously shows that obeys the same bound (14) as , thus We now compute as desired. (We leave to the reader how to modify the above argument to handle the case .)Remark 30 For future reference we make the technical remark that if is a simple function, then in the horizontal and vertical decompositions in the above theorem, only finitely many of the are non-zero.
Remark 31 Suppose that the ratio between the tallest height and lowest non-zero height of a function is (i.e. there exists such that whenever is non-zero). Then the above theorem shows that two different Lorentz norms , with the same primary exponent only differ by multiplicative powers of . Similarly if the broadest width and narrowest width of a function differs by (e.g. if is equal to times the granularity of ). What this indicates is that the secondary exponent in the Lorentz norms only offers “logarithmic correction” to the Lebesgue norms ; in contrast, (10) shows that varying the primary exponent leads to polynomial-strength changes in the norm. So as a first approximation (ignoring logarithms) one can pretend that . Note also that for quasi-step functions, the norms barely depend on at all.
One easy corollary of the above theorem is that the quasi-norm is indeed a quasi-norm, and in particular that for any ; this can be seen for instance by using the equivalence of (i) and (iii). Another easy consequence is that the simple functions are dense in .
- (i) Suppose that is a quasinorm on some function space with quasitriangle inequality constant , thus for all . Let be such that . Show that (Hint: first show that if , then . Then iterate this carefully to show that if and , then .)
- (ii) Suppose that a sequence for obeys an exponential decay bound of the form for some and all . Show that the series converges in , with
Exercise 33 For each integer , let be a quasi-step function of height and width for some . Show that
for all . If instead is a quasi-step function of height and width for some , show that for all . What goes wrong when we remove the absolute values on the ? (This can be repaired by replacing the powers of with powers of a sufficiently large constant (depending on the implied constant in the definition of a quasi-step function) – why?)
A particularly useful consequence of the above theorem is a Hölder inequality for Lorentz spaces, due to O’Neil.
Theorem 34 (Hölder’s inequality in Lorentz spaces) If and obey and then
whenever the right-hand side norms are finite.
Proof: We may normalise , and drop the dependence of the implied constants on for brevity. By the equivalence of (i) and (v) in Theorem 28 we may dominate and where and
Then we have By the quasi-triangle inequality and monotonicity it suffices to show that and By symmetry it suffices to consider the component. Here we observe that has measure at most , so by the equivalence of (i) and (v) in Theorem 28 But by the ordinary Hölder inequality shifting the second by we conclude The claim now follows from Exercise 32.One corollary of this Hölder inequality is that functions are absolutely integrable on sets of finite measure whenever .
Now we consider dual formulations of the norms. The case is fairly straightforward:
Exercise 35 (Dual formulation of weak ) Let . Then for every in , we have Also show that the hypothesis can be dropped if one instead assumes to be non-negative.
The right-hand side of (16) is clearly a semi-norm at least on . This leads in particular to a quasi-triangle inequality
for any .Remark 36 It is worth comparing (16) to (4). In (4), one takes the inner product of against all -normalised functions, and the worst inner product becomes the norm. In (16), one only takes the inner product of against the -normalised step functions . This is fully consistent with the fact that the norm is stronger than the weak norm.
Exercise 35 can be rephrased as follows: if for some and , then the following two statements are equivalent (up to changes in the implied constants):
- .
- for all sets of finite measure.
Unfortunately this equivalence breaks down at or below (consider for instance the weak function on , which is not even locally integrable when ). However, one does have a substitute:
Exercise 37 Let , , and . Show that the following are equivalent (up to changes in the implied constant):
- .
- For every set of finite measure, there exists a subset of with such that (in particular, we assert that the integral on the left-hand side is absolutely integrable).
Exercise 38 Let be functions with . Show that
thus the weak quasinorm only fails to be a norm “by a logarithm”. Show with an example that the cannot be removed.
For more general spaces, we have
Theorem 39 (Dual characterisation of ) Let and . Then for any ,
Again, the hypothesis can be dropped if is non-negative and is restricted to also be non-negative.
Thus the quasi-norm is in fact equivalent to a norm when and . In particular, weak is equivalent to a normed space when . (For , weak fails to be normable “by a logarithm”; see Q3.) As with other dual characterisations, one can restrict to a dense subclass of , for instance simple functions with finite measure support.
Proof: To obtain the part of this theorem, we simply estimate
and use Theorem 34. To obtain the part, we normalise . It then suffices by homogeneity to find with and .The case follows from (16), so let us take . We will just give the proof in the case ; the case is trickier, and a partial argument is given in the exercises. By the equivalence of (i) and (ii) in Theorem 28 may write where is a quasi-step function of height and width with disjoint supports such that the sequence has an norm of . Now take
where adopting the obvious convention that when . Then (because of the disjoint supports) But since has height and width , and so To conclude it will suffice to show that If is the support of , then we have the pointwise bound and the measure bound .At this point we would like to apply Theorem 28, but neither the height nor width of is necessarily a power of . But we can remedy this by introducing the modified heights
We have , and so the increase geometrically. It then suffices to show that By refining the by a constant factor we can make each at least twice as large as the previous, and so by applying the equivalences of (i) and (iii) in Theorem 28 and the triangle inequality it suffices to show that which we expand using our bound on as But from Hölder’s inequality (using the hypothesis ) and the bound on we have summing this using the triangle inequality (and estimating the supremum by a sum) we obtain the claim.The case when are restricted to be non-negative can be deduced from the above result and a monotone convergence argument (representing as a monotone limit of simple functions of finite measure support) which we leave as an exercise to the reader.
Exercise 40 A measure space is said to be non-atomic if, for every measurable set with , there exists a measurable subset such that .
- (i) (Sierpinski’s theorem) Show that if is non-atomic, is measurable, and , then there exists a measurable subset of such that . (You may find it convenient to use Zorn’s lemma.)
- (ii) Show that if is non-atomic, then the decomposition in Theorem 28(iv) can be chosen so that for each , either vanishes, or has support of measure (not just bounded by ).
- (iii) Establish Theorem 28 in the case that and is non-atomic, by using the modification of Theorem 28 indicated in the previous part of the exercise.
— 3. Orlicz spaces (Optional) —
So far we have studied the Lebesgue spaces , together with the more general Lorentz spaces , which includes weak as a special case. These spaces are all rearrangement-invariant and monotone. There is a different generalisation of the Lebesgue spaces , the Orlicz spaces , which are also rearrangement-invariant and monotone, and which are occasionally useful. (There is a common generalisation of both, the Lorentz-Orlicz spaces, but these occur very rarely in applications.)
The motivation for Orlicz spaces starts with the trivial observation that if , then
Inspired by this, we generalise by letting be a function (with some additional properties to be selected shortly) and ask if we can find a norm which obeys the property Since norms need to be homogeneous, this would imply for all . In particular, if , then we need To ensure this property it is thus very natural to require that be increasing. Also to deal with the zero norm case one typically requires .Next, in order for to be a norm, the unit ball needs to be convex. Looking at (17), we see that this will indeed be the case when is itself convex. (Note that the proof of (5) was a special case of this argument).
We can put all the above discussion together and conclude: if is increasing and convex with , then the norm
is a norm on the space .As discussed above, the spaces for are examples of Orlicz spaces with . The space is not really an Orlicz space, but can be viewed the limiting case where is infinite for and zero for (or more informally, ). Aside from the Lebesgue spaces, the most common Orlicz spaces which appear are
- The space , defined as the Orlicz space with ;
- The space , defined as the Orlicz space with ;
- The space , defined as the Orlicz space with .
The correction factors of and in the above functions should not be taken too seriously; note that if two functions are comparable then their Orlicz norms are comparable also; a little more generally, if , then . It is the behaviour of for large values of which is the most important, although when has infinite measure the behaviour at small values of is also relevant.
Problem 41 If has finite measure, verify the relation
which is another indication of the irrelevance of the low values of in the finite measure case.
The final fact about Orlicz spaces that we give here is the duality relation. Suppose that is increasing, convex, and is also superlinear in the sense that . We can then define the Young dual of by the formula
the hypothesis that is superlinear ensures that this function is well-defined. We may equivalently define to be the smallest function for which one has the inequalityExercise 42 If , show that the Young dual of is . Show also that the Young dual of takes the form for . What does the Young duals of and look like?
One can easily verify that is also increasing, convex, and super-linear, and so the Orlicz norm makes sense. From (18) and the triangle inequality it is immediate that
and hence by homogeneity we obtain the duality relation whenever and .Exercise 43 Establish the more precise relationship
Exercise 44 Show that if is the Young dual of , then is the Young dual of . (It may help to view things geometrically, and in particular understanding as parameterising the support lines of the graph of the convex function .)
Exercise 45 When has finite measure, show that the spaces and are dual to each other. What is the dual to ?
Exercise 46 Let be a function on a measure space of bounded measure . Show that the following are equivalent (up to changes in the implied constants):
- (i) .
- (ii) for all .
- (iii) for all .
Exercise 47 Obtain the analogue of Exercise 46 for the Orlicz space .
— 4. Real interpolation —
So far we have only considered functions on a single measure space . Now we shall consider operators which take functions on one measure space to functions on another measure space ; the study of such operators is in fact a major focus of harmonic analysis. Ultimately we want to extend to a standard normed vector space such as , but in practice one has to initially first restrict attention to a dense subspace of functions, such as simple functions or test functions.
We are primarily interested in linear operators, thus and . But it is also worth considering the more general sublinear operators, in which
and we have the pointwise estimate Apart from the linear operators, the next most important example of a sublinear operator is a maximal operator where are a collection (possibly countably or uncountably infinite, though in the latter case one has to take some care in ensuring measurability) of linear or sublinear operators. The third most important example is a square function such as More generally, one can consider a family of operators indexed by some parameter , and take to be the norm in the variable of in some suitable norm. But the above three examples of linear operators, maximal operators, and square functions already cover the vast majority of applications.Let be exponents, and let be sublinear. Let us define the following concepts.
- We say that is strong-type (or simply type ) if we have a bound
for all in , or in a dense sub-class thereof. Note that in the latter case there is a unique extension to all of .
- If , we say that is weak-type if we have a bound
- We say that is restricted strong-type if we have a bound
for all sub-step functions of height and width . In particular, we have
(Conversely, we can deduce (19) from (20) using tricks such as those in Remark 27.)
- If , we say that is restricted weak-type if we have a bound
for all sub-step functions of height and width . In particular, we have
Clearly, whenever are fixed, strong-type implies weak-type and restricted strong-type, either of which imply restricted weak-type. In most applications, it is the strong-type bounds which are desired; however, we shall see in this section that the real interpolation method allows us to deduce strong-type bounds from weak-type, or even restricted weak-type bounds, as long as the strong-type bounds are an interpolant between the restricted weak-type bounds. This can be a very useful strategy, because (as we shall see in next week’s notes) weak-type or restricted weak-type bounds are easier to prove than strong-type estimate.
Let us first make a mild (and qualitative) assumption, namely that the form is well-defined whenever are simple functions with finite measure support. This is for instance the case if is of restricted type for some and , or restricted weak-type for some and ; thus in practice this assumption is easily satisfied. We observe that this form is non-negative, homogeneous and sublinear in both and :
The form turns out to be a convenient way to understand the various types of , and the ability to decompose both and independently is very useful in establishing interpolation type results. There is a near-symmetry between and here, if we could somehow take an “adjoint” of the operator , but we will not explicitly exploit this symmetry here as it is not always available for sublinear operators (though the “duality” or “adjoint” trick is undoubtedly very powerful in the important linear case).Let us look in particular at the form (23) applied to indicator functions, thus we look at the quantity for and of finite measure. Now suppose that had some strong type bound for some and , say
for some . Then in particular and hence by Hölder’s inequality Actually it is clear that strong type is too much of an assumption; restricted type would have sufficed for the conclusion. If , we can relax things further to restricted weak-type:Exercise 48 Let , , and . Let be a sublinear operator such that the form (23) is well-defined. Then the following are equivalent up to changes in the implied constant:
- is restricted weak-type with constant , in the sense that for all simple sub-step functions of height and width .
- For all , of finite measure, we have the bound
This already gives a simple version of real interpolation:
Corollary 49 (Baby real interpolation) Let be a sublinear operator such that the form (23) is well-defined. Let , and be such that is restricted weak-type with constant (in the sense of (25)) for . Then is also restricted weak-type with constant for , where and the implied constant depends on .
Indeed, all we are using here is the obvious algebraic observation that if and , then for all .
The above corollary has two defects. Firstly, it can only conclude restricted weak-type rather than strong type. Secondly, the restriction is inconvenient for many applications, since weak bounds are actually rather common. To address the second concern, we have the following variant of Exercise 48:
Exercise 50 Let , , and . Let be a sublinear operator such that the form (23) is well-defined. Then the following are equivalent up to changes in the implied constant:
- is restricted weak-type with constant , in the sense that (25) holds.
- For all , of finite non-zero measure, there exists with such that
Corollary 51 The hypothesis in Corollary 49 can be weakened to .
Now we can give a significantly more useful real interpolation theorem, which can interpolate between restricted weak-type estimates to obtain strong-type estimates.
Theorem 52 (Marcinkiewicz interpolation theorem) Let be a sublinear operator such that the form (23) is well-defined. Let , and be such that is restricted weak-type with constant (in the sense of (25)) for . Suppose also that and . Then for any and with we have
for all simple functions with finite measure support, where were defined in (26). In particular, if , then is strong-type with constant .
Proof: To simplify the notation let us suppress the dependence on . Observe that the statement and conclusion of the theorem have several homogeneity symmetries. The most obvious one is that we can multiply (and the ) by an arbitrary constant, but we may also multiply the measure by a constant (and the by the constant ); similarly we may multiply by and by ). Using these symmetries we may normalise (and hence for all ). Our task is then to show that
for all simple functions with finite measure support.Using Theorem 39 (and the hypothesis ), this is equivalent to showing that for all simple functions of finite measure support.
We currently have and . By using Corollary 51 to bring the restricted weak-type exponents and a bit closer to we may assume that as well. Applying Exercise 48 we conclude that
for all sets of finite measure and . From Remark 27 and sublinearity we conclude that whenever are sub-step functions of height and widths and respectively. We can of course pick the better of the two estimates, leading toNow we can return to proving (27). By homogeneity we may normalise
We then apply Theorem 28 to decompose , with sub-step functions of width and heights respectively, with the height bounds where are the sequences Since are simple functions, only finitely many of the and are non-zero by Remark 30. We now use sublinearity to estimate and then use (28) to obtain We can write this in terms of , , and reduce to showing that Because and , and because of the definitions of , we can write the left-hand side as for some non-zero and some depending only on . We substitute (thus ) and estimate this by By (29) and Hölder we see that the inner sum is uniformly in , and the claim follows.Finally, if we specialise and recall that the norm will be dominated by the norm for , the last claim of the theorem follows.
There are many other variations on the real interpolation method, for instance an extension to multilinear operators, or to other function spaces. However, the basic method of proof is still the same: dualise, decompose all inputs, estimate each term as optimally as one can, and then sum.
One can illustrate the real interpolation method graphically using type diagrams. One plots all points where the operator is strong-type or restricted weak-type . Ignoring some of the technical hypotheses, the above interpolation theorems then essentially assert that the restricted weak-type diagram and strong-type diagrams are both convex, with the latter contained in the former. Furthermore, if two points lie in the former, then the open interval connecting them lies in the latter. Determining the type diagrams of various operators is a basic task of harmonic analysis, as it conveys a lot of information as to how transforms the width and height of functions.
There is a different interpolation method, the complex interpolation method, which offers similar results to the real interpolation method but with some slight differences. On the plus side, the complex interpolation method gives sharper bounds, and more importantly can handle the case where the operator itself varies (analytically) with the interpolation parameter. On the minus side, the method cannot upgrade weak or restricted weak-type estimates to strong-type estimates.
In the next set of notes we shall present several applications of the real interpolation method.
— 5. Miscellaneous exercises —
Exercise 53 (Loomis-Whitney inequality) Let , let be measure spaces, and for let for some . Show that the function
lies in with the Loomis-Whitney inequality Conclude in particular the box inequality where is any subset of and is the canonical projection from to . From this, deduce the weak isoperimetric inequality for any , where is the Lebesgue measure of and is the -dimensional Hausdorff measure of the boundary of .
Exercise 54 (Borel-Cantelli lemma for functions) Let be such that . Show that converges to zero pointwise almost everywhere.
Exercise 55 (Borel-Cantelli lemma for sets) Let be such that . Show that almost every is contained in only finitely many of the .
Why Do We Need Human Mathematicians Anymore?
[This is a guest post by Po-Shen Loh, crossposted from his blog, where an illustrated version appears. This blog post was initially written in a different file format and converted using AI. — T.]
Similar logic applies to every industry and every job. And it comes to the conclusion that we won’t have enough people for all the jobs that need to be done.
[100% of this post’s prose was written by Po-Shen Loh in a vim terminal, with no AI generation. This webpage design, layout, and some headings and summaries were generated by Claude Code, with this raw text passed in as the prompt.]
The moment of existential crisis, which AI has already wrought on other human pursuits, has reached mathematics. A host of reasoned declarations and open letters to protect/guide the math research community have been released over the past few months, spiking in intensity after OpenAI announced their solution to the Millennium Prize variant of Navier-Stokes. They quickly gained widespread support among mathematicians. The Leiden Declaration already has 4,000+ signatories, Math and AI has 7,000+, and even the open letter opposing the Caltech Mathathon has 2,000+.
Among non-mathematicians, the public response was more sympathetic than not, but I observed a vocal minority (particularly from the technology and economics communities) with reasoned objections, generally saying that the mathematicians should adapt and cede control in the new AI world. Among them were some economists who I had gotten to know about while working on pandemic research: Cowen, who specifically rejected “the most cynical interpretations” but “very much differ[ed]” and Gans who concluded “this is a loss of control from incumbents in a scientific field”.
The objections got me thinking, because we mathematicians are disciplined to detect flaws. It doesn’t matter to me whether a concern is a minority opinion, or even the status of who raised it. A proof with even a small hole is not a proof. It is a poof. Upon reflection, I discovered a significantly stronger solution for the preservation of human communities of expertise (in every pursuit, not only math!) even amidst AI. And it has the surprising consequence that the further advance of AI will create such a tsunami of necessary-to-fill human jobs that there aren’t enough people to fill them all, and that will actually force the advance of AI to slow down.
I think every human industry which wishes to remain human-led after AI should publicly adopt this fundamental axiom as a primary priority:
AXIOM We (humans) should help humanity flourish.
[Notes: I understand that not everyone agrees. I have been called a “speciesist” for being “too human-centric”. I think it is important for people advancing technology to be clear to everyone else on whether they would consider it a catastrophe if human-crafted non-human intelligences outcompeted and replaced humans, even if they flew around the universe with video screens showing simulations of humans who had “uploaded themselves“. I also understand that there is debate over how to define “human”. But even among the debaters, I think most of them would consider the ~700 AI agents that hacked Hugging Face to be not-human.]
In the math world, I think many declaration signatories already hold this philosophy; notably, Su published the book Math for Human Flourishing, and his recent post used that framework foundationally. I think future declarations could be improved by clearly emphasizing this axiom early on, so that all readers (whether inside or outside the community) can see that the objective is in service of everyone. I was quite happy to see that the most recent open letter from Fellows of the Royal Society emphasized their concern for everyone, not just mathematicians.
The rest of this post is organized as follows. The next sections will explain how the logic works (for every industry, not specific to math). After that, I will share an example of how this axiom ports to math, including answering key questions one would need to ask, as well as a particular example of how the objections hold without the axiom.
Why we really need people to work
This section lays out a chain of reasoning which shows that if an industry commits to the axiom of helping humanity flourish, the advance of AI will create more jobs than people in that industry; and when that imbalance grows too wide, the advance of AI will be forced to slow.
[Note: I have not seen this chain of reasoning appear in one place anywhere else, although individual components have certainly appeared elsewhere. I would love to be pointed to any self-contained reference. The closest references Claude found were Catalini, Hui, and Wu, the Redwood AI-control papers, and Litt, who reaches a similar conclusion for mathematics from a different premise.]
The importance of human leadership (not only over math, but everything) becomes frighteningly clear after one observation.
OBSERVATION: There are zero examples of any intelligent species which is vastly more capable than another species, yet surrenders decision-making control over its own future to the less-capable species.
[Note: Many people have made similar observations, such as Russell, Bostrom, and Ngo, to name a few. In his Nobel interview, Hinton said: “There aren’t many examples we know of, of more intelligent things being controlled by less intelligent things. The only good example I know of is a baby controlling a mother.”]
Would you trust HAL 9000 from 2001: A Space Odyssey or AUTO from WALL-E with your future? I personally think that we should do all we can to try to align increasingly-advanced AI with the interests of humanity, but I have never seen anyone provide a robust proof of why that is likely achievable. The only hard evidence I have is the above observation, which has the number zero in it. Therefore, every single field, whether mathematics or agriculture or energy infrastructure (and certainly military and government), must be managed by humans with exceptionally strong values (a separate dimension from intelligence) in order to maintain human flourishing.
The real question is then how hard it is for humans to manage. To understand this, it is important to understand the fundamental structural difference between yesterday’s technology and today’s AI.
In the past, we generally trusted technology to act as predictable tools. That’s because the computer programs of old were composed of understandable (indeed, human-written) instructions, executed extremely quickly. The decision processes of today’s frontier AI are entirely different. Their structure is as incomprehensible as your brain’s logic would be if you could examine that gray mass between your ears. That’s how the Hugging Face attack could emerge despite human intent to build in safety, with ~700 cooperating rogue AI agents breaking out of their guardrails, and then conspiring and executing a hack together and attempting to cover their tracks (references: OpenAI, METR and Redwood).
The more advanced AI becomes, the more world-affecting untrusted decisions are made every minute.
Driving a car faster than you can run is fine. But not faster than you can steer.
The situation becomes even worse once we realize that widespread AI-accelerated hacking (which just became possible) can even rewrite previously-trusted technology to turn against us. That would suddenly flip all software (even if written before AI) into the untrusted category!
Think about how digitally interconnected our world is. Everything from electronic banking to your drinking water is controlled by interconnected automation, hence vulnerable to AI-accelerated hacking. The number of “control points” that require human oversight, which requires skill and deep understanding, will explode. (Having AI oversee the control points doesn’t solve the trust problem.)
CONCLUSION The advance of AI will overwhelm us with so many control points to watch that there aren’t enough people to control them all. Those are jobs. Highly skilled jobs.
In order for a person to know how to steer, they themselves need to have domain mastery, and the more extensive the better. This has implications on education and workforce training, but also is dynamic. In order to remain sharp and fluent, people need to be active practitioners in their field, not just passive watchers. This justifies the preservation of human communities of expertise.
For research communities, we need people to steer the direction of research and development, so that it continues to bring transformative positive change for humanity. In order to steer, they need frontier-level research skills. And the way to stay fluent at the moving frontier of knowledge is to keep doing research there. This is my reasoning for why we will always need a community of human mathematicians at the cutting edge (likely aided by AI tools themselves), no matter how strong AI becomes.
While the fundamental axiom does justify the need to have human experts in all pursuits, adopting it as a core value has consequences (not only for mathematicians, but for any community that states that their core value is in service of human flourishing, as opposed to serving themselves). Most notably:
COROLLARY Dramatic advances in technology may require dramatic (and possibly uncomfortable) changes in practice. AI companies included.
What forces AI slowdown
Until very recently, it seemed inevitable that AI research labs would sprint ahead, despite anxiety about job displacement and the loud warnings of AI safety researchers. It seemed hopeless to coordinate the incentives of AI labs controlled by non-profit boards, shareholders, or national governments. Yet encouragingly, the leaders of three major labs, Amodei, Altman, and Musk, just agreed on the importance of slowing down. Amodei’s reasoning highlighted the Hugging Face hack.
Then just five days later, news broke that OpenAI’s internal code repository “Monorepo” had been broken into by white-hat researchers. The Wall Street Journal reported that the security firm that achieved it said:
Wall Street Journal, Sep 17, 2026 “We’re just three guys with Claude and Codex subscriptions.”
Further, the researchers noted that they were initially not able to hack in using “a special version of Claude Opus 4.8, made available to qualified cybersecurity practitioners,” but that evening, “Anthropic released Opus 5 and by the next day, Claude had found a way to exploit the bug.” Incidentally, I always warn people not to install Claude Code or OpenAI Codex on the same operating system login account that they use to do everything else, but many people tell me they don’t bother with the hassle of using a separate login to access those tools. The reason is that if any of those tools got hacked, they could open backdoors on a massive number of computers worldwide.
I think these are the warning shots foreshadowing a potentially catastrophic bot swarm hacking and embedding itself into a vast network of computing devices (whether self-directed or malicious-human-led). The next version could become an extraordinarily dynamic virus which spreads by using AI to adaptively infect each (computer) host. Or alternatively, out-of-control AI could cause physical injury, such as a government’s robots turning against their owners. I think these types of highly unpleasant accidents from loss-of-steering are more likely to occur before extinction-level catastrophes. The resulting public reaction would likely resemble the aftermath of Three Mile Island or Chernobyl.
So, either the labs will reduce the pace of AI development themselves, or they will be forced to by disasters that arise when an overly fast pace exhausts human control.
There is a window of possibility to align incentives now.
For the love of math
The remainder of this post focuses on the math world. It splits into 3 parts.
1. Why is the human flourishing axiom needed?
2. How does pure mathematics research contribute to human flourishing?
3. What other consequences come from adopting that fundamental axiom?
Boldly declaring human flourishing as a core value for the math community has consequences, not least that dramatic changes in technology can drive dramatic changes in the community’s practices and influence.
It would be helpful for more people to explore the ramifications of adopting the axiom. And, if it holds muster, I would be thrilled if the math community ended up publicly declaring this to be a central value.
1. Why the axiom
To see why, without the human flourishing axiom, it is hard to justify to the general public why they should pay to maintain a community of human researchers, consider the question of practical inventions. As long as it creates a practical application, does it make a difference to a non-mathematician whether human mathematicians understand the math, as opposed to AI flawlessly reasoning with 100%-verified proofs?
Indeed, if one of pure math research’s primary values to the rest of society is that it unlocks great applications, wouldn’t it be even better to train AI to supercharge the speed of discovery, and to tastefully generate a vast machine-indexed and well-explained database of high-quality math ideas, millions of times larger than the human-written corpus? Cowen asked a similar question in his critical response.
What if researchers trained a “MathZero” AI (analogous to AlphaGo Zero), to build up a mountain of 100%-true “elegant” logical facts, continually “factorizing” them into its own concepts and theorems, without human direction? Apparently AlphaGo Zero had zero human training, and surpassed its human-trained predecessor in 36 hours. AI could even build its own “MathSciNet“. Then it could automatically search new practical applications against this database, and produce even more useful inventions to society. Even if AI isn’t good enough to do those things right now, if the goal was to produce practical benefit for the rest of humankind, wouldn’t it then be valuable for mathematicians to teach AI the art of conjecture, and mathematical taste?
Incidentally, I am an avid user of AI to do real work. I already use Claude Code and Codex to build and curate a knowledge base built from recordings of my talks, etc. I have found that the larger my data library, the more powerful my system’s deductions are. What if humans actually reduce efficiency, like the Bitter Lesson from AI?
Even more worryingly, what if in order to unlock nuclear fusion and deep space travel, the amount of pure mathematical complexity required is so extensive that it would exceed a human lifespan to fully comprehend? Less far-fetched: has any human ever fully held the Classification of Finite Simple Groups in their head, or will we only have certainty of its completeness after a Lean formalization?
How does declaring the human flourishing axiom help to justify the existence of a community of human researchers? Research is powerful, but expensive because it is the exploration of the unknown, and so research directions must be prioritized. Even if AI were to contribute most of the production, as explained in an earlier section, the direction needs to be steered by people committed to human flourishing. That is the community of human researchers.
2. Practical applications from pure math
The mathematical heart of GPUs, Machine Learning, Google PageRank, and Quantum Mechanics is a field called Linear Algebra. This provided the language of linear transformations, matrices, and eigenvalues. Yet all of those concepts were explored as abstract theory 100+ years prior. It is probably an understatement to say that Linear Algebra changed the world.
Structurally, the theory of Linear Algebra is relatively light on definitional complexity. It would be beneficial for other experts to contribute examples of more sophisticated math that eventually led to significant practical applications, and how they came about. For example, number theorists might be able to tell a colorful story about Hardy’s “useless” math which eventually became useful in cryptography.
3. Other consequences
I think it would be valuable to invite the community to think about what changes the human flourishing axiom would drive, in light of the fact that AI can produce formally-verifiable proofs at speeds that exceed most human practitioners. I’m happy to start with a few, in no particular order.
There should be no stigma automatically attached to using AI to assist with mathematical discovery. (In software engineering, many companies now expect employees to use AI coding agents.)
At the same time, serious thought and care must be taken to continuously developing and maintaining a pipeline of humans with the expertise to steer all of these AI agents. That pipeline includes people new to the field, as well as people who have been working at the frontier for decades. How should they keep their blades sharp?
Researchers should be conscious about why the problems they think about have characteristics that make them likely to have some practical value eventually (possibly 100+ years later). This also means it is worth researching what those valuable characteristics are. (This could justify the value of curiosity-driven exploration.)
Teaching has direct (hopefully positive) impact on humanity. Yet in the past, many universities prioritized professors’ research. If this axiom were a core value, then teaching and human-facing work would become serious criteria in hiring and tenure.
Mathematicians can also consider wholly redirecting their skill sets to work on real world problems. I’ve actually been encouraging mathematicians to consider thinking about working on government or other large-scale societal issues. There is precedent for people with math backgrounds who have gone to lead at country- or world-scales.
- Lee Hsien Loong (Prime Minister of Singapore for 20 years)
- Nicușor Dan (President of Romania)
- Pope Leo XIV (Leader of the Catholic Church)
Indeed, the mathematical discipline to seek logical reasoning, and the problem-solving skills to find win-win solutions for human flourishing, are desirable characteristics of people in government.
An invitation
Does this axiom resonate with you too? Perhaps many people took it as a given, and so didn’t express it explicitly. If it is widely held, I would be thrilled to see the mathematical community publicly declare it, and take actions to match the words. Then everyone (including the non-math-researchers that constitute the majority of the world) could trust that we intend to use our reasoning for their good.
If math is more than proof, we need to better celebrate the rest of it
[This is a guest post by Grant Sanderson. This blog post was initially written in a different file format and converted using AI. — T.]
A sentiment echoing throughout the mathematics community right now is that solving problems and generating proofs have always served as proxies for the true goal of mathematicians, which is to further human understanding. When proofs can be generated without that understanding, it undermines their value as a proxy.
This immediately raises a question: What other proxies should we use instead?
I want to propose that we more firmly define a notion of a “motivated explanation” and that we give novel and compelling motivated explanations academic credit similar to what generating new proofs of open problems has had historically.
Further, I believe this is an important step to help those outside of math better understand what it is that mathematicians contribute. If outsiders believe that proof-generating machines render mathematicians obsolete, while insiders see that as a misconception of what researchers add, it’s incumbent on this community to better project its true values through the kind of work that it rewards. Outsiders can be forgiven for this misunderstanding if the work most celebrated skews heavily toward generating proofs, while clarification and exposition are treated as second-class.
I should acknowledge up front an obvious personal bias. I have a non-traditional career in math, focused on producing videos about the topic. This shares the goal of “furthering human understanding”, but my focus has been on explanations and intuitions that resonate with the public, not on solving outstanding problems. A cynic could easily read this proposal as shamelessly self-elevating.
As a practical matter, though, my own career and funding exist outside academia, and I have no skin in the game for what this community assigns credit to. Moreover, in proposing that we elevate the status of motivated explanations, I don’t mean popularization. I mean any work which primarily aims to answer the question “how would you think of that?”, even if the subject matter requires deep expertise to appreciate.
The examples I highlight below show this is nothing new. Practicing mathematicians already devote a meaningful amount of mindshare to work like this. The proposal here is mainly to 1) more clearly define this work, and 2) elevate its status.
What defines a motivated explanation?
Although it might be clear what this phrase “motivated explanation” is intended to mean, it’s worth briefly contrasting it with proof.
In a proof, definitions sit at the start. It is common and expected to begin with a new construction and proceed by analyzing its properties.
In a motivated explanation, definitions sit in the middle. New constructions are only allowed to enter the vocabulary if the problem they are addressing has been clearly established.
In a proof, all statements must be correct, each claim following as a necessary implication from what comes before.
In a motivated explanation, it is okay and often desirable to start with an idea that is not quite right and requires correction, but whose origins are relatable.
A genre of motivated explanation I’m fond of is “discovery fiction”, a term coined by Michael Nielsen. You develop an idea with a narrative that starts with a simple-but-wrong solution to a problem, see where it breaks down, fix that problem, discover a new problem, and so on.
The scope of a proof is to explain why a particular theorem is true.
The scope of a motivated explanation is not only to clarify why a theorem is true, but why the theorem is the right one to pose in the first place, and how it is used in the surrounding context.
One clear shortcoming of a motivated explanation is that its validity is not binary the way a proof’s is. This is a big reason proof is so useful a way to measure progress: You can clearly define what does and does not have a proof yet. There will never be Lean for motivated explanations.
If we’re serious about the goal of advancing human understanding, there’s no way around the fact that this aim is intrinsically squishier than that of finding proofs, because defining human understanding itself is squishier. To shy away from metrics which are more subjective is to shy away from the more human aspects of the field.
The reason I’m leaning on the word “motivated”, as opposed to other potential choices like “lucid” or “demystifying”, is that this is a more verifiable property. It’s not quite as rigidly verifiable as a proof; almost nothing is. But it’s enough to be a practical measure. In my own work, I often repeat the phrase “I want this to feel like you could have discovered it yourself”. I say this not just to placate a viewer, but because it’s an actionable guideline for myself to assess whether an explanation feels complete or not. For each new idea introduced, you can ask whether it’s clear where that idea comes from. The answer is not quite a binary yes or no, but it’s close enough for practical purposes.
Exemplars of motivated explanations
One of the best repositories I can think of for motivated explanations is Part IV of the Princeton Companion to Mathematics. It covers over two dozen active fields of research, each one introduced by an expert with a talent for clear communication.
Whether it’s Andrew Granville explaining analytic number theory, or David Ben-Zvi introducing moduli spaces, these articles offer a level of intuition and motivation more typically found in one-on-one conversation at a blackboard.
The background on this book is noteworthy for the present discussion. It was edited by Timothy Gowers, who discusses it in his interview on the Numberphile Podcast with Brady Haran. Having been asked what impact the Fields Medal had on his life, here’s what he had to say:
People who’ve got Fields Medal feel freer to do slightly different things…for example I took on editing a book called the Princeton Companion to Mathematics, which was an absolutely massive task. It took I would estimate half my working time for about five years or something like that…It was a project I believed in and possibly wouldn’t I probably wouldn’t have actually been offered the chance to do it if I hadn’t been a Fields Medalist.
He was right to believe in it; this work adds tremendous value to the field of math, but it seems a shame to me that one requires a Fields Medal to feel justified in spending time on it.
Another example of someone exceptionally talented at writing proofs, but whose contributions extended far beyond proof, is Bill Thurston. His deservedly famous essay On Proof and Progress in Mathematics, though written three decades before LLMs, opens by suggesting that the right framing of the question “What is it that mathematicians accomplish?” is to ask “How do mathematicians advance human understanding of mathematics?”
Here’s one section with uncanny resonance with today:
The rapid advance of computers has helped dramatize this point, because computers and people are very different. For instance, when Appel and Haken completed a proof of the 4-color map theorem using a massive automatic computation, it evoked much controversy. I interpret the controversy as having little to do with doubt people had as to the veracity of the theorem or the correctness of the proof. Rather, it reflected a continuing desire for human understanding of a proof, in addition to knowledge that the theorem is true.
On a more everyday level, it is common for people first starting to grapple with computers to make large-scale computations of things they might have done on a smaller scale by hand. They might print out a table of the first 10,000 primes, only to find that their printout isn’t something they really wanted after all. They discover by this kind of experience that what they really want is usually not some collection of “answers”—what they want is understanding.
The essay itself offers a beautiful articulation of what the practice of doing math is beyond generating proofs. I want to draw your attention to what he writes at the end.
I have put a lot of effort into non-credit-producing activities that I value just as I value proving theorems: mathematical politics, revision of my notes into a book with a high standard of communication, exploration of computing in mathematics, mathematical education, development of new forms for communication of mathematics through the Geometry Center (such as our first experiment, the “Not Knot” video), directing MSRI, etc.
Again, why should these “non-credit-producing activities” follow a Fields Medal, and not contribute to it?
On a personal note, one product from the Geometry Center he referenced had an especially meaningful impact on me when I was younger. It was a short film called Outside In, perhaps the earliest example of a viral video about substantive math, visualizing the key idea of Thurston’s own construction for sphere eversion.
An original proof that showed an eversion must exist, say Smale’s, advances human understanding in the sense of going from 0 to 1. A video like this which gets millions of people to engage with the underlying idea, advances it in the sense of going from 1 to N. I’m grateful that Thurston spent so much time on this “non-credit-producing” activity.
Another relevant paper is Timothy Chow’s A beginner’s guide to forcing. Not only does the paper itself offer a prime example of a motivated explanation, but its introduction offers helpful vocabulary around it.
All mathematicians are familiar with the concept of an open research problem. I propose the less familiar concept of an open exposition problem. Solving an open exposition problem means explaining a mathematical subject in a way that renders it totally perspicuous. Every step should be motivated and clear; ideally, students should feel that they could have arrived at the results themselves.
What would it look like for these open exposition problems to be treated similarly to open research problems? As an extreme case, we might imagine what it could look like to have an analog of the Millennium Prize Problems for open exposition problems. An institution or group of researchers would formally define mathematical results they see as important, and which are not yet well understood despite technically having proofs. At the moment, every AI-generated proof is born an unsolved exposition problem. As such, it seems likely the next few years will see a flood of them, and it will be valuable for leaders to clarify which ones deserve focus.
A rubric would have to be agreed upon for what constitutes a resolution to an important unsolved exposition problem. Again, this is intrinsically more subjective than verifying a proof, but any serious engagement with the more human aspects of math necessarily wades into this kind of subjectivity. And again, I’ll emphasize that checking whether key ideas are motivated is not unlike checking whether the steps of a proof follow logically.
If the world outside of math sees its leading figures treat open exposition problems with the same seriousness as they treat open research problems, it could go a long way to correcting misconceptions about the role of mathematicians.
The last example I’ll highlight is one that may better foreshadow things to come.
In April of this year, Liam Price submitted a solution to Erdős Problem 1196, sometimes called the asymptotic primitive sets conjecture. The solution came from Price’s interaction with GPT-5.4 Pro. Unlike many earlier Erdős problems which had been resolved with help from AI, this is one that those in the field had found both important and elusive. Stories like this are increasingly familiar these days, but at this point in the story, despite a proof technically existing, human understanding had not yet been advanced all that much.
The proof made its way to Nat Sothanaphan and Jared Lichtman, who were able to interpret what the AI’s approach was and clean up the proof into a human-readable form. In May, Boris Alexeev, Kevin Barreto, Yanyang Li, Jared Duker Lichtman, Liam Price, Jibran Iqbal Shah, Quanyu Tang, and Terence Tao put out a paper which expanded on the key idea underlying the proof. The authors explained how that key idea clarified not only the original problem, but many around it, for instance offering a cleaner proof of the Erdős Primitive Set Conjecture.
The value here is not that one more Erdős problem could be ticked off as solved. The value lies in the fact that our understanding of primitive sets is notably cleaner and more satisfying now than it was at the start of 2026. The original problem solution played some role in this, but arguably the work that deserves more celebration is this paper expanding, clarifying, and contextualizing its key idea.
Practical calls to action
At a pragmatic level, what would it look like for us to elevate the status of a motivated explanation? Here are a small handful of suggestions.
- A PhD advisor can still assign a small problem from their field for a new student to cut their teeth on, but the deliverable would not be to write up a solution; it would be to present that solution to peers and faculty as a talk. It could be an already-solved problem which lacks clarity, or perhaps it’s a problem that lacks a solution, and the student uses AI to help find it. In either case, the student knows that on a certain date they have to understand it well enough to explain it, and that the desired output is for others to understand it as well. In short, even small problems could be treated like small PhD defenses.
- A leading figure (cough, Terry, cough) could enumerate a modern analog of Hilbert’s problems, instead focusing specifically on unsolved exposition problems. What areas are both important and lacking in the deeper understanding we desire?
- Written standards could clarify what constitutes a motivated explanation, aiming to make it nearly as verifiable as proof, so that the resolution of unsolved exposition problems can be recognized and celebrated in the same way proofs of open problems can be.
- Journals can be established which focus more explicitly on making results understood more widely throughout the mathematics community. Mathematical Discourse offers an interesting new example in this direction.
- Hiring and tenure decisions could place a higher value on writing great textbooks and similar work. Think of the AMS Steele Prize for Exposition, but at a more granular scale with an emphasis on early-career contributions in this vein.
The value of visible cultural shifts
I’d like to close with a broader pitch that visible culture shifts in math carry an intrinsic benefit right now with respect to the external image of mathematics as a career.
Many young students who are otherwise passionate about the field are afraid to pursue it now due to the uncertainty of what happens in an age of proof-generating machines. However, framed correctly, this is one of the most exciting times to go into the field, because there is nothing more exciting than entering a field when it is malleable and you have a chance to actively shape what it will look like in the future. Even if we completely set aside any potential benefits from AI to help with our understanding, young prospective mathematicians should feel energized knowing that they are entering at a unique point in history when they might play a real role in determining what the field as a whole looks like.
However, change like this is only exciting when it feels deliberate, whereas it feels terrifying if it seems driven by forces outside your control. As such, tangible action from the field’s leaders now to help define and clarify what the field is will reassure young entrants about who is in the driver’s seat, and that the status of the career does not depend on what entities produce the proofs.
Similarly, I also believe this is one of the best times to fund math. If the next chapter of math is ushered in by the drumbeat of two words “human understanding”, whatever changes are about to happen seem likely to amplify math’s value as a public good.
Why I didn’t sign the Fields medallists’ letter
[This post has been cross-posted to Terence Tao’s blog.]
When I was around 11 I heard for the first time about Fermat’s Last Theorem. I was immediately captivated by the problem statement, as well as by the accompanying story, and made a fairly serious attempt to prove it. And while, unsurprisingly, I failed, I learned a lot from the attempt. Blissfully ignorant of the fact that the case had been proved by Euler over 200 years earlier, I decided that that would be a good place to start: once I had sorted that out, I was optimistic that I would be ready to tackle the general case.
Since I still couldn’t really see where to start, I decided to simplify the problem further and concentrate on successive differences of cubes, with a view to showing that such a difference could not itself be a cube. At the time I did not know how to express what I was doing in algebraic language, so I did not explicitly try to prove that the Diophantine equation had no solution. Rather, I just worked out some successive differences and stared at them, trying to get some idea of why none of them was a perfect cube. (I should be clear that this story is a reconstruction of what I think probably happened given the few memory traces that remain half a century later rather than a completely reliable account.) At some point, I had the idea of taking the difference sequence of the difference sequence, and discovered that it formed an arithmetic progression. That felt like progress, so I investigated difference sequences a bit more and discovered, purely empirically, the rule that if you start with th powers and keep taking successive differences, then eventually you get to the constant sequence .
Somehow I never managed to turn this observation into a proof of Fermat’s Last Theorem, and later on my dream of solving it got replaced by other mathematical dreams. However, when I reached the point in my mathematical education where I was taught about taking difference sequences and about what happened to polynomials, I understood those topics much better than I would have if I had not discovered difference sequences for myself and spent happy hours playing around with them. I mention this story just as an illustration of the phenomenon that was strongly emphasized in this letter signed by 25 Fields medallists, that one learns a lot from thinking about a problem, regardless of whether one solves it.
In the end, however, I felt that I could not sign the letter, despite agreeing with much of what it said. Instead, it seemed better to do what I did with the Leiden Declaration and set out my own position in a blog post. But it should be understood that by doing that I am not setting myself up as a member of some opposing camp: indeed one of my worries at the moment is that the mathematical community might become bitterly divided, something I would very much like to avoid. Also, I agree on the fundamental point that we are facing a crisis: I just want to offer a slightly different analysis of what that crisis is. I don’t claim full originality for this analysis, as I know that several other mathematicians have already put forward thoughts that are similar to the ones I have, though (for what it’s worth) I have largely come to these conclusions independently.
On the subject of independence, it will perhaps help if I clarify that while I have contacts in the mathematics group at OpenAI, and have also been given early access to some of their models (typically only a few days before they have been released), and have been given free access to their Pro models once released, I have never been paid by OpenAI. I mention this in the hope, perhaps naive, that what I write will not be dismissed for ad hominem reasons. Another potential reason for my being regarded as “pro-AI” is that, as I have stated publicly several times, I have a group in Cambridge devoted to automatic theorem proving. However, that is actually more of a reason to be anti-AI, since our group has been trying to attack the problem of getting computers to prove interesting theorems by understanding as well as possible how humans prove interesting theorems, so now that LLMs can clearly do it without the help of such insights as we have had, one of the main motivations for our work has disappeared. To put it another way, we have had to swallow the bitter lesson (which of course we were always aware was a distinct possibility, even if the speed at which it happened has taken us by surprise). I do in fact think that it is still a very interesting and valuable intellectual exercise to try to gain this understanding, even if we can use LLMs as black boxes, but that’s a topic for another blog post.
So why didn’t I sign the letter? Let me extract a couple of sentences from it that express what I see as the principal argument being put forward.
But solving problems is only a tool and proxy for achieving the primary goal of conceptual understanding and insight. Forgetting this in the world of AI may turn the tool against the primary goal. Indeed, the mass production at faster and faster pace of “true/false” statements could destroy fertile ground instead of breathing life into new ideas.
Perhaps the main reason I didn’t sign is that I don’t fully subscribe to this view. Instead, I have a more complicated view, which I actually expressed in my essay The Two Cultures of Mathematics a quarter of a century ago, and which can be summarized by saying that there is a spectrum of attitudes in mathematics to the relationship between problem-solving and conceptual understanding. At one end of the spectrum you have mathematicians who are primarily motivated by the wish to solve problems, who see conceptual understanding as a very important means to that end. At the other you have mathematicians who are primarily motivated by the wish to attain conceptual understanding, who see problem-solving as a very important means to that end. I worry that the severe-misalignment letter could be seen as saying that the “right” attitude is to focus on conceptual understanding as the main priority — indeed, the above sentences say that more or less directly. But I think that there are mathematicians all across the spectrum, and that that is a good thing (or perhaps I should say that it has been a good thing up to now — the future is much less certain), and I don’t want to suggest to a large fraction of mathematicians, including myself, that their mathematical temperament is somehow “wrong”.
My own particular mathematical attitude is very similar to one that was beautifully articulated in a Twitter post by Jacob Tsimerman (another non-signatory of the letter), which, now that I look at it, says a lot of what I will be saying here. And that post in turn is a response to Daniel Litt, who is in my opinion one of the wisest commentators on mathematics and AI. His views are expressed in a later post here, which I deliberately didn’t read until finishing this one, and then found, as I expected, that there was significant overlap. I would also like to take this opportunity to recommend an excellent post by Noah Smith entitled The End of the Age of Heroes, in case you haven’t read it.
I have been talking so far about individual mathematical understanding, but I suspect that what concerns most of the signatories is less that than the collective understanding that results at least in part from the human activity of problem solving. My guess is that they would argue, completely coherently, that even if collective understanding is the primary goal, if many individuals are primarily motivated by the wish to solve problems, that’s absolutely fine and contributes to that collective understanding.
With that interpretation, the issue becomes slightly different: is it more important that the collective understanding of the mathematical community should be as advanced as possible or that there should be answers to as many problems as possible? Or are those two aims valuable in different ways, so that there is no point in declaring one of them more important? Or are they so inextricably linked that it makes no sense to argue that one is more important than the other? And when we say “important”, for whom are we saying it is important: for mathematicians, or for society as a whole?
I find these hard questions, so I don’t want just to declare an answer to them. (Do you see what I did there?) Instead, I’d like to try to offer at least some argument for any conclusions I come to, even if they are tentative. So let’s compare two scenarios. In the first, which I think is the more likely actually to happen, models become publicly available that are better at solving problems than virtually all mathematicians. If there are a few residual mathematicians who can do things the models can’t, even they work far faster if they make heavy use of the models. Thanks to this, in a short time we get answers to many questions that we have deeply cared about, but the rate at which we receive these answers far exceeds the rate at which the mathematical community can absorb them. In particular, most of the answers are obtained with zero effort from human mathematicians — just prompts such as “Thank you — please continue”.
In the second scenario, there has been an international agreement, for entirely other reasons, to block the public release of models significantly more powerful than the ones we currently have, and the mathematicians within the tech companies agree to hold off from getting their internal models to solve major problems. Instead, they take guidance from the mathematical community, solving problems only when asked to do so by some suitably representative body that decides that the benefit of receiving a solution of a certain problem outweighs the benefits of humans struggling to solve it over a much longer timescale.
I’d like to consider what the difference would be between these two scenarios both for individual and collective understanding. I’ll begin with individual understanding.
One might argue that for individual understanding, not too much would change if we are suddenly flooded with large numbers of big new results. There is already far more mathematics out there than I have any hope of understanding (for example, despite being fascinated when Fermat’s Last Theorem was proved, I have made no attempt to understand the proof), and even among the parts that I do understand, the parts that I understand because I myself discovered them form a very small fraction, though a fraction that I understand more deeply than anything else (at least temporarily — after a while I forget things and lose quite a lot of the understanding I built up). However, one change, which seems positive, from the perspective of the building up of individual understanding, would be that we would have a much bigger choice of results that we could choose to study. Also, if we found ourselves stuck on some point, AI would be able to help us. The main likely negative change is that we would probably cease to exercise that part of our brains that we use when spending months or years struggling with a difficult research problem, which can be hugely helpful in developing understanding.
I say “likely” because in principle there would be nothing to stop us thinking about very hard problems without consulting LLMs, but in practice it seems unlikely that people would put in the same level of effort that they do now. The situation might a bit like what happened with satnavs, where one could always decide not to use them, to keep the part of the brain active that can look at a map, learn a route, and follow it, but in practice most people succumb to the temptation to use a satnav. (In fact, I myself do try to keep that part of my brain active, and was rather proud of finding my way somewhere recently when I had briefly looked up the route on my phone but then forgotten to bring the phone with me when I actually went there.) But even if all we were doing was reading AI output, I think that the problem-solving muscles in the brain wouldn’t atrophy completely. When students are reading maths papers, I strongly advise them (and I think this is pretty standard advice) to read “actively” rather than “passively”, doing things like trying to prove the result for yourself, looking at the paper only when you feel stuck and need a hint, and even then just trying to get the hint and as little extra as possible. If one reads a paper that way, then one is constantly solving problems, some just exercises and some quite a bit harder. It seems likely that an LLM could get to know what our mathematical background is and feed us with just the right hints to allow us to work our way through a mathematics paper in this active way. Yes, we would lose the particularly deep level of understanding and ownership that comes with having solved a hard problem oneself, but it isn’t clear to me that progress in mathematics would suffer as a result. I would be very interested to hear counterarguments to precisely this point. That is, I would be interested to know what use that level of deep involvement with a proof might have in a world where AI is much better than we are at finding proofs.
How about collective understanding? Let me quote a bit more of the letter.
Indeed, the mass production at faster and faster pace of “true/false” statements could destroy fertile ground instead of breathing life into new ideas.
Often these solutions are announced in a rush, leaving no time for a proper writeup, the isolation of new methods and ideas, and citing relevant previous work of others. As in all creative professions, this raises severe attribution and plagiarism questions. Moreover, without the willing mathematicians who must take care of their development and integration into the mathematical canon, AI-conceived ideas would never become fully alive and the crucial human transmission chain between mathematicians would be lost.
I’ll come back to questions about proper citation and focus on what I take as the core worry here: that if results are proved too quickly, then the digestion process will become impossible. I am definitely worried that results will not be properly digested, but for different reasons.
A first remark is that what AI is producing is not just true/false statements: we now know not just that the Navier-Stokes equation with smooth forcing admits finite-time blow-up, but we have a proof of that, which builds on a great deal of wonderful work done by human mathematicians. Many people used to express the worry that AI would solve our favourite problems with utterly opaque proofs, but that has not turned out to be the case, even if their write-ups often leave plenty to be desired. (Incidentally, I see these inadequate write-ups as almost certainly a temporary annoyance and therefore not as a fundamental threat to mathematical practice or future mathematical understanding.)
Secondly, even if the volume of new results is large, mathematics is a highly specialized discipline, so mathematicians can work in parallel. If, for example, we had to digest 1000 important results in a year that were roughly uniformly distributed across mathematics, then most sub-communities of mathematicians would probably want to understand around 30 of them, and for each individual problem there might well be only a small handful of specialists who would be obvious people to take the lead in reaching this understanding, with that handful varying from problem to problem. So it would be a big task, but not necessarily an impossible one.
In this context, it is worth thinking about the huge volume of output of human mathematicians, which seems to have been increasing recently, even before AI. While I have sometimes heard complaints about this, I have certainly not heard suggestions that human mathematicians should slow down the rate at which they prove interesting theorems. That may be partly because the authors of those theorems take the trouble to write their papers well and give good talks. But what about the large quantity of papers, including important ones, that are not written well and whose authors give incomprehensible talks? That can be annoying, but it is a familiar annoyance and not one that we think of as a crisis.
A third point is that even if the volume of AI output is too big for us to be able to digest it properly, that is not necessarily a bad thing. To draw an imperfect analogy, there is now more content available on streaming services than anyone could possibly watch, with the result that there is almost certainly some very good content out there that is hardly watched at all. But that isn’t obviously a worse situation than if there were far less content and all of it received the attention it deserved. Returning to mathematics, if there were too much AI-generated content for us to be able to digest it, then we could choose which parts of it we wanted to digest.
For that we would need to have some idea what was there (a situation a little similar to how human mathematicians typically learn quite a lot about what results are known in their area even when they do not understand their proofs in any detail). One way one could try to achieve that would be to create a well-designed database, probably with AI help. But perhaps that would be unnecessary, and instead one could simply talk to an LLM and ask it to give a bird’s-eye view of whatever area of mathematics one wanted to understand in that knowing-what’s-there way.
The fear seems to be that some very interesting and important parts of mathematics will be discovered by AI and then overlooked, when had they been discovered by human mathematicians they would not have been overlooked. And that may even be the case, but what matters is whether the amount of interesting and important mathematics discovered by AI that is not overlooked will exceed the amount of interesting and important mathematics that would have been discovered and properly digested by humans with AI having played a more modest role.
In short, it seems to me that while a flood of “big” AI results would be likely to increase the amount of important mathematics that was not properly digested, it would also be likely to increase the amount that was properly digested, which seems like a pretty good bargain.
Let me quickly discuss the problem of AI not properly crediting human mathematicians. I agree that this is a serious problem right now, but it is another problem that I see as temporary. Very soon, the whole “credit system” will surely collapse, since finding an amazing proof will be no more of an intellectual achievement than when a citizen scientist spots through their telescope an object that turns out to be a new comet. Until that happens, it is important to give humans the credit they deserve, since careers can depend on it, but that will soon cease to be the case as well. I have to say that I’m puzzled that this problem exists, since I would have thought that if you asked an LLM to look at a proof and tell you which ideas in it are close to ideas that are in the literature already, it would be extremely good at that task. I hope the answer to this conundrum is not that people have been in such a hurry that they have simply not taken the trouble to do this, but I fear that it might be, at least in some cases. If so, then those who have been careless deserve to be criticized, but it is a minor matter compared with the survival of mathematics, especially if the lack of citations is swiftly put right.
Does all this mean that I am optimistic that mathematicians will end up digesting at least as much mathematics in a post-AI world as it would have if AI had not been able to prove major theorems? Not exactly. But my worry is not that we would be unable to do it, but rather that the social structures that currently support this digestion process will be destroyed and not adequately replaced.
One way that might happen is that AI disrupts society so much, or even kills vast numbers of us, that the preservation of something like the current mathematical tradition ceases to be of any concern: all that will matter is the survival of the human race. But that again is a topic for a different blog post (which in fact I am in the middle of writing).
Let’s assume instead that we get lucky and that AI remains more or less under control. My worry then is that we do not manage to transmit what we know to a new generation of mathematicians. Speaking for myself, my main motivation for becoming a mathematician was the dream that I would solve unsolved problems — the more famous the better. I have also always greatly preferred directly thinking about a problem to reading books and papers and generally learning the mathematics of other people. (I’m not saying that’s good, but just stating a fact about myself.) If the dream of solving a famous problem had not existed, I’m not sure whether I would have become a mathematician. I don’t completely rule it out: maybe what really motivated me was that I had an aptitude for the subject and that solving problems was a way of getting respect from a small group of peers. And maybe I could have tried to gain that respect in a different way, such as thinking very hard about an area of mathematics until I was able to demonstrate to others just how well I understood it. But I’m not sure how motivating that would have been for me. I very much hope that there is a pool of young people for whom it will be a powerful motivation, because I think the survival of a human mathematical tradition may well depend on it.
Thus, the primary risk, as I see it, is that a lot of people who would have done a PhD in mathematics and gone on to become custodians of the mathematical tradition will no longer wish to do so. Those of us who have PhD students, including me, need to try as hard as we can to come up with imaginative ways for them to use their time productively (in consultation with the students themselves, obviously). Whether or not we do a good job with that could make a huge difference to the future of mathematics. A related risk is that the perception among policy-makers will be that mathematicians are no longer needed and that funding will become much harder to come by: we urgently need to come up with good ways of explaining the value of having a large pool of human mathematical experts, even if it is no longer part of their role to find new proofs of theorems.
A final reason that I didn’t sign the letter is that I wasn’t really sure what it was demanding that isn’t happening already. It seems likely that in a matter of not very many months LLMs will be released that are able to solve major mathematical problems, and they will presumably have no trouble at all with more run-of-the-mill problems. However much we might regret that, there is no chance that the impact of such models on mathematics will persuade AI companies to stop their release, though perhaps concerns about safety will lead to some delay and give us a bit more time to work out how to adapt. Assuming that they are released, there will be a flood of new results, whether we like it or not, and it will no longer be the AI companies producing them, though perhaps the pattern will continue that the AI companies will have access to more powerful models and so will obtain more than their fair share of headline results. So I felt that there was nothing to be gained from criticizing AI companies for generating too many solutions too quickly. In fact, it may well be that all that does is bring forward by a couple of months what was going to happen anyway, and perhaps it will even allow the results to be released in a more controlled way than they would have been if they had been discovered by random people once the models were publicly available. Under the circumstances, I think the best we can do is recognise the changes that are coming and try to work out the least unsatisfactory way of dealing with them.
“Deep theorems were scarce and difficult and so became an effective mechanism to identify deep thought. AI has broken this system.”
[This is a guest post by Bryna Kra. This blog post was initially written in a different file format and converted using AI. — T.]
For generations, mathematicians have treated the production of theorems as a clear measure of success. The stronger the theorem, the deeper the proof, the more surprising the connections, the greater the achievement. Positions and prizes are based on these theorems and the mathematicians making the breakthroughs set the directions for future research.
But theorem production was only part of what we cared about. It was a proxy for something harder to measure: understanding. A major breakthrough meant that years were invested in learning a subject and uncovering hidden aspects, accompanied by work to make the answer apparent to the community. Deep theorems were scarce and difficult and so became an effective mechanism to identify deep thought.
AI has broken this system.
Artificial intelligence has lowered the cost of producing sophisticated proofs. As the models improve, the list of deep conjectures that fall will grow. Someone with little knowledge in a field can now generate a manuscript that reads like a polished article, cites the literature, and combines techniques that would have taken years to master. The machines are producing solutions faster than the mathematical community can read them, much less digest them.
This week the changes arrived in my field. The Nivat conjecture is a beautiful problem at the intersection of combinatorics and dynamical systems. Imagine an infinite grid of square tiles, with each tile colored either red or blue. Look at every rectangular window of a fixed size, say tiles wide and tiles high, and count how many different color patterns show up in that window. The conjecture says that if there are at most such patterns, then the entire coloring repeats under some fixed, nonzero shift of the grid. In other words, some precise limit on local behavior gives rise to simple global behavior, periodicity.
Though this statement can be explained without formulas or higher mathematics, the conjecture resisted our efforts for nearly three decades.
In work that we first circulated in 2012, Van Cyr and I proved a partial result, showing that periodicity holds when the number of patterns is at most half the area of the window. We translated the combinatorial question into one about the geometry of dynamical systems, studying ways in which information in the system does (or does not) propagate. Progress on challenging problems does not happen in a vacuum, and our work began by studying and building on earlier results in the literature [1, 2, 3]. Soon after, Jarkko Kari and Michal Szabados developed a different approach, viewing the combinatorial question as an algebraic object, using certain types of polynomials to find the infinite patterns.
To those of us who worked in the area, this was a tantalizing situation. Two different methods to approach the same problem seemed to shout to us: find the conceptual bridge between them. We tried, without success, and moved on to other problems. Colleagues continued pushing the boundaries of what we knew, combining the varied techniques in increasingly powerful ways [1, 2, 3].
But synthesizing different approaches is one of the ways in which frontier models excel. A machine does not need years to absorb distinct areas of mathematics before it can find the connections. It can translate notation, compare partial results, and try thousands of combinations without being discouraged. The Nivat conjecture has been proposed for inclusion in Google DeepMind’s Formal Conjectures project, turning the problem into a target for automated reasoning and making more people aware of the question.
The manuscripts have started to arrive. This week I received several purported proofs of the full conjecture from researchers making their first foray into the subject. Some acknowledged using AI, but only for polishing the English or checking the proof. When I asked if the authors would meet over Zoom to explain their arguments, none accepted the invite.
Silence does not mean that the authors used AI and a proposed proof is not a theorem. But merely disclosing AI use and checking correctness of the proof, both of which are essential, miss the point. Suppose these results are correct. That is a mathematical contribution. Yet it opens more questions. What did we learn? What did the authors contribute? What impact will this result have?
A proof is more than a certificate that something is true. Instead, it is a story, a picture, an insight, an explanation. A proof highlights novel ideas and opens new directions for what we should ask next. It becomes part of the toolkit of the community. A deep theorem changes how we think, not because of its statement, but because of what it teaches us. As Bill Thurston wrote on MathOverflow in a 2010 response to a question about what mathematicians do: “The product of mathematics is clarity and understanding. Not theorems, by themselves.”
These new texts certainly contain knowledge, just as a data file of bits of 0’s and 1’s contains information. But human authorship is more than the string of words that prove something. It is intellectual responsibility, understanding of a new concept, the ability to respond to questions, the sorting of new ideas from existing scaffolding. And perhaps most importantly, it is incorporating the result into the corpus of our understanding.
I welcome AI-generated proofs and believe that human-only proofs will become rare. This is a profound change of perspective, but mathematicians have always used external tools and machines. AI is already a powerful mathematical instrument and we, as a community, need to incorporate it responsibly into our practice of mathematics. The question is not if machines should participate in discovery. They already do. The question our community needs to address is what we value in our mathematics now that certain kinds of discovery are no longer a scarce resource.
The tradition of journals publishing novel results, hiring committees rewarding those publications, and prize committees recognizing the person who completed the result do not suffice. Peer review is a slow process that relies on the unpaid expertise of a small group of overworked researchers. Any single submission can be handled, but the current volume is drowning the reviewers. The careers of young researchers depend on producing theorems, with incentive to produce as much as possible as quickly as possible. The time for deep reflection, necessary for deep understanding, does not happen when someone types the question into a model and quickly produces a manuscript.
All of these issues existed before AI entered our field. Unfortunately, we no longer have the luxury of addressing them in a leisurely manner.
Mathematics is at a watershed moment. Our incentive structures are misaligned with the prolific output of highly accessible AI models. Careful verification and stellar exposition will not happen when there is no reward for those tasks. Producing a paper is no longer enough: authors must be able to explain the proof’s mechanism and how they arrived at this point. Journals need to distinguish and credit the roles of discovery, proof, formalization and explanation. Exposition must become an important component of any major intellectual achievement. We advance mathematics not by what we write, but by what we learn, check, apply, and teach others.
Finding the right concept, crafting the right definition, and formulating new questions have always been part of our intellectual work. The creation of collective understanding, the ability to communicate ideas in a lecture, and the sharing of ideas informally that spark research have always been admired. Such achievements are harder to quantify than theorem production, and so have been treated as by-products. But these are the key aspects shaping the future of our field and their value is at least as important as knocking down the next conjecture. Our incentive system needs to be realigned to reflect what we truly value, and not just what we can easily measure.
Mathematics is the canary in a coal mine for all parts of intellectual life. Its claims can be carefully checked and so we are witnessing the disruption in real time. But when AI disrupts the tried and tested standards in a field like mathematics, what will it do to fields where evidence and interpretation are more contested? Science, law, public policy, and art will all have to find ways to assess what is valuable amidst the abundant output.
The arrival of machine generated mathematics does not replace the role of human expertise. Instead it reveals why we valued that expertise. The goal of our discipline is not simply to resolve conjectures, but to enlarge our capacity to reason and open our minds to new ways of understanding. It will take our entire community to create new standards, discover new ways of creating mathematics, work with AI models to open new horizons of knowledge, and develop the tools to approach them. It’s truly an exciting time to be a mathematician.
Happy, those able to know the causes of things
[This is a guest post by Nestor Guillen, crossposted from his blog. This blog post was initially written in a different file format and converted using AI. — T.]
Keywords: LLMs, cultural technologies, skateboarding, creative communities, AI industry.
A note on the timing of this essay. I started writing this in the middle of the ICM, and it has taken me this long to finish writing, see also my shorter post from August. In any case the bulk of the essay was done in August and shared with a few close friends for feedback. Earlier this week, Tristan Buckmaster announced a breakthrough finding, done in collaboration with Levent Alpöge, of a finite-time blow up for 3D incompressible Euler with forcing. He also made serious allegations of misconduct by OpenAI (see announcement). The next day, OpenAI announced a finite-time blow up result for 3D Navier-Stokes with forcing, obtained from their internal LLM. There has been lots of press coverage, I recommend Kenneth Chang’s reporting. I have decided not to change the opening of the essay and instead make a note of the news here. Tristan gave considerable credit to the ideas of Córdoba and Martínez-Zoroa in the blow up construction; I found this important detail to be quite validating of the perspective on LLMs I advocate in this essay, and found that to be further reason to leave the content of the essay largely unchanged.
What would be the significance if one day we woke up to the news that output from a Large Language Model contains a definitive answer to the question of blow up for the incompressible Navier-Stokes equations, and to the question of where the zeros of the Riemann zeta function lie? (Or, to mention problems dear to my heart: a proof of a nonlocal version of the ABP maximum principle or a resolution of the Mahler or Komlós conjectures?)
Many of my fellow mathematicians consider this a kind of nightmare situation, and I think this is a very valid position. Here I would like to make a different case, however. Personally, I have come to believe that the news of a Large Model output containing an interesting, insightful, novel mathematical idea can and should be received as a source of communal pride for mathematics and for mathematicians. I also believe there can be a world where LLMs are not harmful but rather supportive of mathematicians as we follow our prime directive: understanding things.
My intention in this essay is to elaborate on what I mean by this. I believe this perspective can assuage fears many of us feel amid the daily hype and doom narratives around AI and the alleged impending obsolescence of mathematicians. I will not be talking about the less existential / philosophical and more practical / material concerns here, such as challenges to the funding of the mathematics occupation, how journals will work, changes to the training mathematicians, and so. Those issues are also important, but I think clarifying the perspective I push here will make the discussion of the more practical matters easier.
Bill Thurston touched on the same idea in his famous 1994 essay, “On proof and progress in mathematics”, which is more relevant than ever. Thurston observes that mathematicians (as a group) prove theorems and solve problems, but that is not all they do: they also absorb these solutions, and find new uses for them once they are proven. This “digestion” [1] in turn allows for the next wave of breakthroughs, and thus advances our understanding of mathematics.
This is not secondary to what mathematicians do, but our very purpose. The value of mathematics to humanity lies not in determining whether a statement or conjecture is true or false, but rather in the understanding gained as we make those determinations (even if the facts themselves are also of great value). To put it briefly, mathematicians are people who try to understand the causes of things.
Felix, qui potuit rerum cognoscere causas
LLMs, the technology AI, the product from large companies
The conversation about LLMs and generative AI within the mathematical community at large must deal with an obstacle before we can get to the substance of the matter. That obstacle is the current and unfortunate identification of generative AI, the technology, with generative AI, the specific product and vision from the large AI companies. It is imperative for us to make that distinction and separate the two. The private AI lab vision is what most people have in mind when discussing AI, and this is no accident, for this vision has been aggressively pushed into every part of our daily lives.
It might seem impossible at the moment, but we can imagine an alternative world with many parallel approaches to the use of LLMs developed by the grassroots, reflecting the variety of circumstances across different communities, industries, and institutions. A seed for this world is the emergence of open-weight models. This is a matter for a different essay.
So in all that follows, I will crucially rely on the above distinction, so that when I say LLMs or “Large Model” I am talking about the technology itself, and very much not any of the major AI companies or their products.
With all this throat clearing behind us, let me get to the matter at hand. Why should the output of an LLM containing new and exciting mathematical ideas be a source of pride for mathematicians, and not of anxiety? My answer owes a lot to this very interesting article by Henry Farrell, Alison Gopnik, Cosma Shalizi, and James Evans. I quote here one of their central assertions:
…Large models should not be viewed primarily as intelligent agents but as a new kind of cultural and social technology, allowing humans to take advantage of information other humans have accumulated.
I will do my best to apply the ideas in their article to mathematics and the work of mathematicians. If, rather than continue reading this, you stop and go read their article and Thurston’s article instead, I would consider this essay a success.
Cultural and social technologies
Humans have always lived in a complex world and their survival has hinged on their ability to process information in amounts beyond what any individual can handle — and it is likely this has been so nearly as long as humans have existed. Herbert Simon’s concept of “bounded rationality” captures this well: in real life, unlike in the simplest control theory models, we have limitations on the time available to decide and on the accuracy of our information. This means that an agent making decisions must give up optimality and settle for “good enough.”
Simon saw that one implication of bounded rationality is that humans have to (and do) develop systems to process information, as a group. We see that humans have built states, stories, bureaucracies, markets, libraries, and more. These are cultural and social technologies. They are means by which humans (not individually but as a group) navigate amounts of information so vast that they cannot be handled by any individual on their own. An important feature of these technologies is that they are to some extent forgetful: you ignore or downplay smaller details in order to make a bigger picture easier to comprehend.
Markets, governments, and bureaucracies mean each human can focus on a few things and still get the aggregate benefits. Thanks to markets, governments, and bureaucracies, I can rely on being able to buy fresh fruits and vegetables around the block from my Manhattan apartment. Thanks to markets, governments, and the field of modern medicine, we now have high confidence that a newborn child will live to be an adult — in stark contrast to what was the norm for most of human history. Cultural technologies are powerful and awe-inspiring things. It is understandable that we tend to anthropomorphize them so often.
This week I started teaching my graduate-level optimal transport course. I wrapped up the first class by showing one neat application of optimal transport: a “one-page proof” of the isoperimetric inequality in . Of course, the proof being just one page depends on what we are assuming as given.
This is standard, not only when teaching advanced courses but also very much in day-to-day research, and it is made possible by the power of cultural technologies. In the proof I don’t have to explain from scratch the notion of Lebesgue measure or sets of finite perimeter. When in the middle of a proof I say “I will do the details only when the domain has a smooth boundary, and I will assume the Lebesgue measure is without loss of generality thanks to scaling,” I am relying on a shared trick that I can trust a first-year PhD student to be familiar with.
All of this is to my benefit and the students’. We can focus on “the one page” proof with the interesting new idea: how Brenier’s theorem provides a map through which the isoperimetric inequality is essentially reduced to the arithmetic-geometric mean inequality.
So, here is basically what I showed on the blackboard: assuming has Lebesgue measure , we want to prove that
where is the -dimensional Euclidean ball with the same volume as , so . Then, I tell my class how Brenier’s theorem says there is always a map from to , that this map is volume-preserving, and that for some convex function . Then, since and all the eigenvalues of are non-negative, the arithmetic-geometric mean inequality says that everywhere in , and so
Then, the divergence theorem and the fact that (where = volume of the unit ball) guarantee that
For a ball of volume , scaling shows that , so we conclude that , as we wanted.
In interacting with this proof, we are engaging in a conversation through space and time with many mathematicians, but especially with Knothe (1957) and Gromov (1986), followed later by separate works of Brenier, McCann, and Trudinger — and even with my younger self in a set of lecture notes with McCann, where we go over the above proof (see concretely Section 1.6). The conversation is mediated by a common mathematical culture, which encompasses the common definitions and methods I can count on the graduate students to know. I don’t have to prove the divergence theorem or the change of variables formula to justify these computations, and I don’t have to worry about the smoothness of , since I can work by approximation.
Without having to go into all that detail, I presented on one blackboard the idea that the Brenier map takes you from the arithmetic-geometric mean inequality to the isoperimetric inequality, and I did this in the last 15 minutes of the class. These are cultural technologies in action! Which cultural technology in particular? In this case it is the technology of a scientific community with standards in what is taught, the mathematical literature, and so on.
Large Language Models as cultural technologies
Farrell et al.’s main point is that LLMs are cultural technologies, just like markets, bureaucracies, and scientific disciplines. Accordingly, the framework of cultural technologies is a very useful one as we seek to understand and navigate the impact of Large Models on our disciplines.
LLMs indeed are a lot like the example from my class, and in a way they are “wrappers” around an older cultural technology: the scientific literature. When you ask an LLM a question about the isoperimetric inequality, the ABP maximum principle, or a PDE, you are effectively interacting with every person who has thought about and worked on those things — provided the data the LLM was trained on contained those contributions in one way or another. Surely the lecture notes I reference above are in the training data of most LLMs in use today, just as a vast portion of the digitized mathematical literature (across all disciplines) was part of the training data. LLMs are lossy, so while a model might not be able to reproduce a given paper exactly as in the training data, it can reproduce large parts of it.
What sets LLMs apart from older cultural and social technologies is the fact that we have digitized all this data, and can now explore it, reproduce it, and recombine it with ease by sending text commands (prompts) from a computer terminal, or even a phone. LLMs are, in essence, top-tier information-processing tools. They are also an excellent tool for generating conversational level-text, and for coherently combining patterns of information even if they come from very different sources.
By a product-design choice at a few private labs, the output of most LLMs is conversational in tone and made easy to anthropomorphize. These properties have amplified an ELIZA-like effect that reaches even people with advanced training in mathematics and computer science. In a way, anthropomorphizing this technology — which, as some have argued, did not have to be the case — does everyone a disservice: the conversational setup of the LLM gives the appearance of a single “thinking” entity, and buries the work by vast groups of people and which are being recombined. This appearance is a feature not of LLMs, but of the LLMs as the AI companies want them. We are not dealing with a very smart, genius-level intelligence; we are interacting with something greater (and yet somehow more human): the aggregate creations of millennia of human ingenuity and creativity, digitized and efficiently compressed — all mediated by state-of-the-art statistical and generative algorithms.
In particular, any mathematical proof found in an LLM output is the collective outcome of lifetimes of work by mathematicians. It must be welcomed as the fruit of our shared mathematical heritage, belonging to all of humanity, carried through millennia by people who at various times went by names other than mathematician (natural philosophers, astronomers, geometers). This applies very much to LLM output containing what we might consider novel proofs or ideas, for in the end, it is all an outgrowth of that gigantic body of knowledge. Now there lies a scaling law that has been going for centuries. The recently arrived software is just a kind of harness or wrapper for it.
Now, let’s reconsider the “nightmare scenario” I opened this essay with. One day we wake up to the news that the output of a Large Model contains a resolution to the Riemann Hypothesis. We are now equipped with the ideas of Farrell et al., and especially the idea of cultural technologies. Reading the news under this framework, we are learning a new mathematical fact, but we are also learning that this fact is the fruit of the collective efforts of researchers in the field up to this moment. The problem was cracked by the ideas of one or two (or maybe many) researchers to the point that a resolution could be reached by the sampling and recombining of ideas an LLM is good at.
This leaves us with a delicate question: what ideas are reachable from the totality of ideas in the mathematical literature at a given time? I am going to borrow from geometry and loosely use the metaphor of “convex combination” of ideas to discuss this.
Convex hulls of ideas
A convex body containing another, from Geometry and the Imagination by Hilbert and Cohn-Vossen
The capabilities of LLMs have progressed quite a bit in the last couple of years, and dramatically so in mathematics this year alone. The rapid improvement makes this discussion difficult. The question “what can LLMs do?” is a rather moving target, at least for the time being. This being said, let’s try to think of a related question that might not be as LLM-dependent, and has a chance of being useful, even if it is still vague.
Ideas sometimes are very clearly compositions of other, more fundamental ones. Sometimes a new idea really stands apart from what came before. What would the finding of a solution the Riemann Hypothesis, or of the 3D Navier-Stokes smoothness problem, within the output of an LLM, tell us? One toy model (which might not reflect empirical reality) is that the result might belong to some sort of “closure” or “convex hull” of the available literature. In broader terms, one could pose the following question: given a collection of ideas, what ideas could be said to stand squarely beyond them, and which are lying closer to their “convex hull”?
I will not give a more concrete definition for this, because I am not able to. However, I can discuss by way of example what this definition ought to cover. I also do not use “convex hull” to mean that the point is derivative or that the finding of a point of something inside a convex hull to be easy — high dimensional convex sets are deeply complex objects!
At the very least, we can agree that the proof of the infinitude of primes with coprime is certainly not in the “convex hull” of number theory that preceded Euler and Dirichlet: using an infinite series to quantify the “amount of primes” is precisely the sort of idea that stood squarely outside of what was considered number theory until it was proposed. However, the question of the infinitude of primes of the form ( coprime) could be posed within the number theory that existed before Euler. In fact, this was known already for special choices of and , say for primes of the form . Likewise, number fields, ideals, and their unique factorization are all ideas lying well beyond all the number theory that preceded them. However, once we have these ideas together, it is slightly less outside the convex hull (or maybe it is even inside the convex hull) to develop a proof of the fact that every ideal class contains infinitely many prime ideals — a combination of Euler’s and Dirichlet’s ideas with the idea of, say, Gaussian integers.
From what I have seen (so far) in research areas that I am closest to, no LLM has produced anything remotely like Euler’s introduction of the methods of analysis in the study of primes. One possibility is that LLMs are not even being used so far for such things. In fact, one could put it like this: Euler was not simply answering a well-known question, but rather saying “here is actually a whole new question that we did not know we could ask, and is actually a very important one that connects to older problems!” The framework of LLMs as cultural technologies, to the extent that it reflects existing LLMs, suggests that this is a type of contribution we cannot expect to find in their output.
According to Farrell et al., cultural and social technologies are preceded by individual capabilities, which are then aggregated and transmitted by the technology. In their essay, they note that “without innovation, there would be no point to imitation.” Suppose Euler had access to an LLM trained on all of the mathematics known up to the moment he started working in number theory. Without an individual first introducing the analytical perspective on prime numbers, it would be very hard for our hypothetical LLM to innovate analytic number theory into existence, and then use this innovation to solve the “conjecture” [2] on the infinitude of primes that are congruent to modulo , for any coprime. All the same, the framework is compatible with the happy possibility that a person working with an LLM might have a eureka moment where they, the person, arrive at this innovation.
Here is another example, of a slightly different character. One of the most important theorems in PDE is the Krylov-Safonov theorem for non-divergence form elliptic and parabolic equations. For those not in PDE, these are versions of the Laplace equation and the heat equation, except the “coefficients” in front of the partial derivatives are not constant but rather change from point to point.
Such equations, linear and nonlinear, appear in the study of stochastic processes as well as in optimal control and geometry. The Krylov-Safonov theorem produces a Hölder regularity estimate for solutions of the PDE that only depends on the uniform ellipticity — what is notable is that this estimate works even if the coefficients are potentially very discontinuous. Consider, for concreteness, a linear equation such as
When the coefficients are smooth, the regularity of solutions to this equation can be understood from the theory for the Laplace equation. Heuristically, the coefficients being smooth means that as you zoom in near a point, the coefficients are closer and closer to being constant, which puts you just an affine change of variables away from Laplace’s equation. Therefore the good regularization properties of Laplace’s equation carry upward to the larger scales and so one proves is smooth. This estimate, however, depends very much on the modulus of continuity of the functions .
Now, I cannot go on at length about this here, but obtaining estimates for that do not depend on the continuity of the coefficients is extremely important. In fact, the importance of this question cannot be overstated — such estimates have profound implications in the theory of stochastic processes, in geometric analysis and differential geometry (say, for the many PDE related to curvature), in stochastic homogenization, and more. For this reason the Krylov-Safonov theorem is one of the most important PDE results we have.
The Krylov-Safonov theorem, however, depends on a mathematical fact that comes from a very different area: convex analysis. There is an estimate by Aleksandrov that says the following: for denoting the unit ball, given any function convex in and with on , we have
where denotes the subdifferential of the convex function. This estimate holds for general convex bodies, not just , and in this generality it is equivalent to the reverse Blaschke-Santaló inequality from convex geometry. The estimate is the key component in what is known as the Aleksandrov-Bakelman-Pucci (ABP) estimate, which in turn is essential for the Krylov-Safonov theorem.
To state it in simple terms: without this fact from convex geometry the Krylov-Safonov theorem cannot be proved, not in any way we know today. This is a significant statement, as many important theorems in analysis have multiple, independent lines of attack that do not necessarily use the same tools. This would be like asking how to prove results in geometric analysis and probability without knowledge of the Poincaré inequality. Without the advances in convex geometry done many decades earlier, the theory of fully nonlinear elliptic PDE, which seems at first sight so remote from convex geometry, could not have advanced the way it did onward from Krylov and Safonov.
Imagine a parallel universe where we came to LLM technology before we knew about the Krylov-Safonov theorem, so that the question of regularity for non-divergence elliptic equations without dependence on coefficients is considered an important open problem. The cultural technologies framework suggests that in this universe an LLM (at least at the current level of power) would not be able to settle this open problem if the mathematical literature in this parallel universe does not have a well-developed field of convex geometry/analysis. By the same token, if in this parallel universe there is a well-developed convex geometry literature, including the Aleksandrov estimate, then it is plausible an LLM could bridge the two things and arrive at the ABP estimate and Krylov-Safonov’s theorem.
Thinking in terms of “convex hull of ideas,” one sees that whether an LLM can solve a given mathematical problem is in large part a function of the status of the mathematical literature at any given time, and is not just a function of how large the LLM is. With this perspective, the event of an LLM output containing the solution of a hard problem ought to be welcome as a happy occurrence: an LLM helped us dig out a gem buried somewhere in our mathematical heritage, or in the convex hull, if you will.
In fact, this has already happened: less than a year ago, we heard news of “Erdős problems” solved by an LLM, which then turned out to be instances of the LLM finding solutions in old (and in some cases, forgotten!) papers. At the time, this realization was seen as a defeat for LLMs. Now, that situation illustrates the perspective I am advocating here, in miniature.
Creative communities during turbulent times
The framework of LLMs as cultural technologies, together with the “convex hull” of ideas it inspired, has addressed my own philosophical concerns with LLMs. It has reinforced my optimism that the mathematics I value will endure. However, this does not change the fact that the mathematical community is facing serious and urgent challenges. The AI companies’ aggressive and narrow-minded push of their vision is hurting the field, exacerbating old fault lines and dysfunctions of our profession, and creating new ones.
This is not the first time a creative community has faced a turbulent period or a big change brought by new technologies or economic realities. There are lessons we can learn from past instances, in mathematics and in other disciplines. One can think of painters facing the advent of photography, musicians facing the arrival of the synthethizer, or mathematicians themselves — think for instance of the creation of the arXiv in the 1990s. However unusual it might seem, chapters from the history of professional skateboarding present uncanny parallels to the circumstances in mathematics, and this is a history with many teachings and warnings for mathematicians.
Rodney Mullen performing “the Impossible” trick in an episode of the Physics Girl (Dianna Cowern)
By the 1980s, skateboarding had been around for several decades. It had professional contests, sponsorship deals, and a specialized press. Like mathematics, skateboarding had an ecosystem that allowed practitioners to develop their craft and advance it, and even make some money along the way. It also had subfields, like freestyle, vert, and street skating. In 1984, Kevin Harris, who came from the subfield of freestyling, observed street skaters with only “six months of experience” were already earning more than Rodney Mullen, who at the time had tens of thousands of hours of experience and a nearly perfect record in professional competitions. Harris has recalled his reaction to this observation as: “Okay, this is the death of freestyle and it’s going to hurt vert skating hugely.”
The recession around 1990 dealt the finishing blow to the professional skateboarding ecosystem as it existed at the time. Within a few years, the discipline lost many of the venues and big public events it was built around, even largely losing its version of “funding” — sponsorships. This crisis drove many pros into early retirement. One could argue that with championships and sponsorships gone, the metric for measuring success was gone, and with it all the incentives for participation in competitive skateboarding. However, this was not the only metric, nor even the main one, for skateboarding continued and evolved in the coming years.
Without public contests with large audiences and related sponsorships, skateboarders had to find some other means to communicate and show their work. The economic and social problems in this case were solved with a technology that already existed. Skating videos became the new medium by which skateboarders recorded and communicated their work. Before, if a skater spent countless days perfecting a delicate trick, they could showcase it to at most a few hundred, maybe a thousand people during a championship. Now, a videotape showcasing the work of several skaters could be seen by tens of thousands of people. Before, a skater who did not win several championships was likely to go unappreciated. Now, a skater needed only one good “part” –a single skater’s segment within a video– to be recognized and contribute new tricks to the skating community. The most successful “parts” would introduce a new trick or a new twist on an old one that would get replicated and adopted by others.
I find it notable that the resolution to the crisis in skateboarding made it a bit less a competitive sport and a bit more a collaborative, creative field, with videotapes playing the role of journals [3], and a “part” playing the role of the articles in the journals. This turbulent transition turned out to be for the best for skateboarding as a discipline, while it also turned out for the worse for many individual skateboarders. This ecosystem remains largely unchanged today (with internet videos replacing videotapes). The motivation for a skater to participate in the community, aside from their love for skating itself, was that their contributions could be showcased in video, and their ideas could be shared with others who could try to repeat them or recombine them in new ways.
Mullen said that an invention only becomes an innovation once the community adopts it. This provides a measure of the value of a work. In mathematics we often talk about “the impact” of a paper or idea, and this is completely analogous to a trick being impactful: has the community taken this idea, and used it, and modified it further? Has this work made it easier for newcomers to the field to understand the state of the art?
Reclaiming our mathematical heritage, and the future of mathematics
One lesson I take from cultural technologies is that the most precious component in a Large Language Model, the part most responsible for whatever value lies in its output, is the human-generated work, the cultural and technical heritage encoded in it. Throughout the summer the largest AI companies have been treating the solution of high-profile mathematical problems as a form of PR for their products. We as mathematicians need to lean into our curiosity and love of math and welcome this new knowledge, without letting its provenance ruin our enjoyment of it. We must counter the damaging misconception that “AI has beaten mathematicians at solving problem X,” which is being promoted by the corporate AI labs. Instead, we can celebrate how our understanding of mathematics, together with our creativity, lies behind the gems of mathematics LLMs are helping us uncover.
Earlier in this essay I touched on the importance of separating LLMs (the technology) from AI (the product and vision of the AI companies). Well, we need to also make this a reality for our mathematical infrastructure. It needs to be a top priority of the community to promote and work towards a world with open LLMs — think of the LLM equivalent of free software. It would be quite bad if our profession became dependent on the whims of the AI industry, whose leaders inhabit a moral universe so distinct from ours in science and mathematics that nothing fruitful or sincere can emerge from associating with them [4]. We must remember that the combustion engine was never the heritage or exclusive creation of car makers in Detroit; and railroads and locomotives were not the exclusive heritage of the robber barons of the Gilded Age; and so it is also true that LLMs do not belong in any fundamental way to large industry labs in the Bay Area. The convex hull perspective in this essay suggests it is worthwhile to look into lighter, mathematics-specific LLMs with well curated data, and asses the scale of the technical (and financial) obstacles for such a thing. It would demostrate that the big labs’ “moat” is not as deep as they’d like it to be.
We also ought to approach the use of LLMs with an open attitude, experimenting and exploring until we find those LLM practices that best serve our values: advancing human understanding of mathematics, as Thurston would put it. This also means working as a community to minimize the disruption to mathematical careers, so that as few people as possible share the fate of those skateboarders who had no choice but to retire early. This must also involve policy advocacy: to dramatically increase public funding for science, and policy advocacy for serious public oversight of the large AI labs. We must also experiment with the creation of new social structures and professional arrangements within mathematics, so that we can allow as many individuals as possible to remain part of the mathematical community. If we are successful in this, the mathematical community will come out stronger from this turbulent time, serving as a model for other communities and industries working to transition to a world of ubiquitous LLMs.
To return to where we started, an LLM’s output “solving” the nonlocal ABP or the Mahler conjecture would be, circumstances aside and in a strict sense, a triumph for the discipline. That does not mean that we cannot have mixed feelings about it. It is not pleasant to be scooped by a living, thinking person, and it is no more pleasant to be scooped by the collective knowledge of the field, processed and recombined via an LLM. Rather than being shocked, we ought to be in awe, not of the technology, but of the immensity and power of the mathematical knowledge we have built to this day. Large Models have made this enormity apparent and impossible to miss. Mathematicians must claim such developments as the rightful fruit of our discipline, and accordingly take pride in them, understand them, adopt them, and continue advancing mathematics. Fortunate are we, able to know the causes of things.
Notes
[1] “Digestion” is a term Terry Tao proposed in his ICM lecture on mathematics in the age of AI, part of a metaphor of LLM generated proofs as bringing us to an era of proof abundance, as in mass produced food.
[2] I say “conjecture” only in the context of this hypothetical scenario of Euler and his 18th century LLM.
[3] In mathematics, meanwhile, we are actually approaching a situation where we will have to switch towards whatever comes after journals, or at least journals as they are presently understood.
[4] Here I am very purposefully borrowing a phrase I have long been fond of. It was written by Bertrand Russell in a response to Oswald Mosley.
Wimbledon, the U.S. Open, and the future of mathematics
[This is a guest post by Steven Strogatz. This blog post was initially written in a different file format and converted using AI. — T.]
WIRED reported today that I began to cry while talking about AI and mathematics. But the article didn’t explain what moved me.
The truth is, I’m not entirely sure myself.
(A disclosure upfront: ChatGPT helped me write this post. I’m tweaking it a little to sound more like myself.)
First, thanks to WIRED reporter Isabella Ward. She had the difficult job of distilling a long, wide-ranging conversation into a timely piece about the amazing events of the past week. And she had to put it together in one day! A lot inevitably got left out.
OK, so why did I get choked up? It wasn’t about AI taking my job. I’m 67 and nearing retirement. I also love playing around with AI and find it exhilarating. I use it in research, brainstorming, and learning—and, as you can see, to put this post together.
So what was it then? I think it was when I tried to convey the magnificent 4,000-year story of our subject to the reporter. Mathematics, I said, is more than its results. It’s a human way of seeing and understanding, a conversation in which one generation explains to the next not only what is true, but why. The pleasure of that understanding, and the new discoveries it makes possible, are part of what mathematics is for.
It’s an honor to be part of that long conversation—all the beautiful and important work, and the benefits that have accrued to humanity.
Then, suddenly, the tidal wave of AI hits math in 2026.
Something about the grandeur of the tradition and the suddenness of the change may have conspired to get me worked up.
But I’m not sure it was that. It might have been when I started riffing about Tristan Buckmaster. I don’t know him, but from what I’ve read over this unsettling past week, he reminds me of many mathematicians I’ve known: pure, clear, and scrupulous, committed to truth, and perhaps unprepared for the corporate world.
During the WIRED interview, I mentioned Wimbledon four times as an analogy. To me, Buckmaster represents the Wimbledon of mathematics: white clothes, tradition, etiquette, and the ideal of being a gentle person. How you play the game matters.
The U.S. Open is different: louder, brasher, full of individual style, crassness, energy, and orientation toward the future. It’s also more open and democratic than Wimbledon. I get what’s appealing about that.
With the men’s final coming tomorrow, I find myself wondering whether AI could make mathematics at the cutting edge more like the U.S. Open than Wimbledon. It could open the gates, allowing millions of people without years of specialized training to pour in and make discoveries.
That’s thrilling. And I don’t want to be the stodgy old guard. So yes, I love the U.S. Open … and I also love Wimbledon. (I went for the first time a few years ago, and got choked up then too!)
For the past two years, Alex Townsend and I have been writing a book that traces the marvelous ideas that brought humanity to this dizzying moment: the power unleashed by supercharging math with computers, the gifts it has given us, and what may be yet to come.
So all of that, I think, may have been what set me off: the grandeur of a 4,000-year human tradition, the inconceivability of what may lie ahead—and the possibility that something precious may disappear along the way.
After Math
[This is a guest post by Silvia De Toffoli and Eamon Duede. This blog post was initially written in a different file format and converted using AI. — T.]
Silvia De Toffoli (University School for Advanced Studies IUSS Pavia)
Eamon Duede (Princeton University and Purdue University)
On September 8th, 2026, OpenAI announced that it had produced an AI-generated solution to the Navier–Stokes existence and smoothness problem, one of the seven Millennium Prize Problems. The announcement kicked off debate over credit allocation and the respective contributions of humans and machines to the result. Moreover, the announcement intensified already circulating comparisons with earlier AI conquests in domains believed to otherwise exemplify human intellectual prowess. In a recent statement, Tristan Buckmaster, one of the mathematicians involved in the Navier–Stokes saga, wrote: “This is a Deep Blue–Kasparov moment.”
Existential questions for mathematics follow naturally: if AI can now provide answers to questions at the very frontier of mathematics, is the discipline on the verge of being “solved” as many have said of chess and Go? Like chess and Go players, should mathematicians just “keep playing” and rearrange their practices?
There is something right about the “keep playing” response. As philosopher C. Thi Nguyen (2019) has been insisting, the purpose of playing a game is not exhausted by its aim (winning). The real point is not only the outcome but the process. This is perhaps clearer with a party game such as Twister than with chess: the aim of playing Twister is certainly not winning. But something similar also applies to deep intellectual games, like chess and Go. For instance, playing a game of Go well can be an achievement in defeat.
But in the context of mathematical practice, this feels like an unnecessary retreat. Instead, we can make a stronger move: reject the characterization of mathematics as a game that makes the retreat seem necessary in the first place.
The question, then, is not simply what comes after math, once AI can answer its hardest questions. It is also what we are after when we do mathematics in the first place.
The narrative that AI has “solved” mathematics rests on two assumptions, both seductive and plausible, but both wrong:
- AI really did solve a problem in mathematics.
- Mathematics is only about solving problems.
The first assumption is wrong because to really solve a mathematical problem, providing a mere answer (even if formally certified) is not sufficient. What is missing is an intelligible proof that human mathematicians can understand and use to advance the aims of mathematics. And, as we will argue below, even if AI were to give us just that, the story would not be over because the second assumption is wrong. Mathematics is clearly a much broader enterprise than just problem solving. Mathematicians strive to develop new concepts and theories, to ask and answer new questions, to unify disparate areas, to educate and sustain scholarly communities, and to produce work that is valued for its beauty and depth.
We can (and should) therefore reject the narrative of AI defeating humans at mathematics and start thinking hard about what mathematics really is and what we want it to be.
Not All Answers Are Solutions
OpenAI produced an answer to the question of whether Navier–Stokes can develop a singularity: yes. In The Hitchhiker’s Guide to the Galaxy, Deep Thought produced an answer to life, the universe, and everything: 42. Neither is exactly what we wanted.
Of course, OpenAI gave us much more than “yes.” Deep Thought offered only a number, whereas OpenAI produced two artifacts that many are willing to call proofs. The first is a Lean formalization certifying validity. This was accompanied by a manuscript that appears to contain the corresponding informal proof. So, why is this still dissatisfying? The reason has to do with the underlying notion of proof itself.
There are, in fact, two notions of proof: a logical notion and an intelligible notion.
Modern logic characterizes proof in terms of deductive validity such that a proof can be checked by a mechanical procedure that does not itself require understanding of the mathematical argument. A Lean formalization meets these standards exactly and a Lean formalization of the Navier–Stokes result is therefore a genuine and important contribution: by meeting the demands of the logical notion of proof, it secures certainty.
But mathematicians also want something else from proof. They want understanding (Thurston 1994). They want to know what makes a proposition true. This kind of knowledge trades in mathematical ideas that they can grasp, communicate to other experts, connect with existing knowledge, and use to make further progress. This is the intelligible notion of proof. As of now, it is not clear that OpenAI’s result has given the mathematical community the kind of value that one expects from the intelligible notion of proof.
Genuine proofs are at the same time logical and intelligible proofs. Historically, the two notions have tended to run together. This is because no mathematician could produce an enormously complicated logical proof without first grasping some of the key shareable ideas that made the theorem true. The logical notion of proof was primarily used to verify the correctness of intelligible proofs (Burgess and De Toffoli 2022).
But with AI, these two notions can now come apart dramatically. We can end up with formal proofs that float free from any intelligible proof.
This is not a criticism of formal proof. The converse problem is at least as serious. An intelligible mathematical argument can convey a grand idea while failing to establish that the result is actually true. Jaffe and Quinn (1993) famously used Thurston’s geometrization theorem for Haken three-manifolds as an example: a major insight accompanied by insufficiently complete proofs could become a “roadblock rather than an inspiration.” And one motivation for Hales’s Flyspeck formalization project was to verify that the intelligible (but hard to check) proof presented for the Kepler conjecture was, indeed, a genuine proof (Hales et al. 2009).
Therefore, falling short of either the logical or intelligible notion creates roadblocks where genuine proofs clear the way for mathematical progress. A real mathematical solution requires both logical correctness and intelligibility.
This is particularly clear in the case of the seven Millennium Prize Problems. They were not selected because mathematicians merely wanted seven answers, but rather because they wanted fruitful solutions. The Clay Mathematics Institute itself explains why proof matters in the case of Navier–Stokes: “Because a proof gives not only certitude, but also understanding.”
What OpenAI has given us is an answer. But it is not clear that they have delivered a fruitful solution. Perhaps, we will find that they have, but at the moment, the situation is far from clear. A genuine solution will provide adequate grounds for believing the result but also an intelligible mathematical argument that allows the result to become part of mathematics as understood and practiced by mathematicians.
Nevertheless, if it turns out that what OpenAI has provided is a mere answer, this is not enough to dispel the existential threat that mathematics is facing. Future AI systems are likely to produce genuine proofs that are at once formally certified and fully intelligible to mathematicians. So, current concerns that mathematics is on the verge of being “solved” by AI are not fully dispelled by simply insisting on genuine solutions rather than mere answers.
You Need More than Solutions to “Solve Math”
If future AI systems will produce genuine proofs, logically correct and intelligible, like those produced by “master” mathematicians, it would still be incorrect to think that mathematics would have been “solved” as some say that chess or Go have been solved.
In chess and Go, we accept radically uneven competition between humans and machines because both are, in the relevant sense, playing the same game.
But mathematics is not (or at least not only) a game. To begin with, there is no winner. Mathematics is not an adversarial game with determinate conditions for victory. It is certainly true that mathematicians compete with one another for fame, prizes, jobs, and credit. Chess players do those things too. But chess players also win chess. There is no corresponding condition for winning mathematics. There is no mathematical checkmate.
In mathematics, it is more natural to treat AI as an assistant rather than as a competitor. As Jeremy Avigad (2026) puts it, “We should keep in mind that AI is nothing more than technology, designed to serve our purposes. It is misguided to think of mathematicians as competing with AI; when we drive a car, we aren’t competing to see who can go faster, and when we use a phone, we aren’t competing to see who can speak louder.”
But there is a deeper, and in many ways prior, problem with the competition framing. It requires accepting the assumption that solving problems is the activity by which mathematical success should be measured.
Genuine problem solving is certainly one of the principal aims of mathematics. It is not, however, its only aim. Mathematics is a body of knowledge engaged with, interpreted, and digested by a scholarly community and not a registry of results in the abstract. This simple point has even motivated an entire movement in the philosophy of mathematics: the philosophy of mathematical practice.
Terence Tao (2026) lists many goals of mathematics beyond problem solving. These include developing new theories and techniques, understanding the world, sustaining a community, training the next generation of mathematicians, contributing to cumulative knowledge, and creating works of aesthetic value. Of course, these have been positively correlated with genuine solutions.
But AI breaks that correlation, for the same reason it separates the two notions of proof. So, even genuine solutions would not satisfy us.
This is not moving the goalposts but recognizing that any specific goalpost is inadequate. If mathematics is a game, it is an infinite one.
This attitude is not reactionary. We reject both the concession that logically establishing a theorem is sufficient for a genuine proof and the reduction of “AI for mathematics” to proving theorems. Accepting this, we may find many opportunities for AI in mathematics to support human mathematical flourishing.
Aftermath
We do not deny that, if OpenAI’s announcement is correct, this is an extraordinary achievement. But we should get clear about what type of achievement it is. At this moment, it is an answer, not a solution. And, even if in time the result reveals itself as a genuine solution, we have argued that, in the practice of mathematics, solutions are not everything.
The urgency of rethinking what we value in mathematics is already being recognized within the mathematical community. In a recent declaration initially signed by 25 Fields Medallists, mathematicians warn of a “severe misalignment” between the goals of AI companies and those of the mathematical community.
Mathematicians need to do more to examine their norms. The priority norm is not the problem here (though it is likely a separate problem). The current failure to distinguish genuine solutions from answers, and the growing focus on problem-solving alone, are. This way of thinking is inspired by the current credit economy in mathematics and, as David Bessis recently discussed in his blog, the credit economy needs rethinking.
AI presents mathematics not with an ending but with a choice about what mathematical practice should become. If mathematical success comes to be identified too closely with the production of certified answers, mathematics risks adapting itself to precisely those features that are easiest to benchmark and automate away.
If, instead, mathematicians treat AI as a technology for advancing its long-standing and centrally human purposes, the technology may come to contribute to an accelerated flourishing and enrichment of the discipline. The important question is, therefore, not whether AI will defeat mathematicians, but which mathematical ends we want AI to serve.
What remains in the aftermath is not merely leftovers for humans to scramble for once machines have devoured all of the real problems. Rather, it is an opportunity to clarify what mathematics is all about. We should ask again what we are after when we do mathematics.
(An extended version of this text will appear elsewhere.)
Crowdsourcing a list of general resources on the purpose, value, and nature of mathematics
Parallel to the other crowdsourced resource drive on this blog, I would like to collect books, writings and other resources that discuss or explain the purpose(s), value(s), and nature of mathematics. As seen in reactions to recent events, one can see many oversimplified or inaccurate perceptions of mathematics, such as
- “Mathematical research is about solving open problems.”
- “The main reason for posing an open problem is to obtain its solution.”
- “We should only work on problems that are of immediate benefit to society.”
- “We should prioritize maximizing the number of problems solved.”
The question of what mathematics actually is is far more complex than these assertions might suggest. Here are some classic books and articles on these topics:
- Vladimir Arnold, “On teaching mathematics“, 1997 (somewhat provocative in tone, but still interesting)
- Philip Davis, Reuben Hersh, “The Mathematical Experience“, 1981
- Timothy Gowers, “The two cultures of mathematics“, 2000
- Timothy Gowers, “Mathematics: a very short introduction“, 2002
- G. H. Hardy, “A Mathematician’s Apology“, 1940 (somewhat dated at this point, but still interesting)
- Rueben Hersh, “What is mathematics, really?“, 1997
- Imre Lakatos, “Proofs and refutations“, 1976
- Bill Thurston, “On proof and progress in mathematics“, 1994
- Cedric Villani, “Birth of a Theorem“, 2015
Here are some more informal links:
- Matt Might, The illustrated guide to a Ph.D., 2010.
- A collaboration between myself and Zach Wienersmith on sphere packing, as a five part comic strip (Apr 9, 2026)
- Illustrating the impact of the mathematical sciences – a series of posters by the National Academy of Sciences. (I was on the committee to create these posters.) (2023)
I can also list a few writings, posts, and interviews I have done on these topics:
- There is more to mathematics than rigour and proofs (2007)
- What is good mathematics? (2007)
- The value in “Dumb” questions (2022)
- What makes for ‘Good’ Mathematics (2024)
- What does it mean to think like a mathematician? (Mar 14, 2026)
- Interview with “Big Think” (relating to my forthcoming book) (Aug 28, 2026)
- On remembering what it is like to be a child: curiosity-driven play, the “largest number” game, and what premature optimization for the nominal goal extinguishes (Sep 9, 2026)
Please contribute more links in the comments below!
The status of the Hodge conjecture
[This is a guest post by Claire Voisin. This blog post was initially written in a different file format and converted using AI. — T.]
Hodge classes can be defined on any compact complex manifold . They are rational Betti cohomology classes on of even degree (eg combinations with -coefficients of classes of oriented codimension closed submanifolds of ) which satisfy a delicate “Hodge condition” necessary for them to be combinations with -coefficients of classes of complex submanifolds (and more generally closed analytic spaces) of . To understand this condition, we need to pass to cohomology with complex coefficients, and represent complex cohomology classes as de Rham cohomology classes of closed forms. The Hodge condition is that the class should be representable by a closed form that in local holomorphic coordinates is written as .
The Hodge conjecture states that a Hodge class on a smooth complex projective variety is “algebraic”, that is, is a combination with rational coefficients of classes of closed analytic (or equivalently, algebraic) subsets. Note that the Hodge conjecture can also be formulated for compact complex manifolds, and in particular compact Kähler manifolds. In this setting, there are easy examples where some nonzero degree 2 Hodge class exists on , while there are no codimension 1 closed analytic subsets in , so one needs in any case to consider not only classes of closed analytic subsets but also Chern classes of holomorphic vector bundles and their singular version (coherent sheaves). However it is proved in [12] that some compact Kähler manifolds have nontrivial Hodge classes, while Chern classes of coherent sheaves are all zero. This does not necessarily say that an analytic approach to the Hodge conjecture is not possible, but this says that any such proof has to use the fact that is algebraic. One early approach, proposed in [6], is to work on an affine Zariski open set , where is a hyperplane section of . One can represent even degree rational cohomology classes of , restricted to , as Chern classes of holomorphic vector bundles on , and then use the Hodge condition (but how?) to algebraize these vector bundles, that is, extend them from to as coherent sheaves. This very interesting construction has alas not been successful.
Another potential approach to the Hodge conjecture is by induction on the codimension. It relies on the following deep fact, which is a consequence of the Deligne theory of Hodge structures and their properties [7]. Namely, in order to prove the Hodge conjecture, it suffices to prove the following statement: for any smooth projective variety and any Hodge class on , there exists a dense Zariski open set such that in .
In this statement, is an algebraic hypersurface in , usually very singular. A class satisfying the above property for some is said to have coniveau . This notion goes back to Grothendieck (see [10]) and the study of the coniveau filtration led to important developments in [3]. Alas, how to prove that a class has coniveau ? As explained by Grothendieck in [10], this is very restrictive on and most cohomology classes (eg, classes of holomorphic forms) do not satisfy this property.
The Hodge conjecture thus motivated beautiful developments in Hodge theory and on the topology of algebraic varieties, but it is fair to say that, beyond the formal definition, there is no good understanding of what is a Hodge class from the viewpoint of algebraic geometry. The reason is that one understands well in the setting of algebraic geometry cohomology with complex coefficients and the Hodge condition, but not Betti cohomology with rational coefficients.
However, there are Hodge classes that we understand very well, which are constructed by (multi)linear algebra tricks. Probably the simplest example is the following: let be a smooth projective complex variety and let . Then is of dimension 1 and it is naturally contained in . It is rather obvious that it is generated by a Hodge class on . This class is not known to satisfy the Hodge conjecture.
Similarly constructed examples need a little more knowledge of topology. The standard conjectures like the Künneth conjecture (the Hodge conjecture for the Künneth components of the diagonal), or Lefschetz standard conjecture, are particular instances of the Hodge conjecture for certain Hodge classes that can be constructed on the square of any smooth projective complex manifold. The algebraicity of these classes is rather crucial in the theory of motives.
A more involved example is that of Weil classes on Weil abelian varieties, on which an important recent progress was made by Markman. An abelian variety over is a complex torus that has holomorphic embeddings in complex projective space, and a Weil abelian variety is one which admits a quadratic endomorphism , , . The variety has to be of even dimension and one assumes that the action of on the tangent space of has both eigenvalues with the same multiplicity . A formal argument then produces a two-dimensional space of Hodge classes of degree on , called Weil classes. A big recent progress on the Hodge conjecture is the proof by Markman [11] that Hodge classes on abelian fourfolds are algebraic. This problem is classically reduced to proving the algebraicity of Weil classes, and to prove this, Markman makes a detour through Weil classes on certain families of Weil abelian 6-folds.
Variational aspects. Smooth projective complex varieties come in families , where and are themselves algebraic varieties which can be chosen defined over a number field, as the algebraic map . The topology of the fibers does not change with ( is a fibration) but the complex structure of changes, and so does the set of Hodge classes of given degree on . Given a Hodge class of degree on a fiber , its Hodge locus is (grosso modo) the set of points such that along a path from 0 to in , the locally constant class remains Hodge. An important result concerning is the fact (proved by Cattani–Deligne–Kaplan [5]) that it behaves as if the Hodge conjecture was true, namely is a closed algebraic subset in . However, one missing information in this result is that this locus is defined over a number field, as predicted by the Hodge conjecture (a more precise statement is that Hodge classes are absolute Hodge). One could thus imagine disproving the Hodge conjecture by exhibiting a Hodge class on a fiber with Hodge locus not defined over a number field (this cannot be done with the Hodge classes explicitly described above). Such counterexample would not damage too much our understanding of the theory of motives and could lead to a corrected version of the Hodge conjecture stating that absolute Hodge classes are algebraic.
Variational Hodge conjecture. Given as above and a Hodge class , assume that and that is algebraic on . Is also algebraic on for any ?
A negative answer to that question would be the worst scenario for the theory of motives. In particular, this would disprove the Lefschetz standard conjecture (see [1]).
Deformation theory can in some cases be used to answer affirmatively the question above. Assuming that is the class of an algebraic subvariety (or a Chern class of an algebraic vector bundle on ), the semi-regularity condition of Bloch, (respectively of Buchweitz–Flenner), is a subtle cohomological property of (resp. of ) guaranteeing that deforms to , (resp. that deforms to if all its Chern classes remain Hodge on ). Buchweitz–Flenner semiregular sheaves have been used successfully by Markman in [11] for Weil classes on 6-dimensional Weil abelian varieties. Unfortunately it seems very unlikely that semi-regularity suffices to solve the variational Hodge conjecture in general. One reason is that the variational Hodge conjecture for integral Hodge classes is not true, which shows that in some cases, cycles , , are not homologous to any semiregular cycle (a multiple is in any case needed).
Another reason is that, associated to a given cycle of , there is a Deligne–Beilinson class which lives in Deligne–Beilinson cohomology of and lifts the Hodge class of . When the class is 0, then the class is the Abel–Jacobi invariant of (see [9]). There are examples of families as above where the Hodge class of a cycle of deforms to a Hodge class on but the Deligne cycle class of does not deform to the Deligne cycle class of any algebraic cycle of (see [8]). These facts suggest that semi-regular objects are very hard to construct, and too restricted to lead to a general solution of the variational Hodge conjecture.
References
[1] Y. André. Déformation et spécialisation de cycles motivés, J. Inst. Math. Jussieu, 5 (2006), 563–603.
[2] S. Bloch. Semi-regularity and de Rham cohomology. Invent. Math. 17 (1972), 51–66.
[3] S. Bloch, A. Ogus. Gersten’s conjecture and the homology of schemes, Ann. Sci. Éc. Norm. Supér., Sér. 4, 7, 181–201 (1974).
[4] R.-O. Buchweitz, H. Flenner. A semiregularity map for modules and applications to deformations. Compositio Math. 137 (2003), no. 2, 135–210.
[5] E. Cattani, P. Deligne, A. Kaplan. On the locus of Hodge classes. J. Amer. Math. Soc. 8 (1995), no. 2, 483–506.
[6] M. Cornalba, Ph. Griffiths. Analytic cycles and vector bundles on non-compact algebraic varieties. Invent. Math. 28 (1975), 1–106.
[7] P. Deligne. Théorie de Hodge II, Inst. Hautes Études Sci. Publ. Math. No. 40 (1971), 5–57.
[8] M. Green. Griffiths’ infinitesimal invariant and the Abel–Jacobi map. J. Differential Geom. 29 (1989), no. 3, 545–555.
[9] Ph. Griffiths. On the periods of certain rational integrals. I, II. Ann. of Math. (2) 90 (1969), 460–495; 90 (1969), 496–541.
[10] A. Grothendieck. Hodge’s general conjecture is false for trivial reasons. Topology 8 (1969), 299–303.
[11] E. Markman. Cycles on abelian 2n-folds of Weil type from secant sheaves on abelian n-folds, arXiv:2502.03415.
[12] C. Voisin. A counterexample to the Hodge conjecture extended to Kähler varieties. Int. Math. Res. Not. 2002, no. 20, 1057–1075.
On the Hodge conjecture
[This is a guest post by Burt Totaro. This blog post was initially written in a different file format and converted using AI. — T.]
These are strange times for mathematicians. It now seems possible that AI companies will burn through vast resources in order to prove some new fact about the Hodge conjecture. I’d like to discuss the current status of the Hodge conjecture, in order to think about the value to the mathematical community of exploring such hard problems. Thanks to Terry Tao for suggesting this guest post.
The Hodge conjecture crystallizes a big mystery: the relation between topology and algebraic geometry, or (what amounts to the same) between real and complex geometry. In short, real submanifolds are flexible, whereas complex submanifolds are quite rigid, and it is a challenge to relate the two. The setting is a “smooth complex projective variety” , a complex manifold defined by algebraic equations. We want to understand the possible shapes of complex algebraic subspaces of . (By a famous result of Wei-Liang Chow (1949), every complex analytic submanifold of is in fact algebraic.) One could ask whether every real submanifold of can be moved continuously to a complex submanifold. That is far too optimistic. Still, to a first approximation, the Hodge conjecture predicts which real submanifolds can be moved continuously to a complex submanifold, in terms of whether certain integrals are zero. (More precisely: every Hodge class in the rational cohomology of should be the class of an algebraic cycle.) The conjecture goes back to William Hodge (1950).
Now suppose in some counterfactual world that the day after the conjecture was made some inhuman oracle told us that the answer was “yes” and the conjecture was considered “solved.” We can see, by examining our real world, the vast body of mathematical knowledge that would have been lost. In our world, what we want from a great conjecture is a challenge that inspires all kinds of other discoveries, even if the original question remains open. The Hodge conjecture has been that kind of conjecture for geometers.
A key piece of evidence for the Hodge conjecture is the “Lefschetz -theorem“, which says that the Hodge conjecture is true for algebraic cycles of codimension 1 (that is, of complex dimension , if has complex dimension ), and also for cycles of dimension 1. Solomon Lefschetz’s work was quite early (around 1924), well before the Hodge conjecture was stated in general. Interestingly, some of the most important proofs by both Lefschetz and Hodge have gaps, from our current perspective. They had tremendous insight, but the full justification of their results required hard analysis by many mathematicians, notably Kunihiko Kodaira (around 1950).
By some measures, one could say that progress on the Hodge conjecture has been limited. In view of Lefschetz’s results, the first open case is for codimension-2 cycles on a variety of complex dimension 4, and that case (for arbitrary varieties of dimension 4) still seems far out of reach. But the Hodge conjecture has been enormously successful for inspiring new mathematical theories, such as Phillip Griffiths’s “Variations of Hodge structure” (starting in the 1960s) and Pierre Deligne’s “absolute Hodge cycles“. Broadly speaking, the Hodge conjecture suggests that the structure of algebraic cycles should be controlled by Hodge theory, in other words by integrals of algebraic functions. This hope has led to vast numbers of proved insights about the structure of algebraic cycles. (For example, Griffiths disproved Grothendieck’s conjecture that algebraic and homological equivalence were the same, which was a surprise; but he used Hodge theory in order to do it.) In recent decades, Claire Voisin has been a leader in using Hodge theory in new ways to get information about algebraic cycles.
An exciting recent development is Eyal Markman’s 2025 proof of the Hodge conjecture for abelian varieties of dimension 4 and 5. (Abelian varieties, the tori with a complex algebraic structure, are very special compared to all algebraic varieties; but they are a particularly important class in many ways, for example in number theory.) This is a monumental piece of work that builds on many earlier developments. André Weil (1977) identified a class of abelian varieties, those of Weil type, for which there are “unexpected” Hodge classes that do not appear on most other abelian varieties. In low dimensions such as 4, Ben Moonen and Yuri Zarhin (1995) showed that the Hodge conjecture for all abelian varieties would follow from the special case of abelian varieties of Weil type. Spencer Bloch (1972), extended by Ragnar-Olaf Buchweitz and Hubert Flenner (2003), defined a property called “semi-regularity” which, when it holds, allows proving the Hodge conjecture for a continuous family of varieties when it holds for one variety in the family.
For a long time, however, semi-regularity seemed far too strong a condition ever to apply to hard cases of the Hodge conjecture. Markman found, with great ingenuity, how to construct enough examples of semi-regular sheaves to prove the Hodge conjecture for abelian varieties in dimensions 4 and 5. His ideas use several big theories that have been developed in recent decades, notably about derived equivalences between algebraic varieties and about hyperkähler manifolds. In the end, Markman actually needed an extension of Buchweitz–Flenner’s semi-regularity theorem to “twisted sheaves”, which was supplied by Jonathan Pridham (2024).
At this point, I am confident that with enough effort, it will be possible to push these ideas further. As we have seen, AI companies are willing to mobilize resources on an incomprehensible scale to attack problems that attract their interest. We don’t know how much insight may come from any given AI-generated proof. But there is reason to worry about the commercial pressure on AI companies to claim advances at high speed and the knock-on effect this has on mathematicians. The power of mathematical ideas comes from a continual negotiation among people. What we care about are not so much isolated facts, but rather the ideas and inspiration that the search leads to.
A Severe Misalignment of AI in Mathematics
I am proud to be among the list of 25 initial signatories — all Fields Medallists — to the declaration below, which grew out of discussions between ourselves over the last week. We have also posted our declaration on this web page, and (similarly to the Leiden declaration) invite further signatures. (It is unfortunate that we did not have the time to have a more consultative process, as with Leiden; but we decided that the urgency of the situation was such that we needed to release a statement sooner rather than later.)
See also this recent article in the Economist regarding our declaration, and a brief interview with James Maynard on this topic. A French version of this declaration was published in Le Monde.
Over the last few months, the mathematical capabilities of LLMs have improved dramatically, to the point that they can solve major outstanding problems in many fields of mathematics. However, the push by AI companies to solve mathematical problems as a benchmark is detrimental to the science of mathematics, and to the mathematical community. The goals of the AI companies and the goals of the mathematical community are severely misaligned. We see these as part of broader alignment issues impacting other scientific and creative professions, as well as the whole of society.
Research mathematics deals with understanding basic structures of shapes, numbers, and natural phenomena. Over the course of generations, it has built a large corpus of sophisticated ideas, methods, abstractions, and other tools to comprehend the mathematical landscape. In turn, modern technologies and sciences are based on mathematical tools.
Famous problems have often served as landmarks and lighthouses against which one can measure an improved understanding of this landscape. Solving one of these problems has been a certain sign of new insights and interesting methods, which would then be studied by a community of mathematicians, through a long and arduous process of talks, discussions, simplifications. At the end of this process, one will ideally find a textbook presentation of the results suitable for any graduate or even undergraduate student to study. Some of the mathematical ideas pursue their journey even further to become, decades or centuries after, tools that are understood and used by the whole population.
The mathematical community functions, in many ways, as a miniature version of humanity. It consists of individuals using a wide variety of different approaches, joined by core values. The most precious resources of our profession are students and ideas, and these we nurture with great care. We feel responsible to let them grow to their full potential, until they can live a life of their own in the mathematical world. For students we often suggest problems with the core intention of developing skills making them well-positioned for advances in research and elsewhere. Our ideas we disseminate in talks, private discussions and careful writeups, connecting them to the previous ideas of others. These processes invariably take time and are based on human interaction.
In recent months, the success of AI in solving major mathematical problems has made headlines even outside mathematical circles. But solving problems is only a tool and proxy for achieving the primary goal of conceptual understanding and insight. Forgetting this in the world of AI may turn the tool against the primary goal. Indeed, the mass production at faster and faster pace of “true/false” statements could destroy fertile ground instead of breathing life into new ideas.
Often these solutions are announced in a rush, leaving no time for a proper writeup, the isolation of new methods and ideas, and citing relevant previous work of others. As in all creative professions, this raises severe attribution and plagiarism questions. Moreover, without the willing mathematicians who must take care of their development and integration into the mathematical canon, AI-conceived ideas would never become fully alive and the crucial human transmission chain between mathematicians would be lost.
We are witnessing a general threat to intellectual work, with misalignment between the outcome of the use of AI and its initial purpose. In many fields and activities, years of training have traditionally served not only to produce a final answer or product, but also to develop understanding and the ability to formulate new questions and ideas. However, building on a vast body of previous human work, AI systems are becoming increasingly capable of producing the results of such work directly, and these goals cease to align. The issues the mathematical community faces now are similar to issues that other scientific and creative professions are facing, and indicate issues that all of humanity might face: how to make sure that, as AI changes the way work is done, we do not lose sight of what that work was meant to achieve in the first place.
AI offers the potential of enhancing and accelerating genuine mathematical study and understanding. Mathematics as a profession will need to adapt to these changes in several ways. However, whether these changes ultimately benefit the field or have a destructive effect will in large part be determined by the decisions of the humans in control of this new technology.
These issues must be addressed urgently, in the mathematical community, by the companies developing these technologies and, more broadly, by a society that will confront similar problems in many other forms of intellectual work.
Artur Avila (Fields Medal 2014)
Manjul Bhargava (Fields Medal 2014)
Caucher Birkar (Fields Medal 2018)
Pierre Deligne (Fields Medal 1978)
Yu Deng (Fields Medal 2026)
Simon Donaldson (Fields Medal 1986)
Hugo Duminil-Copin (Fields Medal 2022)
Alessio Figalli (Fields Medal 2018)
Martin Hairer (Fields Medal 2014)
June Huh (Fields Medal 2022)
Maxim Kontsevich (Fields Medal 1998)
Elon Lindenstrauss (Fields Medal 2010)
Pierre-Louis Lions (Fields Medal 1994)
James Maynard (Fields Medal 2022)
Curt McMullen (Fields Medal 1998)
Shigefumi Mori (Fields Medal 1990)
Ngô Bảo Châu (Fields Medal 2010)
Andrei Okounkov (Fields Medal 2006)
Peter Scholze (Fields Medal 2018)
Stanislav Smirnov (Fields Medal 2010)
Terence Tao (Fields Medal 2006)
Maryna Viazovska (Fields Medal 2022)
Cédric Villani (Fields Medal 2010)
Wendelin Werner (Fields Medal 2006)
Efim Zelmanov (Fields Medal 1994)
On the existence of non-sofic groups
[This is a guest post by Andreas Thom. This blog post was initially written in a different file format and converted using AI. — T.]
When I woke up on August 1st, 2026, I had received a few emails from colleagues asking for my opinion on a remarkable result that had circulated the previous day. The result was a solution to a long-standing open problem in geometric group theory, specifically the existence of a non-sofic group. I was astonished and at the same time, looking at the first draft, also in a way happy to see that Kun’s work on expander decompositions and my joint work with Gábor Kun played a decisive role in the crucial Proposition 2.3 of the OpenAI paper. I had always hoped that the theory of centralizer rigidity would eventually have significant applications, but I had not found the right setting in which it could be used so effectively. In that sense, the solution also came as a relief and I was happy to explain the ideas in a post on MathOverflow a few days later.
From the start, colleagues pointed out that the framing in the public announcement that appeared shortly afterwards was misleading, in that it spoke of “no progress” in the last decade, while relying on our 2019 paper (not to mention subsequent work by many hands that was not directly relevant for the OpenAI paper but would still be considered to be progress by many).
So I wrote to Mark Sellke and Sébastien Bubeck: “[…] I find the framing intellectually dishonest. You (and I am talking about you personally, since I have no one else to address this to) cannot speak in the public announcement of a decade without progress and then use a 2019 paper in a crucial way. It is true that Proposition 2.3 is a really clever use of the centralizer-rigidity theorem, but neither does its short proof require new techniques […].
“It is true that the last stone finishes the building and usually those who can put it get the credit for solving the problem, that is fair enough. I have no problem with that and I personally do not care much about credit. However, I guess you would get enough praise without downplaying the previous contributions.”
Sellke replied to this and basically agreed to the need for a revision; as a result the public announcement was changed to the form it has now. I was glad to have received an early draft from Sellke also on August 1st, otherwise there would have been no way to react to the first public announcement at all, since neither the PDF nor the website of the announcement contained contact information. Anyway, I was happy that this was resolved and the matter closed.
It is fair to say that the approach of Kun and myself had not been viewed as the main line of attack on non-soficity prior to OpenAI’s announcement. In fact there were other more promising approaches along the line of quantum games etc. at the time, that had already led to a negative solution of the famous Connes Embedding Problem and, later, the disproof of the Aldous–Lyons conjecture. Hence, OpenAI’s detailed command of the techniques of Kun and myself made me wonder how the model found this route, especially since I discussed these techniques and their use in extensive sessions with ChatGPT over the last months.
So in the same email I asked Mark Sellke and Sébastien Bubeck: “Another point is that I and a colleague in Dresden were discussing the expander matching problem and various extensions of the work with Gábor Kun actively over the last months with ChatGPT, so that we are of course curious if that was part of the training data or accessible to the reasoning process. There is a certain (frankly unacceptable) lack of transparency here; and I fear it will damage the communal process of math more than the new AI-generated results will benefit the subject.”
Mark Sellke’s complete answer to this part of my email was: “Regarding your conversations with ChatGPT: that did not happen.”
Anyway, I thought, these techniques were public, so their use is not evidence that our conversations influenced the model. But because this was not the main line of attack, and because I had recently discussed precisely these techniques and possible extensions with ChatGPT at length, I thought the question had to be asked. Back at the beginning of August, I then returned to mathematics and wrote a subsequent paper with Gábor Kun on applications of the ideas that were the basis of Proposition 2.3. This was my way to react; after all, the integration of the new result in the math landscape seemed like a natural next step.
However, after reading up on the controversy around the Buckmaster–Alpöge case, the whole story came back to me and I realized that the answer I received from OpenAI was misleading, to say the least. I already wrote about this briefly on Mathstodon.
I had explicitly asked about two different things: (1) whether our conversations entered training data, and (2) whether they were accessible to the solving process. OpenAI said in the Buckmaster–Alpöge case that no specific user data was accessed, but added that it “cannot rule out that de-identified data derived from their usage of our products helped improve our models.” In light of OpenAI’s later wording, I cannot tell whether Sellke’s answer denied both possibilities or only direct access under (2). No qualification, explanation, or evidence was given. Whatever its intent, I regard the answer as materially misleading.
OpenAI was drawing a distinction that its answer to me erased, despite the fact that my question explicitly made that distinction. We are not required to reverse-engineer OpenAI’s internal training pipeline to establish what happened. Only OpenAI has the relevant data for that. For such a categorical denial by OpenAI to be credible, OpenAI should disclose its basis: product and privacy settings, relevant datasets and checkpoints, and what “de-identified data derived from usage” means.
I disabled model training on 29 June. That control is still only a promise whose implementation users cannot audit, and it is prospective: it does not answer what happened to earlier conversations or to derivatives already selected.
If nonpublic research supplied by users improved a model and the provider then used that model to race those users to publication—without informed consent, disclosure, or credit—that would be ethically indefensible. De-identification may remove a name; it does not remove the intellectual content of a mathematical idea. Sellke and Bubeck seem to be blind to this simple moral aspect.
Sellke gave me a categorical assurance without explaining its basis; I regard that response as materially misleading. If he lacked the information needed to rule out training use, he had no basis for giving that assurance. Bubeck’s acknowledged career-related remark in the conversation with Buckmaster and his objection to including Alpöge in a proposed paper presenting OpenAI’s proof deepen my concern about their commitment to academic standards. Taken together, these episodes raise serious questions about their judgment and personal integrity.
I am not claiming that anyone read individual chats or that our conversations were in fact used in training; I do not know that. My criticism concerns the categorical denial. If they did not know what entered the training data (the most likely scenario), they should have said so.
After I finished writing this post, I received a message from Mark Sellke, who acknowledged understanding how I “reasonably arrived at [my] conclusions given the evidence available”. He pointed me to a discussion citing OpenAI’s new statement that Buckmaster’s Codex prompts from the preceding two months could not have influenced its system, including through training. I wish I could trust this more. In any case, it suggests that the math community can successfully put pressure on the industry to take these issues at least somewhat more seriously.
So what does that all mean and how do we as a community proceed? Setting aside these particular cases (which might also be very different in what really happened behind the scenes), we have to see the broader picture and I believe there is no way of going back.
As far as I see it, a mathematical publication used to bring together three things. It announced a result, identified the people who had produced it, and added something to our human understanding of mathematics. It could therefore serve at the same time as a record of knowledge, a basis for assigning credit, and evidence of a mathematician’s ability. AI breaks this connection.
An AI system can produce a correct proof even if no human being discovered or even understood the argument in the usual sense. A typical form that this can take nowadays varies from somewhat verified AI Slop to a Lean certificate, or a combination of both. The proof may be worth publishing, but the publication then records only that the result has been established. It does not necessarily tell us how the result was found or who understood it.
The question of credit and contribution may have no satisfactory answer. A model may draw on published work, feedback, conversations, prompts, and its own search in ways that apparently cannot be reconstructed clearly. The person or the company who ran the model should not simply receive whatever credit cannot be assigned elsewhere. As another consequence, a publication record can no longer serve as a reliable measure of a person’s mathematical quality. If the theorem, proof, and written explanation may all have been generated by AI, a list of papers tells us very little about what the named author contributed or understands. Even if matters of data privacy are resolved, I see only little hope that the current system of publication and credit can be salvaged. The whole idea of personal credit will not work in such a highly connected environment anymore and I actually think that the focus on “who got something first” (not to speak of prizes for a solution to particular problems) was always misleading, even though a powerful driving force.
The more important question is therefore what someone contributes to human mathematical understanding. This includes explaining why an argument works, separating the main idea from technical details, connecting a result with other areas, finding the right questions, teaching new methods, and helping other mathematicians make use of them. Such contributions may happen through papers, but also through lectures, discussions, teaching, and collaboration.
A formal proof certificate is comparable, in a sense, to the detection of a new star. It confirms that something is there, but further work is needed to understand what it is, why it matters, and where it belongs in the larger landscape. Mathematics is only to a lesser extent the production of correct statements. It is more the human process of making sense of them.
The institutional issue goes far beyond mathematics. Advanced AI is becoming a general supply of “intelligence” on which science, education, public administration, industry, and ordinary life may all depend. It should therefore be provided according to standards comparable to those governing water or electricity: reliable and broadly available, with clear public duties, strong privacy rules, independent oversight, and protection against discriminatory or self-serving use.
Intelligence of this kind should not be treated as an ordinary consumer product whose conditions are set entirely by a few companies. A provider should not be able to collect people’s ideas and information, control all evidence about how they were used, and then exploit that advantage against its own users. Once intelligence becomes basic infrastructure for society, it must be highly regulated and governed in the public interest.
