Essay

On Mathematicians Disbelieving Straight Lines on Graphs

The Fields medalists are worried about mathematics. Terry Tao and many of his Fields medalist colleagues have written A Severe Misalignment of AI in Mathematics. Like anything Tao writes about mathematics, it’s worth reading in full. They argue that among other things the ability of AI to generate context-free solutions of difficult problems in minutes will be detrimental to the growth of the next generation of mathematicians.

It’s a real concern. Humans learn through work. It may be fun or it may not, but it’s difficult to learn anything substantive as a spectator. It’s hard to learn to catch a baseball by watching someone else catch; you generally need to go outside and try to catch it yourself. In math and science, I have often thought “Ah, I understand this” after a lecture or reading a chapter, only to discover in the problem set (or worse, the exam!) that I haven’t actually done it and therefore didn’t know what I didn’t know.

The problem, in his view, goes deeper.

We are witnessing a general threat to intellectual work, with misalignment between the outcome of the use of AI and its initial purpose. In many fields and activities, years of training have traditionally served not only to produce a final answer or product, but also to develop understanding and the ability to formulate new questions and ideas. However, building on a vast body of previous human work, AI systems are becoming increasingly capable of producing the results of such work directly, and these goals cease to align.

This is a slightly different objection, not just to the effects of AI on learning, but on the discipline of mathematics as a whole. Leave aside his point for a moment and look at this sentence, which has been surprisingly sparse in the discourse following OpenAI’s probable solution of the Navier-Stokes existence and smoothness Millennium Prize problem.

AI systems are becoming increasingly capable of producing the results of such work directly

This admission was essentially absent from the Leiden Declaration just a few months earlier in June. The Leiden Declaration focused on five problems, only one of which was about AI technology in itself:

1. Current automated techniques can produce plausible but unreliable (or even incorrect) arguments which are difficult to distinguish from correct mathematical proofs.

This was dubious when it was written three months ago. Today it’s almost entirely wrong. AI systems can and do still make mistakes in mathematics, but it’s increasingly rare and increasingly easy to detect. Formal verification by proof systems like Lean is fundamental in verifying the results and in training the models in the first place.

The other four objections are that AI could:

  1. Undermine traditional attribution for professional mathematicians,
  2. Disturb funding for professional mathematicians,
  3. Take professional mathematicians out of the verification loop,
  4. Make it more difficult for the profession of mathematics to control its own destiny.

It is essentially a complaint that AI could be perceived as making their profession obsolete. I’m not unsympathetic to this, unsympathetic though I’ve made it sound. It’s coming for the rest of us soon enough. My native profession is just one step to the left:

Comic strip titled “Fields arranged by purity,” with an arrow reading “more pure” pointing right. Six stick figures stand in a row: a sociologist, a psychologist, a biologist, a chemist and a physicist, then, far off to the right, a mathematician. The sociologist says sociology is just applied psychology; the psychologist says psychology is just applied biology; the biologist says biology is just applied chemistry; the physicist says “which is just applied physics. It’s nice to be on top.” The mathematician, well apart from the others, says “Oh, hey, I didn’t see you guys all the way over there.”
“Purity,” xkcd 435, by Randall Munroe. CC BY-NC 2.5.

Where my sympathy diminishes is the sleight of hand in what the Declaration means when it says mathematics. It uses the word mathematics to mean the field of human endeavor where Man ventures into the Platonic realm and emerges transformed, holding a gleaming shard of logical Truth. We have been doing it for thousands of years for reasons both sacred and profane and it is one of the triumphs and glories of our species as a whole and of any individual human who thinks about numbers and wonders “Why?” But the declaration also uses the word to mean “The profession practiced by a small group of people who sit in university halls and turn coffee into theorems.” It’s a beautiful, dignified profession with a history that itself goes back thousands of years to people like Pythagoras. It has buyers and sellers, customers, products, professional associations, and job and salary prospects that ebb and flow with the economy. I love that profession and I wish it the best, but it’s not the same thing as mathematics in that first sense.

And for the last very few years we have witnessed a surprising blind spot of mathematics in that second sense failing to recognize one of the oldest and most fundamental parts of mathematics in the first sense: straight lines on a graph.

The Line

Scatter chart of Epoch Capabilities Index against date, 2023 to 2028. Frontier models trace a straight line rising from Claude 2 at about 118 in mid-2023 to GPT-6 Astra at about 166 in mid-2026, then continue as a dashed projection. Horizontal rules mark the level the line had reached when each milestone fell: grade-school math (GPT-4, March 2023), AIME problems (o1, September 2024), olympiad problems (IMO gold, July 2025), Erdős problems (first full proofs, January 2026) and a Millennium problem (Navier–Stokes, September 2026), with a dashed rule near 183 where the projection reaches October 2027.

The very first text-predictive models like GPT and GPT-2 could autocomplete “1 + 1 = ” with the correct answer. This wasn’t surprising to anyone. Of all the times “1 + 1 = ” has been written, “2” has been the next character most of the time.

By 2023, GPT-4 could do grade school math at a grade school level. “If Tommy has two toys and Susie has three toys, how many toys do they have altogether?” You could still write it off to an extent, although that specific sentence may or may not have actually appeared in the training corpus. Nonetheless it could do these problems, not impressive problems and not very well, but it was an inkling that the technology was not strictly limited to copying from what it had seen before.

By the fall of 2024, models could do some AIME problems. These are starting to be genuine challenges. They vary from relatively easy to relatively hard. You couldn’t hand one to the average adult on the street and expect an immediate solution. Here’s an example of a fairly easy one from 2013:

The AIME Triathlon consists of a half-mile swim, a 30-mile bicycle ride, and an eight-mile run. Tom swims, bicycles, and runs at constant rates. He runs five times as fast as he swims, and he bicycles twice as fast as he runs. Tom completes the AIME Triathlon in four and a quarter hours. How many minutes does he spend bicycling?

Or if you’d rather a harder one from that same year:

For πθ<2π, let

P=12cosθ14sin2θ18cos3θ+116sin4θ+132cos5θ164sin6θ1128cos7θ+

and

Q=112sinθ14cos2θ+18sin3θ+116cos4θ132sin5θ164cos6θ+1128sin7θ+

so that P/Q=22/7. Then sinθ=m/n where m and n are relatively prime positive integers. Find m+n.

I think I would have a chance at solving that pencil-and-paper, but not easily and not quickly.

It was at this point people—mathematicians in particular—should have started paying attention. I am sorry to say that I didn’t. I believed, incorrectly although not unreasonably, that there were theoretical limitations to LLM performance due to the nature of training to predict the next token from an internet corpus. I didn’t fully grasp the difference that reinforcement learning (especially in verifiable domains like math) would make. And I didn’t fully grasp the difference that tool use harnesses like Claude Code would make, although I have more of an excuse there because Claude Code didn’t exist until early 2025.

By mid 2025, AIs were solving International Mathematical Olympiad (IMO) problems. These are intended to challenge the best high-schoolers going into college, and I expect even at the height of my pencil-and-paper math days I would have scored a zero. Example:

Let denote the set of positive integers. A function f: is said to be bonza if

f(a)dividesbaf(b)f(a)

for all positive integers a and b.

Determine the smallest real constant c such that f(n)cn for all bonza functions f and all positive integers n.

IMO 2025, Problem 3

It was at this point that I revised my opinion of LLM-based AI as “stochastic parrots.” They were solving hard problems, that I myself could not solve, and that in fact weren’t in the training data because the problems had been written for this competition specifically.

By early 2026 the AIs were beginning to polish off Erdős problems, unsolved problems in research mathematics that were famously collected by legendary 20th century Hungarian mathematician Paul Erdős. They saturated the FrontierMath benchmark run by epoch.ai, and began making a dent in its successor FrontierMath Open Problems. Terry Tao was actually a fairly prominent early adopter of AI, experimenting with it in his own work.

In the summer of 2026 famous open problems began falling. The Jacobian conjecture. Non-sofic groups. A new bound on the zeros of the Zeta function.

And then the bomb—an unreleased OpenAI model resolved a Clay Mathematics Institute Millennium Prize problem. The Millennium problems are a list of seven of the most famous and difficult open problems in all of research mathematics. Since the list was published in 2000, only one had been solved (the Poincaré conjecture, resolved by Grigori Perelman in 2002–2003). The Navier-Stokes equations describe idealized fluid flow and are probably the most directly physical problem of the Millennium problems, and leaving aside the details for the moment the question is essentially “Does this equation always work or does it sometimes have no reasonable solution?” OpenAI’s model found that the answer was no. To be clear, the timeline and details are disputed1, but it’s not disputed that an AI did in fact solve a Millennium problem.

That straight line on the graph? Follow it up and to the right. It was about 14 months between AI solving hard high-school problems and AI solving one of the hardest and most famous problems in the entire field. This seems to have been entirely unanticipated by the Leiden Declaration’s authors. What is certainly unanticipated is what happens if that straight line continues. In 14 more months—in the fall of 2027—AI will be at a level of capability that makes a Millennium problem look like a high school math problem. A problem for exceptionally bright high schoolers, but high schoolers nonetheless. And if it keeps developing beyond that?

I don’t know. We would be in a world where “type question, get theorem” is something that anyone could do with no interaction with the community of professional mathematicians at all.

A Failure of Imagination

[T]he push by AI companies to solve mathematical problems as a benchmark is detrimental to the science of mathematics, and to the mathematical community. The goals of the AI companies and the goals of the mathematical community are severely misaligned.

That’s Tao, in his blog post signed by many Fields medalists. It’s AGI-pilled in the Mowshowitz sense, in that it seems to grasp that the capability is real. It’s not ASI-pilled, in that it fails to contemplate the onrushing day when “AI companies” can be crossed out and replaced with anyone with curiosity and an internet connection.

Leiden wasn’t even AI-pilled, in the sense that it doesn’t really admit that AI could ever do what AI can in fact do now:

Don’t believe the hype[.] There is currently a strong commercial incentive on the part of the technology industry to overstate the capabilities of their products. Consult with experts, including mathematicians, in forming policy decisions rather than relying on press releases or popular reporting of mathematical results.

This is simply wrong. AI’s present capabilities are not hype. It can really do what is being claimed. Open up Claude Fable 5.1 or GPT-6 and see for yourself. You need not take anyone’s word for it.

This probably extends to the non-proof aspects of mathematics as a profession. Tao and colleagues argue quite reasonably that automated AI results-generation will break the traditional role of mathematicians in simplifying, writing, canonizing, and teaching results. This is a real risk. It’s also true that all these tasks are straightforwardly subject to automation themselves. If the AI has no idea what a beautiful proof is—and to be clear AI proofs at the frontier level do tend to be fairly abstruse—this is exactly in the sweet spot for tasks that AI is likely to rapidly improve at. You can go on sites like Midjourney and A/B test pairs of images and have a custom style generated for you in minutes. Might this be the case for proofs? Not today, but the line tolls for thee in these respects too.

This failure of imagination by mathematicians surprises me. They see abstractions within abstractions to a level that I admire and envy. They are only now starting to see the straight line. I hope they extrapolate it beyond the present. There’s no guarantee that extrapolation will continue, but historically AI has rarely or never stopped improving at human level, especially not in domains like mathematics where verification can be automated. Mathematicians should prepare for this.

The Last Year

If I were a bookie at a mathematical casino I’d probably set the over-under on Millennium problems solved by AI before the end of 2027 at about 3.5, counting Navier-Stokes. There are some problems in particular that I think are hard even by those rarified standards. The Riemann hypothesis has survived well over a century of truly concentrated effort by the best minds in the world, and if the hypothesis is in fact true then there’s no clever counterexample to be found by the kinds of conceptually short leaps that AI is now especially good at. The P = NP question is probably harder still. No one even has a good idea on where to start, much less what deep connections might be needed. Neither humans nor AI nor AI-assisted humans (or vice versa) may be able to address those two in particular.

But in a year, following that straight line, I don’t think manual theorem-proving will exist anymore. Perhaps in recreational or competitive manual theorem-proving, as in chess which is thriving despite the fact that computers are on a completely different planet in terms of skill level. I could be wrong, but Tao’s worry is a simple fact. Mathematicians need to find a way to deal with this and embrace the good they can make of it.

Confessions of a Former Athlete

In some sense I myself am like a college baseball player. I was good in high school, one of the best of my peers. In grad school, probably the loose equivalent of baseball at a big state school, I was middling. Could I have gone pro? Maybe I could have been a minor-leaguer in the postdoc world, but I wasn’t going to the big leagues. I was certainly never going to be a Charles Townes or a Willis Lamb, much less a Maxwell or an Einstein. So, like most college athletes and in fact most STEM Ph.D.s, I left the game and found gainful employment outside the university world.

So I’m not a man with tenure telling the up-and-coming students not to be discouraged while sitting in my ivory tower. I know what it’s like to love a game where I was never the star athlete.

There’s a question we have to ask ourselves: do we love the game, or do we only love playing the game?

I loved both. But I still love the game. I follow physics research, I apply my physics and mathematics skills to the cutting-edge engineering I do now, and I embrace AI in doing so. Do I regret that those still in the game are about to experience the fact that the AI field’s triumph will end thousands of years of human dominance in unassisted communion with the infinite? I do. Do I regret that my own hard-won skills, that I still rely on every day, are no longer so special? I do.

But I’m consoled by the fact that I will know so much more than I ever would have thought. I will learn mathematical truths humans would never have found in my lifetime. I will stand on the shoulders of mechanical giants.

Whether that’s enough, I don’t know. I will not presume to tell anyone what this should mean for them. I can only promise that today, of all the times in intellectual history, will be the most interesting.

Matthew Springer

  1. Which is certainly not historically an uncommon phenomenon in human science and mathematics.