Though it’s already been overshadowed in the impossibly-fast news cycle by Anthropic defector Jacob Coxon’s ludicrous pronouncements about the end of the world (which I hope to get into at some point soon), OpenAI recently made headlines with their announcement that they had solved the Navier-Stokes problem: a longstanding, high-profile problem in fluid dynamics, one of seven math problems that the Clay Mathematics Institute deemed a $1,000,000 Millennium Prize Problem in 2000.
The controversy surrounding this announcement – involving, among other things, a parallel discovery by another team of researchers – has been pretty wild. In this letter, I hope to give a rundown on what we know so far, as well as to reflect on some of what I think we should take away from it all.
Before we get into the main newsletter, I’d like to make clear how I’m thinking about the actual findings. I’m not a mathematician, so I can’t weigh in on the quality of the work. As best as I can tell, LLMs genuinely can be useful to mathematicians. And given the way both LLMs and modern mathematics research work, this makes sense. 1
But it’s also worth noting that, while impressive, this utility does not tell us anything in particular about anything besides its capacity to generate material useful in the context of seeking proofs or counterexamples to specific pre-articulated problems – itself only one aspect of the overall process of doing math, as we’ll get into later in the letter. Humans were necessary at every stage of the research.
In any case -- here is part one of this write-up. Part two (the conclusion) will be coming out on Friday. Thanks for reading! More soon -- Theodora
A premillennial dispensation
Several weeks ago, a before OpenAI’s announcement that they had solved the Millennium Prize Navier-Stokes problem, a math professor at NYU named Tristan Buckmaster announced that he and Levent Alpöge, a mathematician working for the AI company Anthropic, had found a proof which solved a smaller Navier-Stokes problem.
To be clear, Buckmaster and Alpöge hadn’t solved the Millennium Prize problem. But they had made substantial, novel progress toward a solution.
This announcement was unusual not merely for the significance of the solution. After describing the interpersonally collaborative, years-long, AI-assisted process which resulted in the proof, Buckmaster takes the remarkable step of apologizing for the manner in which he’s making his announcement:
There is another part of this story, and one that, honestly, I very much wish I did not have to be concerned with. I am not happy about the presentation quality in these papers. Ideally, we would have preferred to spend weeks turning the LLM generated proofs into something readable from the very first page. This level of care is what these problems and the community devoted to these problems deserves.
He continues: “The Euler writeup, in particular, can only be described as AI slop. I am sorry for this. The reasons are below, and they involve our being pressured by outside factors.”
What happened here?
Warning signs
While Buckmaster’s account has been ably summarized in venues such as TechCrunch (and was usefully elaborated upon in interviews with the New York Times), it's worth running through the basics:
After a long stretch of slow progress, Buckmaster and Alpöge arrived at a verified proof of this smaller Navier-Stokes problem on August 22nd.
Soon after this, they heard through the grapevine that someone had tipped OpenAI off to the fact that progress had been made on the Navier-Stokes problem. Apparently, however, OpenAI believed that a team working at rival Anthropic had just solved a Millennium Prize. Buckmaster wrote to OpenAI to clarify what they had done, and to make clear that theirs was not, in fact, an institutionally-affiliated effort. (Though Alpöge works for Anthropic, his involvement in this proof was effectively a side project.) Buckmaster was not receiving funding from any AI company, relying instead on commercially-available subscriptions – including, as will soon become important, a subscription to OpenAI’s own ChatGPT.
OpenAI quickly followed up asking for more information, then began pressuring Buckmaster into meeting as soon as possible. After asking for more time several times, Buckmaster eventually agreed. In his meeting, he was informed by Sébastien Bubeck, leader of OpenAI's math team, that OpenAI had independently reached a proof for the Navier-Stokes Millennium Prize.2
As Bubeck described the solution, Buckmaster noticed something suspicious. He and Alpöge had taken a unique angle on the problem, one he knew nobody else was working on. OpenAI’s sudden new proof followed this same idiosyncratic approach to its conclusion: namely, the Millennium Prize.
The idea that his work had been stolen was further supported by the fact that, over the course of the meeting, OpenAI revealed that they had first prompted their model after receiving the email from Buckmaster clarifying his findings:
I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
When he asked how they’d arrived at it, Buckmaster was told that they’d reached the solution through a simple one-off prompt. (It reminds me of the joke I’ve heard about AI startups: “Claude, make me a billion-dollar company.”) This was a lie, given that, in reality, “an entire team had been working on the problem.”
While Buckmaster explicitly refuses to accuse OpenAI of anything, OpenAI, according to the Guardian, “could not rule out that data from the pair’s use of their products ‘helped improve their models.’” (They later denied the claim more strongly to the New York Times.)
Grand theft theorem
The story after this gets even stranger, but I’d like to pause here to reflect on this possible theft.
Companies like OpenAI and Anthropic have already shown that they are uninterested in observing basic citational practices. Given the fact that LLMs require massive violations of consent (and possibly copyright law) to procure the data necessary for training, the potential theft in this case comes as no surprise.
Still, its brazenness is both bracing and clarifying. Among other things, the blow-up around the Navier-Stokes problem has clarified for us just how Faustian the bargain of relying on commercial AI products actually is.
In order to use ChatGPT, users must, of course, commit to a Terms of Service agreement, including a privacy policy. This policy states that:
We collect Personal Data that you provide in the input to our Services (“Content”), including your prompts and other content you upload, such as files, images, audio and video, Sora characters, and data from connected services, depending on the features you use. Some of our Services allow you to interact with other users, such as post, comment, or send messages, and we treat those interactions as Content, too.
What are they allowed to do with that data? Among other things, they can use it “to improve and develop our Services and conduct research.”
In their announcement of the Navier-Stokes solution, OpenAI states that they released their results simply in order “to report on the substantial progress of [their] AI models,” and that they don’t intend on claiming the Millennium Prize money.
This comes across as appropriately detached, even noble. But I think it’s also language which carefully lays the groundwork for an argument that they haven’t violated their own privacy policy. If it’s just research into the capacities of AI, it’s fair play, right? They’re not even interested in the money! (A claim which loses some of the moral high ground when you remember that they probably spent at least twenty times the prize money simply finding the solution.)
This is the deal that researchers are making with AI companies when they use commercial AI products. Researchers take advantage of these powerful tools at the expense not merely of “externalities” such as emissions and labor rights violations, but with the knowledge that they are opening their lives’ work up to the whims of fundamentally malign corporations. If they decide that your “training data” is needed for their “research,” they will take it and use it as they see fit.
And while the Buckmaster example is particularly flamboyant3, it’s worth noting that ordinary users are just as subject to this. LLMs require titanic quantities of human-authored text to train: they’re almost certainly using your prompts for at least general-purpose training. And given that they have been repeatedly shown to be capable of regurgitating massive chunks of copyrighted text, there’s nothing intrinsic to the technology keeping it from spitting out your work wherever it sees fit.
All of which is to say: to the so-called frontier AI companies, you are not just a customer using the tools they create. You exist, in their eyes, to provide yet more grist for the mill.
An offer you can’t refuse
As disgusting as the idea that OpenAI trawled a competitor’s chat logs to snipe a high-profile math proof is, I was even more taken aback by the next part, where the company starts behaving like straight-up mafiosi:
According to Buckmaster, at the meeting he was rushed into attending, Bubeck (OpenAI’s mathematics team lead) attempted to pressure Buckmaster into either not releasing his findings on the Navier-Stokes problem or attributing the findings to an “internal OpenAI model.” In the same conversation, OpenAI asked him to take Alpöge, who worked for their competitor Anthropic, off the paper’s byline, despite the fact that he had been Buckmaster’s most important collaborator. Buckmaster continues:
I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”
The New York Times cites another mathematician’s response to the situation:
Peter Sarnak, a mathematician at the Institute for Advanced Study in Princeton, was among those whom Dr. Buckmaster phoned on Sunday. “This kind of bargaining I’ve never heard of,” Dr. Sarnak said. “I admire him for behaving the way mathematicians behave.”
Part of the equation
Often in conversations about AI, skeptics are asked to bracket pesky externalities like “labor rights” and “the continued existence of the ecosphere.” Setting that aside, aren’t these tools still useful for this or for that?
What the kerfluffle around the Navier-Stokes theorem reveals is that this separation is impossible, and not just for idealistic reasons.
On September 11, twenty-five Fields medal-winning mathematicians published an open letter titled “A Severe Misalignment of AI in Mathematics.” This letter begins:
Over the last few months, the mathematical capabilities of LLMs have improved dramatically, to the point that they can solve major outstanding problems in many fields of mathematics. However, the push by AI companies to solve mathematical problems as a benchmark is detrimental to the science of mathematics, and to the mathematical community. The goals of the AI companies and the goals of the mathematical community are severely misaligned. We see these as part of broader alignment issues impacting other scientific and creative professions, as well as the whole of society.
One of the signatories of this letter, Terence Tao, has been making a number of interesting posts on the social media platform Mastodon about this. Tao argues, among other things, that the solution to a given problem isn’t the part that pushes the field forward. Progress comes through digressions and false starts and the accidental discovery of rabbit-holes mathematicians hadn’t considered exploring. By overemphasizing results, mathematicians put the entire collaborative process at risk.
Or, as the open letter puts it: “Solving problems is only a tool and proxy for achieving the primary goal of conceptual understanding and insight…the mass production at faster and faster pace of ‘true/false’ statements could destroy fertile ground instead of breathing life into new ideas.”
And none of this is to mention AI’s pedagogical impact. Teaching isn’t incidental to research; it’s what keeps the very prospect of research alive. Like so much cultural knowledge, a field as complex and cumulative as pure math is ultimately more precarious than we’d like to believe. Unless the understanding necessary to grasp and build on them is transmitted from one generation to the next, mathematicians’ insights will not endure, because there will be nobody left to build on them. And whatever its utility for advanced mathematicians, generative AI is turning out to be wildly deleterious for learning.
And like all academic communities right now, mathematicians are already under tremendous strain. As I mentioned in Part 1, Fortune Magazine reports, “Even before AI threatened their work, federal funding for mathematics research has fallen roughly 72% under the Trump administration’s cuts to the National Science Foundation.”
Means to an end
LLMs only exist thanks to the work of human communities. Human beings created all the worthwhile data in their training sets; human beings are responsible for endlessly laboring to fine-tune their output (sometimes under trauma-inducing conditions, for $2 an hour). Claude or ChatGPT would not be able to puke up a new mathematical counterexample if it hadn’t been trained on the vast history of human-authored mathematical writings and results. As so many capitalists have done in the past, so-called frontier AI companies are now bulldozing the communities whose work their models have gorged on.
All of the mathematicians in this story were trained at conventional universities, amidst communities of professors, graduate students, and undergraduates. The mathematicians in the employ of companies like OpenAI are the direct beneficiaries of these communities. By lending their labor to corporations like Anthropic and OpenAI, they are taking the gifts given to them by the community of mathematicians and redistributing them to future shareholders. As the mathematician Michael Harris puts it in a sweeping, beautiful essay for the Boston Review, “Mathematics only makes sense as a gift economy.”
And there are basic functionality reasons why companies like OpenAI would want to support these communities, too. Human beings were the ones who figured out which questions to ask the chatbots, which approaches to prioritize whether or not their outputs were correct. The need for human input and direction isn’t going away. Without new, non-AI generated material, models get worse; this is a well-documented problem known as “model collapse.”
Which is to say that, were these companies concerned about the long-term utility of their product, speed would be secondary to care and concern. They would be working to truly be of use to mathematicians, rather than attempting to beat them to the punch or kick researchers’ names off of publications.
Math is a social activity, and AI, both in practice and in theory, is intrinsically, brutally antisocial. LLMs are built on stolen training data, exploited labor, and wildly gratuitous expenditures of energy; the corporations responsible for it take advantage of the fact that we are social animals, primed to recognize minds like our own behind language, and use these tendencies to trick us into emotional attachments that will never be reciprocated – all in order to boost their subscription counts.
They aren’t doing this out of their desire to contribute to the progress of human understanding, or because they want to see the field of math pushed forward. They're not doing it because there's big money in pure mathematics: they’re losing money even on their most popular products. And as this all makes clear, they're certainly not doing it because they want to help mathematicians.
So why math?
As the technologist and thinker Tante puts it in a blog post:
It’s not like [math is] a use case actual people or businesses have in their everyday life [...] they are burning a lot of money doing something that many people can’t relate to.
Which probably is the whole point: Math is scary and complex and hard to grasp for many if not most – especially when going into more abstract directions. But everyone had to go through some math in school or in their other education, and everybody kinda knows that math is important. And math is often – in the public eye – still seen as a space for geniuses: Math geniuses who just can think about these things differently from everyone else. Math not as a collaborative endeavor of a group of trained professionals who slowly work towards bringing light in the dark but as a sequence of sparks of insight by individuals.
The focus on mathematical proofs therefore seems to mostly try to place stochastic language models in the real of “genius intelligence” therefore imbuing these models with agency and trying to legitimize pushing these systems into more and more parts of our lives.
These corporations care about math for the same reason they care about cybersecurity: because they think we’re fools and want to impress and scare us, because they know we don’t like their sloppy products, because they’re about to IPO despite still being wildly unprofitable, and because grand pronouncements about their robot sons solving big problems on their own is great marketing for the executives still desperate, against all evidence, that they will be able to replace their workers with sycophantic plagiarism machines.
LLMs as such may wind up being a remarkable tool for practicing mathematicians, but in their current form as “AI,” commercial chatbots are a nightmare of appropriation and pillage, the extrusion of a small handful of unhinged plutocrats trampling anything and everything that stands in the way of conquest. They are what you come up with when you see everything -- knowledge, art, people themselves -- as means to an end.
This is why, no matter what tasks the computer becomes capable of automating, we must resist the corporations attempting to shove this janky, toxic technology down our throats. All the lip service they give to concepts like “alignment,” “safety,” and “community” &endash; functionally, these are lies. Incidents like the Navier-Stokes debacle continue to prove to us who and what they really are. Like their mindless product, these AI companies don’t understand anything about labor, or passion, or learning, or culture, or care. As Buckmaster said to a New York Times interviewer, “They couldn’t understand […] why I wouldn’t throw Levent under the bus.”
Epilogue: At what cost
On the one hand, I see no reason to doubt mathematicians’ claims that they are truly remarkable automation tools, capable of accelerating progress on long-standing, deeply difficult problems. Unlike what I’ve observed from many other fields, even the strongest statements against AI from within the community of mathematicians seems to leave room for the idea that they will ultimately have a place in the practice of pure mathematics.
On the other hand, it’s really worth pausing to appreciate just how unbelievable the spending on these proofs has been. Based on OpenAI’s token-based pricing – prices, that is, based on the cost weaker models; prices which nonetheless still may be subsidized by venture capital investments – solving this problem probably cost OpenAI more than twenty million dollars.
On the level of the individual researcher, the more affordable models Buckmaster himself used are shockingly well-subsidized, because AI as a consumer product remains massively unprofitable. On the highest personal Claude subscription level, $200 a month can get you upwards of $10,000 worth of tokens. Inference is getting cheaper, but it’s nowhere close to profitable, especially given how many of the “advances” in the technology are a result of simply throwing more tokens at the problem. (OpenAI allegedly ran 10,000 instances of their fanciest model simultaneously to calculate the solution.)
There do exist freely-available, locally-hosted models, but their performance tends to lag far behind so-called “frontier” models, given the enormous budgets devoted to training and maintaining proprietary commercial models. (If you’re wondering why these companies are suddenly making grand pronouncements about the importance of slowing down the training process, look no further.)
And even these open-source models are unbelievably expensive. The costs are genuinely hard to grasp. Running a single instance of Kimi K3 at full speed, one of the latest open-source models you can run on a computer you own, requires 1.5 terabytes, or 1500 gigabytes, of VRAM. VRAM is typically used for computer graphics such as those in videogames; the highest-end videogame console available right now, the PlayStation 5 Pro, uses 16 gigabytes. Not only was the model OpenAI used to solve this problem was probably much more expensive to use than Kimi K3, they solved the problem by running 10,000 instances at once.
All of which is to say: LLMs seem to be pretty helpful for brute-forcing math problems. But if there is one thing we know for certain at this point about “AI” technology, it’s that it comes at a cost.
Suggested reading/watching:
-
A Severe Misalignment of AI in Mathematics: I’ll get into this more in the next letter; for now, it’s worth noting that this is a lucid and powerful statement from twenty-five Fields Medal-winning mathematicians on the conflict between AI companies and the discipline of math.
-
One of those signatories, the mathematician Terence Tao, has a lot of very thoughtful posts about the implications of LLMs on math research on his Mastodon blog. Here is one about how fetishizing proofs’ solutions might actually be counterproductive to the overall collective project of mathematics.
-
The computer scientist and author Cal Newport is always good on breaking down the hype around LLMs and “AI” technology; his background in math gives him particularly helpful perspective on an earlier moment of asking this question in his video “Did AI Just “Solve” Math? (Let’s Take a Closer Look).”
-
A darkly funny short story about the absurdity of excluding “externalities” like climate catastrophe and labor rights from the pursuit of computer-assisted math research: Brown University mathematics professor Rich Schwartz’s scathing “The Fate of the Reimann Hypothesis.”
-
I’ll get more into this one in the next letter, too, but my personal favorite thing I’ve read in the course of my research for this piece is Michael Harris’s moving essay on the disjunction between the joys of math (and learning more broadly) and the brutality of AI corporations, “Knowledge Collapse,” for the Boston Review.
The Genuine Intelligence Project is an initiative from the North Carolina branch of the American Association of University Professors. Check out our website, and follow us on Instagram and BlueSky!
-
LLMs train and work best when (a) the language they are training on is highly structured, like math is; (b) there exists a large pre-existing body of training material; (c) they can quickly receive binaristic yes-or-no answers, such as those given by pre-existing tools such as the proof assistant Lean; and (d) when the output won’t cause significant problems if it’s faulty. Reasons (a), (b), and (c) are the reasons LLMs can whip up computer programs that run: computer programming languages are designed to be unambiguous and structured in a way human language will never be, and the question of whether or not a program runs is an easy yes/no test. (Plus, the massive repository of publically-available open-source software on which LLMs are trained (often in violation of licensing agreements) provides a nice data set for LLMs to rip off, sometimes at length. ↩
-
While at Microsoft, Bubeck contributed to the highly influential, intellectually worthless “Sparks of Artificial General Intelligence” paper, notable for, among other things, sourcing its definition of intelligence from an op-ed defending the racist work of pseudoscience The Bell Curve. This was not an accident, as the entire field is lousy with eugenicists; the work of Timnit Gebru and Émile P. Torres is essential reading on this front. ↩
-
This is the longest footnote I’ve ever written; while it’s not worth putting in the main piece, I had to put it somewhere, so here it is:
It’s worth taking a second here to make a note about Alpöge, who has wisely elected, it seems, to sit this press cycle out. Alpöge, an Anthropic employee, has been in the limelight for LLM-assisted math research before. In particular, he made news for his Twitter post announcing that he’d found a counterexample to the Jacobian conjecture. While the finding was inarguably significant, one reason for the press is that he eschewed conventional means of announcing such a finding by electing instead to post it directly on Elon Musk’s social media platform in the distinctive Quirky Millennial™ all-lowercase aw-shucks house style preferred by fellow posting luminaries like Sam Altman. (If you couldn’t tell, I’m perhaps disproportionately irked by this; I have been since I first saw it.)
With the anthropomorphism characteristic of Anthropic’s delusional workforce, Alpöge refers in the same breath to "my friend fable" and "my friend akhil"; in a longer piece about the relationship of AI to mathematics, the actual human friend cited here, University of Chicago mathematician Akhil Mathew, describes the situation to Fortune Magazine with ambivalence: as “a very rapid and very unsettling change…especially for junior mathematicians.”
I’m not sure if Mathew thinks he is referring simply to mathematicians in the earlier stages of “ordinary” careers, but for anyone who knows anything about academia, “junior mathematicians” are the math scholars most vulnerable to the precarious higher-ed labor market, a labor market our current presidential administration has only made worse: according to that same Fortune piece, the Trump administration has cut funding for math research by 72%, even as the Brookings Institute reports that the potential value of Department of Defense AI contracts in 2026 has increased by 1,605% over the last year, to $90.7 billion.
The business pitch of modern AI is the replacement of labor as such. Though the evidence strongly suggests that this will never happen, I am of the opinion that most, if not all, of those whose work lends so-called frontier AI companies power or public legitimacy should be regarded as antagonists of labor.
Given the ambivalence, even anguish, with which other mathematicians are speaking about this current moment – and given the fact that, again, the Trump administration has cut funding for math research by 72% – I struggle to find Alpöge’s smarmy twee brashness anything other than a sign of how fundamentally isolated and anti-solidaristic the culture of Silicon Valley is. And while I remain more sympathetic to Buckmaster, the presence of Alpöge in the kerfluffle is a good reminder that, again, these are Faustian bargains. If you, dear reader, want to avoid getting caught in the crossfire between two megacorporations, I suggest refusing to collaborate with a prominent hype man for one of them. ↩