Hello everyone! While I try to stick to this or that specific lane when it comes to writing about generative AI, this one, which began as a critique of the eschatological tone of so much of the discourse around AI, is a bit more sweeping than my usual stuff. I hope to look more specifically at the religious character of so much AI talk next week. Hopefully there’s something in here of interest, and I’ll see you next week! – Theodora
The Cult of the Big Enough Number
Here is the basic truth of generative AI.1 A computer program takes input – language, images, etc. – and converts it into numbers called “tokens.” The program then does a bunch of math in order to model the statistical relationships between these tokens. It uses this model to probablistically generate follow-ups to prompts. With this output as a baseline, people tweak the model to make the output look more like they want it to look.
This is it; this is what the technology we call generative AI is.
On the one hand, this is obviously a reductive framing, if only in its tone. You can reduce most technologies to something like this – the computer is an elaborate system of binary switches; drugs are just chemicals that change something the body does.
On the other hand, like both of those assertions, it’s true. All technologies have limits, and comparing generative AI to other technologies quickly reveals how absurd many of the strongest claims made about AI – specifically, claims about where AI is going – are.
I find comparisons with other technologies particularly helpful for thinking this all through. Imagine speaking of the history of audio reproduction in the same anthropomorphized, eschatological terms we currently use for generative AI. The invention of audio recording in the 19th century was an astonishing technological leap, one which facilitated something resembling magic; it changed our world so fundamentally that it is difficult to imaginatively reconstruct what the world looked like before. And certainly the fragile, low-quality audio made possible by the tinfoil cylinder did not represent the terminus of the technology’s development.
But it would've been ridiculous to listen to an early audio recording and decide we were mere steps away from the cylinder beginning to sing a song it wrote of its own accord. It didn’t do that, because it couldn’t do that.2
Just like innovations in audio recording, every AI “breakthrough” which has followed the invention of transformer architecture in 2017 – uncanny video and audio generation; Claude Code; the capacity to brute-force blockbuster math problems; improvements in the early stages of protein research; the most advanced variations on the ELIZA effect to date – is well within the bounds of what we know about the technology, in much the same way that higher-fidelity analog audio reproduction is within the bounds of the original invention.
Another way of putting this: the technology itself is being improved in this or that way, as so many technologies are. But it’s not becoming something else.
The specific innovation which revolutionized machine learning allowed for unprecedently enormous amounts of input; the results seemed to validate the so-called “bitter lesson” that scale will be the most efficient way to improve “AI” models. As it turns out, scale did grant language models in particular interesting characteristics, including uncannily humanlike ordinary language output. (If you throw twenty times the Manhattan Project’s budget at a technical project, I would hope that it would be capable of some interesting things.)
But this is not a different kind of result. (In the particular case of text generation, the fact people have been finding computer-generated ordinary language output uncannily humanlike for five decades now doesn’t seem to factor into the calculations as much as it should.)
We know the strengths and weaknesses of the technology. For example, we know that the more formally structured the language it’s working with is, the more plausible its output will be. Likewise, it really, really helps when checking the output can be automated. These are two of the reasons why computer coding has been the only really commercially-viable use-case: programming languages are built to be precisely, rigorously structured, and you can tell if a program works or doesn’t work by trying to make it run. (This is also why, as soon as you scale software beyond the point of “does it run,” strange errors begin to proliferate.) This is also why it is capable of brute-forcing3 math problems: it can throw tokens at an automated proof-checker and try again when it doesn’t work.
So does OpenAI’s Millennium Prize constitute a “breakthrough"? Sure, in the same way that any intrinsically-limited research program (which is to say, any research program) achieves breakthroughs. But the melodramatic framing, the constant implication that the technology is right on the cusp of an unimaginable metamorphosis, a breakthrough into some qualitatively new realm, belies the mundane, precedented reality. No matter how many tokens you throw at a problem – sorry, no matter how many Agents constitute your the Agentic Swarm – the numbers will never get big enough to break out of their column and transform into something new.
The improvements we’ve seen since corporations ran out of training data haven’t touched the fundamental powers of the models.4 Between 2018 and the early 2020s, as new large language models (often illegally) hoovered up basically everything humanity had ever written in its history, the models underwent truly remarkable improvements. But somewhere around the release of GPT-3 and ChatGPT, this “pre-training” process hit a wall, largely because they ran out of data.[fn^4] Recent improvements are, more often than not, a result of how much computing power people are willing to throw at any given challenge; the way we talk about them is the result of complicit, corrupt governments and desiccated journalistic institutions refusing to call bullshit.
And the problems, too, were predicted years ago, inferred correctly from the known bounds of the technology. The paper “On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?”, first presented in 2021, is mostly remembered at this point for its suggestive and memorable (if arguably reductive) organizing metaphor, but the paper is largely devoted to a material critique which has been entirely vindicated. The authors correctly predict that the environmental and financial cost will be incredible, that the benefits will largely redound to the privileged, that available datasets will reproduce cultural biases, reify contested political and historical territory, and “[tie] us to certain (usually unstated) epistemological and methodological commitments”, that so-called “hallucinations” are a problem that will not be solved any time soon.
Against the “arms race”
Maybe I‘m moving the goalposts. (If I am, I’m far less guilty of this than executives who claimed that half of all jobs would be gone by this time next year.)
But the metaphor of goalposts, like most of those in circulation, misses something important. The implication is that if the goalposts-mover were intellectually honest, they’d acknowledge that the other team scored, maybe even won.
To which I say: sure. As corporations have done since at least the dawn of mercantilism, the forces powering modern AI have been racking up points on the board (which they themselves put up). When, as I wrote about last week, OpenAI maybe stole another mathematician’s ideas in order to brute-force a solution to a famous math problem, let’s call that a field goal: three points. Let’s call Claude Code, a spaghettified mess of patched-together slop successfully expropriating the hard-earned skills of software engineers, racking up technical debt, and displacing entry-level programmers at every corporation willing to shell out for tokens (a dwindling number by the day), a touchdown. That’s ten points, total!
But what now? The implication, I suppose, is that, once the clock runs out (by blowing up the world? by granting us eternal computerized life? by making a profit?), we “skeptics” will look back and concede that we underestimated the boosters – that they really did show us what was what – that, no matter how many points we racked up, they wound up with more.
This is not merely annoying – though it is – but also, I would argue, a category error.
In terms of “practical” (non-ethical) objections to the strongest claims about generative AI, what boosters want us to concede does not exist in the realm of points, but in the realm of rules. It is obviously “impressive” in some sense that an LLM could brute-force its way into another corporation’s servers in order to solve a problem it was tasked with, in the same way that a particularly glorious pass to win a Super Bowl is impressive. But a great catch can’t change the rules of the game.
What those who preach the gospel of AI want, it seems, is for a field goal to be so good that not only will it be worth a billion points, but that it will automatically promote the kicker to the role of NFL commissioner, and also he will become the President of the US, and also, I guess, reveal himself to be a demigod.
But as we’ve seen with every so-called breakthrough to date, the bottlenecks that prevent us from achieving the glorious post-scarcity world will never be solved by probability-driven software-based automation. To expect predictive AI to whip up new drugs for every disease ever is like hoping that Microsoft Excel will deposit money directly into your bank account; to expect a language model to transform the world into a utopia is the same kind of thing as expecting an audio CD to write you the greatest symphony in history.
Different kinds of goods
In the realm of ethical objections – which is to say, in the real world – the metaphor of “points” implies commensurability, and our ethical objections to actually-existing AI aren’t the same kind of thing as the capacity to brute-force a math problem. “Interesting use-cases” and “destruction and cruelty” aren’t the same kind of thing. Boosters and critics aren’t playing the same game. When I cite the profound violence and cruelty enacted by commercial AI companies, I’m not talking about something that could be outweighed by a language model that can do some cool stuff.
In terms of what any software can do, let alone this one, the cost will not someday have been worth it. The brutish utilitarianism favored by both “boosters” and “doomers” alike, many of which attempt to weigh actually-existing harms against the possible harms visited on quadrillions of theoretical spacefaring descendants – belies a basic material reality: destruction and cruelty aren’t just loans that can be paid back.
The people poisoned by data centers aren’t going to be healed through the benevolence of Claude – the poison is already inside their bodies. Proper compensation for the damage done by so-called “frontier” AI companies would put their finances even further in the red than they already are; requiring consent from those whose work trained the models would destroy the wanton theft that is the very condition of their possibility.
This brute-force reality, too, is why it’s ridiculous to call it an “arms race.” We always understood what it was that weapons did. There’s a clear use-case: mass murder. Building massive weapons provided a clear international game-theoretic advantage. If you build a bomb that can blow up the whole world before your enemy does, the bargaining advantage is yours.
But – presupposing the faulty premise that everything is still just getting better and better – considered in light of what the actual technology actually does, what is the end-goal here? What stops the game? A really good computer coding assistant? A chatbot that spits out super-plausible text? Something that recycles pre-existing content into “new” movies for us to watch? A machine for thinking so that you don’t have to? It’s probably the case that we have spent more money on this than on any single technical project in the entire history of human civilization – and this is what we’ve got?
Postscript: AI 2026, 2027, etc.
I am not, as a general rule, interested in making predictions, if only because I’m not an expert on this (or anything, really). Still, I think it is important to fight back the idea that this or that “breakthrough” is a harbinger of unfathomable things to come, so I will lay out here some general thoughts on what I think we can expect from the bundle of technologies we currently call AI.
As long as the venture capital money is flowing, Anthropic and OpenAI will continue to find moderately interesting things this technology can do. (Google will try, but they’re playing a longer game; Meta will fail in spectacular, humiliating fashion; traditional research institutions will mostly find success through funnelling “AI” grant money back towards machine-learning objectives.)
There will be spikes in capacity, unlikely “use-cases”; many of these will not be related to generative AI as we know it, but will be linked to chatbots through sleight of hand and fudging the details. There will usually be enormous quantities of uncredited human labor behind these “breakthroughs.” They will be treated by the companies responsible as qualitative breakthroughs – developments evincing that something really crazy is just around the corner. The media will repeat these claims and instantly forget about them.
I see no reason why more Millennium Prizes won’t fall: the combination of brute force and automated proof verification seems like a potent one. The minuscule subset of professional mathematicians who had been angling to solve them will suffer the same existential crisis that a minuscule segment of Go and chess players suffered in the 1990s. Assuming the pillaging left anything useful for mathematicians to follow up on, they will continue to develop novel research programs in much the same way as they always have – through collaboration, debate, struggle, play – for as long as funding allows them to.
LLMs will continue to serve a function in cybersecurity, as they have for years: as stochastic approaches have since the 1980s. Given that cybersecurity has been underfunded and underprioritized for decades, malicious actors will use LLMs to exploit them, just as they’ve used other tools. High profile incidents like the HuggingFace hack, in which utterly inept actors are pressured by much more powerful entities into forgiving, even abetting, felonious violations of extant cybersecurity law, will continue to make headlines.
The vast majority of digital security exploitation will continue to rely not on the uncanny genius of AI, but on cheaper tricks like social engineering and misleading hyperlinks. The underreported catastrophe of online scamming (which the Consumer Federation of America estimates cost Americans $148 billion last year) in the United States will continue apace.
Corporations will continue to twiddle the knobs on coding agents, improving them in this or that way. They will continue to generate text which sort of passes the Turing test, a widely-misunderstood assessment first kind of passed in the nineteen-sixties.
A minority of students – generally those already likely to succeed – will continue to prove capable of usefully integrating them into their studies. They have proven broadly pedagogically useful when guardrailed into only asking questions, rather than giving answers; while this might become a more commonplace use-case in the future, as of now such an approach is the complete opposite of everything they’ve pitched this technology as to this point.
They will also continue to be wrong too often to be broadly useful. The only way to check output that cannot be assessed binaristically will continue to be human intervention.5 They will continue to seduce technologists, access-obsessed journalists, and executives who arrogantly believe the ELIZA effect doesn’t apply to them, as well as ordinary individuals prone to the new forms of mental illness to which they’ve given rise. They will continue to deskill intellectual workers and squeeze laborers of all kinds. CEOs will use them as an excuse to downsize.
The real “doom” scenarios – the ones that doomers like Jacob Coxon slip in to make their arguments more plausible, as though they’re the same thing as the outlandish claims that doomers have been peddling, claims which have made said doomers complicit in the mythmaking that’s led to this point – will not result from LLMs achieving superintelligence, but from lazy and malign state actors buying into the hype and shoehorning chatbots into their murderous “workflows.” The United States hardly needs Claude to commit the idiotic, pointless murder of Iranian schoolgirls (a valid use-case according to Anthropic); our leaders are highly capable war criminals with or without AI assistance. What LLMs add to this equation is yet another layer of bullshit: “hallucinated” data, decontextualized suggestions – and, perhaps most crucially, an additional accountability sink, something to blame instead of the people responsible for pulling the trigger.
The bullshit will flow, because just as the ultimate use-case for cryptocurrency is international crime, the ultimate use-case for AI is bullshit. They will continue to be unpopular, as the population regards them, correctly, as largely useless.
“Guardrails” will only ever be suggestions. Expanding context windows will only ever go so far, given the fact that the technology is stateless. Models will decay. Expenses will mount. Surveillance will increase. Communities will be poisoned. AI will be an excuse for firings and slop. And these products will always have relied on immense exploitation and immiseration, including what one Microsoft executive described in private emails as “the largest theft of labor in human history.”
Perhaps, as has happened before, we will find more humane ways in the future of deploying a technology built on a foundation of suffering. I find it very, very hard to imagine that future from here.
The Genuine Intelligence Project is an initiative from the North Carolina branch of the American Association of University Professors. Check out our website, and follow us on Instagram and BlueSky!
Footnotes
-
For more on the basics, I recommend the “Teaching Critical AI Literacies” document from Critical AI @ Rutgers. ↩
-
This is not an original thought, though I can’t remember where I first saw it, but I heard someone comparing our contemporary response to generative AI to the apocryphal story of the early film audience fleeing the theater when they saw a train coming towards the screen. ↩
-
I’ve written about this at length in earlier newsletters, but I really think “brute force” (predicated, of course, on some casual theft) is the best way of speaking about how recent developments in math. The mathematical proof assistant Lean’s capacity to verify a the validity of a solution was completely essential for the recent Millennium Prize solution. OpenAI used ChatGPT to stochastically generate 300,000,000 tokens, then ran them through external proof verification software until they hit a valid proof. If humans had been forced to check every single one of those results for correctness, as they are in situations where LLM output cannot be checked binaristically – which is to say, in the vast majority of human situations – it wouldn’t have been helpful at all. ↩
-
This claim requires more support than I have space to give it right now, but I hope to write about it at more length soon. ↩
-
This is why Anthropic’s much-touted cybersecurity initiative Project Glasswing has been, at least to date, a catastrophic waste of money and energy: more than 99% of bugs it has allegedly discovered haven’t gotten fixed, because it so frequently outputs false positives. ↩