Hi everyone! Hope your school years are starting off well! This letter is the first in an occasional series I’m going to write about a few strange ideologies and their influence on the political landscape of AI, though I’ll continue to write about other topics alongside this one. More soon! – Theodora
Who Evaluates the Evaluators?
A couple of letters ago, I cited a study by the organization METR, which is short for “Model Evaluation & Threat Research.” (For the record, they're the ones who omitted a noun like “Institute” at the end of their name.)
To be clear, the study seems solid, and it’s responsible for one of the most striking and illustrative visuals I’ve found for illustrating the relationship between the subjective perception of AI’s utility and the evidence of what it actually does.
But in a footnote, I alluded to the “highly dubious milieu” from which METR emerges. I think understanding this milieu is important for understanding the material forces shaping the culture of AI.
METR and their fellow travelers work in a field known broadly as the field of “AI Safety.” One would think that that people working in a field called “AI Safety” would be allies in the struggle against, for example, administrations at academic institutions forcing unsafe and inaccurate chatbots into the working environments of students and teachers. But a look at METR’s mission statement paints a different picture, largely for what it leaves out.
Their mission statement does not mention the environmental and human harms caused by the data center buildout; it does not mention anything about the “technical sweatshops” in which AI data workers labor or the widescale theft of creative works to serve as training data; it does not mention cognitive deskilling, the mental health risks associated with chatbot usage, or LLMs’ fundamental unreliability as an information source – just to name a few. No, METR is exclusively and specifically concerned with “the autonomous capabilities of AI systems” – with “advanced AI systems [pursuing] goals that are at odds with what humans want.”
This, and only this, is the purview of “AI Safety”: not the actually-existing crises caused or exacerbated by what people today call “artificial intelligence,” but one specific hypothetical future crisis: namely, “rogue AI” taking over the world and threatening us all.
AI “Safety” and AI Marketing: Two Sides of the Same Coin
If this seems far-fetched, it’s because it is. As I discussed in my last letter, LLMs are nowhere close to becoming “autonomous.” LLMs receive inputs, convert them into tokens, calculate the next likely token, and convert them back. They cannot “break containment” unless you run them through a separate application (a “wrapper”) designed to give them access to your computer. An OpenAI chatbot in a Pentagon web browser is not capable of leaping through the network and setting off the nukes. Every time you hear about an AI “breaking containment” and doing something it shouldn't, you’re not seeing the first stirrings of Skynet – you’re seeing something roughly equivalent to the time the so-called “Director of Alignment at Meta Superintelligence Labs” told an AI “agent” to clean up her email inbox and it deleted them all.
The problem is not that computers have evolved “superintelligence” and gone rogue. The problem is that people are trusting probabilistic, error-prone computer programs with responsibilities they will never be good at handling, because they cannot understand “responsibilities” like we can. If you’re concerned with human safety in the age of big tech, your priorities should include both ongoing real-life harms and the broader economic and social structures incentivizing risky behaviors. (For a much more specific look at what it is that “AI Safety” leaves out, see Emily Bender's Medium piece on the so-called “schism” between “AI Safety” and “AI Ethics.”)
Ok: so they’re cranks. But there are lots of cranks out there. Why care in particular about this set?
For one thing, they’re far more influential and powerful than most of us outside of Silicon Valley would believe. “AI Safety” constitutes one public-facing front of a much larger movement, a movement emerging from a complex of ideologies that the scholars Timnit Gebru and Émile P. Torres call “the TESCREAL bundle.”
I’ll unpack these ideologies as this series moves forward, but it’s worth noting that the worldviews emerging from this ferment have been espoused by figures like Anthropic’s billionaire CEO Dario Amodei, disgraced crypto investor Sam Bankman-Fried (who at one point was probably the most visible exponent of “effective altruism,” itself maybe the most influential of the TESCREAL ideologies), Elon Musk and Peter Thiel, and others. As Karen Hao documents in her authoritative Empire of AI, OpenAI CEO Sam Altman was briefly removed from his post as president of the organization for failing to live up to TESCREALers’ standards; still, four years later, he said that Eliezer Yudkowsky, the highest-profile crank in AI Safety world, might deserve a Nobel Prize.
And as is so often the case with ideological struggle in the United States, the university is a privileged site of contestation. Though neither has been affiliated with an actual philosophy department in a decade, William MacAskill, Nick Bostrom, and Toby Ord, perhaps the three most influential exponents of effective altruism, used their postings at the Elon Musk-funded “Future of Humanity Institute” at the University of Oxford to lend institutional credence to their fringe philosophical work. One new “advocacy organization” working to prevent “the development of superintelligence that poses existential risks to humans and the planet” is offering $2000 a week to college organizers. While the Center for AI Policy think tank (whose website proclaims they are “bolstering government so we can manage powerful AI”) closed down last May, they continue to sponsor AI Safety groups at a number of universities.
This situation is only set to get worse. Both OpenAI and Anthropic are anticipating massive IPOs this year, and many are anticipating that some of this windfall will trickle down to TESCREAL-affiliated organizations.
Groups and activists bound up in these beliefs – beliefs which, as Gebru and Torres argue, are ultimately founded on eugenics and so-called “race science” – will never lead us to a truly democratic or egalitarian society. The “safety” they espouse won’t protect us from the real harms we’re already seeing from AI. As the economic bubble nears bursting and public opinion continues to turn against these technologies, it will be even more important to keep our attention trained on what we’re actually opposing: not the hypothetical “superintelligence” of the science-fictional future, but the legislators, institutions, and capitalists working to sustain and reinforce today's unequal structures.
Further Reading
I know I make some big claims in this piece about some philosophies that possibly sound good on their face, and I hope to support these claims in emails to follow. In the meantime, here are a few pieces to check out if you’d like to go down the rabbit-hole early.
Here’s an article about Effective Altruism in particular:
Effective Altruism’s Bait-and-Switch
When Effective Altruists talked in public about “doing good,” “helping others,” “caring about the world,” and pursuing “the most impact,” the public understanding was that it meant eliminating global poverty and helping the needy and vulnerable. Inward, “doing good” and the “most pressing problems” were understood as working to mainstream core EA ideas like extinction from unaligned AI.
In the communication with “core EAs,” “the initial focus on global poverty is explained as merely an example used to illustrate the concept – not the actual cause endorsed by most EAs.”
“Longtermism” is another ideology foundational to the broader AI Safety movement. Here is an excellent rundown by the excellent critic (and former adherent) Émile P. Torres:
The Dangerous Ideas of “Longtermism” and “Existential Risk”
What I find most unsettling about the longtermist ideology isn’t just that it contains all the ingredients necessary for a genocidal catastrophe in the name of realizing astronomical amounts of far-future “value.” Nor is it that this religious ideology has already infiltrated the consciousness of powerful actors who could, for example, “save 41 [million] people at risk of starvation” but instead use their wealth to fly themselves to space. Even more chilling is that many people in the community believe that their mission to “protect” and “preserve” humanity’s “longterm potential” is so important that they have little tolerance for dissenters. These include critics who might suggest that longtermism is dangerous, or that it supports what Frances Lee Ansley calls white supremacy (given the implication, outlined and defended by Beckstead, that we should prioritize the lives of people in rich countries). When one believes that existential risk is the most important concept ever invented, as someone at the Future of Humanity Institute once told me, and that failing to realize “our potential” would not merely be wrong but a moral catastrophe of literally cosmic proportions, one will naturally be inclined to react strongly against those who criticize this sacred dogma. When you believe the stakes are that high, you may be quite willing to use extraordinary means to stop anyone who stands in your way.
Here’s a great op-ed about how AI Safety “doomer” rhetoric serves the corporations they’re ostensibly opposed to:
Dear AI Companies, the Doom Trolling Needs to Stop
There are really only two options for the intentions of A.I. companies when they engage in doom trolling. The first is that they actually believe that the systems they’re building have a nontrivial chance of producing hugely disruptive events — from destroying the economy in the best case to wiping out our species in the worst. If this were true, every reasonable ethical system would argue that there is only one acceptable response: to immediately stop working on any product that might accelerate such a future, and lobby with all of your resources to help force other A.I. companies to do the same. From a moral perspective, any other reaction would be monstrous.
The second option is that these A.I. companies aren’t really concerned about these risks, and that they’re injecting these doses of unresolvable doom for other reasons. They might want to amplify the perceived power of their technology at a time when they’re setting up their initial public offerings. Or they hope their performative reports and somber interviews will help them compete for top engineering talent coming from a Silicon Valley culture that’s steeped in this type of doomerism. The venture capitalist and A.I. adviser David Sacks recently suggested that Anthropic was using fear-mongering tactics as a method of “regulatory capture,” which can impede upstart competitors. Any of these reasons would mean that these companies are laundering the anxiety of millions to improve the financial fortunes of a vanishingly small number of major stockholders. This cynicism would be equally monstrous.