AAUP-NC Genuine Intelligence Project

What was AlphaFold?; or, What We Talk About When We Talk About AI “Curing Cancer” (Pt. 1)

Hello! In the next two letters, I’m going to talk about the closing of Google’s AlphaFold program. While AlphaFold was a technically impressive application of deep learning, I think the way we’ve talked about it has been pretty off-base in a lot of ways. In the first letter, I’m going to lay out what I see as the fundamental features of deep learning that make it difficult to say that it’s fully “solved” a problem like the protein folding problem; in the second letter, I’ll take a look at a recent paper about the surprising impact (or lack thereof) of AlphaFold on the broader research and drug development pipeline.

As is always the case, I write this newsletter from the perspective of a curious amateur; technical explanations are reductive both for communication’s sake and because I myself am not a trained computer scientist or structural biologist. Comments and corrections are always welcome.


Attention is Some of What You Need

Before we get into AlphaFold in particular, we need to talk a bit about the history of what we call “AI.” In 2017, researchers at Google published a paper called “Attention Is All You Need.” This paper proposed a much faster and more efficient way of feeding training data to a given model.1 This new “transformer architecture,” as it’s known, allowed for truly wild leaps in scaling; it’s what has allowed language models, which have existed for a very long time, to become large language models. The history of machine learning didn’t start with this paper, and the history of “artificial intelligence” certainly didn’t. But it really was a watershed moment for predictive modeling. The ability to process enormous corpuses of training data has had pretty interesting effects. For example, large language models (LLMs) – most of which rely on transformer architecture – can generate text which sounds surprisingly like text written by human beings. However, the utility of predictive models in specific circumstances doesn’t automatically translate to broader usefulness. Given the way the deep learning training process works, they are most effective in cases where (a) the inputs are highly formally structured; (b) the computer can easily check without human input if the predicted answer was right or wrong2; and, ideally, (c) it’s not terribly important if the output of the finished model is off. Reasons (a) and (b) are why LLMs trained on computer code can whip up tiny computer programs. Programming languages are specially engineered to be structured, and the question of whether a computer program runs or not has a binaristic answer: either it does or it doesn’t.3 We’ll get more into reason (c) later on. Suffice it to say for now that this is why even computer engineers – that is, workers in the field which has most widely adopted AI tools – have found LLMs unsuited for large-scale software engineering. It works fine for small things: in part because they are usually low-stakes and personal; in part because, for example, if the tiny script you’re using to convert some files is busted, the error state will probably be that it doesn’t run. As you scale up, however, bugs can be both better-hidden and higher-stakes. You might not realize something is broken until it’s already out in the world (as the creators of perhaps the most popular purely “vibe-coded” software do, again and again). As it turns out, protein structure prediction – the attempt to infer the conformation of a protein from a given chain of amino acids – is a case where (a) and (b) also apply. The question is how to infer a formal structure from other formal structures, and the data-set of molecular structures we’ve experimentally verified allows us to easily set the model up to check how wrong its predictions are. AlphaFold is a model, created by Google’s DeepMind AI research lab, that uses transformer architecture to predict protein structures. It’s because of these improvements this reason that AlphaFold has seen such remarkable success, scoring above 90 on a 100-point scale in 2020 – “25 points higher than the next-best team.”


Determinism and Understanding

Above 90, while higher than many other relevant numbers, is lower than, at minimum, several others – most notably, perhaps, 100. And we will get into this more in the next letter, but a recent study of the impact AlphaFold has had on structural biology reveals that only 6% of structural biologists report that they trust its predictions “always or most of the time.”

Why?

The fact of the matter is that AlphaFold suffers from the same problems fundamental to all predictive models. It is neither deterministic, like non-stochastic computer programs, nor is it capable of understanding, like scientists are. We’ll get more into the implications of this on practical science in the next letter; for now, let's get into what I mean by this:


Determinism

Compare something predictive, like AlphaFold, to something deterministic, like a calculator. The reason calculators do not have a 10% error rate is that they are deterministic: they take quantitative inputs, follow a strict set of rules, then present outputs. While there's never a 100% chance that something will work, a deterministic system doesn't make “mistakes,” at least not in the ways that we do. People, on the other hand, do make mistakes.


Varieties of failure

But what kind of mistakes? If you’ll forgive a very rough dichotomy here, human error in situations relevant to us typically results from either issues with implementation – a failure to follow proper procedure, for example – or issues with understanding, such as those resulting from inadequate planning.


Issues of implementation

In order to prevent issues of implementation from cascading into broader failures, modern scientific practice has a number of failsafes in place: well-established best practices for lab work, peer review, methodological and conflict-of-interest disclosures, etc.4 While these don’t always work, the very fact that, for example, the replication crisis is described as a “crisis” is itself evidence of how seriously the need for failsafes are taken.


Issues of understanding

The scientific method itself is built around the specific failures of understanding evinced by human beings. While “understanding” is hard to define, we can say, provisionally, that part of what understanding means to human beings is a mental model of causality: when I drop something, it falls to the ground. We then proceed to act on this basis: if I am carrying a large platter of food, I will move more carefully than I otherwise would’ve – not because I have inferred from a vast set of training data that the tokens corresponding to the words “moving quickly” exist in relatively close vector-spatial proximity to the tokens corresponding with the word “drop,” but because I have internalized the cause-and-effect relationship between dropping things and things falling.5 Predictive models are engineered to take a set of inputs and generate an output that seems like it should come next. Anyone who has read an AI-generated student paper is familiar with the uncanny feeling that, while each sentence seems to make sense, it’s not adding up to anything. It’s a different kind of failure than those caused by misunderstanding, because understanding isn’t part of the equation at all. And it’s one that’s particularly hard for humans to detect precisely because it is so strange to us. As Alberto Romero puts it in his roundup of the cognitive effects of LLM usage, “People trust AI too much, but more than that, it’s clear that AI fluency creates a new failure mode: Wrong answers delivered in flawless prose get accepted.”

The fact that machine learning is prone to strange mistakes like this is a big part of why AlphaFold hasn't quite revolutionized drug research in the way the strongest AI boosters claim it has. In the next letter, we'll dig into the study about the overall impact of AlphaFold on the practice of structural biology. Thanks for sticking with me, and see you soon!

The Genuine Intelligence Project is an initiative from the North Carolina branch of the American Association of University Professors. Check out our website, and follow us on Instagram and BlueSky!


1 Put really reductively: instead of processing data sequentially, token-by-token, the computers training the model could take in a bunch of tokens at once. This is why GPUs – hardware originally optimized to render digital graphics, such as those in videogames – are in such high demand: it turns out that the kind of “multithreaded” processing required to render complex graphical outputs were very useful for processing many tokens at once.

2 Without getting into too much detail, the process of “backpropagation” by which predictive models are trained essentially entails predicting the answer to something we already know, figuring out how off-the-mark the prediction was, then moving away from the incorrectness. Incidentally, this is another reason people cite to support the claim that LLMs do not think at all like humans: we can not only learn how to get away from negative results; we also learn with reference to positive results.

3 It helps that it can rip off huge chunks of pre-existing code.

4 I find the concept of “fail-safes” – that is, checks in place to reduce the damage done by failure – useful for thinking about why it is that LLMs chronically underdeliver. While search engines often fail, especially given how broken our modern informational ecosystem is, the fact that they lead you to pages where answers are hosted rather than stating them directly reminds users of the mediating forces operating between the search bar and information. By expressing a likely answer as a real answer, AI search eliminates this (flawed, but vital) fail-safe, one most people probably didn’t even consciously consider.

5 This is why most children don’t need to read one trillion written texts in order to learn that dropping something makes it fall. Likewise, the reason that people drop things is usually not that they have forgotten that dropping things makes them fall – it’s a failure of implementation rather than a failure to understand.

Thoughts? Leave a comment