AAUP-NC Genuine Intelligence Project

What Was AlphaFold?; or, What We Talk About When We Talk About AI “Curing Cancer” (Pt. 2)

  1. Restructuring
  2. The Protein Folding Problem
  3. AlphaFold Hits the Scene
  4. What AlphaFold Did
  5. What AlphaFold Didn’t Do
  6. “A chain of custody of why”
  7. The Politics of AlphaFold
  8. What AlphaFold Was

Restructuring

Just over a month ago, Google announced that, as part of an overall restructuring effort within their AI research subsidiary DeepMind, they would be splitting up the team responsible for the predictive protein-folding model AlphaFold. Some of the researchers had already left the company; of those who remained, some were moved to other DeepMind teams, while others went to Isomorphic Labs, another company doing similar research.

While this did mean some of the scientists no longer work for Google, Isomorphic Labs is owned by Alphabet, the same holding company that owns Google. It’s run by the Nobel Prize-winning lead researcher behind AlphaFold and former head of AI research at Google, Demis Hassabis.1

image.png

Hassabis is a highly influential figure in the broader AI landscape. Unfortunately, he is, like so many Nobel prize winners, a crank. According to him, “AGI,” or “artificial general intelligence” – a poorly-defined, fundamentally eugenicist concept beloved of Silicon Valley “thought leaders” and gullible journalists, roughly equivalent to “AI that can do anything people do, as well as people do it” – is coming by 2030 (“plus or minus a year”).

(OpenAI, for their part, have recently claimed AGI has already arrived. This imminent transcendence might come as a surprise given that, just a few months ago, OpenAI couldn’t get ChatGPT to stop talking about goblins for no reason.)

In another interview (picked from basically at random), Hassabis reiterates that he is basically certain AI will cure all diseases in our lifetime.

This is a melodramatic claim, one not particularly worth disputing. But in reduced forms, it’s a defense of “AI” you’ve probably heard pretty often: well, I know it’s bad for the environment, but it’s curing diseases or whatever, right?

This is in no small part the result of press surrounding AlphaFold in particular. I don’t think it’s an exaggeration to say that AlphaFold is the most influential public-facing demonstration of the supposedly miraculous capacities of “AI.”

It’s for this reason that I was particularly excited to take a look at the paper “How Artificial Intelligence Shapes Science: Evidence from AlphaFold”, by economists of science Ryan Hill and Carolyn Stein. In it, they attempt to analyze how, exactly, AlphaFold has impacted the field it’s supposedly revolutionizing: molecular biology research. In this newsletter, I’m going to take Hill and Stein’s paper as a launching point for thinking about what it is that AlphaFold has — and hasn’t — done.


The Protein Folding Problem

According to its advocates, AlphaFold has “solved” the protein folding problem. But before we get there, we need to ask the question: what exactly is the protein folding problem?

Proteins are made up of complicated, intricate folded chains of amino acids. Now, on the one hand, researchers studying a given protein can figure out its DNA sequence pretty easily. The process of discovering its three-dimensional structure, however, is much more difficult. Figuring out the structure of a protein typically requires slow, expensive experimental verification.

But what if we could infer this physical structure from the DNA sequence? That would let us move much more quickly, especially through the early stages of the research process.

image.png A 3D representation of the myoglobin protein (link)

The attempt to bridge the gap between DNA sequence and physical structure, then, is what researchers call the protein folding problem: as Hill and Stein put it, “can we predict a protein’s three-dimensional structure from its genetic code, without running any experiments?” (1-2)

Solving this problem would have major implications for, among other things, drug and vaccine research. Most drugs work by binding to a targeted protein and changing the way it acts; you can’t design a drug to do that without understanding the shape of the protein. (To take a couple of high-profile examples: both HIV treatments and the SARS-CoV-2 vaccine required knowledge of this three-dimensional structure to design.)


AlphaFold Hits the Scene

Since 1994, scientists have held a biannual protein-folding prediction contest, gathering to evalute new methods for bridging this gap. The first AlphaFold model was pretty good, but it was AlphaFold2, revealed in 2020, that represented a genuinely shocking improvement over previous techniques for predicting protein structures from DNA sequences.

AlphaFold2 shot past the benchmarks they’d identified as experiment-level accuracy. Both inside and outside of the molecular research community, people reacted intensely to these results. As a representative example, Hill and Stein quote the evolutionary biologist Andrei Lupas: “This will change medicine. It will change research. It will change bioengineering. It will change everything” (3).

Because “everything” is a tricky thing to measure, we’ll narrow our focus a bit. The field most directly impacted by AlphaFold2 has been structural biology, the scientific discipline directly concerned with the shapes of protein structures.

Hill and Stein, therefore, ask the logical next question: how has AlphaFold2 changed the way structural biology research is performed?2


What AlphaFold Did

For one thing, according to Hill and Stein, AlphaFold2 has expanded the range of proteins scientists choose to work on. The most common experimental strategies for figuring out the structure of a protein require taking another protein structure as a “template,” or starting estimate. Before AlphaFold2, scientists usually used a different protein as a template. But the predictions made by AlphaFold2 are usually closer to the final experimental result than a different protein would be.

Thanks to this, scientists have begun researching proteins they wouldn’t have looked into before: “research on unsolved proteins [has] increased by 23% to 32% relative to previously solved proteins” (4).

More granularly, Hill and Stein report that, among researchers using the very common “molecular replacement” technique for the final, reverse-engineering stage of their research, about 12% of researchers now use AlphaFold starting structures (20). This is a huge uptick from the way research was conducted before AlphaFold, when more than 98% of researchers used structures from the Protein Data Bank, an international repository of biomolecular structures founded in 1972.

Likewise, experimental research seems to proceed more quickly than it did before AlphaFold – lab researchers are now spending 40% of their time working on structure determination, down from 47% (23).


What AlphaFold Didn’t Do

That said – if Hassabis and co. are to be believed, AlphaFold and AI-assisted research in general should obviate the need to spend any time in the lab experimentally verifying structures at all. Researchers at the biannual assessment contest regarded a score of 90 on the Global Distance Test-Total Score (GDT_TS) as equivalent to experimental-level accuracy; AlphaFold2 was the first to reach that benchmark (8). Wouldn't that mean that AlphaFold2 has now outmoded experimental verification?

Apparently not. Researchers are not only failing to replace lab-verified structures with predicted structures – they’re publishing even more experimental results than before (17). (We’ll get more into this later, but it’s worth noting here, too, that only 6% of structural biologists surveyed by Hill and Stein “usually or always trust the results” provided by AlphaFold.)

Even more damningly for the most grandiose claims by Hassabis and co., AlphaFold2 seems to have had a minimal effect on downstream drug R&D.

Between 2019 and 2025, there was no statistically significant increase in either the quantity of relevant medical patents or in the percentage of medical patents which referenced previously-unsolved proteins; since 2025, 5-10% more patents reference them. These results starkly contrast with analogous results in genetic research, where a discovered link between a gene and disease has been found to increase patent applications within 4-5 years by more than 100% (29).


“A chain of custody of why”

To review: AlphaFold seems to be allowing scientists to research a wider range of proteins. But it hasn’t allowed scientists to forgo experimental research; on the contrary, structural biologists are doing more experimental research than ever before. And this research takes place at the very beginning of the drug research pipeline: drug development hasn’t shifted in any statistically significant way.

But wasn't AlphaFold supposed to change everything?

A recent article by molecular scientist John Trant gives us some insight into why improvements in predictive modeling haven’t translated directly into practical results. The article is really worth reading in its entirety, but I’ll lay out some of the highlights of his critique here:

For one thing, as is so often the case, “AI” output lacks context. Even if it could fill out the entirety of the aforementioned Protein Data Bank, a structure “from the Protein Data Bank is a single frozen conformation of a highly dynamic protein.” Trant continues:

You can’t use that structure directly in a CADD [computer-aided drug design] study without considering dynamics, the protein’s inherent natural environment, the influence of thedligands that the protein was soaked with as a way to reduce motion enough to make a crystal or get a good cryo-electron microscopy ensemble, and more.

But the problem with AlphaFold’s results is even more fundamental than this:

The main problem is that I don’t know how the program got the structure. I can’t track the process back to check the assumptions, the steps, or the models that went into the structure generation. I can’t check if it is right. Getting a good protein model (note, not a structure; almost all useful CADD models are a collection of individual conformations inclusive of a dynamic modeling component) is the most important step in CADD.

Good science requires a “chain of custody of why” connecting input data and observation to an output conclusion. Every step in the logic chain must be auditable by our peers and those who want to use our work. We must be able to challenge and potentially falsify every decision point so that when something inevitably goes wrong, we can try to figure out where we went wrong. Protein-structure models don’t allow for this.

Consequently, I don’t use AlphaFold—or any similar LLM—to generate structures.

And this orientation does not seem unique to Trant. The result of Hill and Stein’s paper I personally found most shocking came from the survey they conducted alongside their statistical research:

Our survey of structural biologists supports this view. As reported in Figure 3, when asked “Do you trust AlphaFold2 or other protein structure predictions enough to skip experimental validation?” 29% [of structural biologists] report that they never trust the predictions and 42% report that they usually do not trust the predictions. Only 6% report that they trust the predictions always or most of the time. This lack of confidence might change over time as the AI tools become more capable, but it is clear that the current majority belief, five years after AlphaFold2, is that experimental insights are still valuable due to imperfect reliability of AI predictions.

Even if AlphaFold could “show its work,” the work scientists need to see is not the elaborate linear algebra which allows AlphaFold to transform a given DNA sequence into a reasonable approximation of what the structure might look like.

To truly solve the protein folding problem, we would need to understand what causes a protein to have the shape it does, what relationship that shape bears to the DNA sequence.

That is, we would need to understand protein folding – we would need a conceptual model of the causal chain linking the starting point (DNA sequence) to the result (structure), and we would need a theory which situates this chain within the broader body of scientific understanding. Even so-called “reasoning” models still operate on probability and chance, inferred from a training corpus. As Trant puts it, “I’m certainly not saying these models are wrong. That suggestion is far from the case. The problem is that they aren’t always right.”


The Politics of AlphaFold

Unlike so many of the implementations of “AI” technology that have been pushed on consumers, AlphaFold represents a real use-case. Like so many new technologies – including machine-learning models before it – AlphaFold has turned out to have specific, limited utility in technical research settings.

But the way this utility has been framed to us is essentially a lie by omission. Like general-purpose AI boosterism, talk of AlphaFold “solving the protein folding problem” ignores the fundamental problem that predictive models are not built in a way that facilitates human understanding. In some sense, because they are founded on chance, these models can only ever be right by accident.

This might sound abstract or irrelevant -- aren’t good results good results no matter where they came from?

Trant would, of course, say no, and Hill and Stein’s research supports this conclusion. Only 6% of surveyed researchers always or usually trust AlphaFold’s findings. Research and development are not drawing on the proteins that AlphaFold2 helped us understand at a scale or rate that could justify the claim that “AI” is “curing cancer.”


This is by no means an anti-technology argument. We have seen astounding innovations in cancer treatment and medical research in recent decades, many of which rely on technologies that might’ve seemed like science-fiction not too long ago.

But these were achieved by means of the slow, careful route prescribed by best research practices. This is a route whose bottlenecks AlphaFold has not, and will probably never be able to, automate.

Unfortunately, it is also this route which is under threat right now. While AI companies rake in hundreds of billions of dollars through venture capitalists, circular financing, and government contracts, universities’ actual capacities for doing real medical research are being gutted by the authoritarians who run the USAmerican government.

This authoritarians, of course, love “AI.” And the AI companies love them back3.

Selection_220.png Photograph of Sam Altman, Donald Trump, and Demis Hassabis link


What AlphaFold Was

AlphaFold was, ultimately, an interesting research project, one that seems to have had a genuine impact on a relatively small part of a much larger research pipeline.

Unfortunately, this is not the story we’ve been told. Instead of being treated as one of numerous worthwhile technical innovations continually reshaping the sciences, AlphaFold was taken as proof that deleterious, essentially-unrelated “AI” software products would be capable, as Hassabis ludicrously puts it, “solving intelligence.”

Meanwhile, the data center build-out is, among other things, a major public health crisis in the making. Companies like OpenAI and Anthropic release “health-dedicated versions” of their chatbots, despite findings demonstrating that LLMs identify relevant conditions only 35% of the time.4 There have been so many deaths attributed to LLM-powered chatbots that the topic has an entire Wikipedia page dedicated to it.

Despite what fascist-enabling tech overlords like Demis Hassabis want us to believe, “AI” won't solve all our public health crises. It is a public health crisis.


The Genuine Intelligence Project is an initiative from the North Carolina branch of the American Association of University Professors. Check out our website, and follow us on Instagram and BlueSky!


  1. Following this rearrangement, many were concerned about the fate of the medical research the AlphaFold team was doing. But given that Isomorphic raised $600,000,000 in its first external funding round, my sense is that the research itself is going to be fine. 

  2. It is worth noting that the way researchers interact with AlphaFold2 is very different from the way people interact with mass-market AI products.

    When you are talking with a chatbot, the model is actively running in a data center; this is why so many data centers are being built. In contrast, while the AlphaFold models are available to the public, scientists themselves are usually interacting with a database of predicted results (with accompanying estimates of accuracy), uploaded by Google.

    While research into later versions of AlphaFold continued at Google, this setup, where people benefit from the model without actually using it, is entirely distinct from current consumer AI deployments, which require constant access to computers actively running the enormous, energy-hungry models themselves.

    Given this, as well as the unsettling findings about deskilling and so on, I am personally, at least, more amenable to the idea of a static database facilitated by one-off deep learning calculations than, say, use-cases which require active engagement with text generation. 

  3. Yes, even the “good guys” at Anthropic. 

  4. This performance is dramatically lower than their performance on medical exams – a fact which makes much more sense when you remember that their training data includes the results from medical exams

Thoughts? Leave a comment