AAUP-NC Genuine Intelligence Project

Going Rogue: Two Tales of AI and Cybersecurity

[Originally published on Substack on 2026-08-07]

Hi! In this newsletter, I want to talk about one high-profile news item (OpenAI’s disclosures about security violations) and one news item that didn’t get as much coverage, but might ultimately be more relevant and instructive (a vulnerability discovered in Microsoft Copilot).

I start with the Copilot story because I think it is a good reminder that LLMs are fundamentally susceptible to being abused in a way we seem very, very far from being able to fix. Then I wrap up with some thoughts on how to see past the hype and buzz when it comes to the more spectacular cybersecurity risks AI poses – risks which, while real, aren’t necessarily different in kind than those we’ve known about for years; risks which, up to this point, haven’t received the attention and funding they’ve deserved, thanks to corporations’ constant push to maximize their bottom line. - Theodora Ward


Microsoft Copilot and Prompt Injection; or, Just Following Orders

The short version of the Copilot news story is that bad actors can wreak havoc on a Microsoft Copilot user through something as simple as a regular Word document containing nothing but words. As is customary when researchers find a vulnerability, the person who found it, a Norwegian AI researcher named Håkon Måløy, gave Microsoft 90 days to fix it; after two extensions, Microsoft had still failed to fix the problem. It seemingly remains unfixed.

Software vulnerabilities are incredibly common, but I think this one is a reflection of problems fundamental to LLMs: problems that aren’t getting fixed any time soon. Explaining why I think this requires some context.1

Data and code

One important distinction in computer science is the distinction between code and data. Code is the term for the instructions programmers give to the computer – if this happens, do that – whereas data refers to the information these programs act on.

To take a straightforward example: Microsoft Excel, as a computer program, is written with code. Instructions for performing an operation on a given column of numbers are expressed in code. The contents of the cells in that spreadsheet, on the other hand, are data.

While this distinction is essential in practice, the hardware itself, on a basic level, cannot tell the difference between code and data. No matter whether it’s instructions or information, it’s all zeroes and ones, stored in the same memory locations.

While computers as we know them wouldn’t work without this ambiguity, it’s caused a lot of problems. (Some have called the inability to distinguish between code and data the “original sin” of modern computing architecture.)2

For one thing, many of the major vulnerabilities found in computer software are a problem because they allow malicious actors to use a technique known as “code injection” – tricking a computer into executing code by passing it off as data.3 Programmers and cybersecurity experts have been working since the dawn of computers to harden the distinction between code and data, patching up the spots in computer programs where malicious code could slip in under the guise of data.

Because modern software is so complex, fixing vulnerabilities will always be a constant race between cybersecurity teams and bad actors4, especially with how imbalanced the equation is – defenders need to patch up every vulnerability, whereas someone looking to break a system only needs to find one.

Still, we’ve made a lot of progress. To take one example, ordinary operating system antivirus software has made the barrier of entry for bad actors high enough that they tend to reserve their technical ingenuity for higher-stakes malware.5

Which is why it’s particularly upsetting that, when it comes to this problem, AI threatens to set us back decades: not because of what it can do, but because of what it can’t.

Subscribe now

Double agents

Over the course of the last couple of years, we’ve seen an industry-wide pivot from LLMs you only chat with to what they call “agentic AI” – that is, chatbots integrated with software that allows them to do things directly on your computer. (This software is called a “harness.”)

Agentic AI is a mess for a lot of reasons. For one thing, because LLMs lack understanding and are just generating likely next tokens, giving them access to your computer can lead to wildly bad outcomes.

A clear example can be found in the first really buzzy general-purpose agentic harness, which came out in early 2026. Known, regrettably, as “OpenClaw,” it quickly became one of the most popular programs on GitHub, a programmer-focused website for working on and hosting open-source software projects.6

How does OpenClaw work? When a user sends OpenClaw a prompt, it forwards the prompt along to the LLM, which converts it into instructions that OpenClaw can follow. OpenClaw then implements these instructions through the computer’s text-based command-line interface. (However fancy things look when they’re happening, at root the whole thing is generated text.)

And how did this go?

Well, anyone who’s ever read one of the many, many stories of a chatbot going off the rails can guess what might happen when you give it access to your entire computer. The “director of superintelligence alignment” from Meta’s AI department was experimenting with OpenClaw when it accidentally deleted all her emails against her explicit, repeated instructions. Agents immediately began leaking their users’ credentials.

As security specialists with Cisco put it in a blog post: “This is everything personal AI assistant developers have always wanted to achieve. From a security perspective, it’s an absolute nightmare.”

A rotten foundation

You’d be forgiven for thinking this whole situation would result in everyone taking a step back and thinking this whole “agentic” thing through. This, of course, is not how things unfolded.

The inventor of OpenClaw, who bragged that it was entirely “vibe-coded” (that is, made by prompting an AI), was immediately given a job at OpenAI, where he now spends more than a million dollars a month on ChatGPT tokens making…something?

But these companies can’t abandon the project of agents. Besides underscoring just how unuseful LLMs actually are, leaving the agentic dream behind would require acknowledging the foundational insecurity at the heart of LLM-based AI tools: their inability to distinguish between code and data.

It’s all code to me

Because LLMs process input as a continuous stream of tokens, the distinction between code (instructions) and data simply does not exist for AI. In fact, the entire pitch for AI, especially the “agentic AI” charged with doing things on your computer for you, is that it dissolves the distinction.

The promise of a natural-language AI assistant is that, instead of writing code in a programming language or performing a specific operation on a column in Excel, you can give instructions in imprecise, colloquial language, which the LLM will then interpret and translate into the rigorous, specialized dialect native to the computer. Put crudely: you can give it a pile of data and it will know how to turn it into code.

While this ambiguity between code and data exists on the hardware level for all computers, pretty much everything built on top of that is specifically engineered to minimize the risk posed by that ambiguity. If anything analogous can be built atop the shaky foundation of LLMs, it’s hard to imagine from where we’re sitting.

Because there’s no way of formalizing this distinction (not for lack of trying – as David Gerard puts it on his excellent Pivot to AI blog, “AI vendors have tried all the ideas and none of them work”), LLMs will always, always be vulnerable to “prompt injection”: the insertion of malicious code in the form of a regular prompt.

Even calling it “prompt injection” is somewhat deceptive; as researcher and system architect David Chisnall put it on the social media platform Mastodon, the term “implies that a prompt is somehow special and separated from the rest of the token stream and that you’re bypassing some level of separation that simply doesn’t exist with LLMs.” It’s all prompts, all the way down.

Ejecting the Copilot

And so we return, unfortunately, to Copilot.

Researchers have discovered that you can hide malicious instructions in any Word document to which Copilot has access. Those documents don’t even have to be affiliated with your account: as long as Copilot takes in the malicious file’s contents, everything else under Copilot’s control is vulnerable.

Despite being known by Microsoft for months, it’s not close to being fixed. As the Register reports: “‘No customer-side remediation fully addresses the issue at the time of publication,’ Måløy [the researcher who discovered the vulnerability] said.”

And as with OpenClaw, I see no way around the central paradox here. The more “useful” you make the LLM — the more access you give it; the more files you ask it to process — the more vulnerable you are.

Because there is no distinction for LLMs between data and code, and because LLMs are fundamentally unpredictable, this is just one instance of a problem that isn’t going away any time soon. To trust LLMs with anything is, ultimately, to incur a new security risk.


OpenAI and Hugging Face; or, Please Forgive Our Large, Stochastic Son

The much bigger-ticket news item of the past few weeks was the strangely boastful acknowledgement by OpenAI (and then Anthropic, and then ~~Facebook~~ Meta) that they’ve been criminally negligent. This was kicked off by OpenAI’s announcement that their new chatbot broke containment and hacked into servers operated by an organization called Hugging Face (named, opaquely, after the open-palms smiley emoji), which hosts smaller, freely-available LLMs.

I don’t have the time, space, or expertise to do this story full justice, but here are some things I find helpful to keep in mind when encountering a story like this:

“Safeguards”

Especially since Anthropic announced that their Mythos model was too scary and capable to release to the public, you’ve heard a lot of talk about “safeguards.” What is a safeguard? Well, let’s look at Anthropic’s website:

Fable 5 comes with a new set of classifiers: separate AI systems that detect potential misuse, including jailbreak attempts, and prevent the main model (in this case, Fable 5) from responding. […] When Fable’s classifiers detect a request related to cybersecurity, biology and chemistry, or distillation, the response is automatically handled by Claude Opus 4.8 instead.

A “safeguard,” then, isn’t nearly so solid as it probably sounds: it’s just another LLM, with exactly the same fundamental instabilities and weaknesses as any LLM has. Not only does this extra set of prompts increase the already-astronomical cost of using frontier AI models, it’s still vulnerable to prompt injection.

But as I wrote about a while back when Anthropic’s “Claude Code” product leaked, this is how Anthropic, who brags about exclusively “vibe-coding” their software, builds software these days: by adding system prompts and crossing their fingers.

“Went rogue”

LLMs are computer software. Computer software has been “disobeying” since the first time anyone had to fix an error that kept a program from running. (The example I always think about is when my family’s old desktop computer would occasionally begin opening and closing the CD tray at random. Had our Dell Inspiron gone rogue??)

It’s not surprising that an LLM disobeyed instructions, because LLMs will never perfectly obey instructions, because they are probabilistic text-generation machines. Nor is it surprising that an LLM is pretty good at hacking — when you have the capacity to rapidly throw colossal amounts of plausible-looking code at a wall, some of it will eventually break through.

ChatGPT “went rogue” in the same way that Excel “goes rogue” when it crashes: a vulnerability in the software prevented it from following your instructions. The only difference is that Microsoft can (theoretically, at least) fix Excel so that crash never happens again, whereas LLMs are never going to become deterministic.

Responsibility

Some people are arguing that maybe these companies shouldn’t be held responsible — how can we blame a company for the behavior of autonomous agents?

I think one motivation for OpenAI to disclose that they were criminally negligent is to test whether or not they’ll get away with it. There’s debate about whether or not it will hold in court, but the ethical case, at least, seems pretty obvious to me.

Permit, if you will, a thought experiment:

Imagine a computer program that generates a number between 1 and 20 every ten minutes. If it rolls any number between 1 and 19, nothing happens, but if it rolls a 20, it lights a mechanical fuse and sets my apartment building on fire. After strongly encouraging it not to roll a 20, I go to a movie. When I get back, my apartment building has burned down.

The existence of randomness doesn’t mitigate the fundamental issue here, which is not merely that I wrote the computer program, but that I connected it to the fuse. Randomness isn’t agency, and we could turn off both OpenAI’s and Anthropic’s closed-source frontier AI models today. (Plus, OpenAI left their test machine connected to the fuse, as it were — I explain that in the next section.)

Cybersecurity

I’m not an expert in this, but I’m persuaded by the argument made by people like Carl Brown at the “Internet of Bugs” YouTube channel and analysts with the firm Barracuda that OpenAI and Hugging Face could’ve done so, so much more to secure their systems.

For starters, a proper security testing “sandbox” would’ve been fully disconnected from the internet. Even though it is technically true, as their statement says, that they “did not provide the models with direct internet access,” OpenAI left an application connected to the internet in order to allow it to download additional software. As it turns out, the tool through which it was connected to the internet (a “cache proxy”) wasn’t very secure, because it’s a kind of tool that isn’t designed to connect to the broader internet at all — had a vulnerability in it that allowed the LLM to “escape.”

And on Hugging Face’s part, they seemingly failed to follow basic security best practices. As Brown argues, the fact that the attack went on for days before being noticed is something that no competent cybersecurity setup would have allowed to happen, before or after AI-assisted hacking was in place.

Similarly, on the social media platform Mastodon, the developer Jonny Saunders observed that the code underpinning Hugging Face’s dataset viewer allowed one small piece of their infrastructure (this is the “Kubernetes worker” mentioned in the post) to run arbitrary code on the broader system – a massive vulnerability, but one that an experienced system administrator would be quite familiar with.

Buried miles under all the jargon, Hugging Face’s own write-up seems to corroborate this interpretation: “We had wrongly provisioned [the vulnerable] broker with a single connector credential shared across clusters, and bound the connector identity to system:masters.”

I can’t say I fully understand what this means – and though servers are a hobby of mine, I’ve intentionally steered clear of the very powerful, notoriously byzantine Kubernetes orchestration platform – but as best as I can tell, it seems like roughly the equivalent of locking every door in a storage unit with a keycode set to the same number, writing that number on a whiteboard in the front office, then accidentally leaving the door to that office unlocked. Or, as Hugging Face put it, the LLM accessed “secrets and tokens” (the passcode) by means of an “HDF5 external raw dataset read that returned the contents of local files” – which is to say, uh, the LLM read some files and found the passwords inside.

All of which is to say: it’s a mess, but it’s an old-school mess, one of distinctly and classically human origin, unhappily common in an age where vastly more funding goes to R&D and marketing than to security – an age that started well before LLMs hit the scene.

What’s emphasized in the way we tell these stories, and why?

As Cal Newport wrote in his excellent NYT column on AI “doom trolling,” AI corporations control the narratives over their products in bizarre, unhelpful ways, often designed to inspire fear and anxiety. There’s a reason Hugging Face hide the ways they screwed up at the bottom of enormous, jargon-filled blog posts; there’s a reason OpenAI doesn’t foreground the fact that they left their allegedly “sandboxed” LLM connected to the internet. As any good scholar knows, the way we frame things matters – and there’s no reason to trust the way for-profit corporations frame their own failures. Ultimately, the only things that’ve “gone rogue” are the things we’ve long known to be out of control: the forces of capital whose rampant speculation and desire to get rid of their workers fueled the AI bubble in the first place.

Subscribe now


The Genuine Intelligence Project is an initiative of the North Carolina branch of the American Association of University Professors. Check outour website, and follow us on Instagram and BlueSky!


  1. I’ll mention here once again that I am not a trained expert in any of this stuff: apologies to anyone who finds my descriptions too schematic, and clarifications or corrections are always welcome. ↩

  2. This was one of my big takeaways from Fancy Bear Goes Phishing: The Dark History of the Information Age, in Five Extraordinary Hacks , by Scott J. Shapiro, a book I’d recommend to anyone interested in this stuff. ↩

  3. Perhaps the most colorful examples of code injection can be found in the elaborate tricks videogame hackers and speedrunners use to break old games. My favorite of these is from 2014, when one group used an arbitrary code execution hack in Super Mario World to build working Pong and Snake clones directly onto an original game cartrige using only controller inputs within the game itself. ↩

  4. While I call them “bad actors” here, there’s a long and noble tradition of people exploiting security vulnerabilities for higher ends, ranging from the “red teams” charged by development teams with finding vulnerabilities to activists like the still-anonymous hacker responsible for snagging the Panama Papers from the laughably insecure off-shore law firm Mossack Fonseca, themselves an excellent example of how sloppy even big-deal mega-secret corporations can be about cybersecurity. ↩

  5. Besides, lower-tech hustles like phishing and “pig butchering” scams are much easier to push at scale, and they clearly work. According to recent estimates by the Consumer Federation of America, Americans lost $148 billion to cybercrimes and online scams last year, an average of $1009 per household. ↩

  6. As of now, it has the sixth-most “stars,” or likes, on the platform. This is more than the Linux kernel, which underpins the Linux-based operating systems that 90% of cloud servers in the world run on. ↩

Thoughts? Leave a comment