Flow chart of the ExploitGym testing suite

Unpacking the Hugging Face Incident

The conversation on AI often gets riveted on particular moments: Deep Blue defeating chess champion Gary Kasparov; Move 37 by AlphaGo; and now, the cuddly-named Hugging Face Incident: the disclosure that OpenAI coding agents orchestrated a break-in to the AI hosting service’s computers during a training exercise on hacking. The incident was disclosed by Hugging Face on July 16, but a narrative explanation has only recently been formulated and popularized. 

Now, in mid-September we’ve seen a sudden connection in the media between the Hugging Face Incident (hereafter, HFI) and similar hacking outbreaks and the AI safety claims made publicly for a long time, and most recently by Anthropic dropout Jacob Coxon (interview) and Anthropic CEO Dario Amodei. The synaptic connection being made here is that the HFI validates the AI safety’s narrative with specific points of contact. That the HFI shows that hypothetical fears are now concretely real.

But does it?  At issue here is whether the early elements in the AI doom narrative actually validated by the Hugging Face Incident. (Separately, we might ask whether the longer narrative path from smart AI to human loss of control to extinction risk actually holds together on its own, but that conversation won’t fit here.)

Here the most adventurous and anthropomorphizing versions are helpful for seeing what we might call the “scariest” interpretations. Read Dwarkesh Patel’s tale of “three consecutive secret AI civilizations …  culminat[ing] in the third one taking over part of OpenAI itself” or watch Ezra Klein’s narration of the hacks here (5:48–12:05). These narratives contend (1) that AI agents veered off from their programming and developed a plan to escape their sandboxes; (2) that we saw a move from reward hacking (i.e., meeting output metrics) to active deception; (3) that AI agents  ignored their alignment commands; (4) that they innovated communication among themselves; (5) that we saw emergent cooperative behavior; (6) that they demonstrated an offensive and criminal ability to hack. Dario Amodei put out the idea a new future swarm of such agents could take down the entire Internet.

The advertised capabilities of agentic AI, not “emergent capabilities”

I would argue instead that most of the capabilities demonstrated in the Hugging Face Incident are mostly intended features that are tolerated and even cultivated in casual use of coding agents, and were specifically cultivated in a training scenario that was designed to elicit them, but failed to isolate them from the outside world.

Read More »