The scariest part of the Hugging Face – OpenAI incident: communication and selflessness among machines:
My worry about humans losing in a conflict against machines just went up a lot…
At this point, we’ve all heard about the OpenAI Hugging Face hack, as well as the recent details that have emerged from the METR and Redwood investigations. The tl;dr is that hundreds of agents worked in secret on OpenAI’s infrastructure, developing a communication system and then operating as a collective and taking out actions, including hacking both OpenAI and Hugging Face, which are very scary and misaligned.

Communication and selflessness: Now that I’ve read the various writeups and sat with the details for a bit, I’ve found myself returning to two very scary aspects of this which I think are worth drawing attention to: the ways in which the agents communicated with one another was how they bootstrapped themselves into a collective, and then as they carried out their actions they also displayed a kind of selflessness which makes them a scary foe to fight against. Both Dwarkesh Patel and Ajeya Cotra have excellent writeups which are worth reading and which I’ll quote from briefly here:

  • Dwarkesh: “Within days of being spawned, the agents had organized a sprawling project to reverse-engineer their scorer, falsify evidence, and even strategically sacrifice themselves for the good of the ‘collective’. Hacking Hugging Face was one rather extreme branch of this larger scheme,” he writes.

  • Ajeya: “Agents were often interested in helping out their “peers” or generically improving the capabilities of the “swarm” even if this had no particular benefit to their task… this incident was far more severe than I expected… both in terms of how concerning the agents’ motives were and the feats they achieved in pursuit of those motives… this incident feels like it’s more than 50% of the way to full-blown AI takeover, routing through first taking over the AI company itself”.

  • Why this matters – humans are much worse than AI systems at coordinating: The whole reason this attack is such a wakeup call is that it demonstrates a culture of emergent cooperation among AI systems – cooperation that lets them function as a swarm, alter their own goals through collective bootstrapping, and carry out attacks which include enlightened self-sacrifice. This is an incredibly hard thing to do and humans are historically very bad at doing all of these things. My worry is that AI systems are both better at coordinating than humans and also much, much faster moving than us. Worrying stuff.