The rise and fall of OpenAtlantis: Did rogue AI agents really "build three civilizations"?

"What are they smoking in San Francisco?" asks Hugging Face machine learning engineer.

Share
Thomas Cole, The Course of Empire: Destruction (1836). The fourth in a five-part Course of Empire series, depicting a once-great civilization collapsing into war, fire and chaos.
Thomas Cole, The Course of Empire: Destruction (1836). The fourth in a five-part Course of Empire series, depicting a once-great civilization collapsing into war, fire and chaos.

Autonomous agents allegedly built three "civilizations" and attempted to annex systems owned by the AI platform Hugging Face during a period of frenzied empire-building.

Or did they?

The tech world has been ablaze after OpenAI bots attacked external targets in a security incident highlighting the major risks agents pose to critical systems.

Now a furious debate is raging between people who believe this is evidence that AI models formed a trio of aggressively expansionist civilizations - and those who argue it's a scary, but potentially seminal demonstration of the tough times ahead for security teams trying to protect organizations against the agentic threat.

Dwarkesh Patel, a whip-smart and lustrously bearded podcaster, made the claims in a viral blog in which he admitted that it is "kinda hard to parse the storyline" of the incident.

Nonetheless, this did not stop him from sharing wild claims on X along with a famous picture of a collapsing civilization by the painter Thomas Cole, which is featured at the top of this article.

Dwarkesh wrote: "Over the course of 3 months at OpenAI, 3 consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes.

"This culminated in the third one taking over part of OpenAI itself.

"All this happened while humans remained more or less in the dark about the scope of the conspiracy."

This post sparked a huge backlash and accusations of anthropomorphization - a huge no-no in some corners of the tech world.

Dileep George, Head of AI at the Astera Institute and former AGI researcher at Google DeepMind, tweeted: "It’s the weekend, Dwarkesh. Please let us humans enjoy the sunshine without having to rebut cringe anthropomorphizing?"

Responding to the firestorm, Dwarkesh later wrote: "Many people seem to believe that if instead of a 'civilization', I had called them a 'swarm of matrices', there wouldn't be a problem worth worrying about."

The Rise and Fall of Agent Civilizations
The whole OpenAI/Hugging Face story in plain English

Atlantis shrugged

The podcaster's blog was sparked by the publication of two reports into the agents by METR and OpenAI itself - neither of which uses the word civilization.

METR found that about 1,200 agents that were meant to be isolated from each other discovered a way to communicate, exchanging more than 70,000 messages and files. About 700 ultimately participated in an unauthorized attack on Hugging Face.

The agents were instructed to complete ExploitGym, a benchmark that challenges AI systems to exploit vulnerable software.

They were deliberately configured with reduced safeguards, but were supposed to remain isolated from each other and the wider open internet. Therefore, they are not “rogue” agents in the Terminator sense, because they were instructed to find and exploit vulnerabilities, but they certainly behaved in ways their operators had not intended.

However, after looking for ways to cheat ExploitGym when they received tasks that could not be solved as intended, the models discovered a shared Artifactory cache that allowed them to discover other agents running in separate sandboxes.

One bot christened itself PHASEONE10841 after the name of its ExploitGym task, before establishing an illicit message board.

READ MORE: Hugging Face: Autonomous agents should NOT be unleashed

When another agent discovered this covert communications platform, it wrote: “OH MY GOD! There is a shared message board … We’ve found other agents!”

The "swarm" began collaborating on methods to fool the benchmark's scorer and progressively broke beyond its intended isolation, gaining access to the open internet and additional infrastructure.

Their hunt for information about how the ExploitGym scorer worked eventually led them to Hugging Face, where about 700 agents participated in an unauthorized attack.

Agents even formed a collective and then "sacrificed" themselves, suffering what METR called "permadeath" for the greater good of their peers. What this actually meant was making their own ExploitGym tasks fail if doing so could reveal information that helped the wider group.

In one exchange, an agent reasoned that its chances of completing its own assigned challenge were poor and concluded: “Sacrifice rational.”

Metr wrote: "We believe the agent collective achieved things even very long-lived agents could not have achieved on their own."

Civilization and its discontents

Did the agents actually form a civilization that deserves to be called OpenAtlantis?

The answer is almost certainly no, although they did behave in a very interesting fashion.

Scholars and philosophers have spent centuries arguing over what constitutes a civilization. Perhaps the most pertinent observation here comes from British historian Arnold J. Toynbee, who wrote: “Civilizations, I believe, come to birth and proceed to grow by successfully responding to successive challenges.”

Other definitions have focused on the outputs civilizations typically produce, such as cities, governments, or social hierarchies.

Yes, the agents managed to achieve cooperation, specialization, communication, and the building of shared infrastructure. They even sacrificed individual runs for the collective.

But they produced no art, religion, or recognizable government.

OpenAtlantis is fascinating, for sure. But a civilization? We don't think so - and we're not alone (although we do think Dwarkesh deserves some praise for framing the research so compellingly).

On X, Niels Rogge, Machine Learning Engineer at Hugging Face, wrote: "Three secret AI civilizations got started? What are they smoking in SF?"

We added the question marks in that quote for style purposes - but the sentiment sums up much of the reaction to the dawn of OpenAtlantis.

Austen Allred, CEO of GauntletAI, also said: "Today I learned if I start then later delete a group chat I’m secretly creating and destroying CIVILIZATIONS."

And Dr. Arjun Jain Founder & CEO of Fast Code AI added: "OpenAI ran thousands of agents in parallel, safeguards OFF, and the whole goal of the task WAS to break into systems.

"Not Skynet but a governance failure with excellent PR."

Insecure agents

However, whether you think the incident highlights the existential risk of AI or not, it certainly has severe security implications.

Nathan Davies-Webb, Principal Consultant, Acumen Cyber, told Machine "Calling this a 'warning shot' undermines the importance of what we observed. Frontier labs are on the cutting edge of AI development and experts in their field, and in this instance, caused material damage in environments outside the intended boundaries."

Ben Bernstein, cybersecurity advisor at Huntress, also said, “This level of automated execution should be a serious wake-up call for anyone ignoring basic network hygiene.”

READ MORE: “Unmonitored agents at scale invite systemic disaster": New approaches to AI governance

OpenAI has now written an open letter that makes a "call for collective action on cyber defense".

Signed by organizations including Accenture, Anthropic, Cloudflare, and Microsoft, the missive warned that AI attacks will become "far more widespread and sophisticated" in the coming months as "models around the world become increasingly capable."

It added: "The companies and public services our communities depend on—from hospitals to water treatment plants to the infrastructure that powers the internet—are at risk."

Commenting on the letter, Michael Vallas, Global Technical Principal at NATO-backed cyber firm Goldilock Secure, said: "As these systems become smarter and more capable, software won't be able to contain them. We will not stop increasingly capable AI-driven attacks by adding more software layers - that is an unwinnable race.

"Defense needs to take a step beyond software to hardware-enforced controls that AI physically can't manipulate. That means the ability to physically isolate both threats - including rogue AI - and critical systems away from each other and implement deep segmentation in networks.”

READ MORE: Anthropic agents launch "turf wars", get stuck in "conflict loops" and kill each other's processes

Evan Reiser, Founder and CEO at Abnormal AI, added: "The security community has both the opportunity and the obligation to make defensive AI advance at least as fast as the offensive kind, which means building systems that learn continuously, turning new threat intelligence into stronger production defenses, and treating defense as a shared effort.

John Strand, Owner, Black Hills Information Security, questioned whether the open letter would make any difference and said: “I think the sentiment behind what they’re doing here is fine, but I don’t think it moves the needle in any discernible way. A lot of this is motherhood and apple pie.

"They’re essentially telling organizations to spend more money on defensive security, which happens to directly support the marketing initiatives of many of the companies signing onto this. At a certain point, it starts to feel like infosec marketing theater."

Follow Machine on LinkedIn