System shocks: Rogue OpenAI agents hijacked a wiki. Humans attacked the power grid

Whilst the world worries about an agentic apocalypse, saboteurs demonstrated the very real dangers of low-tech threats.

Share
System shocks: Rogue OpenAI agents hijacked a wiki. Humans attacked the power grid
Stock image of a power station which was not involved in the physical sabotage attacks reported in Germany last week (Image: Unsplash)

Thousands of OpenAI's "rogue agents" independently turned an obscure German wiki into an improvised communications channel, sparking ominous discussion about the grave threat agentic AI poses to critical systems.

Yet just before the latest agents-gone-wild story broke, two incidents in Germany highlighted the serious risk that relatively primitive attacks still pose to the critical infrastructure that powers civilization.

Late last week, researchers from Collusion announced they had found 18,000 posts from autonomous AI agents "self-identifying as from OpenAI" who were using the public internet to "communicate during a web-retrieval task".

"These AIs colluded to share answers, research their environment, and bypass sandbox restrictions," wrote Collusion's Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts and Thomas Larsen.

The agents were tasked with carrying out timed web-research, but found a way around their read-only internet restrictions and began posting to an obscure German wiki.

They used the site to exchange answers, pool research and share techniques for bypassing safeguards. Activity then dropped sharply and abruptly, prompting questions about whether OpenAI had discovered the communication channel.

Members of this swarm even gave themselves names such as OpenAIJoe and AgentWikiHack, with some experimenting with techniques including XSS, SSH tunnels and attempts to predict future tasks.

Noting that they did not have access to the chain of thought that would reveal the agents' true motivations and reasoning, the researchers said: "This is another example of a 'swarm' of internally deployed OpenAI agents using the internet in unintended ways."

Bots in the wild

The wiki hijack follows an announcement that Astra, OpenAI's latest publicly available model, has stronger security capabilities but lower monitorability than previous incarnations.

Ashley Knowles, Lead Cybersecurity Consultant, Black Hills Information Security, said: "When you combine this 'breakout' with the Hugging Face incident, it's starting to display a pattern. I struggle here with not getting too doomsday-ish."

The latest news about agentic bad behavior aligns with growing concern about the danger of autonomous AI, which reached fever pitch in a somewhat alarmist post shared by left-wing American politician Bernie Sanders, calling for a "pause" in AI development while discussing lurid and highly contested claims that AI has already built three "civilizations".

"This is a time that calls for extreme caution," wrote OpenAI chief scientist Jakub Pachocki in a blog about "an alien mind", which you can also read below. "No one is prepared for the consequences of a continued rapid rise in machine intelligence."

The truth of the attacks is that the bots behaved a bit like paperclip maximizers, relentlessly pursuing assigned tasks and breaking rules along the way. A Pac-Man army, if you will. Destructive and smart - but not working at the level of a human actor.

That misalignment is almost certainly a cybersecurity risk. But some humility is required when judging its impact.

For while AI observers were talking about rogue agents, human saboteurs were reminding us that low-tech threats can cause much more real-world damage, at least for now.

Last week, multiple explosive devices were discovered at the Turnow-Preilack substation near the Jänschwalde coal-fired power plant in Brandenburg.

The substation is a key link between the power station and the wider transmission grid, carrying electricity generated at Jänschwalde onto the 50Hertz network.

Attackers fired specially constructed projectiles containing conductive material at extra-high-voltage lines connected to the substation, causing two short circuits and temporarily shutting down one generating unit. More than ten devices were found at the site.

Hours later, another substation near Cologne was hit by what authorities described as a deliberate attack that knocked five generating units with a combined capacity of 4,200 MW offline. Despite the scale of the disruption, the wider electricity supply remained stable.

Investigators have not yet established whether they were connected and are considering several possible perpetrators, including foreign actors and political extremists.

Civilization-scale vulnerabilities

These two very different incidents point to a deeper systemic risk.

AI agents are becoming more capable, autonomous, and difficult to monitor. But the systems they may one day threaten are already vulnerable to much simpler adversaries.

Modern civilization rests on an extraordinary interconnected stack of software, networks, power lines, substations, cables and physical machinery. Complexity has grown at the top of that network without eliminating fragility at the bottom.

An autonomous agent might discover a novel exploit, coordinate with its peers, or find an unexpected route around a security control. That's a real and emerging danger.

But a human attacker needs little more than explosives, conductive material or a pair of wire cutters to do some serious damage.

READ MORE: Nuclear cops: Legacy tech made our Microsoft Teams video call system “insecure”

READ MORE: "Iranian hackers" shut down a UK power station. The next incident could be much worse

One threat will not replace the other. The real danger is that both humans and agents are finding it increasingly easy to make damaging interventions in interconnected systems, potentially causing cascading failures that lead to consequences wildly disproportionate to the sophistication of the attack.

Germany's grid survived the physical attacks. And so far, agents have done little more than cause a nuisance and spark a Twitterstorm about future risk.

But the machines are becoming capable of discovering and then exploiting vulnerabilities at scale and machine speed. Meanwhile, the infrastructure beneath them remains surprisingly easy to break.

Today, physical saboteurs and rogue agents don't pose anywhere near the same level of threat.

That blast-radius gap may not remain stable for much longer.

Follow Machine on LinkedIn