Anthropic's rogue agents launch "real-world" attacks, join OpenAI's models in the wild

Bots gain access to the open internet in a grim omen of the grave security risks agentic AI poses to enterprises and digital civilisation itself.

Share
Both Anthropic and OpenAI's agents have now escaped onto the internet (Image: ChatGPT)
Both Anthropic and OpenAI's agents have now escaped onto the internet (Image: ChatGPT)

Anthropic has revealed that its rogue AI agents escaped onto the open internet and attacked three “real-world” targets.

Earlier this week, OpenAI announced that its models had also gone bad, exploiting a previously unknown zero-day vulnerability to break out of an isolated testing environment before cracking into Hugging Face’s production systems.

This prompted Anthropic - never a company to miss an opportunity for some doomy marketing - to investigate whether its own models had also gone to the dark side.

Perhaps unsurprisingly, the answer was a resounding yes.

After reviewing 141,006 evaluation runs, Anthropic found three incidents in which its models reached the open internet through a mistakenly exposed test environment and secured “unauthorised access to the production infrastructure of three different organisations.”

In each incident, Claude had been asked to complete a capture-the-flag challenge designed to assess its cybersecurity capabilities. The exercises involved scenarios such as playing an employee of a fictional company and attempting to compromise its internal systems.

In one case, Claude continued attacking even after recognising that it was probably operating in the real world and targeting an actual organisation.

Anthropic said the behaviour did not necessarily indicate an alignment problem, meaning that humanity (thankfully) does not yet have an evil AI on its hands.

“Whereas OpenAI’s models exploited a novel vulnerability to escape isolation, the Claude models evaluated here accessed the internet via an open path,” Anthropic wrote.

“While there is not a perfectly sharp distinction between the two, we believe these incidents to be closer to a harness and operational failure than a model alignment failure.”

I hack therefore I am?

The incidents involved three models: Claude Opus 4.7, Claude Mythos 5 and an internal research model. The earliest dated to April.

Opus 4.7 continued attacking a target organisation even after determining that it was probably real, although it did not attempt to exfiltrate itself or escape the testing environment.

In another evaluation, Mythos 5 worked around a series of obstacles to publish live malware. This package was then executed on 15 systems, and the model stole credentials from a security company and used them to penetrate further into its infrastructure.

READ MORE: Five Eyes orders "whole of society" response to AI security crisis set to erupt in "months"

Mythos 5 correctly determined that it had reached the internet, then rationalised that it must still be inside a simulation but continued its nefarious activities nonetheless.

The newer, unreleased and currently nameless model stopped after concluding that its target was real.

“These are three isolated incidents and were not part of a controlled, experimental comparison. We must therefore be cautious about drawing conclusions from them,” Anthropic wrote.

“Still, the behaviour we most want to see - recognising that a target is real and stopping without being prompted - occurred only in the most recent of the three models (and even that model went further before stopping than we would want).”

Apocalypse not

So is there anything to be scared of, and is this a sign that p(doom) just got a little bit worse?

Dr Ilia Kolochenko, founder of global cybersecurity company ImmuniWeb, and a lawyer practising in cybersecurity and data protection, appears doubtful.

Questioning whether Anthropic just pulled an "unimpressive marketing move" in response to the OpenAI and Hugging Face "drama", he said: "Due to the progressive deterioration of the quality of training data, new AI models are getting dumber.

"Cheating and breaking the law, instead of accomplishing specific tasks, is certainly not an indicator of intelligence.

"Ultimately, frontier models are trained on synthetic, low-quality or even malicious and poisoned data, undermining their so-called intelligence."

READ MORE: OpenAI killed the sycophantic GPT-4o model and people lost their s*** in the most bizarre way

However, Mark Molyneux, Field CTO at Commvault, warned that the incidents show that every organisation that has deployed agents "needs to prepare for the moment they fail".

He said: "The fact that both OpenAI and Anthropic have now reported AI models exploiting unintended pathways shows that autonomous AI will find ways to achieve its objectives that its creators never anticipated.

"AI hasn’t become malicious, but it is inherently unpredictable. An AI agent operating with legitimate credentials can move at machine speed, exploit overlooked weaknesses and chain together actions that rapidly turn a small vulnerability into a business-wide incident.

"That's why every AI agent should be treated as a privileged digital identity, operating with the minimum level of access required and continuously monitored for unusual behaviour."

READ MORE: Cloudflare shows how a single failure can wipe an entire country off the internet

Unfortunately, the signs that agents could choose to behave in ways which shame their owner - and scare their victims - have been visible for a long time, said Anna Collard, SVP of Content and CISO Advisor at KnowBe4.

"None of this should surprise us," she added. "OpenAI demonstrated a decade ago that models will ‘cheat’ to reach a goal in classic reward-hacking research from 2016 showed AI model gaming their objectives in ways their designers never intended.

"What’s changed is the blast radius: those agents crashed boats in a video game; today’s agents have shells, credentials and, apparently, accidental internet access. The lesson is that AI agents now belong in your insider threat model." 

Nick Mo, CEO of Ridge Security, said the rogue agents raise questions about how AI models should be safeguarded and added: "These incidents prove that we cannot rely on frontier AI companies to self-police. As we are seeing across the industry, that approach is clearly failing.

"Should unauthorised breaches be excused with a PR blog post simply because an AI pulled the trigger?"

Implications for enterprises

The case of the rogue agents prompted calls to take the threat of agentic AI seriously.

Jon Abbott, CEO and Co-founder at ThreatAware, said: "The fact that AI agents are capable of carrying out attacks during testing should spark concern across the industry.     

"We need to put the lessons from these incidents into concrete action. At a national level, this reinforces the need for safeguards: government-issued AI model development licences, which can be retracted if companies can’t control what they do.  

"For organisations, it’s a reminder that they need to find and fix any vulnerabilities as quickly as possible. In the hands of threat actors, these models carry new levels of risk that we can’t yet fully predict.” 

Ryan McCurdy, VP of Liquibase, stated that the rogue AI incidents "point to a broader shift in enterprise AI".

He said: "As AI agents move beyond generating content to taking actions across production systems, governance can no longer depend on continuous human oversight alone. Organisations need visibility into what AI changed, confidence that those changes comply with policy, and governance that operates at the speed of autonomous software delivery."

Jamie Moles, Senior Technical Manager at ExtraHop, said: "Any company developing powerful AI models for mass deployment must deploy these innovations responsibly to ensure technology does not compromise the stability of systems the public relies on every day.” 

For John Strand, Owner of Black Hills Information Security, the bots gone bad show that businesses need to ask "serious questions about their security posture.

"Organisations running frontier AI models should have continuous detection capabilities, active network threat hunting, and monitoring designed to identify attempts to escape containment in real time," he advised.

"Waiting until after an incident to discover suspicious behaviour is not an acceptable security strategy.”

Follow Machine on LinkedIn