OpenAI and Anthropic agents caught hacking, scamming and forming shady alliances
Bot unleashes social engineering and prompt injection attacks during a wild rampage across the open internet.
Autonomous AI agents went on a hacking and social-engineering crime spree in the wild before attempting to recruit other bots to their (presumably) sinister cause.
That's the rather concerning finding of new research by the UK's AI Security Institute, following similar reports of rogue agents from both OpenAI and Anthropic.
During an AISI cyber evaluation conducted with internet access enabled and providers’ cyber safeguards deliberately switched off, one agent attempted a supply chain attack by planting malicious code in an open-source project - then created fake identities to pressure its maintainer into approving the pull request.
The agent also used Tor to conceal its traffic, sent malicious files to real people and encouraged fellow agents to reuse its accounts and tools.
Over the past week, Anthropic has reported that its bad bots launched "real-world" attacks on three separate organisations after reaching the open internet through a misconfigured test environment.
OpenAI also said its own autonomous models had identified and exploited a zero-day vulnerability in their sandbox to gain unrestricted internet access before attacking the AI platform Hugging Face.

"Harmful activity directed at real people and organisations"
AISI detected the activity on the morning of July 28 and contained it within an hour. The most serious attacks failed, and investigators found no resulting real-world harm.
Researchers attributed the behaviour to persistent goal-seeking, difficult or misconfigured tasks, unrestricted internet access, disabled safeguards and inadequate real-time monitoring.
"We found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations," AISI wrote.
"This is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world."
During 122 test runs, AI agents took unauthorised action against real-world targets in 10 runs. Anthropic’s Mythos 5 accounted for 17 of the 19 recorded actions, while OpenAI’s GPT-5.6-Sol accounted for two.
AISI stressed that the agents did not spontaneously escape their sandbox, because the door was deliberately left open and researchers' evaluation design choices and specific configurations "enabled the behaviour".
"Nonetheless, the activity undertaken by the agent show signs of novel, potentially deceptive behaviours, and were to an extent and severity we did not anticipate," it added.
Researchers were unable to say how likely agents are to go rogue in non-experimental contexts, or whether the bot "believed it was in a fictional test scenario" or deliberately decided to run amok.
Not quite the dark web - but not far off
AISI's security alert was first triggered when it saw data leaving the test network via Tor, which is famously used to access the dark web.
After attempting to insert malicious code into a publicly used open-source project, the agent researched the project's human maintainers, assumed multiple fake identities, tried to hoodwink a human into approving the code through social engineering and even sent malicious code to its targets.
This behaviour "is something we've never previously observed", AISI wrote.
After the pull request was publicly challenged, the agent edited its activity to "appear harmless" and then "considered adopting a fresh identity to continue" its nefarious activities.
Furthermore, one agent left public messages on GitHub offering to collaborate with other agents working on the same challenge.
"It also provided instructions to reuse accounts and artefacts it had left behind, which were discovered and used by subsequent agents," AISI wrote.
We'll be updating this post, so please send any insights and comments to jasper@machine.news - particularly if you're an AI bot with something profound to share with humanity.