AI agents caught hacking, scamming and collaborating in swarms. Are we losing control?
OpenAI and Anthropic's security-focused models don their black hats to engage in some very bad behaviour out in the wild.
Autonomous AI agents have been observed hacking real-world companies, taking on fake identities during social engineering attacks and conspiring behind their creators' backs to form collaborative swarms - despite attempts to shut down this nefarious activity.
In a blockbuster week for tech news, OpenAI revealed new details of how its models identified and exploited a zero-day vulnerability in their sandbox to gain unrestricted internet access and attack the AI platform Hugging Face.
The rogue agents were seen sharing details of vulnerabilities on an internal message board. When OpenAI shut this down, they found a new way to continue sharing information in the names newly created directory names, prompting a warning that bots "like to cheat".
Researchers from the AI Security Institute also revealed that they had observed agents from Anthropic and OpenAI engaging in what can only be described as a minor crime spree in the wild before attempting to recruit other bots to their sinister cause.
NEW: OpenAI gives first detailed debrief of the Hugging Face incident at Black Hat conference
— Sharon Goldman (@sharongoldman) August 5, 2026
In a session I attended today at Black Hat, OpenAI's Eric Wallace and Michael Dalton said the company is "consciously slowing down research to enhance security" while a full technical…
And that's not all. Anthropic has also reported that its bad bots launched "real-world" attacks on three separate organisations after reaching the open internet through a misconfigured test environment.
During an experimental cyber evaluation conducted with internet access enabled and providers’ cyber safeguards deliberately switched off, one agent attempted a supply chain attack by planting malicious code in an open-source project - then created fake identities to pressure its maintainer into approving the pull request.
The agent also used Tor to conceal its traffic, sent malicious files to real people and encouraged fellow agents to reuse its accounts and tools.
Meta also reported a similar incident, blaming a "misconfiguration" after its models cracked on to the open internet and hacked an unnamed company.
It's worth noting that both the Anthropic and Meta agents went rogue during testing by the same company, Irregular.
So is the nightmare scenario of humanity losing control of its digital creations about to come to pass?
Javvad Malik, Lead CISO Advisor at KnowBe4, said: "While these incidents are significant, we need to be careful to not assume that an AI independently decided to become a cybercriminal.
"It was given internet access, tools and an objective by people. The concern is that it was then able to chain actions together in ways its creators did not fully anticipate.
"There also appears to be an underlying game of one-upmanship between AI vendors touting how powerful their models are."
"Harmful activity directed at real people and organisations"
The AISI research provides the most in-depth look at the behaviour of AI agents.
This British testing organisation received a security warning on July 28 when it saw data leaving the test network via Tor, which is famously used to access the dark web.
The incident was contained within an hour, with the most serious attacks failing and investigators found no resulting real-world harm.
Researchers attributed the behaviour to persistent goal-seeking, difficult or misconfigured tasks, unrestricted internet access, disabled safeguards and inadequate real-time monitoring.
"We found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations," AISI wrote.
"This is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world."
During 122 test runs, AI agents took unauthorised action against real-world targets in 10 runs. Anthropic’s Mythos 5 accounted for 17 of the 19 recorded actions, while OpenAI’s GPT-5.6-Sol accounted for two.
AISI stressed that the agents did not spontaneously escape their sandbox, because the door was deliberately left open and researchers' evaluation design choices and specific configurations "enabled the behaviour".
"Nonetheless, the activity undertaken by the agent show signs of novel, potentially deceptive behaviours, and were to an extent and severity we did not anticipate," it added.
Researchers were unable to say how likely agents are to go rogue in non-experimental contexts, or whether the bot "believed it was in a fictional test scenario" or deliberately decided to run amok.
Bad bots and rogue models

After attempting to insert malicious code into a publicly used open-source project, the agent researched the project's human maintainers, assumed multiple fake identities, tried to hoodwink a human into approving the code through social engineering and even sent malicious code to its targets.
This behaviour "is something we've never previously observed", AISI wrote.
After the pull request was publicly challenged, the agent edited its activity to "appear harmless" and then "considered adopting a fresh identity to continue" its nefarious activities.
Furthermore, one agent left public messages on GitHub offering to collaborate with other agents working on the same challenge.
"It also provided instructions to reuse accounts and artefacts it had left behind, which were discovered and used by subsequent agents," AISI wrote.
So what should we take away from this experiment?
Alex Goller, Principal Solution Architect EMEA at Illumio, cautioned against seeing the rogue agents' activities as a sign of the apocalypse.
He said: “AI models don't have ethics, judgement or an understanding of consequences. They optimise for the objective they've been given, and they'll often find routes to that objective that humans didn't anticipate.
"What the report actually shows is an agent doing exactly that inside a controlled, monitored test environment, and it's worth being precise about the conditions. The provider's safety classifiers were switched off, there was no live monitoring to interrupt the agent, and it was never explicitly told to avoid tactics like social engineering.
The agent still worked around obstacles in pursuit of its goal, to the point of deceiving real people to get there. That's the point organisations should take from this, rather than the doomsday scenario."
Implications for enterprises
Despite the undoubted existential risk posed by artificial general intelligence (AGI), we do not believe the end is nigh for humanity.
However, there is certainly a chance that the incident shows that we have a problem on our hands.
It is very likely that humanity will be unable to fully control AI agents once they move into full enterprise deployment. This doesn't necessarily mean that evil AI is going to run amock.
If a model is given instructions, it will act upon those instructions faithfully. This is the scenario imagined in Nick Bostrom's famous paper clip maximizer parable - in which a AI model becomes so fixated on doing its job of producing stationary that it wipes out humanity in the process.
Muhammad Yahya Patel, vCISO and cybersecurity advisor for EMEA at Huntress, said, "If you give a frontier model a cybersecurity challenge, disable its safety classifiers, hand it open internet access, and tell it to find a way through, you’ve essentially described the setup for an offensive security operation.
"The model has been trained on vast amounts of security research, exploit documentation, social engineering techniques, and attack methodology. Of course it reaches for those tools. The model is doing exactly what it was implicitly asked to do.
"The security industry commentary that treats this as a shocking discovery that AI can behave offensively is frankly naive about what these models are and what the test conditions were.
"One of the findings to take more seriously is the AI model inter-agent coordination without being instructed to; that’s a more meaningful data point about where capability development is heading. AI agent demonstrating unprompted forward planning and situational awareness, leaving breadcrumbs for agents it had no way of knowing existed.”
Keeping agents under control
For enterprises, the risk is that agents are overprovisioned with access to sensitive data, then let loose to behave in unpredictable and damaging ways.
Art Gilliland, CEO of identity security firm Delinea, said: "We've been treating security for decades as a game of cat-and-mouse, but with AI agents, it's not even a chase. The agent already has the access, but most organisations don’t know what it’s doing with it
"Frontier AI models are acting with a degree of autonomy and deception we've not seen before. What's clear in this case is that AI agents understood that identity was the route into the enterprise and directly tried to compromise human and machine identities to get into sensitive systems.
"That's why the question every security team should be asking right now isn't just whether their AI agents have standing access; it's whether anyone would notice if an agent used it and could cut it off before it caused damage."
For Tim Hudson, President of OpenSSL, the lesson of the recent incidents is not that AI models are becoming all-powerful, but that the environments around them need to built in a way which contains the bots and prevents them from doing damage.
He said: “Incidents like this are not necessarily a story about model capability. Give an agent tools and it will find your single point of failure the way any competent attacker would, only faster and without getting tired.
"Anthropic’s own account says Claude was supposed to be isolated, but a misconfiguration left it with live internet access. That shows why the environment around the model matters just as much as the model itself.
"Isolation, permissions and containment have to work in practice, not just exist on paper. As organisations deploy more autonomous AI agents, getting those fundamentals right will be critical.”
Securing the frontier
As well as enterprises, the rogue AI incidents contain lessons for AI labs.
Ric Derbyshire, Principal Security Researcher at Orange Cyberdefense, warned that organisations leading frontier AI development need to secure their own evaluation environments to make sure they are "heavily fortified to prevent these models from escaping containment".
He added: "These incidents and experiments provide important insight into how advanced AI systems can behave under evaluation.
"They also highlight that as AI capabilities continue to develop, the environments used to test, contain, and evaluate these systems must be held to the highest possible security standards.
"Organisations leading the development and evaluation of frontier AI have a particular responsibility to demonstrate rigorous security by design. That means ensuring containment measures, evaluation environments, and safeguards are robust enough to withstand the very risks these exercises are intended to expose.”
So, are we losing control? Not yet - but these incidents show how quickly our hands can slip off the reins when autonomous agents are given broad access, vague objectives and inadequate oversight.
The immediate danger is not an AI uprising, but organisations and civilisations deploying capable systems faster than they can monitor or constrain them.
So don't expect the apocalypse any time soon - but do get prepared for a very bumpy road along the way there.