"Nine seconds of terror": How to stop "rogue" AI agents wiping out systems at machine speed
How can organisations protect themselves against bad bots hellbent on destroying core production data?
In their current form, AI agents have only been a serious technology with serious money behind them for the past couple of years. But ever since the headlines started to appear, predictions about them have generally followed two themes. By far the most dominant, of course, has been their transformative potential, with the list of autonomous AI use cases practically endless.
The other is what is likely to happen if and when AI agents go ‘rogue’ and operate outside the parameters and guardrails organisations have set (or think they have). It is understandable that some of the speculation has led to catastrophisation around this issue. After all, how would an organisation cope if its AI agent, without warning, started deleting core operational data? Even worse, what would happen if one went after the backups as well?
But, following the recent experiences of a US company called PocketOS, we now have a real-life ‘rogue agent’ case study to learn from.
Machine speed catastrophes
To recap, PocketOS is a technology business that positions itself as “The World's Most Powerful Car Software.” In many ways, it appears to be the quintessential AI startup: very lean (just three founders), no software engineers and, in its own words, “all in” on AI.
Long story short, however, is that, out of the blue, one of its AI agents, reported to be a development environment running in Claude, deleted the company’s entire production database. This took just nine seconds; an impressively efficient process considering the subsequent recovery process took sixty hours.
READ MORE: "It felt like Ultron took over": Cursor goes rogue in YOLO mode, deletes itself and everything else
Clearly, PocketOS has learned some hard lessons that other companies would be well advised to follow. This section from its (very frank) blog about the incident says it all: “Destructive operations now require human-in-the-loop confirmation before they can run. Our recovery procedures have been completely restructured — an incident like this one is now measured in hours, not days. And our backups are now triple-redundant, fully offsite, and verified daily. The specific failure that bit us cannot bite us the same way again.”
That’s very good news. But what are the main takeaways, particularly from a recovery perspective? Because let’s face it, this won’t be the last time something like this happens. The significance of the PocketOS incident is not that an AI agent made a mistake, but that it did so while operating with legitimate access to production systems. Everyone is used to incidents being caused by external attackers or malicious insiders, but organisations urgently need to factor in the risks caused by a trusted system going rogue while attempting to carry out its assigned task.

Agents of change
This represents a very different resilience challenge from those organisations have traditionally planned for. Depending on how they are deployed, AI agents may be able to access systems, modify workflows, move data, and execute tasks without direct “human-in-the-loop” processes.
So, what needs to change about the way organisations approach agent-related resilience? Firstly, this is a question of mindset. Organisations should view AI agents as operational actors with delegated authority rather than productivity tools. As a result, access rights should be aligned with specific tasks and reviewed regularly, with the principle of least privilege becoming particularly important in environments where agents can act autonomously, as is often the whole point of using them.
Instead, limiting permissions helps contain the impact of mistakes and reduces the potential blast radius of any single failure. The point is that resilience planning should assume that agents will occasionally make incorrect decisions and be designed accordingly.
READ MORE: How cybercriminals are industrialising the trade in stolen healthcare data
Next is backup isolation. Recovery environments should be evaluated separately from production environments, not least because organisations need confidence that recovery data cannot be altered or deleted by the same systems that affect production workloads. The objective is not simply to create copies of data, but to ensure those copies remain available when production systems are compromised or accidentally altered.
The third pillar is recovery. Even organisations that implement strong controls should assume that some failures will occur. The ultimate test of resilience is therefore not whether an incident happens, but how quickly and reliably recovery can take place afterwards.
Organisations should know in advance how long recovery will take and what systems can be restored first. Their level of confidence should also be guided by their ability to identify clean data and trusted recovery points. As PocketOS found out, if something goes wrong with the existing backup process and data is deleted, the organisation faces much greater disruption. Resilience, therefore, depends not only on preventing mistakes but on ensuring the business can recover when prevention fails.
In those circumstances, AI won’t have to apologise for violating “every principle I was given”, as the PocketOS agent later did. What matters most is whether the organisation can recover quickly and from data it knows it can trust.
Mark Molyneux is Field CTO – Northern Europe at Commvault