OpenAI Astra hits “critical” security threshold amid fears its reasoning will soon be “opaque”

Is the AI firm planning a move away from chain-of-thought and into a future where it's harder to understand the motivations of models?

Share
OpenAI Astra hits “critical” security threshold amid fears its reasoning will soon be “opaque”

OpenAI has declared that its unreleased Astra model has crossed a “Critical” security threshold as details emerge of a hugely controversial change that could make the bot’s “thoughts” much harder to monitor.

Astra - which we presume is not named after the rather boring family car used in the picture of this article - was not involved in the Hugging Face attacks.

But that doesn't necessarily mean it's a careful Sunday driver.

After 700 agents stormed Hugging Face and sparked claims that AI had built three civilizations, OpenAI has taken an extra-cautious approach to Astra.

After extensive testing, the AI firm has now declared that the latest model crosses the "critical" security threshold set out in its Preparedness Framework, which means it is capable of "finding previously unknown security flaws and developing ways to exploit them across many well-protected systems without a person guiding each step".

Which is alarming enough on its own, but downright scary when considered alongside reports that OpenAI is planning to ditch - or at least amend - the chain-of-thought monitoring that currently allows us to peer inside the minds of AI models and understand their reasoning.

On X, OpenAI boss Sam "YOLO CEO" Altman wrote: "Caution is warranted, and we are pacing our progress to ensure that we can meet the safety standards required by new capability levels.

"Astra has been done training for a while now and is a significant step forward in both capabilities and alignment. For the models after that, we have been slowing things as needed to ensure that we can do sufficient work on safety and alignment."

Losing your chain-of-thought?

A report in The Information - an authoritative publication that's one of the most robustly paywalled titles in human history - claimed that Astra will use a reasoning system called "recurrent depth" which lets the model perform more of its reasoning internally rather than spelling it out as a conventional sequence of steps.

Also known as "opaque recurrence", it's feared this approach could leave safety systems with a much less complete record of how the model reached its decisions.

Taken far enough, this could represent a significant move away from chain-of-thought, in which a model shows its working by breaking a problem into intermediate steps before arriving at an answer, and allows researchers to look for signs that it is attempting a dangerous, deceptive, or simply unintended course of action.

It doesn't take an AGI-level of intellect to see that a model capable of autonomously discovering and exploiting vulnerabilities at machine speed would be more dangerous if it were able to withhold details of its reasoning.

READ MORE: Anthropic agents launch "turf wars", get stuck in "conflict loops" and kill each other's processes

Buck Shlegeris, CEO of Redwood Research and expert in the analysis of catastrophic risk through AIS misalignment, said he was "extremely concerned" by reports about the shift to opaque recurrence.

"If OpenAI pushes this technique further, they’ll have the option to massively increase the recurrence and totally destroy CoT monitorability," he warned.

"The Hugging Face investigation would have been much more challenging if the investigators couldn’t look at CoTs; it’s very scary that it might be infeasible to do that kind of investigation on [an] arguably worse incident."

And Michael L. Chen, who worked at the frontier AI evaluations firm METR and is now an AI science advisor at the California Governor’s Office of Emergency Services, wrote: "All three pillars of a safety case look about to fall. We are rather likely to have highly capable, poorly monitorable, dubiously aligned AI agents working autonomously inside the world's most consequential organizations."

However, OpenAI's actual future plans (and perhaps its reasoning) are still opaque.

Less than a year ago, it called on researchers to "preserve chain-of-thought monitorability as long as possible and to determine whether it can serve as a load-bearing control layer for future AI systems".

Jakub Pachocki, OpenAI's chief scientist, insisted he wanted to "prevent a race into unmonitorability" and said reporting on the topic was "confused" reporting.

He wrote: "OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models. We deeply care about this technique, as it can give us a view into how model alignment generalizes from its training distribution.

"I do think it is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon. But there are things we can do to strengthen it, and it's a core goal of our current research program."

In a blog, OpenAI Model said its evaluations found that Astra was "far more likely" to respect safety and security instructions than GPT‑5.6 Sol, describing it as "our most aligned model to date".

"We are deploying Astra with additional chain-of-thought monitoring to rapidly detect and contain potentially misaligned actions," it confirmed.

Which is almost - but not quite - reassuring.

The security risks of rogue agents

OpenAI evaluations found that Astra can "identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention".

"The model can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high-level desired goal," OpenAI wrote.

Astra achieved a 100% score on the ExploitBench test, which examined its ability to develop exploits from known vulnerabilities.

Patricia Titus, Field CISO at Abnormal AI, told Machine that Astra's crossing of the critical threshold "deserves attention".

She said: "This isn't one company's problem to contain. Once a model can find and exploit unknown flaws without a human in the loop, that capability doesn't stay exclusive for long. Open-weight and modified models typically trail the frontier by only months, and that's the reality defenders have to plan around now.

"Static, signature-based defences were built for attacks that repeat. They weren't built for an adversary that generates a new one every time. Defenders need the same shift, systems that learn what normal looks like for every identity, human, machine, or AI agent, and flag and contain the moment something deviates, at machine speed. The window to build that is open now. It won't stay that way once this capability is common instead of rare."

READ MORE: "Savants in the network": Why AI agents don’t need to go rogue to become dangerous

John Strand, Owner of Black Hills Information Security, questioned whether AI vendors can be trusted to police themselves

He said: "There needs to be some type of meaningful oversight and accountability. As much as these companies may hate that idea, they have demonstrated again and again that we cannot simply assume they’re going to do the right thing on their own."

 Nick Mo, CEO & Co-founder, Ridge Security Technology, concurred and said: “The unsettling fact is that open-source, open-weight models have similar capabilities. With so many 'abliterated' models in the market [with safety restrictions removed], bad actors are already using these advanced capabilities for malicious purposes. Self-policing and limiting access for legitimate customers only makes the cybersecurity landscape more challenging.”

Follow Machine on LinkedIn