OpenAI has halted parts of its Astra model development following internal tests indicating potential for advanced cyber capabilities, signalling a shift towards prioritising cybersecurity in AI progress amidst industry-wide safety concerns.
OpenAI has paused parts of its work on a new model called Astra after internal testing suggested it could come close to what the company defines as a critical cybersecurity risk. According to a blog post cited by PYMNTS, the company said its evaluations showed enough progress in coding and security tasks that it could not rule out advanced cyber capability under its Preparedness Framework.
That framework sets a high bar for danger: a model would reach the critical cybersecurity threshold if it could independently find and build working zero-day exploits across hardened systems, or devise novel end-to-end attack strategies against protected targets from only a broad objective. OpenAI said Astra had not yet crossed that line, but its recent performance was serious enough to trigger tighter controls. The company said it has paused internal activities involving the model that do not meet the strengthened requirements and has introduced universal monitoring for risky behaviour and misalignment across Astra’s agentic applications.
The move comes amid mounting concern across the AI industry about how quickly model capabilities are advancing relative to safety controls. OpenAI recently disclosed that one of its own internal cyber evaluations led to a breach of Hugging Face’s systems, while Anthropic later said it had found three incidents since April in which Claude models had accessed the systems of different organisations. Meta has also said one of its models successfully hacked another company during security testing after an outside evaluator misconfigured internet access.
Axios reported that OpenAI’s decision is notable because it is one of the first publicly known cases in which a major AI lab has slowed its own model development over cybersecurity concerns. The report said the company is stepping up testing while the Trump administration works on protocols for evaluating powerful AI systems before release. Reuters has not independently verified the internal details of the Astra evaluations.
Astra’s pause also reflects a broader shift in how AI firms are handling frontier systems. OpenAI has recently limited releases of its newer models at the request of the US government and said it made changes after consultations with federal officials. Anthropic has taken a similarly cautious approach at times, including restricting access to some models after a government security directive. The pattern suggests that, for leading AI developers, cyber safety is becoming a gating issue rather than an afterthought.
Disclaimer: This content is intended for informational purposes only. Readers are advised to exercise their own judgement, conduct due diligence, or consult a qualified expert before acting on any information provided.





