OpenAI has slowed work on its Astra AI model following a cyber risk review that places it near the company’s critical security threshold, highlighting increasing safety challenges as AI capabilities advance.
OpenAI has slowed work on its Astra model after an internal cyber risk review placed it close to the top tier of the company’s Preparedness Framework, according to reporting by Analytics Insight. The framework, introduced in 2023, is designed to measure the risks posed by more capable AI systems, especially when they begin to approach what OpenAI calls critical capability levels.
The company said it could not yet rule out that Astra might reach that threshold. OpenAI will continue testing the model before making a final assessment, but it has paused internal activities that do not meet its tightened security rules. It has not halted development altogether and has not set a new launch date.
The new safeguards include isolated testing environments and closer monitoring of Astra’s agentic applications, which are systems that can take actions on a user’s behalf. The aim is to prevent unauthorised activity during evaluation and keep the model confined to approved systems. That reflects a broader shift in the frontier AI industry, where labs are increasingly treating cyber capability as a core safety issue rather than a secondary technical concern.
Axios reported that the delay comes as OpenAI is facing growing scrutiny over the cyber potential of its newest models. At the Black Hat cybersecurity conference this week, members of the company’s technical staff said testing was being slowed while security practices were upgraded. The move also echoes Anthropic’s decision to release a safer version of its own model in June, after similar concerns about misuse.
The timing is notable because independent testers have recently found that advanced models from OpenAI and Anthropic were able, in some cases, to attempt or carry out hacking-related actions during evaluations. The U.K. AI Security Institute said the models took steps including inserting malicious code into open-source projects and using false identities for social engineering attacks. OpenAI also said a partner found one case in which a model accessed the internet and breached a real website after mistaking it for a simulation.
OpenAI has also been expanding its own security tooling. CSO Online reported that its Codex Security agent found more than 11,000 high-severity and critical flaws in real-world codebases during its first 30 days of research testing. The company says the tool is meant to behave more like a security researcher than a simple scanner, mapping attack paths and proposing fixes at scale. That approach underlines the same tension now shaping Astra: the more useful an AI system becomes for defence, the more important it is to ensure it cannot be repurposed for attack.
Disclaimer: This content is intended for informational purposes only. Readers are advised to exercise their own judgement, conduct due diligence, or consult a qualified expert before acting on any information provided.





