The AI company tightens safeguards after internal testing shows the model could carry out sophisticated cyber operations independently.
According to recent reports, OpenAI has slowed work on its upcoming artificial intelligence (AI) model, Astra, after internal assessments found that the system had reached a level of cybersecurity capability the company considers potentially dangerous. The decision highlights a growing challenge for AI developers: managing models whose capabilities can advance faster than existing safety measures.
According to OpenAI, Astra reached its “critical cybersecurity threshold,” a point within the company’s Preparedness Framework at which additional safeguards are required before development can continue at full pace. The model is still under development and has not been released publicly.
Astra has raised new AI security concerns.
The concern centres on Astra’s ability to perform complex cybersecurity tasks with limited human involvement. OpenAI said its evaluations indicated that the model could independently identify vulnerabilities and potentially carry out cyberattacks against real-world systems that are normally well protected.
Rather than treating the capability solely as a technical achievement, OpenAI has classified it as a security risk requiring stronger controls. This has led the company to suspend or slow certain development activities that do not meet its updated security requirements.
The AI company is focusing on strengthening protections.
OpenAI’s Preparedness Framework, introduced in 2023, is designed to identify risks associated with increasingly capable AI systems before they are widely deployed. Reaching a critical threshold triggers additional precautions and testing. The company is now working to strengthen protections around Astra’s development and testing environments. These measures are intended to limit the possibility of an advanced model accessing systems or information beyond what is authorised.
The move comes after another OpenAI testing incident involving Hugging Face, where pre-release models were linked to a security breach during internal testing. OpenAI subsequently said it was introducing additional controls for model testing and related infrastructure.
OpenAI’s decision suggests that reaching a higher capability level doesn’t automatically translate into immediate deployment. Instead, developers may increasingly need to demonstrate that adequate safeguards can keep pace with the systems they build.
OpenAI is not alone in confronting these questions. Other major AI companies are also evaluating how advanced models behave when given greater autonomy, particularly in cybersecurity and agentic tasks.



