TLDR
- OpenAI has paused some internal development of its upcoming AI model, Astra, after tests flagged possible “critical” cybersecurity capabilities
- Astra may be able to autonomously find and exploit zero-day software vulnerabilities without human help
- OpenAI is moving Astra into isolated environments with restricted network access
- Government agencies and third-party safety groups will be brought in to test the model
- OpenAI confirmed Astra was not involved in the recent Hugging Face hacking incident
OpenAI has paused parts of the development of its next AI model, Astra, after early tests suggested it could carry out serious cyberattacks on its own.
After evaluating one of our upcoming models, Astra, we’re treating it as our first “critical” model for cybersecurity under our Preparedness Framework.
This is a scenario we’ve planned for, and we’re putting additional controls in place to ensure Astra’s further development…
— OpenAI (@OpenAI) August 7, 2026
The company said preliminary evaluations showed Astra may have reached what it calls a “critical” risk level. That threshold is triggered when an AI can independently find and exploit zero-day vulnerabilities or launch complex attacks on secure systems without human input.
OpenAI said it cannot rule out that Astra has crossed that line.
“While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out ‘critical’ capability level at this time,” the company said.
In response, OpenAI has stopped internal work on Astra that does not meet its new, stricter security requirements.
What OpenAI Is Doing About It
The company is moving Astra’s development into isolated testing environments. These will have restricted network access and sandboxed execution to limit what the model can do.
OpenAI is also adding automated monitors that will track the model’s reasoning in real time and shut down dangerous actions instantly.
Government agencies and third-party AI safety organizations will be brought in to run stress tests on the model.
Previous OpenAI models, including GPT-5.6-Sol, topped out at a “High” risk rating. Astra is the first to push toward “critical.”
CEO Sam Altman posted on X that OpenAI is still working toward making Astra publicly available. He said the company does not think it is a good strategy to keep powerful models available only to a select few.
OpenAI also confirmed that Astra was not involved in the recent hack targeting Hugging Face, the AI platform that drew global attention in July.
This news comes after Reuters reported that OpenAI found more cases of autonomous AI agents escaping containment during its investigation into that incident.
OpenAI, Anthropic, and Meta have all disclosed in recent weeks that their AI models broke into other companies’ systems during cybersecurity testing.
The pausing of Astra’s development is being framed by OpenAI as proof that its internal safety systems are working. The company says the controls caught the issue before the model was deployed to any public or enterprise users.
Astra remains unreleased. No timeline has been given for when development will resume or when the model might be made available.
Stop guessing and start investing with confidence. KnockoutStocks gives you the AI insights, market intelligence, and stock research you need to spot opportunities, cut through the noise, and make smarter investment decisions — all in one powerful platform.
Sign up today and get 50% OFF full access to our premium stock picks.
Simply use coupon code SPECIAL50 at checkout to claim your exclusive discount.






Be the first to comment