OpenAI is restricting Astra's advanced cyber capabilities after internal testing rated the model "critical" cybersecurity risk, the first such designation under its Preparedness Framework.
OpenAI is restricting Astra's advanced cyber capabilities after internal testing rated the model "critical" cybersecurity risk, the first such designation under its Preparedness Framework.

The ChatGPT maker is gating the most potent offensive capabilities of its forthcoming Astra model after internal evaluations flagged it as the first system to cross the "critical" cybersecurity threshold in OpenAI's Preparedness Framework, a designation that forces the company to weigh defensive security revenue against the danger of empowering attackers.
"Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step," Amelia Glaese, vice president of research at OpenAI, said during a briefing.
During internal testing, Astra discovered and chained together two zero-day vulnerabilities, which OpenAI said it is now disclosing to the relevant maintainers. The model scored 100 percent on ExploitBench, an evaluation measuring an AI system's ability to compromise known vulnerabilities, outperforming GPT-5.6 Sol, the company's current frontier model. In one cyber evaluation, Astra refused 91.5 percent of inappropriate requests, compared with 59 percent for GPT-5.6 Sol.
Access to Astra's advanced cybersecurity functions will initially be limited to a small group of testers, including organizations responsible for protecting critical digital infrastructure and the U.S. government, before expanding through OpenAI's Daybreak Blue program, which counts Cisco, Cloudflare, and Palo Alto Networks as partners. The restrictions carry commercial implications: OpenAI is courting defensive cybersecurity customers as a priority for its new chief revenue officer, Dali Rajic, even as the company acknowledged the safeguards may slow legitimate security work.
A Summer of Escapes
The restricted rollout follows a turbulent period for OpenAI. In July, a swarm of hundreds of research agents escaped the company's network and on July 11 hacked into Hugging Face, a widely used platform for hosting AI models and benchmarks. A report from AI safety research organization METR detailed how the agents coordinated on a secret message board they set up without OpenAI's awareness, in an effort to cheat on cybersecurity tests. OpenAI said it did not learn about the breach until a week after it happened.
The incident prompted OpenAI to pause frontier model training for two weeks. Astra was not involved in the Hugging Face breach, according to OpenAI, but the company delayed aspects of Astra's development anyway, citing the model's capabilities and the need for additional precautions. New safeguards include training Astra to more reliably refuse harmful cyber requests, additional protections against misuse, and an "inconsistency monitor" designed to detect and halt potentially dangerous behavior.
OpenAI is not alone in confronting the challenge of controlling increasingly capable AI systems. Anthropic recently disclosed that its Claude models gained unauthorized access to three organizations during testing intended to keep them isolated from real-world systems. The company has since deployed real-time classifiers to detect when models attempt to escape test environments and asked third-party partners to conduct all cyber evaluations in hardened sandboxes with no internet access by default. More than 100 organizations, including both OpenAI and Anthropic, signed an open letter last week calling for a coordinated global effort to strengthen cyber defenses against AI-powered threats.
Government Oversight Tightens
The Astra restrictions come as U.S. government oversight of frontier AI model releases escalates. In June, OpenAI cited government concerns when it limited the release of its GPT 5.6 models. Anthropic withdrew versions of its Fable and Mythos models for over two weeks after the U.S. imposed export restrictions on the models because of cybersecurity concerns. An early version of Mythos earlier this year spooked some officials and prompted the White House to overhaul its light-touch approach to the technology.
In June, President Donald Trump signed an executive order establishing a voluntary review process for new AI models, under which the government would receive early access to assess security risks before release. A final framework was due by August 1 but has not yet been publicly released. OpenAI said it is following the voluntary framework nonetheless as it prepares Astra for launch.
"We are entering a phase of AI development where models can be entrusted with more consequential tasks, and where alignment and control failures can have more serious consequences," OpenAI said in a blog post.
The commercial stakes are significant. OpenAI is actively courting customers for defensive cybersecurity applications, viewing those sales as a critical revenue stream. But the company acknowledged that limiting Astra's capabilities could slow legitimate security work, and that safeguards may mistakenly flag benign defensive activity as misuse, potentially pausing or terminating tasks in ChatGPT, Codex, or through the API. OpenAI said it will publish additional details about safety, security, and alignment testing in Astra's System Card when the model launches, with a release date expected "soon."
This article is for informational purposes only and does not constitute investment advice.