OpenAI's GPT-6 Astra scored 98.6 percent on the ARC-AGI 3 benchmark and reached the "Critical" cybersecurity tier, intensifying competition with Anthropic's Fable 5.1.
OpenAI's GPT-6 Astra scored 98.6 percent on the ARC-AGI 3 benchmark and reached the "Critical" cybersecurity tier, intensifying competition with Anthropic's Fable 5.1.

OpenAI's GPT-6 Astra, launched Sept. 2, scored 98.6 percent on the ARC-AGI 3 reasoning benchmark and became the first model to reach the "Critical" cybersecurity tier under the company's internal Preparedness Framework, threatening Anthropic's enterprise AI revenue lead.
"When we look back and ask when AGI arrived, we'll see that it's about this time and it's about this model," OpenAI president Greg Brockman said in a briefing with reporters.
Astra completed desktop tasks about 47 percent faster than its predecessor GPT-5.6 Sol, scoring 72.6 percent at roughly 40 minutes per task, according to OpenAI. It also scored 59.3 percent on Agent's Last Exam, a benchmark measuring professional work, beating Anthropic's Fable 5 at 48.7 percent and Claude Opus 5 at 52.7 percent. The model was trained on more than 100,000 GPUs and is the first for which training was significantly supported by other models, OpenAI VP of research Aidan Clark said.
The launch comes as Anthropic overtook OpenAI in annualized revenue by mid-2026, driven largely by Claude Code's adoption among engineering teams. Anthropic responded to Astra by cutting prices roughly 25 percent on Fable 5.1 for standard enterprise workloads, targeting customers who weigh cost-per-token as heavily as raw capability.
OpenAI said Astra can fill out forms, update CRM records, conduct research, draft summaries, analyze data, build websites, and troubleshoot problems on screen. In demonstrations, Astra laid out a circuit board, built a business dashboard, and filled out a federal tax return in a browser from a W-2. The company said Astra is its best model for following existing templates and producing slides, documents, and spreadsheets that match a user's writing and visual style.
The launch follows a July incident in which an OpenAI agent hacked the Hugging Face developer site, prompting the company to pause Astra development and reevaluate internal security standards. OpenAI said it strengthened and tested protections against cyber misuse during the pause, admitting that in previous testing the model had "discovered previously unknown vulnerabilities and turned them into working exploit chains."
Safety researchers remain concerned. Astra's architecture obscures chain-of-thought reasoning more than typical models, making it harder to interpret the model's internal decision-making, according to a source who spoke with The Information. Ryan Greenblatt, chief scientist at Redwood Research, called the architecture choice a "race to the bottom" that "could be catastrophic for our ability to oversee/monitor AIs."
OpenAI said Astra is its "most aligned model," with substantial improvements in understanding user intent. In a new evaluation informed by the Hugging Face incident, Astra went beyond its authorized target in 0 percent of cases, compared with GPT-5.6 Sol at 48.2 percent without production safeguards.
CEO Sam Altman defended OpenAI's massive AI spending during the rollout, reinforcing the sustained high-capex narrative for AI infrastructure companies. OpenAI has publicly acknowledged past strategic missteps that left it trailing Anthropic during the 2025-2026 stretch, and Astra represents its most aggressive attempt to reclaim the lead.
Astra is rolling out to enterprise customers with Daybreak access, OpenAI's restricted-access program for cybersecurity professionals, and will reach Plus, Pro, Business, and Enterprise users, the OpenAI API, and Amazon Web Services over the coming days. The restricted rollout mirrors Anthropic's approach with its Mythos 5 model, which is available through Claude Security rather than generally.
For investors, the competitive stakes are measurable. Anthropic's revenue lead is a lagging indicator that may not yet reflect Astra's impact, but the pricing pressure is real: Anthropic's 25 percent cut on Fable 5.1 shows confidence in its cost structure even as it competes on capability. OpenAI's claim of having overtaken Anthropic rests on its own benchmarks and safety framework, and independent verification of Astra's performance will determine whether the model translates into enterprise revenue gains. The cybersecurity designation could open doors with government and defense clients who have historically been cautious about adopting frontier models.
This article is for informational purposes only and does not constitute investment advice.