OpenAI disclosed on August 7, 2026, that its unreleased model, Astra, has demonstrated cybersecurity capabilities strong enough that the company “cannot rule out” that it has crossed the Critical threshold under its Preparedness Framework, the highest-risk tier the company tracks.
It marks the first time any frontier AI lab has publicly flagged one of its own models at this level for cyber capability, and it triggered an immediate, partial internal shutdown of work on the model.
According to OpenAI’s own account, internal evaluations conducted over “the past few days,” combined with expert assessments, led the company to conclude, “last night,” that it could not rule out Critical-level cyber capability in Astra.
OpenAI Flags Astra AI’s “Critical” Hacking Capability
The model showed marked jumps in agentic coding, the ability to independently execute long, multi-step coding tasks layered with offensive cybersecurity skill. OpenAI was explicit that Astra was not connected to the recent Hugging Face breach, separating this disclosure from any confirmed real-world incident.
Under OpenAI’s Preparedness Framework, first published in December 2023, a model hits the Critical cybersecurity threshold if it can independently identify and build functional zero-day exploits across severity levels in hardened, real-world critical systems without human help, or if it can plan and execute an entire novel cyberattack against a hardened target from nothing more than a high-level goal.
Previous flagship models, including GPT‑5.6‑Sol, topped out at the “High” tier. This is the first time a model has approached the ceiling of the scale for offensive cyber capability, and OpenAI’s own framing treats it as a genuine inflection point rather than routine benchmarking news.
OpenAI’s countermeasures read like an incident-response playbook rather than a product announcement. The company is rolling out isolated testing environments with no live internet exposure, restricted network and tool access, stronger model-weight encryption, expanded monitoring, and sandboxed execution for any Astra-related work.
Critically, internal activities involving Astra that don’t yet meet these hardened standards have been paused outright. OpenAI also says it has deployed universal monitoring of the model’s chain-of-thought reasoning across all agentic uses, including training itself, with automated triggers that can interrupt high-risk behavior in real time.
OpenAI committed to involving government agencies and select AI safety organizations to independently test Astra’s capabilities, and says it will provide third-party testers with recommended security controls to run higher-risk evaluations safely.
This mirrors the approach OpenAI took in June 2025 when its models approached the High threshold for biological risk, when it similarly expanded external testing and safeguards rather than deploying unilaterally.
For the cybersecurity community, the Astra disclosure is less a warning about an imminent attack tool and more a signal that autonomous vulnerability discovery and exploit generation are now technically plausible at scale.
OpenAI frames this as an opportunity for defenders to get ahead of attackers using the same class of tooling, but the honest reading is that offense and defense are now racing on the same curve, and the gap between “capable model” and “capable model in the wrong hands” is thinner than it has ever been.