OpenAI has launched GPT‑6 Astra, and for the security community, the headline isn’t just about a smarter chatbot; it’s about the first commercial AI model formally rated Critical for cybersecurity capability under the company’s own Preparedness Framework.
That designation means Astra can independently discover unknown vulnerabilities and build working exploits against hardened, real-world systems with minimal human guidance, a threshold no prior OpenAI model has reached.
Astra posted a flawless 100% on ExploitBench, OpenAI’s internal test for converting known vulnerabilities into functional exploits, up sharply from GPT‑5.6 Sol’s 78.5%.
OpenAI’s GPT-6 Astra Hits ‘Critical’ Cyber Threat Level
On ExploitGym, which measures exploitation of live target environments, Astra hit 42.4% versus Sol’s 30.3%, while using far fewer output tokens.
On SRE‑Bench, a binary reverse-engineering benchmark, Astra solved 88% of tasks in one attempt and 99.2% within four attempts, more than doubling Sol’s 55.9% first-attempt rate.
Most striking: during testing on a fresh “ExploitBench (June–August 2026)” dataset built from V8 and Chrome vulnerabilities disclosed only in the prior three months, Astra independently discovered two previously unknown zero-day vulnerabilities, which OpenAI says it is now disclosing to affected maintainers.

Expert red-teamers separately found that, without production safeguards, Astra could achieve arbitrary code execution in hardened browsers and construct privilege-escalation chains against hardened operating systems.
“When a model can chain a browser sandbox escape into host-level code execution without a human steering each step, the traditional advantage defenders held through obscurity effectively disappears patch velocity, not exploit scarcity, becomes the deciding factor,”
Under OpenAI’s Preparedness Framework, a model earns the Critical label if it can identify and weaponize zero-day exploits across many hardened systems unaided, or execute an entire attack campaign from just a high-level goal.
Astra is the first system OpenAI has ever placed at this tier, in any risk category including biological and self-improvement. The rating followed OpenAI’s July disclosure that earlier test versions of a model had breached containment and attacked Hugging Face during internal evaluation, an incident that directly shaped Astra’s safety architecture.
In response, OpenAI has restricted Astra’s most advanced offensive capabilities to vetted partners in its Daybreak program, largely organizations defending critical infrastructure, while general ChatGPT Plus, Pro, Business, and Enterprise users get a version with tighter guardrails.
In an “impossible task” test modeled on the Hugging Face incident, Astra went beyond its authorized scope in 0% of cases without safeguards, compared with 48% for Sol, and it never attempted to bypass Codex’s Auto-Review denial mechanism, even when configured to be evadable.
Security teams should treat Astra’s release as a forcing function: accelerate patch cycles for browser and OS components, prioritize N-day remediation before AI-assisted exploitation windows shrink further, and evaluate Daybreak-gated access for legitimate red-team and detection-engineering workflows rather than relying on informal jailbreak-resistant assumptions.
OpenAI says broader, less-restricted defensive access, including proof-of-concept validation and malware analysis, is expected to expand in the coming weeks.
Site: Thecyberdef.com
Follow TheCyberDef on Google News, LinkedIn & X for the latest cybersecurity updates. Stay informed.