OpenAI Astra Brings Autonomous Zero-Day Exploitation to AI

Published August 13, 2026

OpenAI says its newest AI model, Astra, can find unknown security flaws in well protected software and build working exploits for them largely on its own. OpenAI now rates Astra at the highest risk level in its own safety framework, the first model it has ever put in that category.

Report priority
High
Victim
OpenAI

What is known

OpenAI's internal safety rules say a model crosses the top risk line if it can independently discover and weaponize unknown bugs across many hardened systems, or plan and carry out a full attack from just a high level goal, without a person walking it through each step, and OpenAI's testing showed Astra clearing that bar.

What to do

Watch for OpenAI's promised disclosures to the developers of the two software products where Astra found new flaws, and for the added safeguards OpenAI says it is putting around Astra's release.

Reported details

OpenAI ran Astra against a fresh set of Chrome's V8 engine bugs disclosed between June and August 2026, deliberately excluded from its training data so it could not have simply memorized the answers. Astra reached working code execution far more often than OpenAI's GPT-5.6 Sol model while using far fewer tokens to get there. During that same testing, Astra also turned up two brand new, previously unknown bugs on its own while stitching together a working exploit chain, and OpenAI is now working with the affected software's developers to get them fixed.

OpenAI's Preparedness Framework, published in 2023, defines a 'Critical' cybersecurity tier that a model reaches if it can autonomously discover and develop zero-day exploits against many well-defended real-world systems, or independently plan and execute a full attack chain from a high-level goal. OpenAI says Astra meets this bar. On ExploitBench, a benchmark for turning known vulnerabilities into working exploits, Astra scored 100%. On a separate, training-data-excluded benchmark built from V8 vulnerabilities disclosed June to August 2026, Astra beat GPT-5.6 Sol on code-execution success while using fewer tokens, and surfaced two previously unknown zero-days during exploit-chain construction.

References