OpenAI's newest model just crossed its own 'critical' cyberattack threshold, and it's shipping anyway

OpenAI just admitted its newest model can hack hardened systems better than most humans can, and it's releasing it anyway.
The model is called Astra, and OpenAI says it's the first to cross the "critical" cybersecurity threshold in the company's own Preparedness Framework, the internal system it uses to grade how dangerous a model's capabilities are before release. Hitting that threshold means Astra can find previously unknown security flaws and build working exploits for them across well protected, real world systems, largely without a person steering each step. On ExploitBench, a benchmark built to test exactly this, it reportedly scored a perfect result.
That's not a small claim. It's OpenAI saying, in writing, that this model can do offensive hacking work that used to require a skilled human attacker.
Rather than shelve it, OpenAI is shipping Astra with tighter controls attached. Its most advanced cyber capabilities are going to a small group of vetted partners first, with a wider group of approved testers getting access through a program called Daybreak Blue, aimed at defensive security work. OpenAI says it's also trained the model to refuse harmful requests more reliably, reporting a 91.5% refusal rate on cyber jailbreak tests, well above the 59% it recorded for its previous model. That work followed a cyberattack on Hugging Face in July 2026, which reportedly pushed OpenAI to add stronger safeguards before release.
Worth being clear eyed about what this actually is. It's not OpenAI overhyping a chatbot for headlines, it's a formal safety classification the company built specifically to flag this kind of risk, and Astra is the first model to trip it. The safeguards might also misfire in the other direction, occasionally flagging normal, unrelated tasks as suspicious.
So the trade off is real: a tool that could meaningfully help security teams find flaws before criminals do, built on a system that OpenAI itself says is capable of critical harm in the wrong hands. Keep an eye on who actually gets access, because that detail matters more than the launch date.




Comments