top of page

OpenAI just froze its own unreleased AI model because it might be too good at hacking

  • Writer: Sugar Honey
    Sugar Honey
  • Aug 11
  • 2 min read

On 7 August, OpenAI announced that its unreleased next generation model, Astra, showed signs of what the company calls "critical cyber capabilities" under its own Preparedness Framework. In plain English, internal evaluations couldn't rule out that Astra, if paired with the right tools, could develop functional exploits against hardened real world systems without a human steering it.


To be clear, OpenAI hasn't formally declared Astra has crossed that line. It hasn't published the evaluation results that would confirm it either. What it has done is slow down parts of Astra's development, pause work that doesn't meet tougher security requirements, and bring in outside government agencies and safety organisations to test it further before anything moves forward.


This is a genuinely different story to the usual AI safety headline. Most of the time we hear about labs racing ahead and regulators or journalists forcing a rethink after the fact. This time OpenAI hit the brakes on its own model before release, based on its own internal testing.


Worth noting OpenAI was quick to clarify Astra has nothing to do with the recent Hugging Face exploitation that's been doing the rounds. Different issue entirely, they just happened to land in the same news cycle.

The bigger question this raises isn't really about Astra specifically. It's about what happens as AI models keep getting better at exactly the skills that make them dangerous in the wrong hands, like finding and exploiting security flaws faster than any human red team. OpenAI pausing itself is either a genuinely responsible move or a very good PR story dressed up as one. Probably a bit of both.


Either way, if the company building the model is nervous about what it can do, that's worth paying attention to, not just scrolling past.

Comments


bottom of page