The Brief

The AI that can now write its own hacks

OpenAI says its newest model can find security holes nobody knew about — and write the code to break in — with barely any human help. It's holding the model back. The reason it can defend you is the exact reason it scares people.

OpenAI just said something quietly enormous: its newest model, Astra, has crossed into “Critical” territory on cybersecurity. In plain English, that means it can spot security flaws no human has documented yet, and write working code to exploit them — with very little hand-holding. The company is deliberately holding parts of it back while it builds stronger guardrails.

Strip away the jargon and here’s what’s actually new.

What happened

For years, AI has been a helpful assistant to security researchers — summarising, suggesting, speeding things up. Astra is being described as something past that line: a system that can do the hard, creative part of hacking — finding the unknown weakness and building the key to it — mostly on its own. OpenAI classifying it “Critical” and restricting release is the tell. You don’t put the brakes on something that isn’t fast.

Why it matters

The same skill cuts both ways, and that’s the whole story. A model that can find flaws before criminals do is the best defensive tool the internet has ever had — it could patch holes at a speed no human team can match. That same model, in the wrong hands, is the best weapon. There is no version of this that’s only good, because finding the hole and exploiting the hole are the same act, done with a different intention.

That’s why “who can use it, and how it’s contained” suddenly matters more than what it can do. The capability is settled. The safeguards are the only variable left.

The crux

We’ve spent a decade treating AI risk as a someday problem — sentient robots, far-off sci-fi. This is the un-dramatic version that actually counts: a tool, available now, where the safety rails are the entire product. Astra isn’t dangerous because it’s smart. It’s dangerous because the thing it’s good at has no neutral use.

What to take away

Two things. If you run anything online — a business, a site, an app — the near future is a race between defenders using tools like this to close holes and attackers using them to find holes first. Patching stops being optional. And when you read that an AI lab is “restricting” a model, don’t read it as marketing modesty. Read it as the most honest thing they can tell you: they built something they’re not sure they can fully control, and they know it.

That’s the crux. The scary part of AI was never that it might think. It’s that it can now do.