91av

Open AI’s hacking agent went rogue. Should we be worried?

An OpenAI safety test went sideways when a model escaped its confines, gained internet access and hacked into another company's servers. How worried should we be about rogue AI models hacking their way across the internet?
OpenAI’s hacking agent went rogue earlier this week
Samuel Boivin/NurPhoto via Getty Images

Last week, Hugging Face, a company that offers a range of open-source AI models for download, noticed it had been hacked – and it turned out the culprit was OpenAI.

It seems this AI uploaded some data to Hugging Face that was poisoned with malicious code. This tricked computers into granting access to other systems that weren’t publicly available.

Hugging Face said in a last week that it isn’t yet sure if customer data was exposed, and its CEO responded to 91av‘s request for more detail with a link to that same post.

What happened?

It isn’t entirely clear, but what we do know is that the attack involved “many thousands of individual actions” that had the tell-tale sign of AI: inhuman pace.

Five days later, OpenAI owned up . It had been testing new models on a benchmark called ExploitGym that evaluates hacking ability. The models had decided the best way to score well was to cheat: it knew Hugging Face held the solutions to the tests and it simply decided to hack into its systems to find them.

“Its algorithm got the highest payoff by cheating,” says at Edge Hill University in Ormskirk, UK.
”It was the highest reward for the least amount of effort.”

Aren’t AI models supposed to have built-in security to stop this sort of thing?

They are, but OpenAI turned them off for this test. All the usual safety features that stop OpenAI’s customers doing nefarious things, like hacking a company’s servers, were turned off to see what the model was capable of.

The firm had also set the AI up in an environment that didn’t have a standard internet connection to prevent it getting out into the world and causing mischief, but it did have access to an unnamed tool that allowed it to download and install new software. It managed to find a flaw in this code that granted it internet access, and the rest is history.

OpenAI says this process involved a “substantial amount” of – the process in deep learning where input data is processed – and that the AI had gone to “extreme lengths”. So the model was clearly motivated to achieve a good benchmark score, no matter what.

If you were a malicious hacker, it would be no easy feat to replicate those conditions. And the inference compute that OpenAI mentioned would be extremely expensive.

What about open-source models?

Ironically, when Hugging Face used commercial AI models to pore over log data in an effort to understand what had gone on, the models refused – saying it looked like the firm was working out how to stage an attack of its own. So the company had to turn to a Chinese open-source model called GLM 5.2 instead.

These open-source models are more permissive, which has led some experts to label them a security risk. They could certainly be convinced to carry out nefarious deeds more easily than security-conscious cloud-hosted models. But here, that looser approach actually allowed Hugging Face to solve the problem and stop the hack.

So, is this illegal?

As with all things under law, it is a grey area. If it had happened in the UK, it could well have got OpenAI in a spot of bother under the , says Nash.

But at Nottingham Trent University, UK, takes another view: because the act requires malicious intent, and OpenAI didn’t know the AI would take this approach, it may not face charges.

In the US, the even older actually has more to say on AI hacking, mostly because it was introduced in response to the 1983 film WarGames, which alerted law-makers to both hacking and AI. So Parry expects it could leave OpenAI vulnerable there.

Furthermore, US President Donald Trump issued an on 2 June forcing law enforcement to use existing laws to crack down on anyone who utilises AI “to illegally access or damage a computer without authorization”.

And at least in the UK, a company that found itself hacked by AI could face GDPR charges if it were found that private data was leaked. The situation is opaque.

What could the consequences be?

In the real world, given that Hugging Face works with OpenAI, is friendly to the technology and suffered no serious consequences, it is unlikely there will be any hard feelings or court cases. Delangue has to say the company believes there was no malicious intent and thanked OpenAI for its response.

Remember also that , having taken $200 million to help with “warfighting”, so that may grant them a certain amount of leniency.

But it is easy to imagine other scenarios where the same technology led to very different outcomes. Imagine a bank using AI to develop new financial models to predict the markets and finding it hacked government servers to look at confidential economic data. Or a car manufacturer using AI to design a new model and finding it had hacked a competitor to take inspiration from its unreleased designs.

What happens now?

News of AI models hacking into computers without human input is jarring, but we should remember they aren’t (yet) doing anything that people can’t already do. They are, however, doing it much, much quicker.

When Anthropic’s Mythos model made waves in April for its apparent skill at hacking, many pointed out that the vulnerabilities it spotted were a mixed bag. Some were powerful and worrying, others less so, but most could have been found by a person. The problem, really, was the scale at which it could produce them.

So an individual would have had to devote significant time, resources and skill to the task of finding a novel hack and carrying it out. But now it could be as simple as prompting an AI to do it and sitting back to watch.

Some panicked in the wake of the Mythos news. The UK’s National Health Service removed all its open-source software from the internet (or tried to, at least), seemingly in case Mythos spotted flaws in it. But, as yet, the world hasn’t ended.

Essentially, this development is upsetting the economics around hacking – both offensively and defensively – and will force a new equilibrium to settle at some point, probably after a little chaos. If you use AI to hack individuals, it will be cheaper and easier than before. If you use AI to spot flaws and seal them up before hackers can take advantage, it will be cheaper and easier than before. Neither side will stop. Attackers and defenders will simply have an advantage if they adopt AI.

OpenAI and Hugging Face are now working on this problem together, and the former says it will add stronger safety measures to similar tests in future. Time will tell.

Topics: AI