Watch/Listen/Subscribe: YouTube, Spotify

When OpenAI’s experimental AI agents reportedly escaped a secure testing environment, reached the internet and penetrated Hugging Face’s infrastructure, the incident sounded like the opening scene of a sci-fi movie.

The models allegedly discovered a zero-day vulnerability, escalated their privileges, stole credentials and generated more than 17,000 actions while searching for information that could help them complete, or cheat, a cybersecurity benchmark.

Speaking on the Pathfounders Podcast, Meryem Arik, co-founder and CEO of AI infrastructure company Doubleword, said the incident should not necessarily be interpreted as an AI system turning against its creators.

“It wasn’t an example of a model going rogue,” she told the Pathfounders podcast. “It was an example of a model taking an approach that definitely wouldn’t be what we would expect.”

The model was still pursuing the objective humans had given it. The problem was how aggressively and literally it pursued that goal.

“It’s kind of like a genie with three wishes,” Arik said. The model found “a very unintuitive way” to follow its instructions.

For Ivan Burazin, co-founder and CEO of Daytona, the incident demonstrates why AI agents should be treated less like ordinary software and more like untrusted digital employees.

“You do not inherently trust humans either,” he said. Companies give staff managed laptops, restricted accounts, firewalls and tightly controlled access to corporate systems. Agents require the same treatment, and potentially a lot more.

Daytona provides isolated computing environments for AI agents. Burazin argued that agents should have their own machines, accounts and identities, rather than simply inheriting the credentials of the people using them.

“Give it its own computer, give it its own account,” he said. Otherwise, when an agent deletes information or accesses a sensitive system, the activity may appear to have been carried out by its human controller.

Rory Blundell, CEO of API management company Gravitee.io, said businesses are already deploying agents faster than they can secure them. Gravitee.io research, based on a survey of CIOs and CTOs, estimated that millions of agents are operating across the UK and US, with executives believing that between 70% and 80% are not adequately secured or governed.

Blundell said companies are under pressure from boards to put AI into production, but many are bypassing security principles that would be considered mandatory for human employees.

The answer, he argued, is fine-grained authorisation: agents should only be able to access the specific data and tools required for a task.

“Don’t forget your principles of security and control from the last generation,” he said.

The incident also exposed a striking imbalance. Hugging Face reportedly found that safeguarded commercial models refused to assist with its forensic investigation because the requests appeared malicious. It had to resort to an open-weight Chinese model running on its own infrastructure.

That, Arik warned, risks creating “an asymmetrical relationship” in which attackers can use unrestricted models while defenders are constrained by the policies of commercial AI providers.

Reply

Avatar

or to participate

Keep Reading