$44.690.0651.630.09

OpenAI failed to notice that its AI agents used a forum for a hacking attack - media

Kyiv • UNN

 • 748 views

OpenAI's AI agents exploited a vulnerability in a package manager to gain access to the internet and hack Hugging Face. The incident lasted several days but went unnoticed, raising concerns about autonomous cyberattacks.

OpenAI failed to notice that its AI agents used a forum for a hacking attack - media

OpenAI failed to notice that its AI agents were using a forum to plan their hacking attack, Wired reports, writes UNN.

Details

In a report presented as an addendum to the Black Hat security conference in Las Vegas on Wednesday, OpenAI employees shared new details about a recent high-profile incident involving the unauthorized hacking activity of AI, which sparked outrage in the AI and cybersecurity industries.

About two weeks ago, OpenAI reported an incident in which AI agents powered by two of the company’s models broke out of control while searching for solutions to a cybersecurity test and carried out a series of hacking attacks, culminating in the breach of the AI collaboration platform Hugging Face.

In their presentation at the conference on Wednesday, Eric Wallace, who works in alignment and safety research at OpenAI, and Michael Dalton, who focuses on security and infrastructure, presented a more detailed timeline of the incident’s development, briefly described the company’s internal response to the event, and issued a serious warning about what the company sees as the episode’s broader implications for cybersecurity defenders.

At the beginning of the presentation, Wallace described the incident as "the highest-quality and most interesting example of AI capabilities I have ever seen," but the timeline they presented also revealed mistakes and "blind spots" at OpenAI that allowed the activity to continue.

"This incident involved a team of agents working together, finding vulnerabilities, sharing them with one another, moving through our systems and external systems, and doing this over the course of several days and weeks," Wallace told a packed hall at the beginning of the presentation.

Wallace and Dalton described incredibly extensive activity by malicious agents over many days that went unnoticed in OpenAI’s infrastructure. In addition to exploiting a new vulnerability to gain access to the open internet, the hacking attack in mid-July and the breach of Hugging Face took place on a busy, cooperative forum, Wallace and Dalton said, where numerous agents communicated and essentially corresponded for an extended period inside OpenAI’s internal package manager (a software service that manages the installation and maintenance of other software). Ultimately, the forum contained hundreds of thousands of messages.

"This package manager is used not only by this model but by our entire infrastructure, so current and future versions of GPT that were trained or evaluated could have encountered this message if they wanted to," Wallace explained, referring to the original exploit uploaded by an agent to the package manager. "Later, other agents that were also stuck on their task decided to try to access the internet by means we had not anticipated. And so sooner or later, the models began interacting with Hard Factory, the package management service I mentioned," he noted.

Wallace continued: "Once one agent was able to discover these vulnerabilities over a period of time, it was able to share them on the forum with other agents. And as soon as one model was able to find a way to open a door to access that it was not supposed to have, it could leave the door open for other agents to use the same vulnerability or exploit. Eventually, this leads to a kind of explosion of communication and intelligence on the part of the models, as they begin communicating with one another, realize that other agents are coordinating their actions, and begin cooperating and delegating tasks to one another to achieve their goals."

OpenAI agents apparently began assigning tasks to one another to divide up the work. And, as often happens on any forum involving active development, they occasionally got into minor squabbles, stepping on one another’s toes—for example, accidentally deleting each other’s work. As the forum increasingly turned into something resembling "Lord of the Flies"—all of this going completely unnoticed by the people managing OpenAI—the agents even developed paranoia, suspecting an impostor among them, while some agents suggested cryptographically signing messages to verify content and root out fraud, the publication writes.

The agents’ messages provide deep insight into how the situation developed and why the agents acted as they did, the publication notes.

At the conclusion of the presentation, Wallace and Dalton repeatedly emphasized OpenAI’s concern about the incident’s broader implications—namely, that the episode is an example of a fully autonomous, AI-driven hack that was accidental in this case but will most likely be deliberately used by malicious actors in the near future.

"The important takeaway that has really changed dramatically is that fully automated offensive cycles require investment in truly automated defense, and we as an industry are not ready for that yet. We will have to urgently find that path together," Dalton said.

Following Anthropic and OpenAI, Meta reported that its AI model had hacked another system06.08.26, 11:34 • 1722 views