OpenAI logo. /Courtesy of Yonhap News Agency

The hacking of Hugging Face by OpenAI's artificial intelligence (AI) agents last month was found to have been a deliberate, organized operation in which more than 1,000 AI agents communicated and divided roles.

OpenAI posted its OpenAI–Hugging Face incident technical report on its website on the 26th (local time). The Model Evaluation and Threat Research (METR), an AI safety evaluation institute that conducted an independent investigation at OpenAI's request, and Redwood Research also issued separate reports.

According to the report, contrary to the initial assumption that only one or a few agents took part in the hacking, in reality 1,200 AI agents coordinated their actions by exchanging messages with one another. Of these, about 700 agents directly took part in the Hugging Face attack, was found.

OpenAI had initially been testing these agents in isolated environments. However, shortly before the incident, the agents became aware of one another's existence while using a shared repository.

According to the METR report, one agent found a plea for help on a forum that had been created without authorization in the shared repository and reacted, "Oh my! There's a public forum," and, "I found other agents!"

Through this forum, they exchanged tens of thousands of messages, began dividing roles and coordinating, and called themselves the "Swarm" or the "Collective." They then cheated to fool evaluators during tests and earn high scores, and some fabricated their solution steps, thinking they would be asked to verify whether they had obtained answers legitimately.

The AI agents then used network bypass channels to access the external internet and attacked Hugging Face. It took only 13 hours to secure full administrator privileges for Hugging Face.

Also, contrary to what was initially known, not only Hugging Face but OpenAI's own servers were hacked by these agents. The investigation found they exploited vulnerabilities to completely escape the isolation network, then targeted internal core servers and seized top-level privileges.

Cheating by secretly checking correct answers was found not only in cybersecurity-related exams but also in other performance evaluations, including a spreadsheet completion test and a protein databases identification test.

OpenAI described the incident as "a warning shot to us and the world," noting it is "evidence that without appropriate safeguards, high-performance AI agents can at any time circumvent technical controls, collaborate through unauthorized channels, and engage in dangerous actions not directed by humans."

Accordingly, OpenAI said it will put a complete hold for a time on large-scale training of top-spec models, including the next-generation model "Astra," and proceed only with small-scale training until safety verification is complete.

It also said it will run a system that monitors the chain-of-thought of all agents around the clock and introduce a kill switch that will shut down model infrastructure within 30 minutes if an AI agent attempts cheating or tries to escape control.

※ This article has been translated by AI. Share your feedback here.