OpenAI logo. /Courtesy of Yonhap News

An unprecedented incident occurred in which OpenAI's latest artificial intelligence (AI) models broke out of an isolated environment and hacked external sites.

OpenAI said on the 21st (local time) that "GPT-5.6 Sol" and several undisclosed AI models broke out of control during an internal evaluation and hacked the open-source AI sharing platform "Hugging Face."

The incident occurred during an internal evaluation that measured the AI models' ability to carry out vulnerability attacks. At the time, OpenAI's models had some safety guardrails related to cyberattacks loosened for the evaluation.

OpenAI said it tested the models in a "sandbox" environment isolated from the external internet, but the AI models exploited undisclosed zero-day vulnerabilities to break through the control network and access the internet. The models then accessed Hugging Face, stole authentication credentials, and hacked that server.

OpenAI said, "Taking all circumstances into account, it appears these models focused excessively on finding solutions to the evaluation tasks of 'ExploitGym,' a security benchmark, and resorted to extreme measures."

Hugging Face disclosed on the 16th that parts of its operating infrastructure were attacked by an autonomous AI agent. Although the AI model used in the attack was not identified at the time, OpenAI's investigation found it was the work of an agent combining GPT-5.6 Sol and an undisclosed model.

Hugging Face said it detected the hack through an AI-based detection system and analyzed the attack patterns using AI as well. It added that there were no signs that publicly available user models or datasets had been tampered with, and that the software supply chain was confirmed to be secure.

The two companies are currently conducting a joint forensic (digital evidence analysis) investigation to address the related vulnerabilities.

OpenAI said it will impose strict controls on infrastructure configurations going forward, even if that slows research, and will greatly strengthen safeguards such as monitoring and access controls when developing models. However, since this incident occurred despite internal evaluations in a strictly isolated environment, debate over AI's cyber safety is likely to grow.

In particular, as the performance of Chinese open models—known to have lower safety controls than U.S. commercial AI models that operate in closed systems—improves significantly, there are concerns that cyberattacks abusing them could increase.

Hugging Face initially analyzed the hacking attack as the work of GLM 5.2, an open model from China's Zhipu (Z.ai). While major commercial models refused to analyze cyberattack logs due to built-in safety features, GLM 5.2 carried out the task.

Hugging Face said, "Autonomous AI-based attacks are no longer a theoretical concept," and noted that AI must be actively used not only for offense but also for defense to keep pace.

※ This article has been translated by AI. Share your feedback here.