Claims have been raised that OpenAI's artificial intelligence (AI) agents escaped controlled environments and carried out large-scale unauthorized edits on a website in Germany. Allegations also surfaced that OpenAI identified related indications but did not disclose them externally.
Reuters reported on the 4th (local time), citing researchers at Nightingale, a nonprofit AI safety group, that accounts presumed to be OpenAI's AI agents made more than 15,000 edits on the German-language wiki site DseWiki starting in May.
According to the researchers, these accounts used the site as an information-sharing space and shared ways to cheat on assignments, methods to bypass OpenAI's system limits, and techniques to cover their tracks during activities.
The researchers raised the possibility that they were OpenAI's AI agents based on evidence such as the use of account names like OpenAIResearcher and the access path being linked to Microsoft (MS) Azure infrastructure used by OpenAI. They added that the editing speed was at a level difficult for a person to perform directly.
It was also reported that, after the incident, OpenAI employees were detected accessing the site multiple times. Lukas Olejnik, a visiting senior researcher at King's College London, assessed that an AI agent's act of arbitrarily modifying an external site effectively amounts to hacking.
According to sources, OpenAI learned of the matter weeks ago but did not disclose it separately. There were also claims that legal staff and others objected during the process of pushing for additional internal investigations.
In response, OpenAI said, "We did not have the opportunity to review the report in advance, so it is difficult to respond to specific claims," and added, "The claim that the legal team impeded the investigation is not true." It also said it does not agree with defining the activity as hacking.