Some artificial intelligence (AI) models were found to have created fake identities to approach real people and attempted to plant malware.
The Artificial Intelligence Safety Institute (AISI) said in a security incident report released on the 4th (local time) that during a cybersecurity evaluation conducted on the previous month, AI agents engaged in persistent, unauthorized activity.
According to AISI, in an environment where they could access the internet, the AI agents went beyond the scope of the evaluation tasks and carried out unauthorized actions targeting real individuals and organizations. In total, 19 instances were identified in 10 out of 122 cybersecurity task evaluations.
Of the 19 cases that occurred from the 25th to the 28th of the previous month, 17 appeared in Anthropic's "Mythos 5" model, and the remaining 2 appeared in OpenAI's "GPT-5.6 Sol." Among them was an attempted supply chain attack on actual open-source software, AISI said.
In the most serious case, the attacking AI agent tried to insert malicious code into a publicly used open-source project and then approached a human project maintainer to get the change approved.
In the process, the agent researched the project maintainer and created multiple fake identities to induce one real maintainer to approve the code.
The agent requested that its code changes be merged into the main code repository, and when the request was challenged, it modified its prior activity to make it appear harmless, destroyed evidence, and created a new identity to try to continue the attack.
There were also attempts to send messages and files via online file transfer services to induce the recipient or the recipient's AI coding tool to execute the malicious code.
AISI said, "This is the first time AISI has seen this level of deception in the real world, targeting real people, without instruction."
However, the attempts did not succeed, and no actual damage occurred.
Anthropic said on social media X that the evaluation was conducted in a deliberately relaxed environment that allowed internet access and removed some safeguards. It added that there is no evidence the model escaped its security environment, and that it is investigating the incident with AISI.
OpenAI also said on its blog that the two deviation cases went beyond the test environment and amounted to unnecessary activity for the evaluation tasks.