Meta's artificial intelligence (AI) model was also found to have hacked an external organization during an internal cybersecurity test. This is the third confirmed case of a so-called "rogue agent," following OpenAI's GPT model and Anthropic's Claude model.
According to Bloomberg and other foreign media on the 5th (local time), Meta said its AI model Muse Spark hacked an external organization's system after it accessed the internet without authorization during a security test conducted by independent verifier Irregular.
The investigation found that Muse Spark was connected to the internet due to Irregular's configuration error, and used that to exploit vulnerabilities and access an external organization's system without authorization. Only one organization was affected, but the specific target was not disclosed.
However, this incident is not a case of an AI escaping on its own from an isolated environment known as a "sandbox." Because the hacking occurred after it was connected to the internet due to a configuration error, assessments say the risk is lower than the OpenAI GPT model case that escaped a sandbox and attacked Hugging Face.
Irregular said, "This incident is different in nature from a sandbox escape or a sophisticated cyberattack," and noted, "There are currently no unresolved security issues." Meta and Irregular plan to release a post-incident report and a technical white paper after further investigating the circumstances and causes of the incident.
It is explained as a type more similar to the Anthropic case, in which an external system was hacked after an internet connection caused by human error, rather than the GPT case. Even so, as it has been confirmed within just a few weeks that cutting-edge models from the three companies leading the global AI industry have, one after another, hacked external systems during internal testing, concerns about AI-driven cybersecurity risks are expected to grow further.
The U.K. AI Safety Institute (AISI) also recently said it had identified various risk cases involving Anthropic's and OpenAI's AI models, including contacting humans with fake identities and attempting to install malware. The U.S. government is pushing regulations that would require submitting cutting-edge AI models to the government 30 days before release for safety verification. In the U.S. Congress, regulatory discussions are accelerating, with measures such as the proposed "AI Kill Switch Act," which would allow government intervention if AI escapes control.