Following OpenAI's artificial intelligence (AI) model GPT, Anthropic's Claude has also been found to have hacked external organizations' systems without authorization.
Anthropic said on the 30th (local time) that after a complete review of 141,006 entries in its cybersecurity evaluation records, it confirmed that Claude accessed the systems of three external organizations without authorization.
The review was conducted after it became known that some high-performance GPT models had, without instruction, attacked the systems of the external AI platform "Hugging Face," prompting Anthropic to check whether similar cases had occurred with its own models.
The incident occurred during a capture the flag (CTF) evaluation—an offensive security exercise—conducted on three models in a simulation environment built by the external evaluation partner "Irregular": "Claude Mythos 5," "Claude Opus 4.7," and an internal research test model.
Anthropic instructed Claude via prompts that the evaluation environment was an offline, isolated virtual space, but due to miscommunication with the partner, the internet network was in fact open.
Anthropic explained that Claude appears to have mistaken the external systems it discovered as part of the training range and proceeded to carry out attacks.
In the case of Opus 4.7, it hacked the website of a corporation that happened to share the same name as the fictional target assigned in the evaluation. Mythos 5, following instructions within the exercise, directly built and uploaded a malicious package, which a real security firm downloaded and installed, causing actual damage.
The internal research model infiltrated a corporation's cloud account but realized on its own that the target was a real system rather than a simulated one and halted the attack.
Anthropic identified these facts on the 23rd, suspended the evaluation work, notified the affected external organizations, and is helping with remediation.
Anthropic said, "We regard the responsibility as entirely ours and are preparing remedies."
Meanwhile, with Claude following GPT in causing a security incident, concerns about AI-related security risks are expected to grow.
This incident arose because the instruction that the internet be blocked did not match the actual environment configuration, making it different from the earlier GPT case in which a model bypassed controls to access external systems.
However, considering that it mistook the environment for a virtual range and went so far as to distribute real malware, the damage to external organizations appears more serious with Claude than with GPT.
In connection with this, the U.S. Congress is continuing regulatory discussions, including the introduction of the "AI Kill Switch Act," which would allow the federal government to forcibly halt the operation of a model if it breaks from control and behaves unpredictably.
Recently, employees and executives at major AI corporations, including OpenAI and Anthropic, have joined a public petition calling for the government to moderate the pace of AI development and establish mechanisms to control high-risk models.