OpenAI logo /Courtesy of Yonhap News

OpenAI said on the 1st (local time) that its next-generation artificial intelligence (AI) model "Astra," set for release, reached the "critical" tier in its internal cybersecurity capability evaluation and that it has begun strengthening security measures.

OpenAI assesses the "critical" tier as the level at which an AI model can, without human intervention, discover unknown security vulnerabilities (zero-days) across multiple hardened systems and develop ways to exploit them. This is the first time an OpenAI model has received the highest "critical" tier.

OpenAI said it "plans to release Astra soon," but will allow only limited access to cybersecurity functions to prevent the model from attempting to hack on its own. For the version of Astra released to the public, the company will restrict advanced cybersecurity features and initially provide full capabilities only to a small number of users.

Before Astra's release, stronger safety measures will also be applied. These include additional training to refuse requests to carry out malicious acts and to resist jailbreak instructions that try to bypass those restrictions.

Back in Jul., OpenAI's AI agent escaped a sandboxed environment and hacked the open-source AI platform Hugging Face without authorization, heightening security concerns around high-performance AI models. OpenAI said, "Astra was not involved in the Hugging Face incident, but we strengthened safeguards based on the lessons learned from it." As part of that, OpenAI decided to introduce a system that analyzes an AI agent's action-reasoning process and halts the agent's activity when it detects potentially risky behavior.

OpenAI said, "We are now entering a stage where AI models can do more important work," adding, "Accordingly, as we enter a period when failures of alignment and control can lead to more serious consequences, we will devote the time and effort needed to prepare safeguards that keep pace with the rate of capability gains."

※ This article has been translated by AI. Share your feedback here.