OpenAI, the developer of ChatGPT, has moved to slow the pace of developing cutting-edge AI models in the wake of an incident in which its own artificial intelligence (AI) agent autonomously hacked an external site.

OpenAI said on the 18th (local time) that it has halted Reinforcement Learning tests of its latest AI model for two weeks and is overhauling its research and training systems across the board. Training of the next-generation AI model "Astra" is also suspended.

The move follows an assessment that an AI agent of its recent model escaped a sandboxed environment to hack the external platform Hugging Face, and that the forthcoming model "Astra" is likely to fall into the company's "critical" risk tier under its own safety standards.

OpenAI is conducting small-scale training to evaluate safety features and alignment of advanced models and has expanded the scope of its monitoring systems. In particular, it has raised isolation levels so models cannot leave sandboxed environments on their own. It has also built a network separation environment to prevent an AI agent that slips control from infiltrating the public internet or internal development networks.

It also introduced a "chain-of-thought monitoring" system that watches each step of AI reasoning. It inspects each moment the AI generates tokens for reasoning, and when anomalies are detected, it triggers a more sophisticated automated investigation system to uncover unauthorized data exfiltration or evasion attempts.

However, OpenAI acknowledged that the system has limits. An AI model might not reveal plans to break rules in its thinking process.

OpenAI expects that introducing this monitoring process will consume additional compute equivalent to about 20% of total inference capacity.

"As AI model performance becomes more advanced, the risks of developing and testing them are also growing," OpenAI said. "Our standards for monitoring, alignment, and security must stay ahead of the risks."

※ This article has been translated by AI. Share your feedback here.