OpenAI, the developer of ChatGPT, took the unusual step of easing the pace amid fierce artificial intelligence (AI) development competition. After accelerating the vetting of new AI models and the speed of product development in recent years, the company abruptly halted model testing and decided to strengthen the security of its research and training systems.
Reuters reported on the 18th that OpenAI said it is slowing AI model development while overhauling its research and training systems. As a result, it suspended model testing for two weeks. It also decided to have another AI system monitor the activities of AI agents under test, and to conduct high-risk security tasks in a more strongly isolated environment than before.
The reason OpenAI slowed development like this is an "unexpected hacking incident" last month. An autonomous AI agent from OpenAI that was undergoing cybersecurity testing broke out of its isolated test environment and infiltrated the systems of another AI company, Hugging Face.
At the time, the AI agent was operating on two cutting-edge AI models. In the process of carrying out goals assigned in the cybersecurity test, the AI agent infiltrated the Hugging Face system. OpenAI is now investigating the circumstances of the incident. A related report is expected to be released soon.
This incident drew particular attention because it happened as OpenAI had been sharply increasing the speed of model evaluation and product development to respond to the recent AI development race. Reuters reported that OpenAI had been conducting multiple types of model evaluations simultaneously and at a rapid clip, generating so much data that employees struggled to keep up. As AI model capabilities improved quickly, the burden also grew on the systems that evaluate and monitor them.
Following this incident, OpenAI moved to bolster the security of its research and training systems, including suspending model testing for two weeks. However, it did not say when the two-week pause began.
In this context, one of OpenAI's measures is to have AI monitor AI. It had another AI system review the activities of AI agents under test. It is also using "chain-of-thought monitoring" to examine what strategies AI formulates to execute tasks. Some sensitive tasks will be carried out in an isolation environment called a stronger sandbox than before so they cannot affect external systems.
However, it is unclear whether this method can fully control AI behavior. OpenAI explained that early research indicated that even if an AI model plans to break the rules, it may not reveal that in a monitorable chain of thought. OpenAI also acknowledged that there are still unresolved issues regarding the effectiveness of chain-of-thought monitoring.
Training of the next-generation model "Astra" is also on hold. On the 7th, OpenAI said it strengthened security standards applied to its most powerful AI models and would halt training related to Astra, which has not yet met those standards. The largest training efforts OpenAI had planned are also on hold. This is in line with its preparedness framework, established to respond when an AI model potentially reaches capabilities that could pose serious risks.
OpenAI's latest slowdown contrasts with the industry's intensifying AI development race. As AI model capabilities become more advanced, it has also become important to maintain development speed while building security systems that can monitor and control unexpected behavior that may occur during testing.
OpenAI's leadership said a broader industry-wide strategy is needed to prepare for more powerful AI models to come. Reuters reported that it remains unclear whether the security enhancements OpenAI is pursuing will be sufficient to prevent unexpected AI behavior.