OpenAI's next-generation voice model GPT Live /Courtesy of OpenAI

OpenAI said on the 8th (local time) that it added a next-generation voice technology, "GPT-Live," to ChatGPT, enabling conversations to continue as naturally as communicating with a person.

The existing ChatGPT voice feature used a turn-based system in which the user's voice was converted to text, the AI responded, and then the user asked another question after the AI finished. If the user hesitated during a conversation, the AI sometimes took it as the end of the question and gave an irrelevant answer, and if the user asked a question while the AI was speaking, the conversation often broke off.

OpenAI said GPT-Live was designed to understand conversational context in real time, significantly reducing the limits of existing voice AI. Users can interrupt while the AI is responding to ask a different question, and even if they hesitate mid-question to organize their thoughts, the AI recognizes that the conversation is continuing and remains in listening mode.

OpenAI said, "Through sophisticated end-pointing (recognition of the end of speech) and streaming technology, we overcame the chronic limitation of voice systems: conversational dropouts."

Real-time interpretation and complex reasoning were also improved. In particular, even when asked questions that require real-time web search or advanced data analysis, it quietly delegates the task to the latest model, "GPT-5.5," without breaking the conversational flow and generates results.

Kundan Kumar, a researcher leading the audio model institutional sector at OpenAI, emphasized that they "chose safety first," noting that if a conversation heads in a dangerous direction, the system can steer away to avoid it.

GPT-Live will be rolled out sequentially to ChatGPT accounts worldwide starting that day. The paid plan will include "GPT-Live-1," and the free tier will include "GPT-Live-1 Mini" by default.

※ This article has been translated by AI. Share your feedback here.