Major AI corporations are pushing to advance voice AI technology on the view that next-generation AI-based devices succeeding smartphones will be operated not by tapping screens with a fingertip but by people's voices. The goal is to make the experience of interacting with AI feel as natural as talking with a real person.
OpenAI, the developer of ChatGPT, is developing a dedicated AI device with a target release next year. According to the AI industry on the 10th, the AI device OpenAI is preparing with Jony Ive, who oversaw Apple's product design, is a donut-shaped portable smart speaker and is expected to become the first hardware product that embodies ChatGPT in a physical form. Equipped with a camera and sensors, the device can recognize its surroundings and can be operated by voice. It is designed to learn from conversations with users and provide increasingly personalized service.
Meta, Facebook's parent company, has singled out AI-based smart glasses as a next-generation device to replace smartphones and has released related products since 2021, all of which run by voice. When a user says "Hey Meta" and asks a question out loud, the AI built into the glasses recognizes it and responds. Google and Apple are also set to launch AI smart glasses.
Corporations are focusing on implementing natural Conversational AI so that AI devices touted as the "next-generation smartphone" can quickly permeate users' daily lives. Until now, voice AI had drawbacks such as a beat-late response, robotic and awkward speech, and poor understanding of context.
GPT-Realtime-2, the latest voice AI model unveiled by OpenAI in May this year, is a voice AI model that can handle complex requests based on GPT-5-level reasoning. Its ability to process and generate voice input in real time sets it apart from conventional voice AI.
Conventional voice AI went through multiple steps from "user voice recognition → speech-to-text (STT) conversion → large language model (LLM)-based reasoning → generating a text answer → text-to-speech (TTS)." As a result, response speeds were slow, and conversations often felt disjointed because users and AI waited for each other to finish speaking in a turn-based format.
GPT-Realtime-2 addresses these issues by applying a full-duplex method that handles speaking and listening in real time. It is designed to respond instantly even if a user interrupts while the AI is speaking to ask a new question or to correct something said earlier, enabling communication as natural as talking with a real person. OpenAI's advanced model assesses the conversation multiple times per second and decides in real time whether to keep speaking, listen, pause, or call a tool.
An OpenAI official said, "More than 150 million people each week use ChatGPT's voice conversation and dictation features," adding, "We are advancing real-time voice AI technology so it can go beyond simple Q&A to listen according to the flow of conversation, reason, translate, take dictation, and perform tasks."
Gemini 3.1 Flash Live, Google's flagship voice AI model, is designed to recognize fine-grained elements such as a user's intonation, speed, pitch, and laughter and continue reasoning and conversation based on them. The company said affective dialog is possible, where the AI detects when a user speaks in an angry tone and changes how it responds.
Gemini 3.1 Flash Live also operates based on an audio-to-audio technology that processes voice in real time without going through speech-to-text conversion, and features multimodal capabilities that recognize and process images, text, and video along with voice.
Meta, Facebook's parent company, has expanded related investments, including acquiring voice AI-focused startups, to strengthen the Conversational AI technologies to be applied to its AI glasses products. In July last year, Meta acquired PlayAI, a startup that creates human-like voices with AI-based speech synthesis technology, and in August it bought WaveformsAI, which develops technology that understands human emotions and reflects them in voice.
With expectations that voice AI will go beyond limited uses such as "AI agents" at corporations' call centers and be embedded in various wearable devices such as smart speakers and smart glasses, funding is also pouring into startups focused solely on voice AI. Five promising voice AI startups—ElevenLabs, Deepgram, Hume AI, Cartesia, and Cesami—have raised more than $1.5 billion (about 2 trillion won) in cumulative funding in recent years.