Kakao said on the 4th that it has advanced the voice generation technology of its in-house developed omni artificial intelligence (AI) model "Kanan-a-o (Kanana-o)."
While conventional voice AI focused on reading text naturally, Kanan-a-o reflects the user's desired speaking style, emotion, and intonation. When a user gives natural-language prompts such as "Read it very fast," "Read it in a low voice," "Read it in a sad voice," or "Read it in a Gyeongsang dialect," the AI generates speech that precisely incorporates speed, volume, pitch, as well as emotion, intonation, and intensity.
In addition to role-based instructions such as "Like a sports broadcast," "Like a news anchor," and "As if reading a children's book," it can also execute composite instructions that combine multiple conditions, such as "Lower the tone and read it quickly in a sad voice."
Kakao said Kanan-a-o scored 94.50 on the InstructTTSEval Korean benchmark, which evaluates compliance with speaking instructions, surpassing GPT-4o-mini-tts (91.10).
Kanan-a-o also improved generation speed and efficiency through its in-house voice tokenizer LM-SPT (LM-aligned SPeech Tokenizer). LM-SPT is a technology that lets AI compress and represent speech with fewer tokens. It reduces the amount of data the AI must process, generating speech faster and more efficiently.
Noh Byung-seok, Kakao unified foundation model performance lead, said, "We will apply the Kanan-a-o model to a variety of services to deliver a more natural and convenient AI voice experience."