Kakao unveiled an adaptive tokenization technology that makes AI video generation 3.2 times faster.

Kakao said on the 9th that at ECCV 2026, Lee Yeon-gyeong, a Kakao research engineer, will present a paper introducing an adaptive tokenization technology that addresses the long-standing challenges of massive computation and high calculation expense in AI video generation models.

/Courtesy of Kakao

ECCV is considered one of the world's three major Computer Vision conferences along with CVPR (Conference on Computer Vision and Pattern Recognition) and ICCV (International Conference on Computer Vision). This year, it is being held in Sweden through the 12th.

Recently in the AI industry, as research expands to high-resolution and long-duration video generation, cutting the exponentially growing calculation expense has emerged as a key task. Conventional video compression (tokenization) compressed data only at a fixed ratio regardless of content, so it generated unnecessarily many tokens even in scenes with static backgrounds or little motion, consuming compute resources of video generation models.

Focusing on the high redundancy between video frames, Kakao researched an adaptive tokenization technology called "KATok (Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation)," in which AI autonomously judges and adjusts the number of tokens needed for computation based on the spatiotemporal complexity of the video. It allocates more tokens to complex scenes with large motion and high information density, while removing redundant tokens in static or simple regions. Even without a person manually specifying token counts, the AI determines the necessary computation based on content and allocates resources far more efficiently.

/Courtesy of Kakao

Experiments showed that KATok reconstructs high-quality videos with significantly fewer tokens than conventional methods, demonstrating outstanding data compression performance. Applied across the entire video generation pipeline, the optimized token representation improved training speed of video generation models by about 6.9 times and generation speed by about 3.2 times compared with previous approaches.

Choi Dong-jin, Kakao's applied AI model performance lead, said, "The more computation surges, as with high-resolution and long-duration videos, the greater the compute-saving effect of KATok's adaptive tokenization will be," adding, "We will verify its generality with various AI video generation models and develop it into a core foundational technology for large-scale video generation that reduces computation while maintaining excellent image quality."

※ This article has been translated by AI. Share your feedback here.