KAIST researchers, faculty startup Panesia, and Meta proposed a next-generation AI data center architecture that consolidates central processing units (CPUs), artificial intelligence (AI) accelerators, and memory as if they were a single system. It is a method that tightly consolidates the computing resources of the entire data center beyond individual server units.
KAIST said on Sept. 9 that Professor Jeong Myeong-su of the School of Electrical Engineering and researchers at Panesia published on Aug. in the international journal Nature Reviews Electrical Engineering a review paper co-authored with Meta researchers that analyzes and proposes a next-generation AI data center architecture.
With the spread of Generative AI and large-scale AI models, cases using hundreds or thousands of AI accelerators simultaneously are increasing. However, as the number of accelerators grows, data transfer latency between devices can act as a bottleneck that determines overall computing performance.
As a solution, the researchers used CXL (Compute Express Link). CXL is an interface standard that consolidates CPUs, accelerators, memory, and more at high speed and allows resource sharing.
While conventional data centers consolidate computing resources per server over a network, the method proposed by the researchers lowers the boundaries between servers and focuses on configuring CPUs, accelerators, and memory into a single consolidation domain. It is the concept of operating the entire data center like one large computing system.
Unlike existing high-speed consolidation technologies such as Nvidia NVLink or UALink, which mainly focus on accelerator consolidation within a rack, the researchers proposed expanding the scope of CXL consolidation to the entire data center, including CPUs and memory.
According to the researchers, in this architecture up to 960 accelerators can be configured in a single consolidation domain. This is about 13 times larger than current NVLink-based racks. The researchers also analyzed that data transfer latency could be reduced from the microsecond (μs, one-millionth of a second) level in existing network-based architectures to the hundreds of nanoseconds (ns, one-billionth of a second) level.
Because CPUs, accelerators, and memory can be configured as separate resources, there is potential to improve operational efficiency. If a specific device fails, only that component can be replaced, or unused compute or memory resources can be allocated to other tasks.
This paper is a review that summarizes the development status of CXL technology and presents design directions for applying it to large-scale AI data centers. Jeong's team at KAIST has researched CXL-based computer systems and semiconductor technologies and has since proceeded with semiconductor implementations of related technologies through the founding of Panesia.
Jeong Myeong-su, a KAIST professor and Panesia CEO, said, "As AI scales up, not only the performance of individual accelerators but also how quickly and efficiently accelerators and memory are consolidated is becoming more important," adding, "We proposed a direction for configuring the entire data center as if it were a single computing system based on CXL."
References
Nature Reviews Electrical Engineering (2026), DOI: https://doi.org/10.1038/s44287-026-00315-5