A concept rendering of an AI-produced liquid-cooling infrastructure. /Courtesy of LG CNS

As the artificial intelligence (AI) data center market grows rapidly, water-cooling technology that cools the heat of high-performance graphics processing units (GPUs) is drawing attention. Existing data centers used air-cooling technology that blows in cold air, but AI data centers that run state-of-the-art GPUs at scale are introducing liquid cooling technology that lowers heat with water because it is difficult to control heat generation with airflow alone.

According to market research firm MarketsandMarkets on the 1st, the global data center cooling market is expected to grow at an average annual rate of 16.1% from $13.23 billion (about 18.14 trillion won) this year to $37.62 billion (about 51.6 trillion won) in 2033. This is because demand for AI and cloud services is driving a surge in data center power consumption and load, increasing the importance of cooling technology.

In particular, advanced GPUs have high power density and heat generation, requiring more stable and efficient cooling technology. If equipment overheats because the heat generated by high-performance AI Semiconductor chips is not cooled in time, chip performance can deteriorate or fail.

Existing data centers were designed around air cooling that cools heat with cold air. Air cooling has the advantage of relatively easy installation and maintenance, but its cooling efficiency falls in AI data centers, where the power density per GPU rack is more than twice that of general data centers. Air cooling can generally handle power densities of 20–40 kW per rack, but the power density of server racks populated with Nvidia's next-generation "Vera Rubin" GPUs reaches 230 kW.

Liquid cooling is in the spotlight as a technology that complements the limitations of such air cooling. A representative method is "direct-to-chip (D2C) cooling," which attaches a cold plate (metal cooling plate) through which coolant flows to AI chips to absorb heat. By using liquid, which has higher heat transfer efficiency than circulating air, the industry says it can recover 65–75% of the heat generated by GPUs and Neural Processing Unit (NPU) devices.

However, because the remaining 25–35% of heat from memory, network equipment, and other components besides GPU servers must be cooled by air, AI data centers designed around liquid cooling ultimately operate in a hybrid structure that combines the two methods.

The industry also expects that "immersion cooling," which removes heat by submerging entire servers in a special coolant, will be commercialized later. For now, the high expense of special coolants and the difficulty of maintaining related equipment are obstacles, but if rack-level power usage in AI data centers rises to the stage of exceeding 1 megawatt (MW), demand for immersion cooling is also expected to spread.

Companies at home and abroad are also making preemptive investments in liquid cooling-related technology and infrastructure. LG CNS will apply "direct-to-chip cooling," a liquid cooling technology, to the Samsong data center in partnership with Naver Cloud. It is also conducting joint research on immersion cooling technology with the Korea Energy Technology Evaluation and Planning (KETEP). AI infrastructure company Elice plans to introduce to a modular data center now under construction a cooling technology that uses hot water of 40 degrees Celsius or higher, instead of chilled water, to efficiently cool the heat generated when running high-power GPU equipment.

In June, Nvidia set the standard for its next-generation Vera Rubin–dedicated AI data center as "100% water cooling," arguing, "Applying a fully liquid-cooled infrastructure can dramatically reduce power consumption." According to Nvidia, when a large-scale 50-megawatt (MW) data center switches to a liquid cooling infrastructure, it can save more than $4 million (about 5.6 billion won) annually in cooling-related energy and water expenses.

Samil PwC said in a recently published report that "the competitive landscape of the AI industry is shifting from securing GPUs to 'thermal management,'" and analyzed that "as the power density and heat generation of AI servers surge, how efficiently limited power and space are used has emerged as the factor that determines a data center's competitiveness."

※ This article has been translated by AI. Share your feedback here.