The demo zone at the AMD Advancing AI 2026 event in San Francisco, United States, on the 22nd (local time)./Courtesy of Hwang Min-kyu, San Francisco correspondent

The days of tossing a single question to a chatbot and being done are over. Now artificial intelligence (AI) is evolving into "agentic AI," which plans on its own, calls up multiple tools, and reviews intermediate results, trading dozens or hundreds of exchanges even for a single request. Unlike the past, when one token could finish one question, far more tokens are now consumed even to handle a single task.

Among them, AMD took a different approach instead of boasting the performance of a single semiconductor chip. It redesigned the entire AI infrastructure—chips, servers, and the network equipment connecting them as a whole—around the benchmark of how cheaply it can produce "one token." The approach starts from the concern that no matter how fast a single chip computes, the overall expense does not fall if the server and network surrounding that chip become a bottleneck.

◇ "For AI services, the key is not 'speed' but the 'price per token'"

On the 22nd (local time) at Moscone West in San Francisco, the United States, AMD unveiled the new AI accelerator Instinct MI455X and Helios, a system that bundles it at rack scale, at its annual event AMD Advancing AI 2026. That day, AMD highlighted as Helios' greatest strength that chips, servers, and networks were designed together as a single set from the start, and emphasized how much the actual operating expense was reduced when running an entire rack end to end.

A token is the smallest unit an AI model uses to generate sentences. When you pose a question to AI services like ChatGPT or Claude, the AI generates an answer one token at a time, and the electricity and hardware expense required to produce these tokens become the operating cost for AI corporations. In particular, as approaches like agentic AI—in which AI handles tasks through multiple self-directed steps—proliferate, the number of tokens consumed to process a single request increases exponentially.

AMD also examined the structural changes underway in the AI industry that day. Andrew Dieckmann, corporate vice president in charge of AMD's data center GPU business, said, "In 2024, 'training' to teach models accounted for 60% of AI semiconductor demand, but in 2026, 'inference,' which delivers answers in actual services, has reversed to 60%," and noted, "Once a model is built well, the main job of AI semiconductors has become delivering answers every day to countless users, and in increasingly complex ways."

In this trend, it is more important how much money the entire infrastructure—chips, servers, and networks integrated as one—spends to create each token than how fast a single chip computes. This is why AMD prioritized the "rack-level economics" over individual component performance in this announcement.

◇ Redesigning beyond chips to make an entire rack a single system

The next-generation AI system Helios unveiled at the AMD Advancing AI 2026 event in San Francisco, United States, on the 22nd (local time)./Courtesy of Hwang Min-kyu, San Francisco correspondent

The concrete outcome of this approach is AMD's next-generation AI solution, Helios. AMD not only improved individual semiconductors but also designed server CPUs (central processing units), GPUs, and the network equipment connecting them together as a single set from the start. Rather than maximizing each component's performance, the approach aims to have the entire system mesh like gears to reduce wasted expense.

AMD also claimed an advantage over rival Nvidia's next-generation rack system, the Vera Rubin NVL72, in terms of performance as well as expense. According to AMD, in its own tests running a real AI model (Kimi K2 Thinking), Helios processed 10%–15% faster than the Nvidia product depending on response-time conditions. Converted into expense, AMD calculates that the same budget can produce up to 30% more tokens.

AMD said it focused on handling more computation for the same expense when designing Helios. Instead of boosting individual component performance, it aimed to raise rack-level output per expense. For AI corporations, the strategy is to help produce more results from AI services on the same budget and expand AMD's influence in the AI ecosystem.

An AMD official said, "Redesigning not individual chips but the entire infrastructure around economics ultimately reflects the judgment AI corporations make when choosing semiconductors—not 'raw compute capability' itself, but 'how cheaply can the service be operated,'" adding, "Large AI corporations such as OpenAI, Meta, Microsoft, Oracle, and Anthropic have already decided to adopt Helios because this productivity logic applied."

※ This article has been translated by AI. Share your feedback here.