AMD unveiled Helios, a rack-scale AI infrastructure that optimizes the entire system—from semiconductors to networking and software—in step with the agentic artificial intelligence (AI) paradigm, along with the sixth-generation server central processing unit (CPU) Epyc Venice.
On Aug. 23 (local time) in San Francisco, AMD Chief Executive Lisa Su introduced the two products during the keynote at the annual Advancing AI 2026 event, emphasizing that in the agentic AI era, what matters is not the performance of a single chip but the performance of the entire AI infrastructure powered by that chip. She said AMD restarted the design from scratch to target an AI infrastructure with world-class efficiency.
Su said, "We moved away from the previous generation approach of optimizing GPUs and CPUs, networking, power, and cooling separately, and chose an approach that treats an entire data center as a single product," adding, "As a result, we can now supply AI chips with world-class efficiency," expressing confidence.
◇ AI accelerator market outlook tripled in a year
Su sharply raised the AI accelerator market outlook from last year. At the same event last year, AMD projected the AI accelerator market would grow to $500 billion by 2028, but this year it reset the figure to $1.4 trillion by 2030. She explained that at that scale, the AI accelerator market alone in 2030 would be comparable to the size of the entire semiconductor market today.
Behind the upward revision is the assessment that the very use of AI is changing. Su said this year is "the first year when the computing used to run already-built models globally will exceed the computing used to train AI models," adding, "About 60% of total AI computing capacity this year will be devoted to inference." As AI, which began with chatbots, has evolved over the past five to six months into agents that plan, use tools, and complete tasks on their own, computing demand is not climbing gradually but leaping in steps, she noted.
She also noted that behind this surge in demand lies a real burden on corporations. Su asked rhetorically, "Aren't all of us on the front lines feeling our AI budgets rising every month?" and stressed that ultimately the key is not performance but cost per token.
◇ The design unit is not a single chip but the entire rack
This shift in demand structure has had a major impact on how AMD designs hardware. In the past, the focus was on making a single high-performance chip; now, the entire rack—including CPUs, GPUs, and networking—has to be designed as one. Meta's Santosh Janardhan, head of infrastructure and a guest on stage, said, "The era of optimizing servers, networking, power, and cooling separately is over, and we have entered a phase where we must design the entire data center as one system together with partners."
Helios, which officially entered mass production that day, is the first result of this shift. It bundles the Instinct MI455X AI accelerator, the sixth-generation server CPU Epyc Venice, and the Pensando networking chip that handles in-rack communications into a single rack, featuring a design that combines GPUs, memory, power delivery, and cooling into one module. Alongside, AMD released performance figures showing it leads in memory capacity and bandwidth compared with rival Nvidia's latest rack-scale product, Vera Rubin (NVL72).
Su said flatly, "Simply put, Helios is the world's best AI rack," and reiterated, "What customers ultimately gain is expense savings, because they can serve more users with the same investment."
◇ "GPUs alone aren't enough"… CPU importance comes to the fore
Su also noted that agentic AI does not run on GPU compute alone. "For an agent to carry out a task, it must go through dozens of steps of inference, tool calls, and data access, and coordinating this entire process falls to the CPU," she said. She explained that demand for server CPUs, overshadowed by the growth of the GPU market, is in fact a completely new growth axis created by agentic AI.
The sixth-generation Epyc server CPU Venice, whose mass production was released that day, features a single architecture segmented by workload. On a single design foundation, AMD implemented a CPU that feeds data to GPUs, a high-density CPU that runs thousands of agents simultaneously, and a general-purpose CPU that handles corporations' routine tasks, each differently.
Su emphasized that this segmentation strategy translated into real performance gaps. Compared with competing x86 processors, the CPU for agent sandboxes can handle more than twice as many agents per watt, and compared with competing ARM-based processors, it widened the rack-level performance per watt gap by up to more than three times. "At data center scale, such gaps ultimately translate into how many agents you can run simultaneously on the same power," she said.
◇ Major customers such as OpenAI, Anthropic, and Meta add support
Representatives from Anthropic, OpenAI, Meta, and Cerebras took the stage in turn to outline how they use AMD's AI systems and their collaborations. OpenAI Vice President of Compute Strategy Sachin Katti said, "The previously announced contract of up to 6 GW is progressing smoothly." He also said, "Starting with AMD's previous-generation GPUs, this time we were the first in the industry to receive Helios racks and are already using them to train large-scale models."
Meta's Santosh Janardhan said, "Over multiple generations, we have used AMD server CPUs at a scale of millions," adding, "We now have such a close relationship that, rather than buying finished goods, we sit down with AMD from the earliest design stages to think together about power, cooling, servers, and communications equipment."
Cerebras CEO Andrew Feldman introduced the company's business of making ultra-large AI chips supplied on-premises and in the cloud, saying it has secured many large customers such as OpenAI and Amazon Web Services (AWS). "Until now, the industry had to choose between 'processing large volumes quickly' and 'delivering instant answers,' but by teaming up with AMD this time, we were able to increase throughput fivefold while keeping response times the same," he said.