Top News

AMD Helios Top Platform Cuts AI Data Center Costs by 30%
Samira Vishwas | September 3, 2026 9:24 AM CST

AMD Helios: What AMD claims now is that the economics of AI infrastructure are changing from server TCO to rack efficiency, when power, cooling, network, memory, and software should be considered together as one system. This change is due to more and more complex AI workloads that may include even agentic AI, when multiple operations like CPU orchestration, database queries, API calls, inference, and tool execution are required, and not only GPU computations.

In their September 2, 2026 thought-leadership post “The Data Center Math Has Changed,” AMD says that enterprises should reconsider their infrastructure, taking into account consolidation, power availability, and total cost of ownership, and not just comparing processors or servers.

This new argument appears at the time when both hyperscalers and enterprises face one main limitation of AI data centers – that of power, cooling, and physical capacity, and not just the number of computations that can be purchased.

The Old Data Center Math Is Dead: Why Racks, Not Servers, Rule AI

Typical enterprise infrastructure calculations start usually with server – what kind of performance does a processor provide, how much the system costs, how many such systems are required?

AMD now claims that this kind of calculation neglects additional costs that appear elsewhere in the infrastructure.

Consolidation of servers can decrease the number of required physical systems for running the workload, thus saving rack space, electrical and cooling power, and possibly decreasing software licensing costs, which depend on sockets and/or CPU cores number. According to AMD, 5th Generation EPYC processors are designed for this model – there are possible configurations with up to 192 cores and a frequency up to 5GHz. And the upcoming 6th Gen AMD EPYC “Venice” family will allow achieving up to 256 cores.

One of the most interesting examples provided by AMD is a comparison of 100 servers with five- to seven-year-old Intel Xeon Platinum 8280 dual-CPU processors.

Representational image based on an official image | News

According to AMD, these systems can be replaced with 14 servers based on AMD EPYC 9965 processors, which will mean an 86% decrease in server count, keeping the same performance level. AMD claims that such a configuration will provide 69% less power consumption and 41% lower TCO for three years.

The important thing is that this is AMD’s modeled example and not the results of independent testing of all possible enterprise environments.

The economics matter most when software is licensed per-core or per-socket. Fewer servers and cores equals reduced licensing costs, if not for users, devices,s or other license metrics.

Agentic AI Shifts the Workload: Why CPUs and Memory Are Back in the Spotlight

The AI infrastructure discussion was fixated on accelerators for too long, but that changed with the advent of agentic AI. An AI agent is not a single inference engine. It has to reason, retrieve information, query databases, call APIs, run tools, and return results before responding to a prompt.

According to AMD, it is a full-stack workload where CPUs, accelerators, memory, and networks are critical. CPUs are becoming important for orchestration in addition to their traditional role as database servers, thus complementing GPUs, FPGAs, and other accelerators.

That is why AMD’s 6th Gen EPYC Venice, Instinct, and Pensando networking combine to create systems capable of delivering enterprise-scale AI performance. The CPU is no longer an auxiliary to the GPU. It is now part of a heterogeneous stack for building AI infrastructure.

That changes the conversation around AI accelerators. For enterprise decision-makers, the economics have shifted. The question is no longer about how many FLOPS one could get for the money. It is about how much work the infrastructure could deliver in terms of tokens per watt, tokens per dollar, rack utilization, power utilization efficiency (PUE), and other factors.

AMD Helios and MI455X: The Blueprint for Rack-Scale Performance

AMD’s approach to this challenge is embodied in its rack-scale AI platform, AMD Helios. The latest reference design, AMD Helios, has 72 AMD Instinct MI455X GPUs, 18 AMD EPYC Venice 6th Gen CPUs, and network switches from AMD Pensando. ROCm software connects and manages this large-scale system. AMD states that the design is based on open industry standards such as Open Rack Wide from Open Compute Project, UALink, and Ultra Ethernet Consortium interconnects.

Perhaps the most interesting part of the design is the memory. AMD is prioritizing memory bandwidth and capacity in MI455X, thus generating more throughput for large language models (LLMs) and reducing memory bottlenecks. Each MI455X GPU has up to 432GB of HBM4 and 23.3 TB/s of memory bandwidth, with 50% more memory capacity than the equivalent Nvidia Vera Rubin GPU.

At the rack-level, AMD Helios has 31TB of memory, 2.9 exaFLOPS of FP4 performance, 260TB/s of scale-up memory bandwidth, and 43TB/s of scale-out memory bandwidth.

It is important to note that for heterogeneous systems, memory bandwidth and capacity are often more important than raw FLOPS. They dictate how much data could be moved between accelerators and other CPUs for parallel processing.

AMD is also prioritizing economics. According to the company, its solution could deliver up to 30% more inference tokens per dollar than the leading competitor’s equivalent configuration. This figure was generated by AMD Performance Lab comparing a Helios configuration against the Nvidia Vera Rubin NVL72 using the Kimi K2 Thinking benchmark. In other words, it is not a third-party benchmark, but a comparison between two configurations at a specific workload generated by Kimi.

Why this matters for enterprise TCO

A simplified AI infrastructure calculation increasingly looks like:

Total AI cost ≈ hardware + electricity + cooling + networking + software + operations ÷ useful tokens generated

That means $/token can become more meaningful for an inference-heavy deployment than the initial purchase price of a GPU.

The Sustainability Equation: 20x Rack-Scale Efficiency by 2030

AMD’s rack-scale strategy also connects directly to its data-center sustainability targets.

The company says it is targeting a 20x improvement in rack-level energy efficiency for AI training and inference by 2030, compared with its 2024 baseline. AMD projects that approximately two 2030 AMD racks could deliver the same compute as 570 racks in 2024, with use-phase electricity consumption reduced by 20x and carbon intensity by 28x.

AMD Performance
Image Credit: Reuters / CGTN

These are roadmap projections rather than completed 2030 results. AMD says the projections are based on its silicon and system roadmap and a methodology validated by energy-efficiency expert Jonathan Koomey.

Nevertheless, the numbers illustrate why rack-scale efficiency is becoming an economic issue rather than merely a sustainability metric.

AMD vs Nvidia: The Competition Moves Beyond the GPU

Nvidia remains the dominant force in AI accelerators, but AMD’s strategy is increasingly aimed at competing at a broader infrastructure level.

Rather than positioning Instinct as an isolated GPU alternative, AMD is combining EPYC CPUs, Instinct MI455X accelerators, Pensando networking and ROCm into a single architecture.

That strategy is already being tested by major infrastructure buyers. AMD and Meta announced a partnership covering up to 6 gigawatts of AMD GPUs, with the first deployment built around an MI450-based custom GPU, 6th Gen EPYC Venice, ROCm and Helios.

OpenAI has also agreed to deploy up to 6 gigawatts of AMD GPUs, while Microsoft is expanding AMD infrastructure on Azure and plans to deploy Helios for frontier-model inference.

These partnerships do not mean Nvidia is being displaced across the market. Instead, they demonstrate that hyperscalers are increasingly interested in multiple compute architectures, supply diversification and workload-specific infrastructure economics.

For AMD, the opportunity is therefore not simply to sell another accelerator. It is to persuade customers that the entire rack should be evaluated differently.


READ NEXT
Cancel OK