AMD has unveiled Helios, a rack-scale AI system designed to make 72 Instinct MI455X GPUs operate as a single coherent compute unit—backed by 31 terabytes of HBM4 memory and an advertised aggregate bandwidth of about 260 TB/s.
Multiple specialized outlets describe the goal in blunt terms: offer hyperscalers an integrated alternative to Nvidia-dominated training and inference racks by bundling GPU, CPU, and networking into one platform. Microsoft is among the first named customers, with deployments expected in Azure in the second half of 2026.
Helios aims to make 72 GPUs behave like one giant accelerator
The core pitch behind Helios is architectural: AMD says it can run 72 Instinct MI455X GPUs as a coherent whole for large-scale AI training and inference. In coverage and AMD presentations cited by the source article, Helios is framed as AMD’s largest single AI compute system yet—built around the idea of treating a rack like one GPU.
The bet is that this reduces the operational and software complexity that shows up when models sprawl across many nodes, where parameter exchanges and synchronization can become bottlenecks.
Memory is central to the positioning. Helios is paired with 31 terabytes of HBM4 that’s described as coherent across the rack, a figure highlighted as a key differentiator versus competing racks. For modern training workloads, memory capacity and bandwidth can dictate batch sizes, parallelization strategies, and how much data must be offloaded to system memory.
Sources cited in the French article emphasize full rack-scale coherence—something that, if it holds up in production, could reduce costly software workarounds such as heavy model fragmentation or more aggressive activation checkpointing.
AMD is also touting an aggregate bandwidth figure of roughly 260 TB/s, aimed at keeping compute units fed so data movement doesn’t become the limiting factor. As FP4 and FP8 formats become more common in inference and training, pressure often shifts toward memory, interconnects, and collective-communication efficiency—areas Helios is positioned to address.
For data center operations teams, the “rack as one system” approach is meant to translate into fewer logical nodes to manage, clearer resource allocation, and a more stable target for AI frameworks mapping large training runs onto shared infrastructure. The source article notes AMD hasn’t publicly detailed every integration parameter in the reporting it cites, but the intent is to feel closer to one very large accelerator than a fragmented cluster.
EPYC “Venice” CPUs and Pensando networking are built into the platform
Helios isn’t presented as a simple GPU pile-up. It’s positioned as an integrated platform combining sixth-generation EPYC “Venice” CPUs with Pensando networking components. The strategy reflects how large-scale AI depends not just on raw accelerator throughput, but on how efficiently data reaches the GPUs.
CPUs still orchestrate parts of the input pipeline, preprocessing, storage interaction, and control tasks. In modern stacks, the stability and availability of those components can directly affect GPU utilization—one of the biggest drivers of real training cost.
The inclusion of Pensando hardware puts the network front and center. With 72 GPUs working together, latency, congestion, and collective-operation quality become critical. AMD is trying to sell an end-to-end solution where interconnect design is treated as a first-class element alongside compute.
Hardware footprint is also part of the story: coverage describes Helios as a rack that’s double the width of a conventional rack. That matters for operators because it can impose constraints on aisle layout, density, cooling, and power distribution—more like planning for industrial equipment than simply adding another server cabinet.
Performance figures reported in the article put Helios at up to 2.9 exaflops for FP4 inference and 1.4 exaflops for FP8 training. The emphasis reflects a broader shift in how vendors market AI systems—highlighting reduced-precision formats that map more directly to real AI workloads rather than FP32 peak numbers.
AMD is positioning Helios against Nvidia’s Vera Rubin NVL72
Helios is described as AMD’s direct answer to Nvidia’s rack-scale approach, with explicit comparisons to Vera Rubin NVL72. In this segment, the fight isn’t just about marginal performance gains—it’s about shipping racks in volume, meeting delivery timelines, ensuring parts availability, and providing a coherent software stack.
The source article points to memory as a key differentiator: Helios is described as matching the 72-GPU form factor while offering 1.5 times the memory. If that holds in final configurations, it could matter for workloads that push capacity limits, such as training very large multimodal models or running inference with extended context windows.
The reported 31 terabytes of coherent HBM4 also signals a software ambition: make memory distribution more transparent. In practice, frameworks must navigate complex topologies, choose communication strategies, and adapt parallelization. A system designed to minimize communication penalties can improve scaling efficiency—and even small utilization gains can translate into major annual operating savings at rack scale.
Still, the competitive backdrop remains tight. Enterprise GPUs sell quickly, and buyers often decide based on availability and ecosystem maturity, not just spec sheets. Helios is AMD’s attempt to prove it can deliver a complete platform that’s credible against the market’s reference point.
Microsoft is expected to deploy Helios in Azure in the second half of 2026
One of the most consequential details in the reporting cited by the French article is Microsoft’s role: it’s presented as a Helios customer, with deployments expected in Azure starting in the second half of 2026. For AMD, a hyperscaler reference can matter beyond sheer unit volume—it serves as platform validation, reassures integrators, and can accelerate interest from other large buyers.
For Microsoft, the logic is described as twofold: diversify suppliers and secure access to compute capacity. Cloud AI has become a race for availability, and the ability to offer high-performance instances depends heavily on supply chains and delivery cycles. Adding Helios to Azure could improve resilience and provide negotiating leverage—an approach commonly used by hyperscalers to avoid single-vendor dependence in strategic categories.
Running Helios in a public cloud also raises the bar: multi-tenant requirements, isolation, observability, billing, and integration into managed services. A fast machine in a lab isn’t enough; it needs allocation tools, access controls, update mechanisms that avoid major disruption, and operational procedures. The article notes that while implementation details aren’t public, Microsoft’s mention as a customer implies those issues are being treated seriously.
What remains unclear is the scale and pacing. “Second half of 2026” leaves room for pilots, limited-region rollouts, and then broader deployments. In projects like this, ramp-ups often happen in stages—validating performance, qualifying monitoring pipelines, and adapting internal software—before a platform becomes widely visible to cloud customers.
FAQ
What is AMD Helios? Helios is AMD’s rack-scale AI system that groups 72 Instinct MI455X GPUs and is designed to run them as a coherent unit, with large HBM4 memory and an interconnect aimed at large-scale training and inference.
What key specs are being reported? The cited sources report 72 GPUs, 31 terabytes of rack-coherent HBM4, and about 260 TB/s of aggregate bandwidth, plus up to 2.9 exaflops FP4 (inference) and 1.4 exaflops FP8 (training).
Why does 31 TB of HBM4 matter? Very large HBM4 capacity can enable larger models and reduce complex partitioning and back-and-forth to other memory tiers, potentially improving efficiency and simplifying deployment depending on the workload.
Is Helios aimed at a specific Nvidia system? Yes. It’s positioned as a direct response to Nvidia’s rack-scale AI systems, including Vera Rubin NVL72.
When is Microsoft expected to use Helios in Azure? The reporting cited says Microsoft expects Helios deployments in Azure starting in the second half of 2026; exact volumes and ramp timing aren’t publicly detailed.
Key takeaways
• AMD Helios groups 72 Instinct MI455X GPUs into a coherent rack-scale AI system.
• The platform highlights 31 TB of coherent HBM4 and about 260 TB/s of aggregate bandwidth.
• Reported performance targets reach up to 2.9 exaflops FP4 inference and 1.4 exaflops FP8 training.
• Helios is positioned against Nvidia’s integrated AI racks, including Vera Rubin NVL72.
• Microsoft is cited as an early customer, with Azure deployments expected in the second half of 2026.
Sources
Techzine.eu; MLQ.ai; Mark Thiele (LinkedIn); L’Usine Digitale; Justo Global
Key Takeaways
- AMD Helios packs 72 Instinct MI455X GPUs into a unified AI rack
- The system highlights 31 TB of HBM4 and bandwidth of around 260 TB/s
- Figures of up to 2.9 FP4 exaflops and 1.4 FP8 exaflops are reported
- Helios is positioned against Nvidia AI racks, including the Vera Rubin NVL72
- Microsoft plans a rollout in Azure in the second half of 2026
Frequently Asked Questions
What is Helios at AMD?
Helios is a rack-scale AI system introduced by AMD that combines 72 Instinct MI455X GPUs with an architecture designed to make them operate as a cohesive unit, featuring massive HBM4 memory and an interconnect built for large-scale training and inference.
What key figures have been announced for Helios?
The cited sources mention 72 GPUs, 31 TB of rack-scale coherent HBM4, and aggregate bandwidth of around 260 TB/s. Performance figures of up to 2.9 exaflops in FP4 (inference) and 1.4 exaflops in FP8 (training) are also reported.
Why is 31 TB of HBM4 memory important?
A very large HBM4 capacity can enable larger models, reduce the need for complex partitioning, and limit data movement to other memory tiers. For AI workloads, this can improve efficiency and simplify deployment depending on the workload.
Is Helios targeting a specific competitor?
Yes. Helios is positioned as a direct response to Nvidia’s AI racks, particularly Vera Rubin NVL72, in the segment of integrated systems designed for large-scale training and inference.
When does Microsoft plan to use Helios in Azure?
According to information reported in the press, Microsoft is planning Helios deployments in Azure starting in the second half of 2026. The volumes and exact ramp-up pace have not been publicly detailed.



