AMD MI350 Instinct: 288GB HBM3e and 1.5 PFLOPS FP8 challenge NVIDIA at datacenter scale
In one sentence AMD launches the MI350 Instinct GPU with 288GB HBM3e memory, double the bandwidth of MI300X, and 1.5 PFLOPS FP8 performance, paired with ROCm 7.0 featuring significantly improved PyTorch compatibility.
Training a large AI model — the kind that powers modern chatbots or image generators — requires an almost unimaginable amount of parallel computation. That work happens inside datacenters filled with specialized graphics cards called accelerators. Until recently, if you needed to train a serious AI model, you bought NVIDIA hardware. Full stop. There was simply no credible alternative at scale.
AMD is now changing that with the MI350 Instinct. This new accelerator packs 288 gigabytes of ultra-fast HBM3e memory onto a single card — more than double what many previous solutions offered — and moves data at twice the speed of AMD's own previous flagship, the MI300X. In practice, this means you can load much larger AI models entirely into the GPU's memory without needing complex tricks to split data across dozens of cards.
The raw compute performance in FP8 — the low-precision number format preferred for fast AI training — reaches 1.5 petaflops per card. That figure puts it squarely in competition with NVIDIA's H200 and GB200 accelerators.
Arguably just as important is ROCm 7.0, the software stack that drives AMD's AI hardware. One reason organizations stuck with NVIDIA even when AMD hardware looked competitive on paper was that frameworks like PyTorch worked poorly or required significant porting effort on AMD GPUs. ROCm 7.0 closes much of that gap, bringing the developer experience meaningfully closer to what CUDA users are accustomed to.
For large research labs and enterprises spending tens of millions on AI infrastructure, a credible second supplier means leverage in negotiations and reduced single-vendor risk — which is a strategic advantage that goes well beyond raw benchmark numbers.
Companies
AMD
Tools
MI350, ROCm, PyTorch
Tags
Sources