AMD Instinct is a brand of data center graphics processing units (GPUs) developed by AMD. Introduced in 2016 as a replacement for the FirePro S line, the Instinct product family is designed to accelerate deep learning, artificial intelligence workloads, and high-performance computing (HPC) applications. Unlike AMD's Radeon brand, which targets consumer and gaming markets, Instinct cards are optimized for compute-intensive tasks such as training neural networks and running large-scale scientific simulations. The brand directly competes with Nvidia's data center GPU offerings and Intel's Xeon Phi and Data Center GPU lines.
The product line was originally marketed as AMD Radeon Instinct, but AMD dropped the Radeon branding before the launch of the MI100 in November 2020. Since then, the Instinct series has evolved through multiple architectures, incorporating advanced packaging and memory technologies to address the growing demands of generative AI and machine learning. As of 2023, AMD-based supercomputers using Instinct GPUs have achieved leading positions on global performance and efficiency rankings.
Early Products (2016-2017)
AMD announced the first three Radeon Instinct products on December 12, 2016, and released them on June 20, 2017. Each card was based on a different GPU architecture, targeting distinct segments of the compute market.
MI6
The MI6 used a Polaris 10 chip with 16 GB of GDDR5 memory and a thermal design power (TDP) under 150 W. It delivered 5.7 TFLOPS of FP16 and FP32 performance, making it suitable primarily for inference tasks rather than training. Its peak double-precision (FP64) performance was 358 GFLOPS. The card was passively cooled, reflecting its low-power design.
MI8
The MI8 was based on the Fiji architecture, similar to the Radeon R9 Nano, with a TDP below 175 W. It featured 4 GB of High Bandwidth Memory (HBM) and achieved 8.2 TFLOPS in FP16 and FP32. Like the MI6, it was oriented toward inference workloads, with a peak FP64 performance of 512 GFLOPS.
MI25
The MI25 utilized a Vega architecture with HBM2 memory. It delivered 12.3 TFLOPS in FP32 and could reach 24.6 TFLOPS in FP16 by leveraging lower-precision arithmetic. The card had a TDP under 300 W with passive cooling and offered 768 GFLOPS of peak FP64 performance at one-sixteenth the FP32 rate. The MI25 was positioned for both training and inference, benefiting from the Vega architecture's improved compute capabilities.
MI50 and MI60 (Vega 20)
The MI50 and MI60, based on the Vega 20 variant of the Graphics Core Next (GCN) 5 architecture, were introduced as the last Instinct cards to carry the Radeon branding. They supported half-rate FP64 performance, a significant improvement over earlier models, and retained the ability to produce display output. These cards were used in early exascale and HPC deployments, bridging the gap between workstation and data center roles.
MI100 Series (CDNA 1)
The MI100 series, launched in November 2020, marked the transition to AMD's CDNA (Compute DNA) architecture, which was specifically designed for compute workloads. The CDNA 1 cards removed all rendering-related hardware, such as rasterization units, and added matrix processing units to accelerate transformer models and other deep learning operations. This architecture laid the groundwork for AMD's subsequent AI-focused accelerators.
MI300 Series (CDNA 3)
The MI300A and MI300X, introduced in late 2023, are data center accelerators built on the CDNA 3 architecture, optimized for HPC and generative AI workloads. The architecture uses a scalable chiplet design, leveraging TSMC's advanced packaging technologies, including CoWoS (chip-on-wafer-on-substrate) and InFO (integrated fan-out), to combine multiple chiplets on a single interposer. AMD's Infinity Fabric interconnect enables high-speed, low-latency data transfer between chiplets and the host system.
MI300A
The MI300A is an accelerated processing unit (APU) that integrates 24 Zen 4 CPU cores with four CDNA 3 GPU cores, totaling 228 compute units (CUs) in the GPU section and 128 GB of HBM3 memory. The Zen 4 cores, built on a 5 nm process, support the x86-64 instruction set, AVX-512, and BFloat16 extensions, allowing them to run general-purpose applications and provide host-side computation. The MI300A achieves peak performance of 61.3 TFLOPS FP64 (122.6 TFLOPS with FP64 matrix operations) and 980.6 TFLOPS FP16 (1961.2 TFLOPS with sparsity), with 5.3 TB/s of memory bandwidth. It supports PCIe 5.0 and CXL 2.0 interfaces for heterogeneous system integration.
MI300X
The MI300X is a dedicated generative AI accelerator that replaces the CPU cores with additional GPU resources, resulting in 304 CUs and 192 GB of HBM3 memory. It is designed for applications such as natural language processing, computer vision, and deep learning. The MI300X delivers 653.7 TFLOPS of TF32 performance (1307.4 TFLOPS with sparsity) and 1307.4 TFLOPS of FP16 (2614.9 TFLOPS with sparsity), with the same 5.3 TB/s memory bandwidth. It supports PCIe 5.0 and CXL 2.0, along with AMD's ROCm software stack, which provides a unified programming model for developing and deploying generative AI applications.
MI350 Series (CDNA 4)
The MI350X and MI355X, announced in 2025, are built on the CDNA 4 architecture, targeting advanced AI training and inference. Manufactured on TSMC's 3 nm (N3) process, they feature a high-performance chiplet design with 288 GB of HBM3E memory and 8 TB/s of bandwidth. CDNA 4 introduces native support for low-precision formats FP4 and FP6, in addition to FP8 and FP16, boosting FP4 compute to up to 9.2 PetaFLOPS on the MI355X. The architecture maintains Infinity Fabric interconnect for efficient data transit. In 2026, AMD released the MI350P as a PCIe card with approximately half the compute capacity and power requirements of the MI350X, offering a more accessible option for certain deployments.
Software Stack
ROCm
AMD's software ecosystem for Instinct products is organized under the Radeon Open Compute (ROCm) meta-project, which provides a collection of drivers, libraries, and tools for GPU computing. ROCm is designed to compete with Nvidia's CUDA platform, offering an open-source alternative for developers.
MxGPU
The MI6, MI8, and MI25 support AMD's MxGPU virtualization technology, which enables sharing of GPU resources across multiple users or virtual machines. This feature is valuable for cloud and enterprise environments where efficient resource utilization is critical.
MIOpen
MIOpen is AMD's deep learning library that enables GPU acceleration of neural network training and inference. It extends the GPUOpen Boltzmann Initiative software and supports popular frameworks including Theano, Caffe, TensorFlow, MXNet, Microsoft Cognitive Toolkit, Torch, and Chainer. Programming is supported in OpenCL and Python, and AMD's Heterogeneous-compute Interface for Portability (HIP) and Heterogeneous Compute Compiler allow compilation of CUDA code to run on AMD hardware, easing migration for developers.
Performance and Market Position
In June 2022, supercomputers based on AMD's Epyc CPUs and Instinct GPUs took the lead on the Green500 list of the most power-efficient supercomputers, holding the top four spots with over 50% lead over any other system. One of these, the Frontier supercomputer, has been the fastest in the world on the TOP500 list since June 2022 and remained so as of 2023. This achievement highlighted the efficiency and compute capability of Instinct accelerators in large-scale HPC environments.
The Instinct line continues to evolve in response to the rapid growth of large language models and generative AI, with each generation increasing memory capacity, bandwidth, and support for lower-precision formats to meet the demands of training and inference at scale.
See Also
- AMD - The parent company
- Intel - A competitor in data center GPUs
- TSMC - Manufacturer of advanced chips
- Artificial intelligence - Primary application domain