MI60 AMD | Alldatasheet

Document overview

  • Manufacturer or author: Provided By ALLDATASHEET.COM(FREE DATASHEET DOWNLOAD SITE)
  • PDF pages: 2

Technical content

AMD RADEON INSTINCT™ MI60 UNLEASH DISCOVERY ON THE WORLD'S FASTEST DOUBLE PRECISION PCIe® ACCELERA TOR1 The AMD Radeon Instinct™ MI60 accelerator based on the world's rst 7nm GPU, is equipped with optimized deep learning operations to drive the latest workows in AI and deep learning. Combined with ultra-fast 32GB HBM2 memory with up to 1 TB/s memory bandwidth, the MI60 brings customers supercharged compute capabilities to meet today’s demanding system requirements of handling large data eciently for training these complex neural networks used in deep learning. For High Performance Compute (HPC) workloads, the AMD Radeon Instinct™ MI60 accelerator delivers up to 7.4 TFLOPS of FP64 performance - the world's fastest double precision PCIe® accelerator 1, allowing scientists and researches across the globe to more eciently process HPC parallel codes across several industries including life sciences, energy, automotive and aerospace, academics, government and more. AMD’s next-generation HPC solutions are designed to deliver optimal compute density and performance per node. Coupled with the ROCm Open eCosystem, the MI60 brings a scalable HPC-class solution that gives scientists and researchers system control down to the metal. PERFORMANCE Compute Units Stream Processors Peak INT8 Peak FP16 Peak FP32 Peak FP64 Bus Interface 4,096 Up to 59.0 TOPS Up to 29.5 TFLOPS Up to 14.7 TFLOPS Up to 7.4 TFLOPS PCIe®Gen 3 and Gen 4 Capable2 MEMORY 32GB HBM2 4,096-Bits

1 GHz

Y es⁵ Linux® 64-bit Ye s Innity Fabric™ Links MxGPU technology OS Support ROCm Compatible Ye s3 Ye s4 RELIABILITY ECC (Full-chip) RAS Support Full-Height, Dual Slot 10.5" Long Passively Cooled 300W TDP Three Y ear Limited⁶ BOARD DESIGN Board Form Factor Length Thermal Max Power Warranty

The ROCm Open eCosystem, is an open-source HPC/Hyperscale-class platform for GPU computing that’s also programming-language independent. ROCm brings choice, minimalism and modular software development to GPU computing. Optimized Deep Learning Operations with exible mixed-precision FP16, FP32, INT8 & INT4 capabilities brings customers supercharged compute performance to meet today’s demanding system requirements of handling large data eciently for training complex neural networks and running inference against those networks used in deep learning. World’s Fastest Double Precision PCIe® Accelerator¹ Ultra-fast double precision performance with up to 7.4 TFLOPS FP64 performance on the Radeon Instinct™ MI60 Compute GPU, allows scientists and researches across the globe to more eciently process HPC parallel codes across several industries including life sciences, energy, nance, automotive and aerospace, academics, government, defense and more. AMD Innity Fabric™ Link T wo Innity Fabric™ Links per GPU for high speed Direct-Connect GPU hives delivering up to 200 GB/s GPU peer-to-peer bandwidth – 6x faster than using PCIe 3.0 alone.⁷ Ultra-Fast HBM2 Memory With 32 GB HBM2 memory utilizing a four-stack memory conguration, the new AMD Radeon Instinct MI60 delivers ultra-high memory bandwidth with up to 1TB/s memory bandwidth. AMD RADEON INSTINCT™ MI60 1.Calculated on Oct 22, 2018, the Radeon Instinct MI60 GPU resulted in 7.4 TFLOPS peak theoretical double precision oating-point (FP64) performance. AMD TFLOPS calculations conducted with the following equation: FLOPS calculations are performed by taking the engine clock from the highest DPM state and multiplying it by xx CUs per GPU. Then, multiplying that number by xx stream processors, which exist in each CU. Then, that number is multiplied by 1/2 FLOPS per clock for FP64. TFLOP calculations for MI60 can be found at https:/ /www.amd.com/en/products/professional-graphics/instinct-mi60 External results on the NVidia T esla V100 (16GB card) GPU accelerator resulted in 7 TFLOPS peak double precision (FP64) oating-point performance. Results found at: https:/ /images.nvidia.com/content/technologies/volta/pdf/437317-Volta-V100-DS-NV-US-WEB.pdf AMD has not independently tested or veried external/third party results/data and bears no responsibility for any errors or omissions therein. performance and features. 3.ECC support on 2nd Gen Radeon Instinct™ GPU cards, based on the “Vega 7nm” technology has been extended to full-chip ECC including HBM2 memory and internal GPU structures. 4.Expanded RAS (Reliability, availability and serviceability) attributes have been added to AMD’s 2nd Gen Radeon Instinct™ Vega 7nm technology based GPU cards and their supporting ecosystem including software, rmware and system level features. AMD’s remote manageability capabilities using advanced out-of-band circuitry allow for easier GPU monitoring via I2C, regardless of the GPU state. For full system RAS capabilities, refer to the system manufacturer’s guidelines for specic system models 5.MxGPU SR-IOV technology for hardware-based virtualized compute deployments only available on specied Radeon Instinct™ server GPU cards (MI50 and MI60) when using KVM as a hypervisor . MxGPU SR-IOV virtualization usage beyond this requires further development with AMD for deployment. Current plans for “Vega 7nm” based product support will include 1 VM per GPU up to 8 GPUs in a VM. service available in the U.S. and Canada only, email access is global. 7.As of Oct 22, 2018. Radeon Instinct™ MI50 and MI60 “Vega 7nm” technology-based accelerators are PCIe® Gen 4.0* capable providing up to 64 GB/s peak theoretical transport data bandwidth from CPU to GPU per card with PCIe Gen 4.0 x16 certied servers. Previous Gen Radeon Instinct compute GPU cards are based on PCIe Gen 3.0 providing up to 32 GB/s peak theoretical transport rate bandwidth performance. Peak theoretical transport rate performance is calculated by Baud Rate * width in bytes * # directions = GB/s per card. PCIe Gen3: 8 * 2 * 2 = 32 GB/s. PCIe Gen4: 16 * 2 * 2 = 64 GB/s. Radeon Instinct™ MI50 and MI60 “Vega 7nm” technology-based accelerators include dual Innity Fabric™ Links providing up to 200 GB/s peak theoretical GPU to GPU or Peer-to-Peer (P2P) transport rate bandwidth performance per GPU card. Combined with PCIe Gen 4 compatibility providing an aggregate GPU card I/O peak bandwidth of up to 264 GB/s. Performance guidelines are estimated only and may vary. Previous Gen Radeon Instinct compute GPU cards provide up to 32 GB/s peak PCIe Gen 3.0 bandwidth performance. Innity Fabric™ Link technology peak theoretical transport rate performance is calculated by Baud Rate * width in bytes * # directions * # links = GB/s per card. Innity Fabric Link: 25 * 2 * 2 = 100 GB/s. MI50 |MI60 each have two links: 100 GB/s * 2 links per GPU = 200 GB/s. Refer to server manufacture PCIe Gen 4.0 compatibility and performance guidelines for potential peak performance of the specied server model numbers. Server manufacturers may vary conguration oerings yielding dierent results. https:/ /pcisig.com/ , https:/ /www.chipestimate.com/PCI-Express-Gen-4-a-Big-Pipe-for-Big-Data/Cadence/T echnical-Article/2014/04/15 , https:/ /www.tomshardware.com/news/pcie-4.0-power-speed-express,32525.html AMD has not independently tested or veried external/third party results/data and bears no responsibility for any errors or omissions therein. *Pending © 2018 Advanced Micro Devices, Inc. All rights reserved. AMD, the AMD Arrow logo, Radeon, and combinations thereof are trademarks of Advanced Micro Devices, Inc. OpenCL is a trademark of Apple Inc. used by permission by Khronos. Other product names used in this publication are for identication purposes only and may be trademarks of their respective companies. ROCm Open eCosystem DEEP LEARNING AND HPC APPLICA TIONS HARDWARE ABSTRACTION OPTIMIZED ECOSYSTEM Deep Learning and HPC Libraries OPEN FRAMEWORKS Energy Life Sciences AutomotiveFinancial Services HPCCloud / Hyperscale