Ampere is the codename for a graphics processing unit (GPU) microarchitecture developed by Nvidia as the successor to both the Volta and Turing architectures. It was officially announced on May 14, 2020, and is named after French mathematician and physicist André-Marie Ampère.

Nvidia announced the Ampere architecture GeForce 30 series consumer GPUs at a GeForce Special Event on September 1, 2020. Nvidia announced the A100 80GB GPU at SC20 on November 16, 2020. Mobile RTX graphics cards and the RTX 3060 based on the Ampere architecture were revealed on January 12, 2021.

Nvidia announced Ampere's successor, Hopper, at GTC 2022, and "Ampere Next Next" (Blackwell) for a 2024 release at GPU Technology Conference 2021.

Details

Architectural improvements of the Ampere architecture include the following:

  • CUDA Compute Capability 8.0 for A100 and 8.6 for the GeForce 30 series
  • TSMC's 7 nm FinFET process for A100
  • Custom version of Samsung's 8nm process (8N) for the GeForce 30 series
  • Third-generation Tensor Cores with FP16, bfloat16, TensorFloat-32 (TF32) and FP64 support and sparsity acceleration. The individual Tensor cores have with 256 FP16 FMA operations per clock 4x processing power (GA100 only, 2x on GA10x) compared to previous Tensor Core generations; the Tensor Core Count is reduced to one per SM.
  • Second-generation ray tracing cores; concurrent ray tracing, shading, and compute for the GeForce 30 series
  • High Bandwidth Memory 2 (HBM2) on A100 40GB & A100 80GB
  • GDDR6X memory for GeForce RTX 3090, RTX 3080 Ti, RTX 3080, RTX 3070 Ti
  • Double FP32 cores per SM on GA10x GPUs
  • NVLink 3.0 with a 50Gbit/s per pair throughput
  • PCI Express 4.0 with SR-IOV support (SR-IOV is reserved only for A100)
  • Multi-instance GPU (MIG) virtualization and spatial GPU partitioning feature in A100 supporting up to seven instances
  • PureVideo feature set K hardware video decoding with AV1 hardware decoding for the GeForce 30 series and feature set J for A100
  • 5 NVDEC for A100
  • Adds new hardware-based 5-core JPEG decode (NVJPG) with YUV420, YUV422, YUV444, YUV400, RGBA. Should not be confused with Nvidia NVJPEG (GPU-accelerated library for JPEG encoding/decoding)

Chips

  • GA100
  • GA102
  • GA103
  • GA104
  • GA106
  • GA107
  • GA10B

Comparison of Compute Capability: GP100 vs GV100 vs GA100

GPU featuresNvidia Tesla P100Nvidia Tesla V100Nvidia A100
GPU codenameGP100GV100GA100
GPU architecturePascalVoltaAmpere
Compute capability6.07.08.0
Threads / warp323232
Max warps / SM646464
Max threads / SM204820482048
Max thread blocks / SM323232
Max 32-bit registers / SM655366553665536
Max registers / block655366553665536
Max registers / thread255255255
Max thread block size102410241024
FP32 cores / SM646464
Ratio of SM registers to FP32 cores102410241024
Shared Memory Size / SM64 KBConfigurable up to 96KBConfigurable up to 164KB

Comparison of Precision Support Matrix

FP16FP32FP64INT1INT4INT8TF32BF16FP16FP32FP64INT1INT4INT8TF32BF16
Supported CUDA Core PrecisionsSupported Tensor Core Precisions
Nvidia Tesla P4NoYesYesNoNoYesNoNoNoNoNoNoNoNoNoNo
Nvidia P100YesYesYesNoNoNoNoNoNoNoNoNoNoNoNoNo
Nvidia VoltaYesYesYesNoNoYesNoNoYesNoNoNoNoNoNoNo
Nvidia TuringYesYesYesNoNoNoNoNoYesNoNoYesYesYesNoNo
Nvidia A100YesYesYesNoNoYesNoYesYesNoYesYesYesYesYesYes

Legend:

  • FPnn: floating point with nn bits
  • INTn: integer with n bits
  • INT1: binary
  • TF32: TensorFloat32
  • BF16: bfloat16

Comparison of Decode Performance

Concurrent streamsH.264 decode (1080p30)H.265 (HEVC) decode (1080p30)VP9 decode (1080p30)
V100162222
A10075157108

Ampere dies

DieGA100GA102GA103GA104GA106GA107GA10BGA10F
Die size826mm2628mm2496mm2392mm2276mm2200mm2448mm2?
Transistors54.2B28.3B22B17.4B12B8.7B21B?
Transistor density65.6 MTr/mm245.1 MTr/mm244.4 MTr/mm244.4 MTr/mm243.5 MTr/mm243.5 MTr/mm246.9 MTr/mm2?
Graphics processing clusters87663221
Streaming multiprocessors12884604830201612
CUDA cores819210752768061443840256020481536
Texture mapping units512336240192120806448
Render output units192112969648323216
Tensor cores512336240192120806448
RT coresN/A8460483020812
L1 cache24MB10.5MB7.5MB6MB3MB2.5MB3MB1.5MB
192KB per SM128KB per SM192KB per SM128KB per SM
L2 cache40MB6MB4MB4MB3MB2MB4MB1MB

A100 accelerator and DGX A100

The Ampere-based A100 accelerator was announced and released on May 14, 2020. The A100 features 19.5 teraflops of FP32 performance, 6912 FP32/INT32 CUDA cores, 3456 FP64 CUDA cores, 40GB of graphics memory, and 1.6TB/s of graphics memory bandwidth. The A100 accelerator was initially available only in the 3rd generation of DGX server, including 8 A100s. Also included in the DGX A100 is 15TB of PCIe gen 4 NVMe storage, two 64-core AMD Rome 7742 CPUs, 1TB of RAM, and Mellanox-powered HDR InfiniBand interconnect. The initial price for the DGX A100 was $199,000.

Starting from P100, to V100, to A100, to H100, to B200 and to R100; the comparison of accelerators used in DGX:

General & Architecture

ModelArchitectureSocketGPUFabrication ProcessTransistor count (billion)Die size (mm2)Launched
P100PascalSXM/SXM2GP100TSMC 16FF+15.3610Q2 2016
V100 16GBVoltaSXM2GV100TSMC 12FFN21.1815Q3 2017
V100 32GBSXM3
A100 40GBAmpereSXM4GA100TSMC N754.2826Q1 2020
A100 80GBQ4 2020
H100HopperSXM5GH100TSMC 4N80814Q3 2022
H200Q3 2023
B100BlackwellSXM6GB100TSMC 4NP208N/AQ4 2024
B200
R100RubinSXM7—N/aTSMC 3N338—N/aH2 2026

Cores, Clock & Power

ModelBoost clock (MHz)#SMCores (FP32 CUDA)Cores (FP64 excl. tensor)Cores (Mixed INT32/FP32)Cores (INT32)TDP (W)
P10014805635841792—N/a—N/a300
V100 16GB15308051202560N/A5120300
V100 32GB350
A100 40GB1410108691234566912N/A400
A100 80GB
H100198013216896460816896N/A700
H2001000
B100—N/a—N/a—N/a—N/a—N/a—N/a700
B200—N/a—N/a—N/a—N/a—N/a—N/a1000
R100—N/a—N/a—N/a—N/a—N/a—N/a2300

Memory & Cache

ModelMemory Type (HBM)VRAM Size (GB)Memory Speed (Gb/s)Bus width (bits)Bandwidth (TB/s)L1 Cache Per SM (KB)L1 Cache Total (KB)L2 Cache (KB)
P100HBM2161.440960.722413444096
V100 16GBHBM2161.7540960.9128102406144
V100 32GB32
A100 40GBHBM2402.451201.521922073640960
A100 80GBHBM2e803.2
H100HBM3805.251203.351922534451200
H200HBM3e1416.361444.8
B100HBM3e192881928N/AN/AN/A
B200
R100HBM4—N/a—N/a—N/a—N/a—N/a—N/a—N/a

Compute Performance, Interconnect & Networking

ModelFP32 (TFLOPS)FP64 (TFLOPS)INT8 dense tensorFP16 dense tensorbfloat16 dense tensorTF32 dense tensorFP64 dense tensorInterconnect (NVLink; TB/s)Networking
P10010.65.3—N/a21.2—N/a—N/a—N/a0.16ConnectX-4 (100 Gb/s)
V100 16GB15.77.8N/A125 TFLOPSN/AN/AN/A0.3ConnectX-5 (100 Gb/s)
V100 32GB
A100 40GB19.59.7624 TOPS312 TFLOPS312 TFLOPS156 TFLOPS19.5 TFLOPS0.6ConnectX-6 (200 Gb/s)
A100 80GB
H10067341.98 POPS990 TFLOPS990 TFLOPS495 TFLOPS67 TFLOPS0.9ConnectX-7 (400 Gb/s)
H200
B100—N/a—N/a3.5 POPS1.98 PFLOPS1.98 PFLOPS989 TFLOPS30 TFLOPS1.8ConnectX-7 (400 Gb/s)
B200—N/a—N/a4.5 POPS2.25 PFLOPS2.25 PFLOPS1.2 PFLOPS40 TFLOPS
R100—N/a—N/a—N/a—N/a—N/a—N/a—N/a—N/aConnectX-9 (1600 Gb/s)

Products using Ampere

  • GeForce MX series GeForce MX570 (mobile) (GA107)
  • GeForce 20 series GeForce RTX 2050 (mobile) (GA107)
  • GeForce 30 series GeForce RTX 3050 Laptop GPU (GA107) GeForce RTX 3050 (GA106 or GA107) GeForce RTX 3050 Ti Laptop GPU (GA107) GeForce RTX 3060 Laptop GPU (GA106) GeForce RTX 3060 (GA106 or GA104) GeForce RTX 3060 Ti (GA104 or GA103) GeForce RTX 3070 Laptop GPU (GA104) GeForce RTX 3070 (GA104) GeForce RTX 3070 Ti Laptop GPU (GA104) GeForce RTX 3070 Ti (GA104 or GA102) GeForce RTX 3080 Laptop GPU (GA104) GeForce RTX 3080 (GA102) GeForce RTX 3080 12GB (GA102) GeForce RTX 3080 Ti Laptop GPU (GA103) GeForce RTX 3080 Ti (GA102) GeForce RTX 3090 (GA102) GeForce RTX 3090 Ti (GA102)
  • Nvidia Workstation GPUs (formerly Quadro) RTX A1000 (mobile) (GA107) RTX A2000 (mobile) (GA106) RTX A2000 (GA106) RTX A3000 (mobile) (GA104) RTX A4000 (mobile) (GA104) RTX A4000 (GA104) RTX A5000 (mobile) (GA104) RTX A5500 (mobile) (GA103) RTX A4500 (GA102) RTX A5000 (GA102) RTX A5500 (GA102) RTX A6000 (GA102) A800 Active
  • Nvidia Data Center GPUs (formerly Tesla) Nvidia A2 (GA107) Nvidia A10 (GA102) Nvidia A16 (4 × GA107) Nvidia A30 (GA100) Nvidia A40 (GA102) Nvidia A100 (GA100) Nvidia A100 80GB (GA100) Nvidia A100X NVIDIA A30X
Products using Ampere (per Chip)
TypeGA10BGA107GA106GA104GA103GA102GA100
GeForce MX series—N/aGeForceMX570(mobile)—N/a—N/a—N/a—N/a—N/a
GeForce 20 series—N/aGeForceRTX2050(mobile)—N/a—N/a—N/a—N/a—N/a
GeForce 30 series—N/aGeForceRTX3050Laptop GeForceRTX3050 GeForceRTX3050TiLaptopGeForce RTX 3050 GeForceRTX3060Laptop GeForce RTX 3060GeForce RTX 3060 GeForce RTX 3060 Ti GeForce RTX 3070 Laptop GeForce RTX 3070 GeForceRTX3070TiLaptop GeForce RTX 3070 Ti GeForceRTX3080LaptopGeForce RTX 3060 Ti GeForceRTX3080TiLaptopGeForceRTX3070Ti GeForceRTX3080 GeForceRTX3080Ti GeForceRTX3090 GeForceRTX3090Ti—N/a
Nvidia Workstation GPUs—N/aRTX A1000 (mobile)RTX A2000 (mobile) RTX A2000RTX A3000 (mobile) RTX A4000 (mobile) RTX A4000 RTX A5000 (mobile)RTX A5500 (mobile)RTX A4500 RTX A5000 RTX A5500 RTX A6000—N/a
Nvidia Data Center GPUs—N/aNvidia A2 Nvidia A16—N/a—N/a—N/aNvidia A10 Nvidia A40Nvidia A30 Nvidia A100
Tegra SoCsAGXOrin OrinNX OrinNano—N/a—N/a—N/a—N/a—N/a—N/a

See also

External links