Overview
NVIDIA A10 Tensor Core GPU delivers 24GB GDDR6 and 250 TOPS INT8 in a passive single-slot PCIe card designed for rack-server inference. At 31.2 TFLOPS FP32 and 125 FP16 TFLOPS with 600 GB/s bandwidth, it's positioned for cost-efficient enterprise AI inference where multiple A10 cards per server provide horizontal scaling. Widely deployed for LLM serving and real-time computer vision.
Key Features
- 24GB GDDR6 — 600 GB/s bandwidth
- 31.2 TFLOPS FP32 / 250 TOPS INT8 for inference
- Passive blower cooling — server rack compatible
- PCIe 4.0 x16, 150W TDP
- NVIDIA Ampere Tensor Cores with INT8 optimization
Ideal For
Enterprises deploying multi-GPU inference servers where cost-per-inference matters — the A10's passive cooling and INT8 acceleration make it efficient for production LLM serving and computer vision at scale.
FP32
31.2 teraFLOPS
TF32 Tensor Core
62.5 teraFLOPS | 125 teraFLOPS*
BFLOAT16 Tensor Core
125 teraFLOPS | 250 teraFLOPS*
FP16 Tensor Core
125 teraFLOPS | 250 teraFLOPS*
INT8 Tensor Core
250 TOPS | 500 TOPS*
INT4 Tensor Core
500 TOPS | 1,000 TOPS*
RT Core
72 RT Cores
Encode/decode
1 encoder2 decoder (+AV1 decode)
GPU memory
24GB GDDR6
GPU memory bandwidth
600GB/s
Interconnect
PCIe Gen4 64GB/s
Form factors
Single-slot, full-height, full-length (FHFL)
Max thermal design power (TDP)
150W
Prices may vary. Verify on vendor site.
Quick Specs
- FP32
- 31.2 teraFLOPS
- TF32 Tensor Core
- 62.5 teraFLOPS | 125 teraFLOPS*
- BFLOAT16 Tensor Core
- 125 teraFLOPS | 250 teraFLOPS*
- FP16 Tensor Core
- 125 teraFLOPS | 250 teraFLOPS*
- INT8 Tensor Core
- 250 TOPS | 500 TOPS*
- INT4 Tensor Core
- 500 TOPS | 1,000 TOPS*
