Skip to product information
1 of 1

Cisco

Cisco HCI-GPU-A30 | A30 24GB HBM2 GPU, PCIe 4.0, Passive, 165W

SKU:HCI-GPU-A30

Stock Status: Enquire

Request Quote
Sale Sold out
Shipping calculated at checkout.

Description

The Cisco HCI-GPU-A30 delivers NVIDIA A30 Tensor Core GPU compute acceleration for AI inference, training, and HPC workloads in mainstream enterprise servers. Built on NVIDIA Ampere architecture with 3,584 CUDA cores and 24GB HBM2 memory, it provides 933 GB/s memory bandwidth and supports Multi-Instance GPU technology to partition into up to four isolated instances. The dual-slot passive design with 165W TDP integrates seamlessly into Cisco HCI and UCS C-Series rack servers.

Features

NVIDIA Ampere Architecture with 3rd Generation Tensor Cores for accelerated AI and HPC
- 24GB HBM2 memory with ECC providing large memory capacity for complex models
- 933 GB/s memory bandwidth for fast data movement in memory-intensive workloads
- Multi-Instance GPU (MIG) technology partitions into up to 4 isolated GPU instances
- Tensor Float 32 (TF32) precision delivering up to 20x speedup over FP32 without code changes
- FP64 Tensor Cores for double-precision HPC applications
- Structural sparsity support for up to 2x performance boost on sparse models
- PCIe Gen 4.0 x16 interface delivering 64 GB/s bidirectional bandwidth
- Dual-slot passive cooling design optimized for data center servers
- Compatible with NVIDIA AI Enterprise software suite and CUDA-X libraries
- Support for VMware vSphere virtualization with vGPU and MIG
- CUDA compute capability 8.0 for broad framework compatibility
- Compatible with PyTorch, TensorFlow, RAPIDS, and NGC containers

Warranty

All products sold by XS Network Tech include a 12-month warranty on both new and used items. Our in-house technical team thoroughly tests used hardware prior to sale to ensure enterprise-grade reliability.

All technical data should be verified on the manufacturer data sheets.

View full details

specs-tabs

Collapsible content

Technical Specifications

FAQs

Technical Specifications

GPU Architecture: NVIDIA Ampere GA100
- CUDA Cores: 3,584
- Tensor Cores: 224 (3rd Generation)
- Memory: 24GB HBM2 with ECC
- Memory Bandwidth: 933 GB/s
- Memory Interface: 3,072-bit
- Peak FP32 Performance: 10.3 TFLOPS
- Peak FP64 Performance: 5.2 TFLOPS
- TF32 Tensor Core Performance: 165 TFLOPS
- System Interface: PCIe Gen 4.0 x16
- Form Factor: Dual-slot full-height full-length (FHFL)
- Cooling: Passive (requires system airflow)
- Power Consumption: 165W typical, 180W maximum
- Power Connector: 1x 8-pin PCIe
- Multi-Instance GPU (MIG): Up to 4 GPU instances
- Compute Capability: 8.0
- Certifications: Compatible with Cisco UCS and HCI platforms

FAQs

Q: What workloads is the Cisco HCI-GPU-A30 designed for?
A: The A30 GPU is optimized for AI inference at scale, AI training, high-performance computing, data analytics, and engineering simulation. It excels in workloads requiring large memory capacity and high bandwidth, including large language models, deep learning frameworks, and GPU-intensive enterprise applications.

Q: What is Multi-Instance GPU technology and how does it work?
A: Multi-Instance GPU allows the A30 to be partitioned into up to four independent GPU instances, each with dedicated high-bandwidth memory, cache, and compute cores. This enables multiple users or workloads to securely share a single GPU while maintaining hardware-level isolation, maximizing utilization in virtualized and containerized environments.

Q: Which Cisco servers support the HCI-GPU-A30?
A: The HCI-GPU-A30 is designed for Cisco HCI and UCS C-Series rack servers. GPU cards must be procured from Cisco with the required unique SBIOS ID for compatibility with CIMC and UCSM management. Consult Cisco compatibility guides for specific slot configurations and server models.

Q: Does this GPU require active cooling or special power requirements?
A: The A30 uses passive cooling and requires adequate server airflow to operate within its thermal envelope. It draws 165W typical power with a maximum of 180W TDP and connects via a single 8-pin PCIe power connector. Ensure your server chassis provides sufficient cooling and power delivery.

Q: How does the A30 compare to other data center GPUs for AI workloads?
A: The A30 offers 24GB HBM2 memory and 933 GB/s bandwidth, delivering 1.5x more memory and 3x more bandwidth than the T4 GPU. While positioned below the A100 in compute performance, the A30 provides excellent price-to-performance for mainstream AI inference and training tasks, with Tensor Core acceleration and support for TF32, FP16, INT8, and sparse model optimization.

Recently Viewed

  • Request a Quote

    Looking for competitive pricing? Submit a request, and our team will provide a tailored quote that fits your needs.

  • Contact Us Directly

    Have a question or need immediate assistance? Call us for expert advice and real-time support.

    Call us Now  
  • Contact Us Directly

    Have a question or need immediate assistance? Call us for expert advice and real-time support.

    Contact us