Skip to product information
1 of 1

Cisco

Cisco CAI-GPU-H200-NVL | H200 NVL Graphics Card, 141GB HBM3e, 600W, PCIe 5.0

SKU:CAI-GPU-H200-NVL

Stock Status: Enquire

Request Quote
Sale Sold out
Shipping calculated at checkout.

Description

The Cisco CAI-GPU-H200-NVL is an NVIDIA H200 NVL Tensor Core GPU built on the Hopper architecture, delivering exceptional performance for AI training, LLM inference, and high-performance computing workloads. Featuring 141GB of HBM3e memory with 4.8TB/s bandwidth and 16,896 CUDA cores, this PCIe 5.0 add-in card is optimized for enterprise data centers requiring powerful AI acceleration in air-cooled rack environments. With NVLink support for 2-way or 4-way GPU interconnect and multi-instance GPU capabilities, it provides scalable performance for generative AI, deep learning, and scientific computing applications.

Features

NVIDIA Hopper architecture with 4th generation Tensor Cores for accelerated AI and HPC workloads
- 141GB HBM3e memory with ECC providing 1.5x more capacity than H100 NVL and 4.8TB/s bandwidth
- 16,896 CUDA cores and 528 Tensor Cores for massive parallel processing performance
- Multi-precision compute support: FP64, FP32, FP16, BF16, TF32, FP8, and INT8 for diverse workload optimization
- Transformer Engine with FP8 precision for up to 2x faster LLM training and inference
- NVLink 4.0 support with 900GB/s per GPU bandwidth for 2-way or 4-way GPU configurations
- Multi-Instance GPU (MIG) technology supporting up to 7 isolated GPU instances
- PCIe 5.0 x16 interface delivering 128GB/s bidirectional bandwidth
- Confidential Computing support for secure AI workloads
- 600W configurable TDP (450W-600W) for flexible power management
- Passive cooling design optimized for enterprise air-cooled data center environments
- Dual-slot FHFL form factor compatible with standard PCIe servers
- Full CUDA, cuDNN, TensorRT, and NVIDIA AI software stack compatibility
- Advanced memory bandwidth optimization for reduced latency in memory-bound workloads
- Support for NVIDIA AI Enterprise software suite (included as 5-year subscription)

Warranty

All products sold by XS Network Tech include a 12-month warranty on both new and used items. Our in-house technical team thoroughly tests used hardware prior to sale to ensure enterprise-grade reliability.

All technical data should be verified on the manufacturer data sheets.

View full details

specs-tabs

Collapsible content

Technical Specifications

FAQs

Technical Specifications

GPU: NVIDIA H200 NVL, Hopper architecture (GH100)
- CUDA Cores: 16,896
- Tensor Cores: 528 (4th generation)
- Memory: 141GB HBM3e with ECC
- Memory Bandwidth: 4.8TB/s
- Memory Interface: 6144-bit
- FP32 Performance: 60 TFLOPS
- FP16/BF16 Tensor Core Performance: 1,671 TFLOPS (with sparsity)
- FP8 Tensor Core Performance: 3,341 TFLOPS (with sparsity)
- TF32 Tensor Core Performance: 835 TFLOPS (with sparsity)
- TDP: 600W (configurable 450W-600W)
- Form Factor: 2-slot Full-Height Full-Length (FHFL) PCIe card
- Interface: PCI Express 5.0 x16
- Cooling: Passive
- Interconnect: NVLink 4.0 (900GB/s per GPU, supports 2-way or 4-way bridge)
- Multi-Instance GPU: Up to 7 MIG instances
- Power Connector: 1x 16-pin (12+4 pin) PCIe
- Dimensions: 267mm (L) x 111.8mm (H) approx. (dual-slot width)
- Process Node: 4nm

FAQs

Q: What is the difference between H200 NVL and H200 SXM?
A: The H200 NVL is a PCIe add-in card with 600W TDP designed for air-cooled enterprise servers, while the H200 SXM is a 700W module requiring liquid cooling and specialized HGX baseboard systems. Both share the same 141GB HBM3e memory and GH100 silicon, but the SXM variant offers slightly higher peak compute performance due to higher power delivery.

Q: Can the H200 NVL be used in standard PCIe servers?
A: Yes, the H200 NVL is designed as a drop-in PCIe 5.0 x16 card compatible with standard enterprise rack servers. However, the passive cooling design requires adequate server airflow and chassis design to handle the 600W thermal load.

Q: What is NVLink and how does it benefit multi-GPU configurations?
A: NVLink is NVIDIA's high-speed GPU-to-GPU interconnect providing 900GB/s bandwidth per GPU, which is 7x faster than PCIe 5.0. The H200 NVL supports 2-way or 4-way NVLink bridges, enabling efficient data sharing for distributed AI training and large-scale inference workloads across multiple GPUs.

Q: What workloads benefit most from the H200 NVL's 141GB memory?
A: The H200 NVL excels at memory-intensive workloads including large language model inference and fine-tuning, generative AI applications with extended context windows, high-parameter-count model training, scientific simulations, and HPC applications requiring large working datasets.

Q: Does the H200 NVL support Multi-Instance GPU?
A: Yes, the H200 NVL supports up to 7 MIG instances, allowing a single GPU to be partitioned into multiple isolated instances with dedicated compute, memory, and bandwidth resources. This enables secure multi-tenancy and maximizes GPU utilization across diverse workloads.

Recently Viewed

  • Request a Quote

    Looking for competitive pricing? Submit a request, and our team will provide a tailored quote that fits your needs.

  • Contact Us Directly

    Have a question or need immediate assistance? Call us for expert advice and real-time support.

    Call us Now  
  • Contact Us Directly

    Have a question or need immediate assistance? Call us for expert advice and real-time support.

    Contact us