Skip to product information
1 of 1

Cisco

Cisco UCSC-GPU-H200-NVL | H200 NVL GPU, 141GB HBM3e, 4.8TB/s, 600W PCIe 5.0

SKU:UCSC-GPU-H200-NVL

Stock Status: Enquire

Request Quote
Sale Sold out
Taxes included. Shipping calculated at checkout.

Description

The Cisco UCSC-GPU-H200-NVL is an NVIDIA H200 NVL Tensor Core GPU built on the Hopper architecture, delivering 141GB of HBM3e memory with 4.8TB/s bandwidth for AI inference, generative AI, and high-performance computing workloads. This PCIe 5.0 accelerator features 600W TDP, NVLink support for multi-GPU configurations, and fits standard air-cooled enterprise rack designs. Ideal for large language model training, deep learning, scientific simulations, and memory-intensive AI applications requiring massive memory capacity and throughput.

Features

141GB HBM3e memory with 4.8TB/s bandwidth - 1.4x more bandwidth and nearly double the capacity of H100
- NVIDIA Hopper architecture with 4th generation Tensor Cores for advanced AI acceleration
- 16,896 CUDA cores and 528 Tensor Cores for massive parallel processing
- PCIe 5.0 x16 interface with 128GB/s host bandwidth
- NVLink 4.0 support with 900GB/s bidirectional bandwidth per GPU for multi-GPU scaling
- Multi-Instance GPU (MIG) technology supporting up to 7 isolated GPU instances
- Support for FP64, FP32, TF32, FP16, BF16, FP8, and INT8 precision formats
- 3,341 TFLOPS FP8 Tensor Core performance with sparsity for AI inference
- 50MB L2 cache for improved data locality and performance
- 600W TDP with configurable power profiles (450W-600W) for flexible deployment
- 2-slot FHFL form factor compatible with standard enterprise servers
- Passive cooling design for air-cooled data center environments
- Compatible with NVIDIA CUDA, cuDNN, TensorRT, and AI Enterprise software stack
- Advanced memory bandwidth optimization for memory-bound AI workloads
- Support for long-context inference with extended KV cache capacity

Warranty

All products sold by XS Network Tech include a 12-month warranty on both new and used items. Our in-house technical team thoroughly tests used hardware prior to sale to ensure enterprise-grade reliability.

All technical data should be verified on the manufacturer data sheets.

View full details

specs-tabs

Collapsible content

Technical Specifications

FAQs

Technical Specifications

GPU: NVIDIA H200 NVL Tensor Core (Hopper GH100 architecture)
- Memory: 141GB HBM3e
- Memory Bandwidth: 4.8TB/s
- CUDA Cores: 16,896
- Tensor Cores: 528 (4th generation)
- FP8 Performance (with sparsity): 3,341 TFLOPS
- TF32 Tensor Core Performance (with sparsity): 835 TFLOPS
- FP64 Performance: 60 TFLOPS
- TDP: 600W (configurable 450W-600W)
- Form Factor: 2-slot Full-Height Full-Length (FHFL) PCIe card
- Interface: PCIe 5.0 x16
- NVLink: 900GB/s bidirectional per GPU (NVLink 4.0, supports 2-way or 4-way configurations)
- Multi-Instance GPU (MIG): Up to 7 instances per GPU
- L2 Cache: 50MB
- Base Clock: 1,365MHz / Boost Clock: 1,785MHz
- Memory Clock: 1,313MHz (effective 5,252MHz)

FAQs

Q: What types of workloads is the H200 NVL optimized for?
A: The H200 NVL is optimized for AI training and inference, large language models like GPT and Llama, generative AI applications, scientific simulations, high-performance computing, and memory-intensive deep learning workloads. Its 141GB HBM3e memory enables processing of massive datasets and large-scale models that exceed traditional GPU memory limits.

Q: How does the H200 NVL differ from the H200 SXM variant?
A: The H200 NVL is a PCIe 5.0 add-in card with 600W TDP, designed for air-cooled enterprise racks and standard server hardware. The H200 SXM5 variant uses SXM5 packaging with higher TDP, typically requires liquid cooling, and is optimized for HGX/DGX systems with NVSwitch fabric. Both share the same GH100 die and 141GB HBM3e memory subsystem.

Q: Can the H200 NVL be used in multi-GPU configurations?
A: Yes, the H200 NVL supports NVLink bridge connectors enabling 2-way or 4-way GPU configurations with 900GB/s bidirectional bandwidth per GPU. This allows up to four GPUs to work together for accelerated large language model inference and HPC applications, providing up to 1.7x faster LLM performance and 1.3x better HPC performance compared to H100 NVL configurations.

Q: Is the H200 NVL compatible with existing H100 or A100 infrastructure?
A: Yes, the H200 NVL PCIe card maintains backward compatibility and is a drop-in replacement for H100 and A100 GPUs in existing systems. It uses the same PCIe 5.0 interface and works with NVIDIA's full software stack including CUDA, cuDNN, and TensorRT, enabling seamless upgrades without infrastructure changes.

Q: What cooling requirements does the H200 NVL have?
A: The H200 NVL is designed for air-cooled enterprise rack environments and operates at 600W TDP with configurable power profiles down to 450W for constrained thermal environments. This makes it suitable for standard data center deployments with proper ventilation, unlike the SXM5 variant which typically requires liquid cooling.

Recently Viewed

  • Request a Quote

    Looking for competitive pricing? Submit a request, and our team will provide a tailored quote that fits your needs.

  • Contact Us Directly

    Have a question or need immediate assistance? Call us for expert advice and real-time support.

    Call us Now  
  • Contact Us Directly

    Have a question or need immediate assistance? Call us for expert advice and real-time support.

    Contact us