Skip to product information
1 of 1

Cisco

Cisco CAI-GPU-L40S | L40S 48GB GDDR6 GPU, 350W TDP, PCIe 4.0 x16, Passive

SKU:CAI-GPU-L40S

Stock Status: Enquire

Request Quote
Sale Sold out
Shipping calculated at checkout.

Description

The Cisco CAI-GPU-L40S is a high-performance data center GPU accelerator built on NVIDIA Ada Lovelace architecture, designed for generative AI workloads, large language model inference and training, 3D graphics rendering, and video processing. Featuring 48GB of GDDR6 ECC memory with 864 GB/s bandwidth and 350W TDP, this passively cooled, 2-slot FHFL form factor GPU delivers breakthrough multi-workload performance for enterprise AI and visualization applications. Optimized for 24/7 data center operations in Cisco UCS C-Series and X-Series servers.

Features

NVIDIA Ada Lovelace AD102 GPU architecture with 18,176 CUDA cores
- 48GB GDDR6 ECC memory for data integrity in mission-critical applications
- Fourth-generation Tensor Cores with Transformer Engine and FP8 precision
- Hardware-accelerated ray tracing with third-generation RT Cores
- Support for structural sparsity and optimized TF32 format for AI training
- NVIDIA DLSS 3 for AI-enhanced graphics and resolution upscaling
- Multi-stream video encoding and decoding engines (NVENC/NVDEC)
- PCIe Gen4 x16 interface for high-bandwidth host connectivity
- Passive thermal design for reliable 24/7 data center operation
- Secure boot with hardware root of trust for enhanced security
- NEBS Level 3 ready for telecommunications and critical infrastructure
- 4x DisplayPort outputs for multi-monitor visualization workflows
- Support for NVIDIA Omniverse, RTX Virtual Workstation (vWS), and AI Enterprise
- Compatible with CUDA, cuDNN, TensorRT, and NVIDIA GPU Cloud (NGC) containers
- 2-slot FHFL form factor for high-density multi-GPU server configurations

Warranty

All products sold by XS Network Tech include a 12-month warranty on both new and used items. Our in-house technical team thoroughly tests used hardware prior to sale to ensure enterprise-grade reliability.

All technical data should be verified on the manufacturer data sheets.

View full details

specs-tabs

Collapsible content

Technical Specifications

FAQs

Technical Specifications

GPU: NVIDIA L40S based on Ada Lovelace architecture (AD102)
- Memory: 48GB GDDR6 with ECC
- Memory Bandwidth: 864 GB/s
- Interface: PCIe Gen4 x16
- Form Factor: Full-Height, Full-Length (FHFL), dual-slot, 10.5 inch
- Thermal Design: Passively cooled
- Maximum Power (TDP): 350W
- FP32 Performance: 91.6 TFLOPS
- Tensor Core Performance: Fourth-generation Tensor Cores with Transformer Engine and FP8 support
- Display Outputs: 4x DisplayPort connectors
- Dimensions: 10.51 x 4.37 x 1.37 inches (26.7 x 11.1 x 3.5 cm)
- Weight: Approximately 3.74 lbs (1.7 kg)
- Compliance: NEBS Level 3 ready, secure boot with root of trust

FAQs

Q: Which Cisco UCS servers support the CAI-GPU-L40S?
A: The CAI-GPU-L40S is compatible with Cisco UCS C845A M8 AI rack servers and can be deployed in configurations supporting 2, 4, 6, or 8 GPUs. It is also supported in Cisco UCS X-Series modular systems via the X440p PCIe Node with X-Fabric technology, enabling PCIe Gen4 GPU fabric connectivity.

Q: What workloads is the L40S GPU optimized for?
A: The L40S is optimized for generative AI and large language model (LLM) inference and training, 3D graphics rendering, video encoding and transcoding, virtual desktop infrastructure (VDI), AI-enhanced content creation, scientific visualization, and mixed AI/graphics pipelines in enterprise data center environments.

Q: Does the L40S require active cooling or special power considerations?
A: The L40S features passive cooling and requires adequate server chassis airflow for thermal management. With a 350W TDP, ensure your server power supply and PCIe slot support the full power delivery requirements. Cisco UCS servers designed for GPU workloads provide appropriate airflow and power infrastructure.

Q: How does the L40S compare to the previous generation L40?
A: The L40S delivers nearly double the TF32 Tensor Core TFLOPS and FP16 Tensor Core performance compared to the L40. It increases maximum power from 300W to 350W and adds fourth-generation Tensor Cores with Transformer Engine support and FP8 precision for enhanced AI inference and training performance.

Q: What software and virtualization platforms are supported?
A: The L40S supports NVIDIA vGPU software for virtual desktop infrastructure (VDI) and application virtualization, NVIDIA AI Enterprise for production AI workloads, and is compatible with major AI frameworks including PyTorch, TensorFlow, and NVIDIA NeMo. It integrates with VMware vSphere, Citrix Virtual Apps and Desktops, and other enterprise virtualization platforms.

Recently Viewed

  • Request a Quote

    Looking for competitive pricing? Submit a request, and our team will provide a tailored quote that fits your needs.

  • Contact Us Directly

    Have a question or need immediate assistance? Call us for expert advice and real-time support.

    Call us Now  
  • Contact Us Directly

    Have a question or need immediate assistance? Call us for expert advice and real-time support.

    Contact us