Skip to product information
1 of 1

Cisco

Cisco UCSX-GPU-L40S | NVIDIA L40S 48GB GDDR6 Graphics Card, 350W, PCIe Gen4

SKU:UCSX-GPU-L40S

Stock Status: Enquire

Request Quote
Sale Sold out
Taxes included. Shipping calculated at checkout.

Description

The Cisco UCSX-GPU-L40S is an NVIDIA L40S GPU accelerator designed for data center AI and graphics workloads. Built on Ada Lovelace architecture with 48GB GDDR6 memory, it delivers powerful performance for generative AI, LLM inference, 3D rendering, and video processing. This full-height full-length dual-slot card features fourth-generation Tensor Cores with FP8 support and third-generation RT Cores for advanced ray tracing capabilities.

Features

NVIDIA Ada Lovelace architecture with enhanced compute capabilities
- Fourth-generation Tensor Cores with FP8 precision support via Transformer Engine
- Third-generation RT Cores delivering 2x ray-tracing performance over previous generation
- 48GB GDDR6 memory with ECC for data integrity
- 18,176 CUDA cores for parallel processing workloads
- Up to 1,466 TFLOPS FP8 Tensor performance with structural sparsity
- PCIe Gen4 x16 interface for high-bandwidth connectivity
- NVIDIA DLSS 3 support for AI-enhanced graphics rendering
- Hardware-accelerated video encoding and decoding with AV1 support
- Passive cooling design optimized for data center rack environments
- DisplayPort outputs for direct monitor connectivity
- Secure boot with root of trust technology
- NEBS Level 3 ready for telecommunications and data center deployments
- Compatible with NVIDIA CUDA, TensorRT, and AI frameworks
- Support for NVIDIA Omniverse for metaverse and simulation applications

Warranty

All products sold by XS Network Tech include a 12-month warranty on both new and used items. Our in-house technical team thoroughly tests used hardware prior to sale to ensure enterprise-grade reliability.

All technical data should be verified on the manufacturer data sheets.

View full details

specs-tabs

Collapsible content

Technical Specifications

FAQs

Technical Specifications

GPU: NVIDIA L40S based on Ada Lovelace architecture
- Memory: 48GB GDDR6 with ECC
- Memory Bandwidth: 864 GB/s
- Memory Interface: 384-bit
- CUDA Cores: 18,176
- Tensor Cores: 568 fourth-generation with FP8 support
- RT Cores: 142 third-generation
- FP32 Performance: 91.6 TFLOPS
- TF32 Tensor Core Performance: 366 TFLOPS (733 TFLOPS with sparsity)
- Interface: PCIe Gen4 x16
- Form Factor: Full-Height Full-Length (FHFL), dual-slot
- Power Consumption: 350W TDP
- Cooling: Passive
- Display Outputs: 4x DisplayPort connectors
- Management: Compatible with Cisco UCS X-Series Modular System

FAQs

Q: What servers is the UCSX-GPU-L40S compatible with?
A: This GPU is designed for Cisco UCS X-Series Modular Systems, specifically the X410c M7 Compute Node with X440p PCIe expansion nodes. It requires appropriate riser cards and power cables that are included with the PCIe node.

Q: What workloads is the L40S GPU optimized for?
A: The L40S excels at generative AI inference, large language model fine-tuning and deployment, 3D rendering and visualization, video transcoding and streaming, virtual workstation applications, and mixed AI-graphics workloads. It's particularly effective for models up to 30-40 billion parameters.

Q: How does the L40S compare to training-focused GPUs like the H100?
A: The L40S uses GDDR6 memory with 864 GB/s bandwidth versus H100's HBM3 at 3.35 TB/s, making it more cost-effective for inference workloads at lower batch sizes. The H100 is better suited for large-scale training and high-concurrency serving, while the L40S offers excellent price-performance for inference and fine-tuning tasks.

Q: Does the L40S support virtualization?
A: Yes, the L40S supports NVIDIA vGPU software and RTX Virtual Workstation (vWS) technology, enabling GPU resources to be shared across multiple virtual machines for remote workstation users and multi-tenant environments.

Q: What precision formats does the L40S support?
A: The L40S supports FP32, TF32, FP16, BFLOAT16, INT8, and FP8 precision formats. The fourth-generation Tensor Cores with Transformer Engine can dynamically switch between FP8 and FP16 to optimize performance and memory usage for AI workloads.

Recently Viewed

  • Request a Quote

    Looking for competitive pricing? Submit a request, and our team will provide a tailored quote that fits your needs.

  • Contact Us Directly

    Have a question or need immediate assistance? Call us for expert advice and real-time support.

    Call us Now  
  • Contact Us Directly

    Have a question or need immediate assistance? Call us for expert advice and real-time support.

    Contact us