Skip to product information
1 of 1

Cisco

Cisco UCSC-GPU-L40S | L40S GPU, 48GB GDDR6, 350W, 2-Slot FHFL, PCIe Gen4

SKU:UCSC-GPU-L40S

Stock Status: Enquire

Request Quote
Sale Sold out
Taxes included. Shipping calculated at checkout.

Description

The Cisco UCSC-GPU-L40S is an NVIDIA L40S GPU accelerator designed for demanding data center workloads including generative AI, LLM inference and training, 3D rendering, and video processing. Built on the Ada Lovelace architecture with 18,176 CUDA cores, 568 fourth-generation Tensor Cores, and 142 third-generation RT Cores, it delivers exceptional performance for multi-workload environments. With 48GB GDDR6 memory, 864GB/s bandwidth, and PCIe Gen4 x16 connectivity, the L40S provides versatile compute and graphics acceleration for enterprise servers.

Features

NVIDIA Ada Lovelace architecture with 18,176 CUDA cores for exceptional parallel processing
- 48GB GDDR6 memory with ECC protection for large models and datasets
- Fourth-generation Tensor Cores with FP8 precision for accelerated AI training and inference
- Third-generation RT Cores with 212 TFLOPS RT performance for photorealistic ray tracing
- PCIe Gen4 x16 interface with 64GB/s bidirectional bandwidth for high-speed data transfer
- Triple NVENC and NVDEC engines with AV1 encoding and decoding support for video workflows
- NVIDIA DLSS 3 support with frame generation for enhanced rendering performance
- Secure Boot with Root of Trust for data center security
- NEBS Level 3 ready for telecommunications and mission-critical deployments
- Passive thermal design for integration into enterprise server environments
- vGPU software support for virtualized desktop and application delivery
- 864GB/s memory bandwidth for demanding memory-intensive workloads
- Up to 1,466 TFLOPS FP8 Tensor performance with sparsity for transformer models
- Multi-workload capability supporting AI, graphics, video, and compute simultaneously

Warranty

All products sold by XS Network Tech include a 12-month warranty on both new and used items. Our in-house technical team thoroughly tests used hardware prior to sale to ensure enterprise-grade reliability.

All technical data should be verified on the manufacturer data sheets.

View full details

specs-tabs

Collapsible content

Technical Specifications

FAQs

Technical Specifications

GPU: NVIDIA L40S based on Ada Lovelace architecture
- Memory: 48GB GDDR6 with ECC
- Memory Bandwidth: 864GB/s
- CUDA Cores: 18,176
- Tensor Cores: 568 (4th generation)
- RT Cores: 142 (3rd generation)
- Interface: PCIe Gen4 x16 (64GB/s bidirectional)
- FP32 Performance: 91.6 TFLOPS
- TF32 Tensor Core Performance: 183 TFLOPS (366 TFLOPS with sparsity)
- FP8 Tensor Core Performance: 733 TFLOPS (1,466 TFLOPS with sparsity)
- Form Factor: Dual-slot Full-Height Full-Length (FHFL)
- Display Outputs: 4x DisplayPort 1.4a
- Max Power Consumption: 350W
- Power Connector: 16-pin
- Thermal Solution: Passive cooling
- Video Encoding/Decoding: 3x NVENC / 3x NVDEC (AV1 support)
- Dimensions: 4.4" (H) x 10.5" (L)
- Security: Secure Boot with Root of Trust
- Compliance: NEBS Level 3 ready

FAQs

Q: What workloads is the UCSC-GPU-L40S optimized for?
A: The L40S is designed for generative AI, large language model (LLM) inference and training, 3D rendering, real-time ray tracing, video encoding and decoding, virtual desktop infrastructure (VDI), and NVIDIA Omniverse applications. Its balanced AI compute and graphics capabilities make it ideal for mixed-workload data center deployments.

Q: What Cisco UCS servers are compatible with this GPU?
A: The UCSC-GPU-L40S is compatible with select Cisco UCS C-Series rack servers such as the C240 M8 and UCS X-Series PCIe nodes like the X440p. It requires a dual-slot full-height full-length (FHFL) PCIe slot and 350W power delivery. Check your server's configuration guide for GPU compatibility and required accessories like air ducts and power cables.

Q: Does the L40S support GPU virtualization?
A: Yes, the L40S supports NVIDIA vGPU software, enabling multiple virtual machines to share GPU resources for virtual desktop infrastructure (VDI) and virtual workstation deployments. This allows efficient GPU utilization across virtualized environments for both AI and graphics workloads.

Q: How does the L40S compare to the previous generation L40 and A40 GPUs?
A: The L40S delivers up to 2x faster AI inference performance compared to the original L40 and up to 5x higher inference performance than the NVIDIA A40. It increases maximum power from 300W (L40) to 350W and nearly doubles TF32 Tensor Core performance, making it significantly more capable for modern AI and rendering workloads.

Q: What display connectivity does the L40S provide?
A: The L40S includes 4x DisplayPort 1.4a outputs, allowing direct monitor connectivity for visualization workloads, real-time rendering, and professional graphics applications. This makes it suitable for both headless data center compute tasks and interactive visual computing scenarios.

Recently Viewed

  • Request a Quote

    Looking for competitive pricing? Submit a request, and our team will provide a tailored quote that fits your needs.

  • Contact Us Directly

    Have a question or need immediate assistance? Call us for expert advice and real-time support.

    Call us Now  
  • Contact Us Directly

    Have a question or need immediate assistance? Call us for expert advice and real-time support.

    Contact us