Skip to product information
1 of 1

Cisco

Cisco UCSC-GPU-L4M6 | L4 24GB AI Inference GPU, 72W, Single-Slot PCIe 4.0

SKU:UCSC-GPU-L4M6

Stock Status: Enquire

Request Quote
Sale Sold out
Taxes included. Shipping calculated at checkout.

Description

The Cisco UCSC-GPU-L4M6 is an NVIDIA L4 Tensor Core GPU designed for energy-efficient AI inference, video processing, and virtual workstation deployments in Cisco UCS C-series rack servers. Built on the Ada Lovelace architecture with 24GB GDDR6 ECC memory and 72W TDP, this single-slot, low-profile GPU delivers up to 485 TOPS INT8 performance while drawing power entirely from the PCIe slot without auxiliary cables. Ideal for edge computing, AI inference at scale, video transcoding, and virtualized graphics workloads where rack space and power efficiency are critical.

Features

NVIDIA Ada Lovelace GPU architecture with 5nm process technology
- Fourth-generation Tensor Cores with FP8, FP16, BFLOAT16, TF32, and INT8 precision support
- Third-generation RT Cores for hardware-accelerated ray tracing
- 24GB GDDR6 ECC memory for data integrity in mission-critical AI workloads
- 300 GB/s memory bandwidth for high-throughput inference
- Hardware-accelerated AV1 encode/decode with 2x NVENC and 4x NVDEC engines
- NVIDIA DLSS 3 (Deep Learning Super Sampling) for graphics and visualization
- PCIe 4.0 x16 interface with 64 GB/s bidirectional bandwidth
- Single-slot, low-profile passive cooling design for maximum density
- 72W TDP enables slot-powered operation without auxiliary power cables
- Up to 120X AI video performance improvement over CPU-only solutions
- 2.5X generative AI performance versus NVIDIA T4 predecessor
- Support for NVIDIA TensorRT, CUDA, cuDNN, and AI frameworks
- Cisco CIMC and UCS Manager integration for centralized GPU monitoring
- Energy-efficient design delivers up to 99% better efficiency than CPU inference

Warranty

All products sold by XS Network Tech include a 12-month warranty on both new and used items. Our in-house technical team thoroughly tests used hardware prior to sale to ensure enterprise-grade reliability.

All technical data should be verified on the manufacturer data sheets.

View full details

specs-tabs

Collapsible content

Technical Specifications

FAQs

Technical Specifications

GPU: NVIDIA L4 Tensor Core (Ada Lovelace architecture, 5nm process)
- Memory: 24GB GDDR6 ECC
- Memory Bandwidth: 300 GB/s
- FP32 Performance: 30.3 teraFLOPS
- Tensor Core Performance: 242 teraFLOPS FP16, 485 teraFLOPS FP8, 485 TOPS INT8
- TDP: 72W (70W variant for UCS configuration)
- Form Factor: Single-slot, half-height, half-length (HHHL), low-profile
- Interface: PCIe 4.0 x16
- Cooling: Passive heatsink
- Video Encode/Decode: 2x NVENC, 4x NVDEC, 4x JPEG decoders with AV1 hardware acceleration
- Power Delivery: Slot-powered, no auxiliary PCIe power connectors required
- Compatibility: Cisco UCS C220 M6, C240 M6, C240 M8 rack servers (requires Cisco SBIOS ID)
- Dimensions: Low-profile single-slot card

FAQs

Q: Which Cisco UCS servers are compatible with the UCSC-GPU-L4M6?
A: The UCSC-GPU-L4M6 is compatible with Cisco UCS C220 M6 and C240 M6/M8 rack servers. It requires Cisco-specific SBIOS ID integration with CIMC and UCSM management, so the GPU must be procured as a Cisco part number rather than using generic NVIDIA L4 cards. Multiple L4 GPUs can be installed depending on available PCIe riser slots.

Q: What are the primary use cases for the NVIDIA L4 GPU?
A: The L4 excels at AI inference workloads including recommendation engines, natural language processing, computer vision, and generative AI at up to 2.5X the performance of the previous-generation T4. It also delivers hardware-accelerated video transcoding with AV1 codec support, virtual desktop infrastructure (VDI) for up to hundreds of concurrent users, and lightweight graphics rendering with ray tracing and DLSS 3 support.

Q: Does the UCSC-GPU-L4M6 require auxiliary power connectors?
A: No, the L4 draws its entire 70-72W power budget directly from the PCIe 4.0 x16 slot without requiring auxiliary 6-pin or 8-pin power cables. This simplifies installation and makes it ideal for dense server deployments and edge locations where power distribution and cooling capacity are limited.

Q: What AI precision formats does the L4 support?
A: The L4 supports multiple precision formats optimized for different workloads: FP32, TF32, FP16, and BFLOAT16 for training and mixed-precision inference; FP8 and INT8 with fourth-generation Tensor Cores for high-throughput inference. FP8 support delivers up to 485 teraFLOPS of AI performance, making it highly efficient for deploying large language models and transformer-based architectures.

Q: Can I install the L4 in a server that wasn't ordered as GPU-ready?
A: Yes, but retrofitting requires the UCSC-GPUKIT-240M8= GPU kit (for C240 M8 servers) which includes low-profile CPU heatsinks, GPU air duct, thermal paste, and air blockers. You only need one GPU kit per server regardless of how many L4 GPUs you add. Note that the L4 does not require the air duct that double-wide GPUs need, simplifying installation.

Recently Viewed

  • Request a Quote

    Looking for competitive pricing? Submit a request, and our team will provide a tailored quote that fits your needs.

  • Contact Us Directly

    Have a question or need immediate assistance? Call us for expert advice and real-time support.

    Call us Now  
  • Contact Us Directly

    Have a question or need immediate assistance? Call us for expert advice and real-time support.

    Contact us