NVIDIA GPUDirect Remote Direct Memory Access for Cloudera AI
NVIDIA GPUDirect Remote Direct Memory Access (RDMA) enables applications to directly access remote GPU memory without routing data through the operating system kernel. By minimizing CPU involvement in network transfers, GPUDirect RDMA reduces latency and overhead while freeing CPU resources for application processing.
This capability is especially beneficial for distributed GPU workloads such as large-scale model training, high-throughput data processing, High-Performance Computing (HPC), and machine learning pipelines. In NVIDIA Hyperscale Graphics eXtension (HGX) and NVIDIA Deep Learning GPU Xceleration (DGX)-based clusters, GPUDirect RDMA improves node-to-node communication performance for workloads that span multiple GPUs and multiple hosts.
Designed for high-bandwidth, low-latency communication, GPUDirect RDMA is well suited for AI, ML, and big data workloads that require efficient GPU-to-network data transfer.
You can configure NVIDIA GPUDirect Remote Direct Memory Access for Cloudera AI workloads, including jobs, sessions, applications, model deployments, and Cloudera AI Inference service.
This feature is supported in both Cloudera Embedded Container Service environments and OpenShift environments for Cloudera AI 1.5.5 SP4 and higher releases.
Supported Remote Direct Memory Access configuration models
NVIDIA GPUDirect RDMA configuration models enable high-throughput GPU networking through SR-IOV, and multi-rail, architectures in OpenShift and Embedded Container Service environments and in shared device architectures in OpenShift environments. Cloudera recommends the SR-IOV method for most production deployments because it provides strong workload isolation and scalable performance.
SR-IOV method
SR-IOV uses hardware-level virtualization to divide a physical network adapter into multiple virtual functions (VFs). Each VF functions as an independent PCIe network device that can be assigned to a workload.
SR-IOV is often the best compromise between performance isolation and multi-tenant scalability for production Cloudera AI workloads.
Multi-rail configuration
In a multi-rail configuration, each GPU node connects to the network fabric through multiple independent network interfaces (rails), typically using high-speed NICs such as NVIDIA ConnectX-7 adapters. By enabling more than one active network port per host, multi-rail configurations increase throughput, improve resiliency, and optimize utilization of available system resources.
The NVIDIA Spectrum-X multi-rail AI-optimized interconnect architecture combines SR-IOV with multiplane load balancing to scale GPU-to-GPU bandwidth efficiently across switch tiers. This architecture enables high-bandwidth, low-latency RDMA and GPUDirect communication between GPU nodes, which is critical for distributed AI and high-performance computing workloads.
To configure a high-performance multi-rail environment, each GPU rail on the compute host must have a dedicated SR-IOV Network Node Policy. Dedicated policies allow Kubernetes to distinguish between physical NICs and assign specific rails to GPU-accelerated workloads, ensuring consistent network isolation, predictable performance, and efficient traffic distribution.
Shared device
The shared device model uses a software-based approach to allow multiple workloads running on the same OpenShift Container Platform worker node to access a single NVIDIA GPUDirect RDMA device simultaneously. This model is supported only in Red Hat OpenShift environments.
Choose this option when you want efficient resource sharing and can tolerate performance interference between workloads.
| Method | Environment | Sharing model | Isolation | Scalability | Typical use case |
|---|---|---|---|---|---|
| SR-IOV |
Cloudera Embedded Container Service OpenShift |
Hardware-partitioned virtual functions | High | High | Performance-oriented, multi-tenant AI and High-Performance Computing workloads |
| Shared device |
OpenShift |
Software-level shared access to one physical RDMA device | Low | Very high | Many pods with moderate sensitivity to interference |
| Multi-rail configuration |
Cloudera Embedded Container Service OpenShift |
Multiple physical RDMA devices aggregated per pod for combined bandwidth | High | High | Large-scale AI training workloads requiring maximized network throughput and redundancy |
Naming conventions for RDMA configurations
| Default name used in configuration | NVIDIA driver resource naming by network configuration |
|---|---|
| default/roce-network | rdma/rdma_shared_device_eth |
| default/ipoib-network:rdma | rdma_shared_device_ib |
| default/sriovib-network | openshift.io/sriovlegacy (OpenShift) |
| default/sriovib-network | nvidia.com/sriovlegacy (Cloudera Embedded Container Service) |
| Default name used in configuration | NVIDIA driver resource naming by network configuration |
|---|---|
| default/roce-network | rdma/rdma_shared_device_eth_123 |
| default/sriovib-network | nvidia.com/sriovlegacy_12 (Cloudera Embedded Container Service) |
For implementing NVIDIA GPUDirect Remote Direct Memory Access (RDMA) on Cloudera Embedded Container Service follow the relevant documentation in: Setting up Single or Mult-Rail cluster.
For implementing NVIDIA GPUDirect Remote Direct Memory Access (RDMA) on OpenShift Container Platform follow the relevant OpenShift documentation: NVIDIA GPUDirect Remote Direct Memory Access (RDMA).
