Setting up Host node

As the driver management is turned off in the Cloudera Control Plane deployment, you must execute the following setup on all worker nodes prior to joining the cluster.

Refer: CUDA Installation Guide for Linux.

For Platform support information (Cloudera deploys gpu-operator 26.3.2): Refer:Platform Support instructions.

  1. To enable Kernel Dependencies and EPEL, run the following commands:
    Worker nodes require kernel development packages matching the exact active kernel version.
    # Register RHEL 9 Repositories (skip on Rocky/AlmaLinux 9)
    sudo subscription-manager repos --enable=rhel-9-for-x86_64-appstream-rpms
    sudo subscription-manager repos --enable=rhel-9-for-x86_64-baseos-rpms
    sudo subscription-manager repos --enable=codeready-builder-for-rhel-9-x86_64-rpms
    
    # Install matching kernel development headers
    sudo dnf install -y kernel-devel-$(uname -r) kernel-headers EPEL
    
    # Add CUDA & Toolkit Repositories
    sudo dnf config-manager --add-repo https://developer.download.nvidia.com/compute/cuda/repos/rhel9/x86_64/cuda-rhel9.repo
    curl -s -L https://nvidia.github.io/libnvidia-container/stable/rpm/nvidia-container-toolkit.repo | \
      sudo tee /etc/yum.repos.d/nvidia-container-toolkit.repo
    
  2. Install the Driver and FabricManager
    You must install drivers natively on the host.
    # Install open-dkms NVIDIA driver & FabricManager
    sudo dnf module enable -y nvidia-driver:open-dkms
    sudo dnf install -y nvidia-open nvidia-fabricmanager
    
    # Enable FabricManager (Mandatory for HGX A100/H100/B200 NVSwitch systems)
    sudo systemctl enable --now nvidia-fabricmanager
    
  3. Link must be up and running. You must run this on every HGX node:
    rdma link show
    
    # Example OUTPUT
    link mlx5_0/1 subnet_prefix fe80:0000:0000:0000 lid 28 sm_lid 28 lmc 0 state ACTIVE physical_state LINK_UP 
    link mlx5_1/1 subnet_prefix fe80:0000:0000:0000 lid 11 sm_lid 28 lmc 0 state ACTIVE physical_state LINK_UP 
    link mlx5_2/1 subnet_prefix fe80:0000:0000:0000 lid 26 sm_lid 28 lmc 0 state ACTIVE physical_state LINK_UP 
    link mlx5_3/1 subnet_prefix fe80:0000:0000:0000 lid 8 sm_lid 28 lmc 0 state ACTIVE physical_state LINK_UP 
    link mlx5_4/1 subnet_prefix fe80:0000:0000:0000 lid 33 sm_lid 28 lmc 0 state ACTIVE physical_state LINK_UP 
    link mlx5_5/1 subnet_prefix fe80:0000:0000:0000 lid 5 sm_lid 28 lmc 0 state ACTIVE physical_state LINK_UP 
    link mlx5_6/1 subnet_prefix fe80:0000:0000:0000 lid 35 sm_lid 28 lmc 0 state ACTIVE physical_state LINK_UP 
    link mlx5_7/1 subnet_prefix fe80:0000:0000:0000 lid 25 sm_lid 28 lmc 0 state ACTIVE physical_state LINK_UP 
    
Set up the Single/Multi Rail for homogeneous HGX cluster