Troubleshooting GPU and Network operators

Review some of the troubleshooting scenarios related to the GPU and Network operators that you might encounter.

Unresolvable CDI devices

After an HGX node reboot, if the GPU operator Pods fail to start due to the following issue (you can locate it among the Pod events):

failed to inject CDI devices: unresolvable CDI devices management.nvidia.com/gpu=all
  1. Run the following command on the affected HGX node where the Pods are running:
    sudo /usr/local/nvidia/toolkit/nvidia-ctk cdi generate \
    --vendor=management.nvidia.com \
    --device-name-strategy=type-index \
    --output=/var/run/cdi/management.nvidia.com-gpu.yaml
    
  2. Later delete the affected Pods on the HGX node.

The /dev/nvidia-uvm is missing

After an HGX node reboot, if the GPU operator Pods are not able to start because /dev/nvidia-uvm is missing, then run the following commands on the affected HGX node.
sudo modprobe nvidia
sudo modprobe nvidia-uvm
sudo nvidia-modprobe -u -c=0

Later delete the affected Pods on the HGX node.

Multi-Node GPUDirect RDMA resource creation and memory registration

You might encounter an error or failure scenario during Multi-Node GPUDirect RDMA resource creation and memory registration step.

Use the following information to set up the Safety Valve with the given values:
  1. From the Cloudera Manager → Containerized Cluster → ECS service → Configuration tab.
  2. Input the Safety Valve name: rke2_agent_systemd_dropin_safety_valve.
  3. Value:
    [Service]
    LimitMEMLOCK=infinity
Setting this safety valve will cause configuration staleness and all the Cloudera Embedded Container Service agent roles must be restarted for these changes to take effect.