Troubleshooting GPU and Network operators
Review some of the troubleshooting scenarios related to the GPU and Network operators that you might encounter.
Unresolvable CDI devices
After an HGX node reboot, if the GPU operator Pods fail to start due to the following issue (you can locate it among the Pod events):
failed to inject CDI devices: unresolvable CDI devices management.nvidia.com/gpu=all
The /dev/nvidia-uvm is missing
/dev/nvidia-uvm is missing, then run the
following commands on the affected HGX node.sudo modprobe nvidia
sudo modprobe nvidia-uvm
sudo nvidia-modprobe -u -c=0
Later delete the affected Pods on the HGX node.
Multi-Node GPUDirect RDMA resource creation and memory registration
You might encounter an error or failure scenario during Multi-Node GPUDirect RDMA resource creation and memory registration step.
Use the following information to set up the Safety Valve with the given
values:
- From the Cloudera Manager → Containerized Cluster → ECS service → Configuration tab.
- Input the Safety Valve name:
rke2_agent_systemd_dropin_safety_valve. - Value:
[Service] LimitMEMLOCK=infinity
