Setting up single or multi rail for homogenous HGX cluster
Setting up the HGX cluster involves multiple operational phases that includes discovering rails, defining YAML configuration, applying the changes, and verifying configurations.
Discovering the rails
- HGX node names to:
<node-0> <node-1> - Username to:
<username>
for NODE in <node-0> <node-1>; do
echo "=== $NODE ==="
ssh -o StrictHostKeyChecking=no <username>@$NODE \
'printf "%-6s %-14s %-14s %-14s %-9s %-9s\n" "GPU" "pf-name" "network-type" "nic-vendor-id" "totalvfs" "state"
ibdev2netdev 2>/dev/null | while read -r m _ _ _ nd _; do
case $m in mlx5_[0-7])
t=$(cat /sys/class/net/$nd/type 2>/dev/null)
v=$(cat /sys/class/net/$nd/device/vendor 2>/dev/null | sed "s/0x//")
vfs=$(cat /sys/class/net/$nd/device/sriov_totalvfs 2>/dev/null || echo "-")
st=$(cat /sys/class/net/$nd/operstate 2>/dev/null || echo "-")
case $t in 32) l=InfiniBand;; 1) l=Ethernet;; *) l=unknown;; esac
printf "%-6s %-14s %-14s %-14s %-9s %-9s\n" "GPU${m#mlx5_}" "$nd" "$l" "$v" "$vfs" "$st"
;; esac
done'
echo
done
# Example OUTPUT per node
=== <node-0> ===
GPU pf-name network-type nic-vendor-id totalvfs state
GPU0 ibp15s0 InfiniBand 15b3 8 up # rail0
GPU1 ibp20s0 InfiniBand 15b3 8 up # rail1
GPU2 ibp28s0 InfiniBand 15b3 8 up # ...
GPU3 ibp102s0 InfiniBand 15b3 8 up
GPU4 ibp155s0 InfiniBand 15b3 8 up
GPU5 ibp170s0 InfiniBand 15b3 8 up
GPU6 ibp187s0 InfiniBand 15b3 8 up
GPU7 ibp218s0 InfiniBand 15b3 8 up # rail7
Refer to the NVIDIA documentation to seek directions for any kind of setup.
YAML configuration templates per rail
The YAML configuration template that must be defined per rail. It is
<highlighted> and # commented what values
need to be substituted from the previous section.
# --- IPPool: allocates IPs for this rail ---
apiVersion: nv-ipam.nvidia.com/v1alpha1
kind: IPPool
metadata:
name: rail<N>-pool # <- index
namespace: network-operator
spec:
subnet: "10.0.<N>.0/24" # <- index
perNodeBlockSize: 20
gateway: "10.0.<N>.1" # <- index
---
# --- SriovIBNetwork: exposes the rail as a Network Attachment Definition ---
apiVersion: sriovnetwork.openshift.io/v1
# Use SriovIBNetwork for InfiniBand. Use SriovNetwork for Ethernet.
kind: <SriovIBNetwork OR SriovNetwork>
metadata:
name: rail<N>-network # <- index
namespace: network-operator
spec:
vlan: 0 # <- Ethernet only — DELETE this line for InfiniBand
linkState: enable # <- InfiniBand only — DELETE this line for Ethernet
networkNamespace: "default"
resourceName: "sriovlegacy<N>" # <- index
ipam: |-
{
"type": "nv-ipam",
"poolName": "rail<N>-pool" # <- index
}
---
# --- SriovNetworkNodePolicy: configures the VFs on the physical function ---
apiVersion: sriovnetwork.openshift.io/v1
kind: SriovNetworkNodePolicy
metadata:
name: rail<N>-policy # <- index
namespace: network-operator
spec:
deviceType: netdevice
mtu: 1500
nicSelector:
vendor: "<nic-vendor-id>" # <- e.g.: 15b3
# for single-rail use the pf-name of GPU0,
# for multi-rail use the pf-name of the corresponding GPU
pfNames:
- <pf-name> # <- e.g.: ibp15s0
nodeSelector:
# e.g.: feature.node.kubernetes.io/pci-15b3.present: "true"
feature.node.kubernetes.io/pci-<nic-vendor-id>.present: "true" # <- e.g.: 15b3
numVfs: <totalvfs> # <- number of VFs to create per PF, e.g.: 8
linkType: <IB OR eth> # <- InfiniBand | Ethernet (from network-type)
priority: 90
isRdma: true
resourceName: sriovlegacy<N> # <- index
Configuring single-rail
- You need to define only one instance of the YAML configuration template from above.
- Save this in a file on the Cloudera Embedded Container Service master node,
and call it
all-rails-combined.yaml. - Remove every
# helpercomment.
Configuring multi-rail
Use the following information:
- Combine each configuration template instance into an
all-rails-combined.yamlfile and insert it on the Cloudera Embedded Container Service master node, add a---separator after each instance. - For every new rail, increment the index
<N>, commence fromN=0. - For the
pf-nameinSriovNetworkNodePolicy, use thepf-namefor the corresponding GPU. - Remove every
# helpercomment.
Applying the defined configurations
Before you start, you must login as root and set
kubectl and KUBECONFIG. Perform the actions in
the order of flow.
kubectl command, and KUBECONFIG. Disabling Operator Drain
Because of the Longhorn’s instance manager disruption policy, it is not
possible to evict the instance-manager Pods from the Multi-Instance
GPU (MIG) capable nodes.
kubectl patch sriovoperatorconfig default -n network-operator --type=merge \
-p '{"spec":{"disableDrain":true}}'
# Example OUTPUT
sriovoperatorconfig.sriovnetwork.openshift.io/default patched
Applying the configuration changes
You must apply the changes:
kubectl apply -f all-rails-combined.yaml
# Example OUTPUT
ippool.nv-ipam.nvidia.com/rail0-pool created
sriovibnetwork.sriovnetwork.openshift.io/rail0-network created
sriovnetworknodepolicy.sriovnetwork.openshift.io/rail0-policy created
...
ippool.nv-ipam.nvidia.com/rail7-pool created
sriovibnetwork.sriovnetwork.openshift.io/rail7-network created
sriovnetworknodepolicy.sriovnetwork.openshift.io/rail7-policy created
Synchronizing the changes
Allow about 15 minutes and run the following command on a periodic basis:
kubectl get sriovnetworknodestate -n network-operator
# Successful reconciliation OUTPUT looks like the following for every HGX node:
# * SYNC STATUS + DESIRED SYNC STATE =
# * InProgress + Reboot_Required
# * Succeeded + Reboot_Required
# * Succeeded + Idle
# * CURRENT SYNC STATE = idle
NAME SYNC STATUS DESIRED SYNC STATE CURRENT SYNC STATE AGE
<node-0> Succeeded Idle Idle 10h
<node-1> InProgress Reboot_Required Idle 10h
Rebooting every HGX Node
- Sequential (safest and
recommended):
- Reboot one HGX node at a time.
- Parallel (only if ALL true):
- No active GPU workloads on HGX nodes.
- No critical Longhorn volumes with replicas only on those nodes.
- Full HGX downtime is acceptable.
- Batched parallel
(middle-ground):
- Split HGX nodes into small batches (for example, two at a time).
- Reboot each batch in parallel. After one batch has finished, verify and then do the same for the next batch.
- After every reboot session, you must perform a verification on the HGX node(s).
NODE=<node-0>
kubectl cordon $NODE
ssh <username>@$NODE 'sudo reboot'
# Wait some time (about 15 minutes) after HGX node reboot...
kubectl uncordon $NODE
Verifying the change requests
Use the following information to verify the changes
kubectl get sriovnetworknodepolicy,sriovibnetwork,sriovnetwork,ippools.nv-ipam.nvidia.com -n network-operator
# Example successful OUTPUT for Infiniband
# For Ethernet, you should see SriovNetwork CRs, instead of SriovIBNetwork
NAME AGE
sriovnetworknodepolicy.sriovnetwork.openshift.io/rail0-policy 44m
...
sriovnetworknodepolicy.sriovnetwork.openshift.io/rail7-policy 44m
NAME AGE
sriovibnetwork.sriovnetwork.openshift.io/rail0-network 44m
...
sriovibnetwork.sriovnetwork.openshift.io/rail7-network 44m
NAME SUBNET GATEWAY BLOCK SIZE
ippool.nv-ipam.nvidia.com/rail0-pool 10.0.0.0/24 10.0.0.1 20
...
ippool.nv-ipam.nvidia.com/rail7-pool 10.0.7.0/24 10.0.7.1 20
Verifying the Network Access Devices
Use the following information to verify the changes
kubectl get net-attach-def -n default -o json | \
jq -r '.items[] | select(.metadata.name | test("^rail")) | .metadata.name as $n |
(.spec.config | fromjson) | "\($n)\t\(.type)\t\(.link_state // "n/a")"'
# Example successful OUTPUT for Infiniband
# For Ethernet, the middle column should show 'sriov'
rail0-network ib-sriov enable
...
rail7-network ib-sriov enable
Confirming that the operator has successfully reconciled
Use the following information to verify the changes
kubectl get sriovnetworknodestate -n network-operator
# Example successful OUTPUT
NAME SYNC STATUS DESIRED SYNC STATE CURRENT SYNC STATE AGE
<node-0> Succeeded Idle Idle 13h
<node-1> Succeeded Idle Idle 13h
Confirming if the SR-IOV operator applied the policy correctly on each node
<node-0>
<node-1>.for NODE in <node-0> <node-1>; do
echo "=== $NODE ==="
kubectl get sriovnetworknodestate $NODE -n network-operator \
-o jsonpath='{range .status.interfaces[*]}{.name}{": numVfs="}{.numVfs}{"\n"}{end}'
done
# Example successful OUTPUT
=== <node-0> ===
ibp15s0: numVfs=8
ibp20s0: numVfs=8
ibp28s0: numVfs=8
ibp102s0: numVfs=8
ibp155s0: numVfs=8
ibp170s0: numVfs=8
ibp187s0: numVfs=8
ibp218s0: numVfs=8
=== <node-1> ===
ibp15s0: numVfs=8
ibp20s0: numVfs=8
ibp28s0: numVfs=8
ibp102s0: numVfs=8
ibp155s0: numVfs=8
ibp170s0: numVfs=8
ibp187s0: numVfs=8
ibp218s0: numVfs=8
Confirming if the Kubernetes scheduler can actually detect and assign the VFs to pods
Before you start, you must substitute the HGX node names to:
<node-0> <node-1>.
for NODE in <node-0> <node-1>; do
echo "=== $NODE ==="
kubectl get node $NODE -o json | \
jq '.status.allocatable | with_entries(select(.key | contains("sriov")))'
done
# Example successful OUTPUT
=== <node-0> ===
{
"nvidia.com/sriovlegacy0": "8",
"nvidia.com/sriovlegacy1": "8",
"nvidia.com/sriovlegacy2": "8",
"nvidia.com/sriovlegacy3": "8",
"nvidia.com/sriovlegacy4": "8",
"nvidia.com/sriovlegacy5": "8",
"nvidia.com/sriovlegacy6": "8",
"nvidia.com/sriovlegacy7": "8"
}
=== <node-1> ===
{
"nvidia.com/sriovlegacy0": "8",
"nvidia.com/sriovlegacy1": "8",
"nvidia.com/sriovlegacy2": "8",
"nvidia.com/sriovlegacy3": "8",
"nvidia.com/sriovlegacy4": "8",
"nvidia.com/sriovlegacy5": "8",
"nvidia.com/sriovlegacy6": "8",
"nvidia.com/sriovlegacy7": "8"
}
Verifying GPU Operator pods
You must verify that every GPU operator Pod STATUS is
Running/Completed:
kubectl get pods -n gpu-operator
Ensuring single/multi rail is functional
Substitute the values highlighted in the code and run the following script:
#!/bin/bash
# ==============================================================================
# multirail_connectivity_test_template.sh
#
# Cross-node multi-rail SR-IOV/RDMA connectivity test.
# Schedules one pod per GPU node, attaches all rails, pings across every rail.
#
# This is the "fill-in-the-blanks" version: the logic is fixed, you only edit
# the CONFIG block below. Every line you must change is marked # <-- CHANGE
# ==============================================================================
# ------------------------------------------------------------------------------
# CONFIG -- edit everything in this block to match your cluster
# ------------------------------------------------------------------------------
# kubectl + kubeconfig. On embedded ECS/RKE2 these are the on-node defaults;
# on a normal kubeconfig set K=kubectl and KC=$HOME/.kube/config.
K=/var/lib/rancher/rke2/bin/kubectl # <-- CHANGE if kubectl is elsewhere
KC=/etc/rancher/rke2/rke2.yaml # <-- CHANGE to your kubeconfig path
# Container image for the test pods. Needs sh + ping + ls (busybox is enough).
# In an air-gapped cluster point this at YOUR registry mirror.
IMG="<YOUR_REGISTRY>/busybox:latest" # <-- CHANGE to a reachable image
# The two GPU nodes to test between (must be different nodes).
# Get names with: $K --kubeconfig=$KC get nodes
NODE_A="<GPU_NODE_A>" # <-- CHANGE to your 1st GPU node name
NODE_B="<GPU_NODE_B>" # <-- CHANGE to your 2nd GPU node name
# Namespace holding the rail NetworkAttachmentDefinitions and where pods run.
NS=default # <-- CHANGE if your rail NADs live elsewhere
# Comma-separated list of rail NAD names to attach.
# List them with: $K --kubeconfig=$KC get net-attach-def -n $NS
NETS="rail0-network,rail1-network,rail2-network,rail3-network,rail4-network,rail5-network,rail6-network,rail7-network" # <-- CHANGE to your rail NADs
# The SR-IOV device-plugin resources each rail consumes, one "name: \"1\"" per
# rail. Find them with: $K --kubeconfig=$KC get node $NODE_A -o json | grep -i sriov
# The number of lines here MUST match the number of rails in NETS above.
RESOURCE_LIMITS=$(cat <<'LIMITS'
nvidia.com/sriovlegacy0: "1"
nvidia.com/sriovlegacy1: "1"
nvidia.com/sriovlegacy2: "1"
nvidia.com/sriovlegacy3: "1"
nvidia.com/sriovlegacy4: "1"
nvidia.com/sriovlegacy5: "1"
nvidia.com/sriovlegacy6: "1"
nvidia.com/sriovlegacy7: "1"
LIMITS
)
# ^-- CHANGE the resource names/count above to match your rails
# Rail count (used only for the ping loop) and the IPv4 prefix your rail IPAM
# hands out. Example: pools 10.0.0.0/24 .. 10.0.7.0/24 -> prefix "10.0" and
# the per-rail octet is the rail index. Adjust the regex if your pools differ.
RAIL_COUNT=8 # <-- CHANGE to your number of rails
RAIL_IP_PREFIX="10.0" # <-- CHANGE to your rail subnet prefix
# ------------------------------------------------------------------------------
# END CONFIG -- no changes needed below this line
# ------------------------------------------------------------------------------
run() {
sudo "$K" --kubeconfig="$KC" "$@"; # drop "sudo" if not needed for kubeconfig
}
ex() {
run exec "$1" -c t -n "$NS" -- sh -c "$2" 2>&1;
}
cleanup() {
run delete pod mr-a mr-b -n "$NS" --ignore-not-found --grace-period=0 --force >/dev/null 2>&1
}
cleanup
for P in "mr-a $NODE_A" "mr-b $NODE_B"; do
set -- $P
cat <<YAML | run apply -f - >/dev/null
apiVersion: v1
kind: Pod
metadata: { name: $1, namespace: $NS, annotations: { k8s.v1.cni.cncf.io/networks: "$NETS" } }
spec:
nodeName: $2
restartPolicy: Never
containers:
- name: t
image: $IMG
command: ["sh","-c","sleep 3600"]
securityContext: { runAsUser: 0, capabilities: { add: ["IPC_LOCK","NET_ADMIN","NET_RAW"] } }
resources:
limits:
$RESOURCE_LIMITS
YAML
done
echo "--- wait for Running ---"
for i in $(seq 1 24); do
a=$(run get pod mr-a -n "$NS" -o jsonpath='{.status.phase}' 2>/dev/null)
b=$(run get pod mr-b -n "$NS" -o jsonpath='{.status.phase}' 2>/dev/null)
[ "$a" = "Running" ] && [ "$b" = "Running" ] && break; sleep 5
done
run get pods mr-a mr-b -n "$NS" -o wide
echo "=== mr-a secondary interfaces (net1..netN) count ==="
ex mr-a "ls /sys/class/net/ | grep -c net"
echo "=== CROSS-NODE PING on ALL $RAIL_COUNT rails (mr-a -> mr-b) ==="
ASTAT=$(run get pod mr-a -n "$NS" -o jsonpath='{.metadata.annotations.k8s\.v1\.cni\.cncf\.io/network-status}')
BSTAT=$(run get pod mr-b -n "$NS" -o jsonpath='{.metadata.annotations.k8s\.v1\.cni\.cncf\.io/network-status}')
for n in $(seq 0 $((RAIL_COUNT-1))); do
# Pull mr-b's IP on rail $n out of its Multus network-status annotation.
bip=$(echo "$BSTAT" | grep -oE "${RAIL_IP_PREFIX}\.${n}\.[0-9]+" | head -1)
res=$(ex mr-a "ping -c 2 -w 6 $bip >/dev/null 2>&1 && echo OK || echo FAIL")
echo "rail${n} -> ${bip}: ${res}"
done
cleanup
# Example successful OUTPUT
NAME READY STATUS RESTARTS AGE IP NODE
mr-a 1/1 Running 0 6s 10.42.1.145 <node-0>
mr-b 1/1 Running 0 6s 10.42.2.121 <node-1>
=== mr-a secondary interfaces (net0..net7) count ===
8
=== CROSS-NODE PING on ALL 8 rails (mr-a -> mr-b) ===
rail0 -> 10.0.0.81: OK
rail1 -> 10.0.1.81: OK
rail2 -> 10.0.2.81: OK
rail3 -> 10.0.3.81: OK
rail4 -> 10.0.4.81: OK
rail5 -> 10.0.5.81: OK
rail6 -> 10.0.6.81: OK
rail7 -> 10.0.7.81: OK
