Setting up single or multi rail for homogenous HGX cluster

Setting up the HGX cluster involves multiple operational phases that includes discovering rails, defining YAML configuration, applying the changes, and verifying configurations.

Discovering the rails

Make sure that before running this command you substitute:
  • HGX node names to: <node-0> <node-1>
  • Username to: <username>
for NODE in <node-0> <node-1>; do
  echo "=== $NODE ==="
  ssh -o StrictHostKeyChecking=no <username>@$NODE \
   'printf "%-6s %-14s %-14s %-14s %-9s %-9s\n" "GPU" "pf-name" "network-type" "nic-vendor-id" "totalvfs" "state"
     ibdev2netdev 2>/dev/null | while read -r m _ _ _ nd _; do
      case $m in mlx5_[0-7])
        t=$(cat /sys/class/net/$nd/type 2>/dev/null)
        v=$(cat /sys/class/net/$nd/device/vendor 2>/dev/null | sed "s/0x//")

        vfs=$(cat /sys/class/net/$nd/device/sriov_totalvfs 2>/dev/null || echo "-")
        st=$(cat /sys/class/net/$nd/operstate 2>/dev/null || echo "-")
        case $t in 32) l=InfiniBand;; 1) l=Ethernet;; *) l=unknown;; esac
        printf "%-6s %-14s %-14s %-14s %-9s %-9s\n" "GPU${m#mlx5_}" "$nd" "$l" "$v" "$vfs" "$st"
      ;; esac
    done'
  echo
done
# Example OUTPUT per node
=== <node-0> ===
GPU    pf-name        network-type   nic-vendor-id  totalvfs state 
GPU0   ibp15s0        InfiniBand     15b3           8        up    # rail0
GPU1   ibp20s0        InfiniBand     15b3           8        up    # rail1
GPU2   ibp28s0        InfiniBand     15b3           8        up    # ...
GPU3   ibp102s0       InfiniBand     15b3           8        up
GPU4   ibp155s0       InfiniBand     15b3           8        up
GPU5   ibp170s0       InfiniBand     15b3           8        up
GPU6   ibp187s0       InfiniBand     15b3           8        up
GPU7   ibp218s0       InfiniBand     15b3           8        up    # rail7

Refer to the NVIDIA documentation to seek directions for any kind of setup.

YAML configuration templates per rail

The YAML configuration template that must be defined per rail. It is <highlighted> and # commented what values need to be substituted from the previous section.

# --- IPPool: allocates IPs for this rail ---
apiVersion: nv-ipam.nvidia.com/v1alpha1
kind: IPPool
metadata:
  name: rail<N>-pool            # <- index
  namespace: network-operator
spec:
  subnet: "10.0.<N>.0/24"       # <- index
  perNodeBlockSize: 20
  gateway: "10.0.<N>.1"         # <- index
---
# --- SriovIBNetwork: exposes the rail as a Network Attachment Definition ---
apiVersion: sriovnetwork.openshift.io/v1
# Use SriovIBNetwork for InfiniBand. Use SriovNetwork for Ethernet.
kind: <SriovIBNetwork OR SriovNetwork>
metadata:
  name: rail<N>-network          # <- index
  namespace: network-operator
spec:
  vlan: 0           # <- Ethernet only — DELETE this line for InfiniBand
  linkState: enable # <- InfiniBand only — DELETE this line for Ethernet
  networkNamespace: "default"
  resourceName: "sriovlegacy<N>" # <- index
  ipam: |-
    {
      "type": "nv-ipam",
      "poolName": "rail<N>-pool" # <- index
    }
---
# --- SriovNetworkNodePolicy: configures the VFs on the physical function ---
apiVersion: sriovnetwork.openshift.io/v1
kind: SriovNetworkNodePolicy
metadata:
  name: rail<N>-policy           # <- index
  namespace: network-operator
spec:
  deviceType: netdevice
  mtu: 1500
  nicSelector:
    vendor: "<nic-vendor-id>"  # <- e.g.: 15b3
    # for single-rail use the pf-name of GPU0,
    # for multi-rail use the pf-name of the corresponding GPU
    pfNames:
      - <pf-name> # <- e.g.: ibp15s0
  nodeSelector:
    # e.g.: feature.node.kubernetes.io/pci-15b3.present: "true"
    feature.node.kubernetes.io/pci-<nic-vendor-id>.present: "true" # <- e.g.: 15b3
  numVfs: <totalvfs>     # <- number of VFs to create per PF, e.g.: 8
  linkType: <IB OR eth>  # <- InfiniBand | Ethernet (from network-type)
  priority: 90
  isRdma: true
  resourceName: sriovlegacy<N> # <- index

Configuring single-rail

Use the following information:
  • You need to define only one instance of the YAML configuration template from above.
  • Save this in a file on the Cloudera Embedded Container Service master node, and call it all-rails-combined.yaml.
  • Remove every # helper comment.

Configuring multi-rail

Use the following information:

You need to define as many instances of the YAML configuration template as many rails the discovery command has identified in the previous section. In this example, the rail count is 8, so the YAML configuration template from above should be repeated 8 times for each rail.
  • Combine each configuration template instance into an all-rails-combined.yaml file and insert it on the Cloudera Embedded Container Service master node, add a --- separator after each instance.
  • For every new rail, increment the index <N>, commence from N=0.
  • For the pf-name in SriovNetworkNodePolicy, use the pf-name for the corresponding GPU.
  • Remove every # helper comment.

Applying the defined configurations

Before you start, you must login as root and set kubectl and KUBECONFIG. Perform the actions in the order of flow.

Login to Cloudera Embedded Container Service master node and set kubectl command, and KUBECONFIG.

Disabling Operator Drain

Because of the Longhorn’s instance manager disruption policy, it is not possible to evict the instance-manager Pods from the Multi-Instance GPU (MIG) capable nodes.

Therefore you must disable automatic draining:
kubectl patch sriovoperatorconfig default -n network-operator --type=merge \
  -p '{"spec":{"disableDrain":true}}'
# Example OUTPUT
sriovoperatorconfig.sriovnetwork.openshift.io/default patched

Applying the configuration changes

You must apply the changes:

kubectl apply -f all-rails-combined.yaml
# Example OUTPUT
ippool.nv-ipam.nvidia.com/rail0-pool created
sriovibnetwork.sriovnetwork.openshift.io/rail0-network created

sriovnetworknodepolicy.sriovnetwork.openshift.io/rail0-policy created
...
ippool.nv-ipam.nvidia.com/rail7-pool created
sriovibnetwork.sriovnetwork.openshift.io/rail7-network created
sriovnetworknodepolicy.sriovnetwork.openshift.io/rail7-policy created

Synchronizing the changes

Allow about 15 minutes and run the following command on a periodic basis:

kubectl get sriovnetworknodestate -n network-operator
# Successful reconciliation OUTPUT looks like the following for every HGX node:
#   * SYNC STATUS + DESIRED SYNC STATE =
#       * InProgress + Reboot_Required
#       * Succeeded + Reboot_Required
#       * Succeeded + Idle
#   * CURRENT SYNC STATE = idle
NAME      SYNC STATUS  DESIRED SYNC STATE  CURRENT SYNC STATE  AGE
<node-0>  Succeeded    Idle                Idle                10h
<node-1>  InProgress   Reboot_Required     Idle                10h

Rebooting every HGX Node

You must review and finalise as to which HGX node reboot scenario to proceed with.
  • Sequential (safest and recommended):
    • Reboot one HGX node at a time.
  • Parallel (only if ALL true):
    • No active GPU workloads on HGX nodes.
    • No critical Longhorn volumes with replicas only on those nodes.
    • Full HGX downtime is acceptable.
  • Batched parallel (middle-ground):
    • Split HGX nodes into small batches (for example, two at a time).
    • Reboot each batch in parallel. After one batch has finished, verify and then do the same for the next batch.
  • After every reboot session, you must perform a verification on the HGX node(s).
The reboot commands to be executed on each HGX node:
NODE=<node-0>
kubectl cordon $NODE
ssh <username>@$NODE 'sudo reboot'
# Wait some time (about 15 minutes) after HGX node reboot...
kubectl uncordon $NODE

Verifying the change requests

Use the following information to verify the changes

kubectl get sriovnetworknodepolicy,sriovibnetwork,sriovnetwork,ippools.nv-ipam.nvidia.com -n network-operator
# Example successful OUTPUT for Infiniband
# For Ethernet, you should see SriovNetwork CRs, instead of SriovIBNetwork
NAME                                                            AGE
sriovnetworknodepolicy.sriovnetwork.openshift.io/rail0-policy   44m
...
sriovnetworknodepolicy.sriovnetwork.openshift.io/rail7-policy   44m

NAME                                                     AGE
sriovibnetwork.sriovnetwork.openshift.io/rail0-network   44m
...
sriovibnetwork.sriovnetwork.openshift.io/rail7-network   44m

NAME         SUBNET        GATEWAY    BLOCK SIZE
ippool.nv-ipam.nvidia.com/rail0-pool   10.0.0.0/24   10.0.0.1   20
...
ippool.nv-ipam.nvidia.com/rail7-pool   10.0.7.0/24   10.0.7.1   20

Verifying the Network Access Devices

Use the following information to verify the changes

kubectl get net-attach-def -n default -o json | \
  jq -r '.items[] | select(.metadata.name | test("^rail")) | .metadata.name as $n |
    (.spec.config | fromjson) | "\($n)\t\(.type)\t\(.link_state // "n/a")"'
# Example successful OUTPUT for Infiniband
# For Ethernet, the middle column should show 'sriov'
rail0-network ib-sriov enable
...
rail7-network ib-sriov enable

Confirming that the operator has successfully reconciled

Use the following information to verify the changes

kubectl get sriovnetworknodestate -n network-operator
# Example successful OUTPUT
NAME       SYNC STATUS   DESIRED SYNC STATE   CURRENT SYNC STATE   AGE
<node-0>   Succeeded     Idle                 Idle                 13h
<node-1>   Succeeded     Idle                 Idle                 13h

Confirming if the SR-IOV operator applied the policy correctly on each node

Before you start, you must substitute the HGX node names to: <node-0> <node-1>.
for NODE in <node-0> <node-1>; do
  echo "=== $NODE ==="
  kubectl get sriovnetworknodestate $NODE -n network-operator \
    -o jsonpath='{range .status.interfaces[*]}{.name}{": numVfs="}{.numVfs}{"\n"}{end}'
done
# Example successful OUTPUT
=== <node-0> ===
ibp15s0: numVfs=8
ibp20s0: numVfs=8
ibp28s0: numVfs=8
ibp102s0: numVfs=8
ibp155s0: numVfs=8
ibp170s0: numVfs=8
ibp187s0: numVfs=8
ibp218s0: numVfs=8
=== <node-1> ===
ibp15s0: numVfs=8
ibp20s0: numVfs=8
ibp28s0: numVfs=8
ibp102s0: numVfs=8
ibp155s0: numVfs=8
ibp170s0: numVfs=8
ibp187s0: numVfs=8
ibp218s0: numVfs=8

Confirming if the Kubernetes scheduler can actually detect and assign the VFs to pods

Before you start, you must substitute the HGX node names to: <node-0> <node-1>.

for NODE in <node-0> <node-1>; do
  echo "=== $NODE ==="
  kubectl get node $NODE -o json | \
    jq '.status.allocatable | with_entries(select(.key | contains("sriov")))'
done
# Example successful OUTPUT
=== <node-0> ===
{
  "nvidia.com/sriovlegacy0": "8",
  "nvidia.com/sriovlegacy1": "8",
  "nvidia.com/sriovlegacy2": "8",
  "nvidia.com/sriovlegacy3": "8",
  "nvidia.com/sriovlegacy4": "8",
  "nvidia.com/sriovlegacy5": "8",
  "nvidia.com/sriovlegacy6": "8",
  "nvidia.com/sriovlegacy7": "8"
}
=== <node-1> ===
{
  "nvidia.com/sriovlegacy0": "8",
  "nvidia.com/sriovlegacy1": "8",
  "nvidia.com/sriovlegacy2": "8",
  "nvidia.com/sriovlegacy3": "8",
  "nvidia.com/sriovlegacy4": "8",
  "nvidia.com/sriovlegacy5": "8",
  "nvidia.com/sriovlegacy6": "8",
  "nvidia.com/sriovlegacy7": "8"
}

Verifying GPU Operator pods

You must verify that every GPU operator Pod STATUS is Running/Completed:

kubectl get pods -n gpu-operator

Ensuring single/multi rail is functional

Substitute the values highlighted in the code and run the following script:

#!/bin/bash
# ==============================================================================
# multirail_connectivity_test_template.sh
#
# Cross-node multi-rail SR-IOV/RDMA connectivity test.
# Schedules one pod per GPU node, attaches all rails, pings across every rail.
#
# This is the "fill-in-the-blanks" version: the logic is fixed, you only edit
# the CONFIG block below. Every line you must change is marked  # <-- CHANGE
# ==============================================================================

# ------------------------------------------------------------------------------
# CONFIG  --  edit everything in this block to match your cluster
# ------------------------------------------------------------------------------

# kubectl + kubeconfig. On embedded ECS/RKE2 these are the on-node defaults;
# on a normal kubeconfig set K=kubectl and KC=$HOME/.kube/config.
K=/var/lib/rancher/rke2/bin/kubectl                  # <-- CHANGE if kubectl is elsewhere
KC=/etc/rancher/rke2/rke2.yaml                       # <-- CHANGE to your kubeconfig path

# Container image for the test pods. Needs sh + ping + ls (busybox is enough).
# In an air-gapped cluster point this at YOUR registry mirror.
IMG="<YOUR_REGISTRY>/busybox:latest"                 # <-- CHANGE to a reachable image

# The two GPU nodes to test between (must be different nodes).
# Get names with:  $K --kubeconfig=$KC get nodes
NODE_A="<GPU_NODE_A>"                                # <-- CHANGE to your 1st GPU node name
NODE_B="<GPU_NODE_B>"                                # <-- CHANGE to your 2nd GPU node name

# Namespace holding the rail NetworkAttachmentDefinitions and where pods run.
NS=default                                           # <-- CHANGE if your rail NADs live elsewhere

# Comma-separated list of rail NAD names to attach.
# List them with:  $K --kubeconfig=$KC get net-attach-def -n $NS
NETS="rail0-network,rail1-network,rail2-network,rail3-network,rail4-network,rail5-network,rail6-network,rail7-network"  # <-- CHANGE to your rail NADs

# The SR-IOV device-plugin resources each rail consumes, one "name: \"1\"" per
# rail. Find them with:  $K --kubeconfig=$KC get node $NODE_A -o json | grep -i sriov
# The number of lines here MUST match the number of rails in NETS above.
RESOURCE_LIMITS=$(cat <<'LIMITS'
        nvidia.com/sriovlegacy0: "1"
        nvidia.com/sriovlegacy1: "1"
        nvidia.com/sriovlegacy2: "1"
        nvidia.com/sriovlegacy3: "1"
        nvidia.com/sriovlegacy4: "1"
        nvidia.com/sriovlegacy5: "1"
        nvidia.com/sriovlegacy6: "1"
        nvidia.com/sriovlegacy7: "1"
LIMITS
)
# ^-- CHANGE the resource names/count above to match your rails

# Rail count (used only for the ping loop) and the IPv4 prefix your rail IPAM
# hands out. Example: pools 10.0.0.0/24 .. 10.0.7.0/24 -> prefix "10.0" and
# the per-rail octet is the rail index. Adjust the regex if your pools differ.
RAIL_COUNT=8                                         # <-- CHANGE to your number of rails
RAIL_IP_PREFIX="10.0"                                # <-- CHANGE to your rail subnet prefix

# ------------------------------------------------------------------------------
# END CONFIG  --  no changes needed below this line
# ------------------------------------------------------------------------------

run() {
  sudo "$K" --kubeconfig="$KC" "$@"; # drop "sudo" if not needed for kubeconfig
}
ex() {
  run exec "$1" -c t -n "$NS" -- sh -c "$2" 2>&1;
}
cleanup() {
  run delete pod mr-a mr-b -n "$NS" --ignore-not-found --grace-period=0 --force >/dev/null 2>&1
}

cleanup
for P in "mr-a $NODE_A" "mr-b $NODE_B"; do
  set -- $P
cat <<YAML | run apply -f - >/dev/null
apiVersion: v1
kind: Pod
metadata: { name: $1, namespace: $NS, annotations: { k8s.v1.cni.cncf.io/networks: "$NETS" } }
spec:
  nodeName: $2
  restartPolicy: Never
  containers:
  - name: t
    image: $IMG
    command: ["sh","-c","sleep 3600"]
    securityContext: { runAsUser: 0, capabilities: { add: ["IPC_LOCK","NET_ADMIN","NET_RAW"] } }
    resources:
      limits:
$RESOURCE_LIMITS
YAML
done

echo "--- wait for Running ---"
for i in $(seq 1 24); do
  a=$(run get pod mr-a -n "$NS" -o jsonpath='{.status.phase}' 2>/dev/null)
  b=$(run get pod mr-b -n "$NS" -o jsonpath='{.status.phase}' 2>/dev/null)
  [ "$a" = "Running" ] && [ "$b" = "Running" ] && break; sleep 5
done
run get pods mr-a mr-b -n "$NS" -o wide

echo "=== mr-a secondary interfaces (net1..netN) count ==="
ex mr-a "ls /sys/class/net/ | grep -c net"

echo "=== CROSS-NODE PING on ALL $RAIL_COUNT rails (mr-a -> mr-b) ==="
ASTAT=$(run get pod mr-a -n "$NS" -o jsonpath='{.metadata.annotations.k8s\.v1\.cni\.cncf\.io/network-status}')
BSTAT=$(run get pod mr-b -n "$NS" -o jsonpath='{.metadata.annotations.k8s\.v1\.cni\.cncf\.io/network-status}')
for n in $(seq 0 $((RAIL_COUNT-1))); do
  # Pull mr-b's IP on rail $n out of its Multus network-status annotation.
  bip=$(echo "$BSTAT" | grep -oE "${RAIL_IP_PREFIX}\.${n}\.[0-9]+" | head -1)
  res=$(ex mr-a "ping -c 2 -w 6 $bip >/dev/null 2>&1 && echo OK || echo FAIL")
  echo "rail${n} -> ${bip}: ${res}"
done
cleanup
# Example successful OUTPUT
NAME   READY   STATUS    RESTARTS   AGE   IP            NODE
mr-a   1/1     Running   0          6s    10.42.1.145   <node-0>
mr-b   1/1     Running   0          6s    10.42.2.121   <node-1>
       === mr-a secondary interfaces (net0..net7) count ===
       8
       === CROSS-NODE PING on ALL 8 rails (mr-a -> mr-b) ===
       rail0 -> 10.0.0.81: OK
       rail1 -> 10.0.1.81: OK
       rail2 -> 10.0.2.81: OK
       rail3 -> 10.0.3.81: OK
       rail4 -> 10.0.4.81: OK
       rail5 -> 10.0.5.81: OK
       rail6 -> 10.0.6.81: OK
       rail7 -> 10.0.7.81: OK