User Tools

Site Tools


edge-node-clustering:sriov-debug

SR-IOV Troubleshooting Commands (EVE-K / KubeVirt)

Source: Slack conversation between Manny Calero and Pramodh Pallapothu
Context: EVE-OS with KubeVirt SR-IOV passthrough to VMs


1. Discover SR-IOV Capable Devices

List all PCI devices:

lspci

Find all PCI devices that support SR-IOV and show the max VFs they can create:

for dev in /sys/bus/pci/devices/*/; do
  if [ -f "$dev/sriov_totalvfs" ]; then
    echo "$dev: $(cat $dev/sriov_totalvfs) VFs max"
  fi
done

Check SR-IOV capability on specific PF (Physical Function):

lspci -v -s 20:00.0 | grep -i "SR-IOV\|Virtual"
lspci -v -s 14:00.0 | grep -i "SR-IOV\|Virtual"

2. Check Active and Max VFs on a PF

cat /sys/bus/pci/devices/0000:20:00.0/sriov_numvfs    # currently active
cat /sys/bus/pci/devices/0000:20:00.0/sriov_totalvfs  # max supported

3. Map PFs to Network Interfaces

Identify which network interface corresponds to each PF:

for pf in 20:00.0 20:00.1 22:00.0 22:00.1 14:00.0 16:00.0; do
  iface=$(ls /sys/bus/pci/devices/0000:$pf/net/ 2>/dev/null || echo "no netdev")
  echo "$pf -> $iface"
done

Example output confirming PF-to-interface mapping:

20:00.0 -> eth2
20:00.1 -> eth3
22:00.0 -> no netdev
22:00.1 -> no netdev
14:00.0 -> keth0
16:00.0 -> no netdev

4. Check VF Driver Binding

Verify which driver is bound to each VF. VFs should be bound to vfio-pci before VM deployment:

for dev in 0000:20:10.0 0000:20:10.1 0000:20:10.2 0000:20:10.3; do
  driver=$(readlink /sys/bus/pci/devices/$dev/driver 2>/dev/null)
  if [ -z "$driver" ]; then
    echo "$dev: NO DRIVER"
  else
    echo "$dev: ${driver##*/}"
  fi
done

Expected output when VFs are ready for passthrough:

0000:20:10.0: vfio-pci
0000:20:10.1: vfio-pci
0000:20:10.2: vfio-pci
0000:20:10.3: vfio-pci
Note: VFs bound to vfio-pci is correct. Once a VM is deployed and takes a VF, the driver will change to the NIC driver (e.g. ixgbe).

5. Inspect PF and VF MAC Addresses

Show the PF interface along with all VF MAC assignments:

ip link show eth2
ip link show eth3

Example output — VF MAC addresses assigned by KubeVirt:

6: eth2: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 ...
    link/ether c4:00:ad:b8:8c:e2 brd ff:ff:ff:ff:ff:ff
    vf 0     link/ether 02:16:3e:44:3c:ca brd ff:ff:ff:ff:ff:ff, spoof checking on, link-state auto, trust off, query_rss off
    vf 1     link/ether 00:00:00:00:00:00 brd ff:ff:ff:ff:ff:ff, spoof checking on ...
Note: VFs with a non-zero MAC are actively assigned to a VM. After a reboot, confirm the same VM gets the same MAC and IP.

6. Check Network Attachment Definitions (NADs)

NADs must exist for each SR-IOV interface before a VM can use them. Run inside the kube container (eve enter kube):

kubectl get network-attachment-definition -A

Expected output for a working SR-IOV setup:

NAMESPACE      NAME                          AGE
eve-kube-app   network-instance-attachment   17d
eve-kube-app   sriov-eth2                    17d
eve-kube-app   sriov-eth3                    17d
Important: If sriov-eth2 / sriov-eth3 NADs are missing, the VM will fail to start with a Multus CNI error. This can be caused by a race condition at startup — upgrading to a patched image resolves it.

7. Fix Missing SR-IOV CNI Binary

If VM deployment fails with:

failed to find plugin "sriov" in path [/var/lib/cni/bin /var/lib/rancher/k3s/data/current/bin]

The SR-IOV CNI binary is missing. Enter the kube container and copy it manually (workaround for older images):

eve enter kube
cp /opt/cni/bin/sriov /var/lib/cni/bin
Note: Newer images include the binary automatically via the kube-sriov-cni-ds-amd64 DaemonSet. If the pod exists but the binary is still missing, verify the DaemonSet ran successfully.

8. Check SR-IOV Device Plugin Logs

The device plugin registers VFs with Kubelet. Check its pod logs to verify VFs are being advertised:

# From inside kube container
kubectl logs -n kube-system <sriov-device-plugin-pod>

Healthy output shows VFs registered per PF:

initServers(): selector index 0 will register 20 devices
device added: [identifier: 0000:20:10.1, vendor: 8086, device: 15c5, driver: vfio-pci]
...
starting eth2_vfs device plugin endpoint at: eve.network_eth2_vfs.sock
Plugin: eve.network_eth2_vfs.sock gets registered successfully at Kubelet

9. Common Error Reference

Error Cause Resolution
failed to find plugin “sriov” in path SR-IOV CNI binary missing from node Copy binary: cp /opt/cni/bin/sriov /var/lib/cni/bin or use newer image
error adding container to network “sriov-eth2” NAD sriov-eth2 missing Upgrade image; race condition prevents NAD creation on older builds
failed to create SR-IOV hostdevices: context deadline exceeded network-info file not populated in time Transient; retry VM deploy or check virt-launcher pod events
VFs have NO DRIVER EVE code should auto-bind to vfio-pci Verify correct EVE-K image with SR-IOV feature is running
VM shows no SR-IOV interfaces NADs not attached in VMI config Check app instance config / UI; NAD must be referenced in the VM network spec

10. General Workflow Summary

  1. Onboard the node to ZED Cloud
  2. Verify PFs are visible via lspci and mapped to interfaces
  3. Confirm VFs are created and bound to vfio-pci
  4. Verify NADs (sriov-eth2, sriov-eth3, etc.) exist in the eve-kube-app namespace
  5. Deploy the VM with SR-IOV interfaces set to AppDirect in the model config
  6. After deployment, verify unique MAC addresses on each VF via ip link show ethX
  7. Test reboot persistence — VM should get the same MAC and IP after reboot

Last updated: 2026-05-23 | Source: Slack DMs with Pramodh Pallapothu

edge-node-clustering/sriov-debug.txt · Last modified: by 172.17.0.1