====== SR-IOV Troubleshooting Commands (EVE-K / KubeVirt) ====== //Source: Slack conversation between Manny Calero and Pramodh Pallapothu//\\ //Context: EVE-OS with KubeVirt SR-IOV passthrough to VMs// ---- ===== 1. Discover SR-IOV Capable Devices ===== List all PCI devices: lspci Find all PCI devices that support SR-IOV and show the max VFs they can create: for dev in /sys/bus/pci/devices/*/; do if [ -f "$dev/sriov_totalvfs" ]; then echo "$dev: $(cat $dev/sriov_totalvfs) VFs max" fi done Check SR-IOV capability on specific PF (Physical Function): lspci -v -s 20:00.0 | grep -i "SR-IOV\|Virtual" lspci -v -s 14:00.0 | grep -i "SR-IOV\|Virtual" ---- ===== 2. Check Active and Max VFs on a PF ===== cat /sys/bus/pci/devices/0000:20:00.0/sriov_numvfs # currently active cat /sys/bus/pci/devices/0000:20:00.0/sriov_totalvfs # max supported ---- ===== 3. Map PFs to Network Interfaces ===== Identify which network interface corresponds to each PF: for pf in 20:00.0 20:00.1 22:00.0 22:00.1 14:00.0 16:00.0; do iface=$(ls /sys/bus/pci/devices/0000:$pf/net/ 2>/dev/null || echo "no netdev") echo "$pf -> $iface" done Example output confirming PF-to-interface mapping: 20:00.0 -> eth2 20:00.1 -> eth3 22:00.0 -> no netdev 22:00.1 -> no netdev 14:00.0 -> keth0 16:00.0 -> no netdev ---- ===== 4. Check VF Driver Binding ===== Verify which driver is bound to each VF. VFs should be bound to ''vfio-pci'' before VM deployment: for dev in 0000:20:10.0 0000:20:10.1 0000:20:10.2 0000:20:10.3; do driver=$(readlink /sys/bus/pci/devices/$dev/driver 2>/dev/null) if [ -z "$driver" ]; then echo "$dev: NO DRIVER" else echo "$dev: ${driver##*/}" fi done Expected output when VFs are ready for passthrough: 0000:20:10.0: vfio-pci 0000:20:10.1: vfio-pci 0000:20:10.2: vfio-pci 0000:20:10.3: vfio-pci > **Note:** VFs bound to ''vfio-pci'' is correct. Once a VM is deployed and takes a VF, the driver will change to the NIC driver (e.g. ''ixgbe''). ---- ===== 5. Inspect PF and VF MAC Addresses ===== Show the PF interface along with all VF MAC assignments: ip link show eth2 ip link show eth3 Example output — VF MAC addresses assigned by KubeVirt: 6: eth2: mtu 1500 ... link/ether c4:00:ad:b8:8c:e2 brd ff:ff:ff:ff:ff:ff vf 0 link/ether 02:16:3e:44:3c:ca brd ff:ff:ff:ff:ff:ff, spoof checking on, link-state auto, trust off, query_rss off vf 1 link/ether 00:00:00:00:00:00 brd ff:ff:ff:ff:ff:ff, spoof checking on ... > **Note:** VFs with a non-zero MAC are actively assigned to a VM. After a reboot, confirm the same VM gets the same MAC and IP. ---- ===== 6. Check Network Attachment Definitions (NADs) ===== NADs must exist for each SR-IOV interface before a VM can use them. Run inside the kube container (''eve enter kube''): kubectl get network-attachment-definition -A Expected output for a working SR-IOV setup: NAMESPACE NAME AGE eve-kube-app network-instance-attachment 17d eve-kube-app sriov-eth2 17d eve-kube-app sriov-eth3 17d > **Important:** If ''sriov-eth2'' / ''sriov-eth3'' NADs are missing, the VM will fail to start with a Multus CNI error. This can be caused by a race condition at startup — upgrading to a patched image resolves it. ---- ===== 7. Fix Missing SR-IOV CNI Binary ===== If VM deployment fails with: failed to find plugin "sriov" in path [/var/lib/cni/bin /var/lib/rancher/k3s/data/current/bin] The SR-IOV CNI binary is missing. Enter the kube container and copy it manually (workaround for older images): eve enter kube cp /opt/cni/bin/sriov /var/lib/cni/bin > **Note:** Newer images include the binary automatically via the ''kube-sriov-cni-ds-amd64'' DaemonSet. If the pod exists but the binary is still missing, verify the DaemonSet ran successfully. ---- ===== 8. Check SR-IOV Device Plugin Logs ===== The device plugin registers VFs with Kubelet. Check its pod logs to verify VFs are being advertised: # From inside kube container kubectl logs -n kube-system Healthy output shows VFs registered per PF: initServers(): selector index 0 will register 20 devices device added: [identifier: 0000:20:10.1, vendor: 8086, device: 15c5, driver: vfio-pci] ... starting eth2_vfs device plugin endpoint at: eve.network_eth2_vfs.sock Plugin: eve.network_eth2_vfs.sock gets registered successfully at Kubelet ---- ===== 9. Common Error Reference ===== ^ Error ^ Cause ^ Resolution ^ | ''failed to find plugin "sriov" in path'' | SR-IOV CNI binary missing from node | Copy binary: ''cp /opt/cni/bin/sriov /var/lib/cni/bin'' or use newer image | | ''error adding container to network "sriov-eth2"'' | NAD ''sriov-eth2'' missing | Upgrade image; race condition prevents NAD creation on older builds | | ''failed to create SR-IOV hostdevices: context deadline exceeded'' | ''network-info'' file not populated in time | Transient; retry VM deploy or check virt-launcher pod events | | VFs have ''NO DRIVER'' | EVE code should auto-bind to ''vfio-pci'' | Verify correct EVE-K image with SR-IOV feature is running | | VM shows no SR-IOV interfaces | NADs not attached in VMI config | Check app instance config / UI; NAD must be referenced in the VM network spec | ---- ===== 10. General Workflow Summary ===== - Onboard the node to ZED Cloud - Verify PFs are visible via ''lspci'' and mapped to interfaces - Confirm VFs are created and bound to ''vfio-pci'' - Verify NADs (''sriov-eth2'', ''sriov-eth3'', etc.) exist in the ''eve-kube-app'' namespace - Deploy the VM with SR-IOV interfaces set to **AppDirect** in the model config - After deployment, verify unique MAC addresses on each VF via ''ip link show ethX'' - Test reboot persistence — VM should get the same MAC and IP after reboot ---- //Last updated: 2026-05-23 | Source: Slack DMs with Pramodh Pallapothu//