====== ZKS GitOps Demo Cluster — Deploy From Scratch ====== This guide covers a full from-scratch deployment of the ZKS demo cluster GitOps stack using Fleet (Rancher Continuous Delivery). It includes MetalLB, KubeVirt, CDI, Ubuntu VM bootstrap, and the UniFi network dashboard. ===== Architecture Overview ===== ^ Layer ^ Tool ^ Purpose ^ | GitOps engine | Fleet (Rancher CD) | Watches Git repo, applies Helm charts | | Load balancer | MetalLB | Assigns LAN IPs to LoadBalancer services | | Virtualisation | KubeVirt + CDI | Runs VMs inside Kubernetes pods | | Network visibility | network-dashboard | UniFi API → topology web UI | | Observability | Grafana | Dashboards backed by PostgreSQL + NFS | **Repo:** ''https://github.com/ManCalAzure/zks.git'' (branch: ''main'')\\ **Chart root:** ''zks/demo/charts/'' ^ Folder ^ Purpose ^ | ''infrastructure/'' | Always-on platform components (MetalLB config, KubeVirt) | | ''active/'' | Application workloads Fleet deploys automatically | | ''inactive/'' | Charts present in Git but not watched by Fleet | To activate a chart move it from ''inactive/'' → ''active/'' and push. To deactivate, move it back. --- ===== Prerequisites ===== * ZKS cluster imported into Rancher / Fleet * Fleet GitRepo pointed at the repo above, path ''zks/demo/charts/'' * MetalLB **operator** already installed on the cluster (CRDs must exist before ''infrastructure/metallb-config'' is applied) * An NFS share at ''192.168.0.101:/volume1/nano-r5s/'' (used by Grafana) * An HTTP file server at ''http://192.168.0.101:1080/'' serving the Ubuntu cloud image * PostgreSQL at ''192.168.0.2:5432'' (used by Grafana) * UniFi controller at ''https://192.168.0.1'' with an API key --- ===== Step 1 — MetalLB IP Pool ===== **Chart:** ''infrastructure/metallb-config'' MetalLB is configured via two CRs: an ''IPAddressPool'' and an ''L2Advertisement''. **IP pool:** ''192.168.0.200 – 192.168.0.225'' ^ Service ^ LoadBalancer IP ^ | Grafana | 192.168.0.205 | | network-dashboard | 192.168.0.207 | This chart is always-on in ''infrastructure/''. Fleet applies it on every sync. No action needed after initial GitRepo setup. **Verify:** kubectl get ipaddresspool -n metallb-system kubectl get l2advertisement -n metallb-system --- ===== Step 2 — KubeVirt Operator ===== **Chart:** ''infrastructure/kubevirt-operator''\\ **Versions:** KubeVirt ''v1.3.1'', CDI ''v1.60.1'' ==== What it does ==== The chart runs a sequenced install via Helm hooks: ^ Hook weight ^ Job ^ What it does ^ | -6 | configmap | Renders the KubeVirt CR (with SR-IOV feature gate) | | -5 | rbac | Creates ServiceAccount + ClusterRole in ''kube-system'' | | 0 | install-job | Applies KubeVirt operator → waits for ready → applies CR → installs CDI | SR-IOV is enabled by default (''sriov.enabled: true''). ==== Verify install ==== # Job should show Completed 1/1 kubectl get jobs -n kube-system | grep kubevirt # KubeVirt operator pods kubectl get pods -n kubevirt # CDI pods kubectl get pods -n cdi # CRDs registered kubectl get crd | grep kubevirt kubectl get crd | grep cdi The ''ubuntu-vm'' chart has a ''dependsOn'' referencing this bundle. Fleet will not apply ubuntu-vm until kubevirt-operator reaches Ready state. Do not remove this dependency. --- ===== Step 3 — Ubuntu VM ===== **Chart:** ''active/ubuntu-vm'' ==== Overview ==== Deploys a KubeVirt ''VirtualMachine'' backed by a CDI ''DataVolume'' that imports the Ubuntu Noble cloud image over HTTP. ^ Setting ^ Value ^ | VM name | ubuntu-vm1 | | Namespace | vms | | vCPUs | 2 | | Memory | 4 Gi | | Disk | 20 Gi (''local-path'' storage class) | | Run strategy | Always | | Network | Pod network (masquerade / NAT) | | Image source | ''http://192.168.0.101:1080/noble-server-cloudimg-amd64.img'' | ==== Cloud-init ==== Cloud-init is stored as a multipart MIME file at ''active/ubuntu-vm/files/cloud-init.txt'' and mounted as a Kubernetes Secret (key: ''userdata''). **What cloud-init configures:** * Hostname: ''VM1'' * User: ''manny'' (sudo, no password, SSH key injected) * Password auth enabled on SSH * Packages: ''wget'', ''curl'', ''traceroute'' * Network: DHCP on ''enp1s0'' To change the SSH key or hostname edit ''files/cloud-init.txt'' and push. ==== KubeVirt volume reference ==== The cloud-init secret is referenced in the VM spec as: cloudInitNoCloud: secretRef: name: ubuntu-vm1-cloudinit The field is ''secretRef'', **not** ''userDataSecretRef''. KubeVirt v1.3.1 strict decoding will reject the latter with a BadRequest error. ==== DataVolume lifecycle ==== ^ Phase ^ Meaning ^ | WaitForFirstConsumer | PVC waiting for a pod to be scheduled (normal for ''local-path'') | | ImportScheduled | CDI importer pod scheduled | | ImportInProgress | Image downloading from HTTP server | | Succeeded | Import complete, VM can boot | The DataVolume stays in ''WaitForFirstConsumer'' until the VirtualMachine object exists and KubeVirt schedules the VMI pod. If you see the DV stuck there, check that the VM object was created: kubectl get vm -n vms kubectl get dv -n vms kubectl get vmi -n vms ==== Multus / SR-IOV (optional) ==== To attach a secondary interface instead of pod NAT, change ''values.yaml'': network: type: multus multusNetwork: "default/eth1-ipvlan" --- ===== Step 4 — Network Dashboard ===== **Chart:** ''active/network-dashboard''\\ **LoadBalancer IP:** ''192.168.0.207'' A Python ''ThreadingHTTPServer'' that queries the UniFi Network API and renders an interactive topology tree. ==== Features ==== * Auto-refresh every 3600 s * Clickable device cards → port drill-down modal (MAC, IP, speed, TX/RX) * WAN stats on UDM devices * Online/offline status with colored indicators ==== UniFi API ==== * Endpoint: ''/proxy/network/api/s/default/stat/device'' * Auth: ''X-API-KEY'' header * API key is stored in a Kubernetes Secret ''unifi-credentials'' (key: ''api-key'') The secret must be created manually before deploying (Fleet will not create it): kubectl create namespace network-dashboard kubectl create secret generic unifi-credentials \ --from-literal=api-key= \ -n network-dashboard ==== Readiness probe ==== No readiness probe is configured. The Zededa ZKS CNI prevents the kubelet from reaching pod IPs for HTTP probes. The ''/health'' endpoint exists in the app but is not wired to a probe. ==== Verify ==== kubectl get svc -n network-dashboard # EXTERNAL-IP should show 192.168.0.207 kubectl get pods -n network-dashboard # Should show 1/1 Running Then open ''http://192.168.0.207'' in a browser. --- ===== Step 5 — Grafana (inactive by default) ===== **Chart:** ''inactive/grafana'' (move to ''active/'' to enable) ^ Setting ^ Value ^ | LoadBalancer IP | 192.168.0.205 | | PostgreSQL | 192.168.0.2:5432 | | NFS server | 192.168.0.101 | | NFS path | /volume1/nano-r5s/grafana | | Plugin | yesoreyeram-infinity-datasource | ==== Storage ==== Uses a static PV/PVC pair (no dynamic provisioning). Both must have ''storageClassName: ""'' to bind correctly: storageClassName: "" # on both PV and PVC If you see ''pod has unbound immediate PersistentVolumeClaims'', verify ''storageClassName: ""'' is set on **both** the PV and PVC. PVC spec is immutable — delete and recreate if you need to fix it on a running cluster. ==== Verify IP after deploy ==== kubectl get svc -n grafana # Look for EXTERNAL-IP column --- ===== Troubleshooting ===== ==== Fleet doesn't apply a chart after moving it to active/ ==== Fleet re-syncs on git push. If a resource was silently skipped (e.g. CRD not ready at apply time), force a re-sync from Rancher UI: **Continuous Delivery → Git Repos → (repo) → Force Update** Or annotate the bundle from inside the cluster to trigger reconciliation. ==== VM object not created despite KubeVirt CRDs existing ==== Fleet applied the chart before KubeVirt CRDs were registered. Manually apply the VM: kubectl apply -f - <<'EOF' apiVersion: kubevirt.io/v1 kind: VirtualMachine metadata: name: ubuntu-vm1 namespace: vms spec: runStrategy: Always template: metadata: labels: kubevirt.io/domain: ubuntu-vm1 spec: domain: cpu: cores: 2 resources: requests: memory: 4Gi devices: disks: - name: rootdisk disk: bus: virtio - name: cloudinit disk: bus: virtio interfaces: - name: default masquerade: {} networks: - name: default pod: {} volumes: - name: rootdisk dataVolume: name: ubuntu-vm1-dv - name: cloudinit cloudInitNoCloud: secretRef: name: ubuntu-vm1-cloudinit EOF Then force a Fleet re-sync so Git remains the source of truth. ==== KubeVirt install Job shows i/o timeout ==== Transient ''dial tcp 10.43.0.1:443: i/o timeout'' errors during the installer Job are normal under load. The Job has ''restartPolicy: OnFailure'' and will retry. Check final status: kubectl get jobs -n kube-system | grep kubevirt # Should eventually show 1/1 ==== Zededa UI tags not syncing to Fleet cluster labels ==== ZedCloud tags are ZedCloud-only — they do not propagate to Rancher cluster labels or Fleet bundle selectors. Use Fleet workspace membership (cluster groups) for coarse targeting, not ''clusterSelector'' based on ZedCloud tags. --- ===== Cluster Group → Fleet Workspace Mapping ===== ^ Zededa concept ^ Fleet concept ^ | Cluster group | Fleet workspace | | Cluster tags | Not synced — ZedCloud only | | Fleet GitRepo | Scoped to a workspace | To target a specific cluster group, scope the Fleet GitRepo to the matching workspace in Rancher. --- ===== Quick Reference — Key IPs ===== ^ Resource ^ IP / Address ^ | MetalLB pool | 192.168.0.200 – 192.168.0.225 | | Grafana LB | 192.168.0.205 | | network-dashboard LB | 192.168.0.207 | | UniFi controller | 192.168.0.1 | | NFS server | 192.168.0.101 | | PostgreSQL | 192.168.0.2:5432 | | Ubuntu image HTTP | 192.168.0.101:1080 |