Table of Contents
ZKS GitOps Demo Cluster — Deploy From Scratch
This guide covers a full from-scratch deployment of the ZKS demo cluster GitOps stack using Fleet (Rancher Continuous Delivery). It includes MetalLB, KubeVirt, CDI, Ubuntu VM bootstrap, and the UniFi network dashboard.
Architecture Overview
| Layer | Tool | Purpose |
|---|---|---|
| GitOps engine | Fleet (Rancher CD) | Watches Git repo, applies Helm charts |
| Load balancer | MetalLB | Assigns LAN IPs to LoadBalancer services |
| Virtualisation | KubeVirt + CDI | Runs VMs inside Kubernetes pods |
| Network visibility | network-dashboard | UniFi API → topology web UI |
| Observability | Grafana | Dashboards backed by PostgreSQL + NFS |
Repo: https://github.com/ManCalAzure/zks.git (branch: main)
Chart root: zks/demo/charts/
| Folder | Purpose |
|---|---|
infrastructure/ | Always-on platform components (MetalLB config, KubeVirt) |
active/ | Application workloads Fleet deploys automatically |
inactive/ | Charts present in Git but not watched by Fleet |
To activate a chart move it from inactive/ → active/ and push. To deactivate, move it back.
—
Prerequisites
- ZKS cluster imported into Rancher / Fleet
- Fleet GitRepo pointed at the repo above, path
zks/demo/charts/ - MetalLB operator already installed on the cluster (CRDs must exist before
infrastructure/metallb-configis applied) - An NFS share at
192.168.0.101:/volume1/nano-r5s/(used by Grafana) - An HTTP file server at
http://192.168.0.101:1080/serving the Ubuntu cloud image - PostgreSQL at
192.168.0.2:5432(used by Grafana) - UniFi controller at
https://192.168.0.1with an API key
—
Step 1 — MetalLB IP Pool
Chart: infrastructure/metallb-config
MetalLB is configured via two CRs: an IPAddressPool and an L2Advertisement.
IP pool: 192.168.0.200 – 192.168.0.225
| Service | LoadBalancer IP |
|---|---|
| Grafana | 192.168.0.205 |
| network-dashboard | 192.168.0.207 |
This chart is always-on in infrastructure/. Fleet applies it on every sync. No action needed after initial GitRepo setup.
Verify:
kubectl get ipaddresspool -n metallb-system kubectl get l2advertisement -n metallb-system
—
Step 2 — KubeVirt Operator
Chart: infrastructure/kubevirt-operator
Versions: KubeVirt v1.3.1, CDI v1.60.1
What it does
The chart runs a sequenced install via Helm hooks:
| Hook weight | Job | What it does |
|---|---|---|
| -6 | configmap | Renders the KubeVirt CR (with SR-IOV feature gate) |
| -5 | rbac | Creates ServiceAccount + ClusterRole in kube-system |
| 0 | install-job | Applies KubeVirt operator → waits for ready → applies CR → installs CDI |
SR-IOV is enabled by default (sriov.enabled: true).
Verify install
# Job should show Completed 1/1 kubectl get jobs -n kube-system | grep kubevirt # KubeVirt operator pods kubectl get pods -n kubevirt # CDI pods kubectl get pods -n cdi # CRDs registered kubectl get crd | grep kubevirt kubectl get crd | grep cdi
<note important>
The ubuntu-vm chart has a dependsOn referencing this bundle. Fleet will not apply ubuntu-vm until kubevirt-operator reaches Ready state. Do not remove this dependency.
</note>
—
Step 3 — Ubuntu VM
Chart: active/ubuntu-vm
Overview
Deploys a KubeVirt VirtualMachine backed by a CDI DataVolume that imports the Ubuntu Noble cloud image over HTTP.
| Setting | Value |
|---|---|
| VM name | ubuntu-vm1 |
| Namespace | vms |
| vCPUs | 2 |
| Memory | 4 Gi |
| Disk | 20 Gi (local-path storage class) |
| Run strategy | Always |
| Network | Pod network (masquerade / NAT) |
| Image source | http://192.168.0.101:1080/noble-server-cloudimg-amd64.img |
Cloud-init
Cloud-init is stored as a multipart MIME file at active/ubuntu-vm/files/cloud-init.txt and mounted as a Kubernetes Secret (key: userdata).
What cloud-init configures:
- Hostname:
VM1 - User:
manny(sudo, no password, SSH key injected) - Password auth enabled on SSH
- Packages:
wget,curl,traceroute - Network: DHCP on
enp1s0
To change the SSH key or hostname edit files/cloud-init.txt and push.
KubeVirt volume reference
The cloud-init secret is referenced in the VM spec as:
cloudInitNoCloud: secretRef: name: ubuntu-vm1-cloudinit
<note warning>
The field is secretRef, not userDataSecretRef. KubeVirt v1.3.1 strict decoding will reject the latter with a BadRequest error.
</note>
DataVolume lifecycle
| Phase | Meaning |
|---|---|
| WaitForFirstConsumer | PVC waiting for a pod to be scheduled (normal for local-path) |
| ImportScheduled | CDI importer pod scheduled |
| ImportInProgress | Image downloading from HTTP server |
| Succeeded | Import complete, VM can boot |
The DataVolume stays in WaitForFirstConsumer until the VirtualMachine object exists and KubeVirt schedules the VMI pod. If you see the DV stuck there, check that the VM object was created:
kubectl get vm -n vms kubectl get dv -n vms kubectl get vmi -n vms
Multus / SR-IOV (optional)
To attach a secondary interface instead of pod NAT, change values.yaml:
network: type: multus multusNetwork: "default/eth1-ipvlan"
—
Step 4 — Network Dashboard
Chart: active/network-dashboard
LoadBalancer IP: 192.168.0.207
A Python ThreadingHTTPServer that queries the UniFi Network API and renders an interactive topology tree.
Features
- Auto-refresh every 3600 s
- Clickable device cards → port drill-down modal (MAC, IP, speed, TX/RX)
- WAN stats on UDM devices
- Online/offline status with colored indicators
UniFi API
- Endpoint:
/proxy/network/api/s/default/stat/device - Auth:
X-API-KEYheader - API key is stored in a Kubernetes Secret
unifi-credentials(key:api-key)
The secret must be created manually before deploying (Fleet will not create it):
kubectl create namespace network-dashboard kubectl create secret generic unifi-credentials \ --from-literal=api-key=<YOUR_API_KEY> \ -n network-dashboard
Readiness probe
No readiness probe is configured. The Zededa ZKS CNI prevents the kubelet from reaching pod IPs for HTTP probes. The /health endpoint exists in the app but is not wired to a probe.
Verify
kubectl get svc -n network-dashboard # EXTERNAL-IP should show 192.168.0.207 kubectl get pods -n network-dashboard # Should show 1/1 Running
Then open http://192.168.0.207 in a browser.
—
Step 5 — Grafana (inactive by default)
Chart: inactive/grafana (move to active/ to enable)
| Setting | Value |
|---|---|
| LoadBalancer IP | 192.168.0.205 |
| PostgreSQL | 192.168.0.2:5432 |
| NFS server | 192.168.0.101 |
| NFS path | /volume1/nano-r5s/grafana |
| Plugin | yesoreyeram-infinity-datasource |
Storage
Uses a static PV/PVC pair (no dynamic provisioning). Both must have storageClassName: “” to bind correctly:
storageClassName: "" # on both PV and PVC
<note warning>
If you see pod has unbound immediate PersistentVolumeClaims, verify storageClassName: “” is set on both the PV and PVC. PVC spec is immutable — delete and recreate if you need to fix it on a running cluster.
</note>
Verify IP after deploy
kubectl get svc -n grafana # Look for EXTERNAL-IP column
—
Troubleshooting
Fleet doesn't apply a chart after moving it to active/
Fleet re-syncs on git push. If a resource was silently skipped (e.g. CRD not ready at apply time), force a re-sync from Rancher UI:
Continuous Delivery → Git Repos → (repo) → Force Update
Or annotate the bundle from inside the cluster to trigger reconciliation.
VM object not created despite KubeVirt CRDs existing
Fleet applied the chart before KubeVirt CRDs were registered. Manually apply the VM:
kubectl apply -f - <<'EOF' apiVersion: kubevirt.io/v1 kind: VirtualMachine metadata: name: ubuntu-vm1 namespace: vms spec: runStrategy: Always template: metadata: labels: kubevirt.io/domain: ubuntu-vm1 spec: domain: cpu: cores: 2 resources: requests: memory: 4Gi devices: disks: - name: rootdisk disk: bus: virtio - name: cloudinit disk: bus: virtio interfaces: - name: default masquerade: {} networks: - name: default pod: {} volumes: - name: rootdisk dataVolume: name: ubuntu-vm1-dv - name: cloudinit cloudInitNoCloud: secretRef: name: ubuntu-vm1-cloudinit EOF
Then force a Fleet re-sync so Git remains the source of truth.
KubeVirt install Job shows i/o timeout
Transient dial tcp 10.43.0.1:443: i/o timeout errors during the installer Job are normal under load. The Job has restartPolicy: OnFailure and will retry. Check final status:
kubectl get jobs -n kube-system | grep kubevirt # Should eventually show 1/1
Zededa UI tags not syncing to Fleet cluster labels
ZedCloud tags are ZedCloud-only — they do not propagate to Rancher cluster labels or Fleet bundle selectors. Use Fleet workspace membership (cluster groups) for coarse targeting, not clusterSelector based on ZedCloud tags.
—
Cluster Group → Fleet Workspace Mapping
| Zededa concept | Fleet concept |
|---|---|
| Cluster group | Fleet workspace |
| Cluster tags | Not synced — ZedCloud only |
| Fleet GitRepo | Scoped to a workspace |
To target a specific cluster group, scope the Fleet GitRepo to the matching workspace in Rancher.
—
Quick Reference — Key IPs
| Resource | IP / Address |
|---|---|
| MetalLB pool | 192.168.0.200 – 192.168.0.225 |
| Grafana LB | 192.168.0.205 |
| network-dashboard LB | 192.168.0.207 |
| UniFi controller | 192.168.0.1 |
| NFS server | 192.168.0.101 |
| PostgreSQL | 192.168.0.2:5432 |
| Ubuntu image HTTP | 192.168.0.101:1080 |
