====== ZKS GitOps Demo Cluster — Deploy From Scratch ======
This guide covers a full from-scratch deployment of the ZKS demo cluster GitOps stack using Fleet (Rancher Continuous Delivery). It includes MetalLB, KubeVirt, CDI, Ubuntu VM bootstrap, and the UniFi network dashboard.
===== Architecture Overview =====
^ Layer ^ Tool ^ Purpose ^
| GitOps engine | Fleet (Rancher CD) | Watches Git repo, applies Helm charts |
| Load balancer | MetalLB | Assigns LAN IPs to LoadBalancer services |
| Virtualisation | KubeVirt + CDI | Runs VMs inside Kubernetes pods |
| Network visibility | network-dashboard | UniFi API → topology web UI |
| Observability | Grafana | Dashboards backed by PostgreSQL + NFS |
**Repo:** ''https://github.com/ManCalAzure/zks.git'' (branch: ''main'')\\
**Chart root:** ''zks/demo/charts/''
^ Folder ^ Purpose ^
| ''infrastructure/'' | Always-on platform components (MetalLB config, KubeVirt) |
| ''active/'' | Application workloads Fleet deploys automatically |
| ''inactive/'' | Charts present in Git but not watched by Fleet |
To activate a chart move it from ''inactive/'' → ''active/'' and push. To deactivate, move it back.
---
===== Prerequisites =====
* ZKS cluster imported into Rancher / Fleet
* Fleet GitRepo pointed at the repo above, path ''zks/demo/charts/''
* MetalLB **operator** already installed on the cluster (CRDs must exist before ''infrastructure/metallb-config'' is applied)
* An NFS share at ''192.168.0.101:/volume1/nano-r5s/'' (used by Grafana)
* An HTTP file server at ''http://192.168.0.101:1080/'' serving the Ubuntu cloud image
* PostgreSQL at ''192.168.0.2:5432'' (used by Grafana)
* UniFi controller at ''https://192.168.0.1'' with an API key
---
===== Step 1 — MetalLB IP Pool =====
**Chart:** ''infrastructure/metallb-config''
MetalLB is configured via two CRs: an ''IPAddressPool'' and an ''L2Advertisement''.
**IP pool:** ''192.168.0.200 – 192.168.0.225''
^ Service ^ LoadBalancer IP ^
| Grafana | 192.168.0.205 |
| network-dashboard | 192.168.0.207 |
This chart is always-on in ''infrastructure/''. Fleet applies it on every sync. No action needed after initial GitRepo setup.
**Verify:**
kubectl get ipaddresspool -n metallb-system
kubectl get l2advertisement -n metallb-system
---
===== Step 2 — KubeVirt Operator =====
**Chart:** ''infrastructure/kubevirt-operator''\\
**Versions:** KubeVirt ''v1.3.1'', CDI ''v1.60.1''
==== What it does ====
The chart runs a sequenced install via Helm hooks:
^ Hook weight ^ Job ^ What it does ^
| -6 | configmap | Renders the KubeVirt CR (with SR-IOV feature gate) |
| -5 | rbac | Creates ServiceAccount + ClusterRole in ''kube-system'' |
| 0 | install-job | Applies KubeVirt operator → waits for ready → applies CR → installs CDI |
SR-IOV is enabled by default (''sriov.enabled: true'').
==== Verify install ====
# Job should show Completed 1/1
kubectl get jobs -n kube-system | grep kubevirt
# KubeVirt operator pods
kubectl get pods -n kubevirt
# CDI pods
kubectl get pods -n cdi
# CRDs registered
kubectl get crd | grep kubevirt
kubectl get crd | grep cdi
The ''ubuntu-vm'' chart has a ''dependsOn'' referencing this bundle. Fleet will not apply ubuntu-vm until kubevirt-operator reaches Ready state. Do not remove this dependency.
---
===== Step 3 — Ubuntu VM =====
**Chart:** ''active/ubuntu-vm''
==== Overview ====
Deploys a KubeVirt ''VirtualMachine'' backed by a CDI ''DataVolume'' that imports the Ubuntu Noble cloud image over HTTP.
^ Setting ^ Value ^
| VM name | ubuntu-vm1 |
| Namespace | vms |
| vCPUs | 2 |
| Memory | 4 Gi |
| Disk | 20 Gi (''local-path'' storage class) |
| Run strategy | Always |
| Network | Pod network (masquerade / NAT) |
| Image source | ''http://192.168.0.101:1080/noble-server-cloudimg-amd64.img'' |
==== Cloud-init ====
Cloud-init is stored as a multipart MIME file at ''active/ubuntu-vm/files/cloud-init.txt'' and mounted as a Kubernetes Secret (key: ''userdata'').
**What cloud-init configures:**
* Hostname: ''VM1''
* User: ''manny'' (sudo, no password, SSH key injected)
* Password auth enabled on SSH
* Packages: ''wget'', ''curl'', ''traceroute''
* Network: DHCP on ''enp1s0''
To change the SSH key or hostname edit ''files/cloud-init.txt'' and push.
==== KubeVirt volume reference ====
The cloud-init secret is referenced in the VM spec as:
cloudInitNoCloud:
secretRef:
name: ubuntu-vm1-cloudinit
The field is ''secretRef'', **not** ''userDataSecretRef''. KubeVirt v1.3.1 strict decoding will reject the latter with a BadRequest error.
==== DataVolume lifecycle ====
^ Phase ^ Meaning ^
| WaitForFirstConsumer | PVC waiting for a pod to be scheduled (normal for ''local-path'') |
| ImportScheduled | CDI importer pod scheduled |
| ImportInProgress | Image downloading from HTTP server |
| Succeeded | Import complete, VM can boot |
The DataVolume stays in ''WaitForFirstConsumer'' until the VirtualMachine object exists and KubeVirt schedules the VMI pod. If you see the DV stuck there, check that the VM object was created:
kubectl get vm -n vms
kubectl get dv -n vms
kubectl get vmi -n vms
==== Multus / SR-IOV (optional) ====
To attach a secondary interface instead of pod NAT, change ''values.yaml'':
network:
type: multus
multusNetwork: "default/eth1-ipvlan"
---
===== Step 4 — Network Dashboard =====
**Chart:** ''active/network-dashboard''\\
**LoadBalancer IP:** ''192.168.0.207''
A Python ''ThreadingHTTPServer'' that queries the UniFi Network API and renders an interactive topology tree.
==== Features ====
* Auto-refresh every 3600 s
* Clickable device cards → port drill-down modal (MAC, IP, speed, TX/RX)
* WAN stats on UDM devices
* Online/offline status with colored indicators
==== UniFi API ====
* Endpoint: ''/proxy/network/api/s/default/stat/device''
* Auth: ''X-API-KEY'' header
* API key is stored in a Kubernetes Secret ''unifi-credentials'' (key: ''api-key'')
The secret must be created manually before deploying (Fleet will not create it):
kubectl create namespace network-dashboard
kubectl create secret generic unifi-credentials \
--from-literal=api-key= \
-n network-dashboard
==== Readiness probe ====
No readiness probe is configured. The Zededa ZKS CNI prevents the kubelet from reaching pod IPs for HTTP probes. The ''/health'' endpoint exists in the app but is not wired to a probe.
==== Verify ====
kubectl get svc -n network-dashboard
# EXTERNAL-IP should show 192.168.0.207
kubectl get pods -n network-dashboard
# Should show 1/1 Running
Then open ''http://192.168.0.207'' in a browser.
---
===== Step 5 — Grafana (inactive by default) =====
**Chart:** ''inactive/grafana'' (move to ''active/'' to enable)
^ Setting ^ Value ^
| LoadBalancer IP | 192.168.0.205 |
| PostgreSQL | 192.168.0.2:5432 |
| NFS server | 192.168.0.101 |
| NFS path | /volume1/nano-r5s/grafana |
| Plugin | yesoreyeram-infinity-datasource |
==== Storage ====
Uses a static PV/PVC pair (no dynamic provisioning). Both must have ''storageClassName: ""'' to bind correctly:
storageClassName: "" # on both PV and PVC
If you see ''pod has unbound immediate PersistentVolumeClaims'', verify ''storageClassName: ""'' is set on **both** the PV and PVC. PVC spec is immutable — delete and recreate if you need to fix it on a running cluster.
==== Verify IP after deploy ====
kubectl get svc -n grafana
# Look for EXTERNAL-IP column
---
===== Troubleshooting =====
==== Fleet doesn't apply a chart after moving it to active/ ====
Fleet re-syncs on git push. If a resource was silently skipped (e.g. CRD not ready at apply time), force a re-sync from Rancher UI:
**Continuous Delivery → Git Repos → (repo) → Force Update**
Or annotate the bundle from inside the cluster to trigger reconciliation.
==== VM object not created despite KubeVirt CRDs existing ====
Fleet applied the chart before KubeVirt CRDs were registered. Manually apply the VM:
kubectl apply -f - <<'EOF'
apiVersion: kubevirt.io/v1
kind: VirtualMachine
metadata:
name: ubuntu-vm1
namespace: vms
spec:
runStrategy: Always
template:
metadata:
labels:
kubevirt.io/domain: ubuntu-vm1
spec:
domain:
cpu:
cores: 2
resources:
requests:
memory: 4Gi
devices:
disks:
- name: rootdisk
disk:
bus: virtio
- name: cloudinit
disk:
bus: virtio
interfaces:
- name: default
masquerade: {}
networks:
- name: default
pod: {}
volumes:
- name: rootdisk
dataVolume:
name: ubuntu-vm1-dv
- name: cloudinit
cloudInitNoCloud:
secretRef:
name: ubuntu-vm1-cloudinit
EOF
Then force a Fleet re-sync so Git remains the source of truth.
==== KubeVirt install Job shows i/o timeout ====
Transient ''dial tcp 10.43.0.1:443: i/o timeout'' errors during the installer Job are normal under load. The Job has ''restartPolicy: OnFailure'' and will retry. Check final status:
kubectl get jobs -n kube-system | grep kubevirt
# Should eventually show 1/1
==== Zededa UI tags not syncing to Fleet cluster labels ====
ZedCloud tags are ZedCloud-only — they do not propagate to Rancher cluster labels or Fleet bundle selectors. Use Fleet workspace membership (cluster groups) for coarse targeting, not ''clusterSelector'' based on ZedCloud tags.
---
===== Cluster Group → Fleet Workspace Mapping =====
^ Zededa concept ^ Fleet concept ^
| Cluster group | Fleet workspace |
| Cluster tags | Not synced — ZedCloud only |
| Fleet GitRepo | Scoped to a workspace |
To target a specific cluster group, scope the Fleet GitRepo to the matching workspace in Rancher.
---
===== Quick Reference — Key IPs =====
^ Resource ^ IP / Address ^
| MetalLB pool | 192.168.0.200 – 192.168.0.225 |
| Grafana LB | 192.168.0.205 |
| network-dashboard LB | 192.168.0.207 |
| UniFi controller | 192.168.0.1 |
| NFS server | 192.168.0.101 |
| PostgreSQL | 192.168.0.2:5432 |
| Ubuntu image HTTP | 192.168.0.101:1080 |