User Tools

Site Tools


zks:zks-demo-cluster-deploy

ZKS GitOps Demo Cluster — Deploy From Scratch

This guide covers a full from-scratch deployment of the ZKS demo cluster GitOps stack using Fleet (Rancher Continuous Delivery). It includes MetalLB, KubeVirt, CDI, Ubuntu VM bootstrap, and the UniFi network dashboard.

Architecture Overview

Layer Tool Purpose
GitOps engine Fleet (Rancher CD) Watches Git repo, applies Helm charts
Load balancer MetalLB Assigns LAN IPs to LoadBalancer services
Virtualisation KubeVirt + CDI Runs VMs inside Kubernetes pods
Network visibility network-dashboard UniFi API → topology web UI
Observability Grafana Dashboards backed by PostgreSQL + NFS

Repo: https://github.com/ManCalAzure/zks.git (branch: main)
Chart root: zks/demo/charts/

Folder Purpose
infrastructure/ Always-on platform components (MetalLB config, KubeVirt)
active/ Application workloads Fleet deploys automatically
inactive/ Charts present in Git but not watched by Fleet

To activate a chart move it from inactive/ → active/ and push. To deactivate, move it back.

—

Prerequisites

  • ZKS cluster imported into Rancher / Fleet
  • Fleet GitRepo pointed at the repo above, path zks/demo/charts/
  • MetalLB operator already installed on the cluster (CRDs must exist before infrastructure/metallb-config is applied)
  • An NFS share at 192.168.0.101:/volume1/nano-r5s/ (used by Grafana)
  • An HTTP file server at http://192.168.0.101:1080/ serving the Ubuntu cloud image
  • PostgreSQL at 192.168.0.2:5432 (used by Grafana)
  • UniFi controller at https://192.168.0.1 with an API key

—

Step 1 — MetalLB IP Pool

Chart: infrastructure/metallb-config

MetalLB is configured via two CRs: an IPAddressPool and an L2Advertisement.

IP pool: 192.168.0.200 – 192.168.0.225

Service LoadBalancer IP
Grafana 192.168.0.205
network-dashboard 192.168.0.207

This chart is always-on in infrastructure/. Fleet applies it on every sync. No action needed after initial GitRepo setup.

Verify:

kubectl get ipaddresspool -n metallb-system
kubectl get l2advertisement -n metallb-system

—

Step 2 — KubeVirt Operator

Chart: infrastructure/kubevirt-operator
Versions: KubeVirt v1.3.1, CDI v1.60.1

What it does

The chart runs a sequenced install via Helm hooks:

Hook weight Job What it does
-6 configmap Renders the KubeVirt CR (with SR-IOV feature gate)
-5 rbac Creates ServiceAccount + ClusterRole in kube-system
0 install-job Applies KubeVirt operator → waits for ready → applies CR → installs CDI

SR-IOV is enabled by default (sriov.enabled: true).

Verify install

# Job should show Completed 1/1
kubectl get jobs -n kube-system | grep kubevirt
 
# KubeVirt operator pods
kubectl get pods -n kubevirt
 
# CDI pods
kubectl get pods -n cdi
 
# CRDs registered
kubectl get crd | grep kubevirt
kubectl get crd | grep cdi

<note important> The ubuntu-vm chart has a dependsOn referencing this bundle. Fleet will not apply ubuntu-vm until kubevirt-operator reaches Ready state. Do not remove this dependency. </note>

—

Step 3 — Ubuntu VM

Chart: active/ubuntu-vm

Overview

Deploys a KubeVirt VirtualMachine backed by a CDI DataVolume that imports the Ubuntu Noble cloud image over HTTP.

Setting Value
VM name ubuntu-vm1
Namespace vms
vCPUs 2
Memory 4 Gi
Disk 20 Gi (local-path storage class)
Run strategy Always
Network Pod network (masquerade / NAT)
Image source http://192.168.0.101:1080/noble-server-cloudimg-amd64.img

Cloud-init

Cloud-init is stored as a multipart MIME file at active/ubuntu-vm/files/cloud-init.txt and mounted as a Kubernetes Secret (key: userdata).

What cloud-init configures:

  • Hostname: VM1
  • User: manny (sudo, no password, SSH key injected)
  • Password auth enabled on SSH
  • Packages: wget, curl, traceroute
  • Network: DHCP on enp1s0

To change the SSH key or hostname edit files/cloud-init.txt and push.

KubeVirt volume reference

The cloud-init secret is referenced in the VM spec as:

cloudInitNoCloud:
  secretRef:
    name: ubuntu-vm1-cloudinit

<note warning> The field is secretRef, not userDataSecretRef. KubeVirt v1.3.1 strict decoding will reject the latter with a BadRequest error. </note>

DataVolume lifecycle

Phase Meaning
WaitForFirstConsumer PVC waiting for a pod to be scheduled (normal for local-path)
ImportScheduled CDI importer pod scheduled
ImportInProgress Image downloading from HTTP server
Succeeded Import complete, VM can boot

The DataVolume stays in WaitForFirstConsumer until the VirtualMachine object exists and KubeVirt schedules the VMI pod. If you see the DV stuck there, check that the VM object was created:

kubectl get vm -n vms
kubectl get dv -n vms
kubectl get vmi -n vms

Multus / SR-IOV (optional)

To attach a secondary interface instead of pod NAT, change values.yaml:

network:
  type: multus
  multusNetwork: "default/eth1-ipvlan"

—

Step 4 — Network Dashboard

Chart: active/network-dashboard
LoadBalancer IP: 192.168.0.207

A Python ThreadingHTTPServer that queries the UniFi Network API and renders an interactive topology tree.

Features

  • Auto-refresh every 3600 s
  • Clickable device cards → port drill-down modal (MAC, IP, speed, TX/RX)
  • WAN stats on UDM devices
  • Online/offline status with colored indicators

UniFi API

  • Endpoint: /proxy/network/api/s/default/stat/device
  • Auth: X-API-KEY header
  • API key is stored in a Kubernetes Secret unifi-credentials (key: api-key)

The secret must be created manually before deploying (Fleet will not create it):

kubectl create namespace network-dashboard
kubectl create secret generic unifi-credentials \
  --from-literal=api-key=<YOUR_API_KEY> \
  -n network-dashboard

Readiness probe

No readiness probe is configured. The Zededa ZKS CNI prevents the kubelet from reaching pod IPs for HTTP probes. The /health endpoint exists in the app but is not wired to a probe.

Verify

kubectl get svc -n network-dashboard
# EXTERNAL-IP should show 192.168.0.207
 
kubectl get pods -n network-dashboard
# Should show 1/1 Running

Then open http://192.168.0.207 in a browser.

—

Step 5 — Grafana (inactive by default)

Chart: inactive/grafana (move to active/ to enable)

Setting Value
LoadBalancer IP 192.168.0.205
PostgreSQL 192.168.0.2:5432
NFS server 192.168.0.101
NFS path /volume1/nano-r5s/grafana
Plugin yesoreyeram-infinity-datasource

Storage

Uses a static PV/PVC pair (no dynamic provisioning). Both must have storageClassName: “” to bind correctly:

storageClassName: ""   # on both PV and PVC

<note warning> If you see pod has unbound immediate PersistentVolumeClaims, verify storageClassName: “” is set on both the PV and PVC. PVC spec is immutable — delete and recreate if you need to fix it on a running cluster. </note>

Verify IP after deploy

kubectl get svc -n grafana
# Look for EXTERNAL-IP column

—

Troubleshooting

Fleet doesn't apply a chart after moving it to active/

Fleet re-syncs on git push. If a resource was silently skipped (e.g. CRD not ready at apply time), force a re-sync from Rancher UI:

Continuous Delivery → Git Repos → (repo) → Force Update

Or annotate the bundle from inside the cluster to trigger reconciliation.

VM object not created despite KubeVirt CRDs existing

Fleet applied the chart before KubeVirt CRDs were registered. Manually apply the VM:

kubectl apply -f - <<'EOF'
apiVersion: kubevirt.io/v1
kind: VirtualMachine
metadata:
  name: ubuntu-vm1
  namespace: vms
spec:
  runStrategy: Always
  template:
    metadata:
      labels:
        kubevirt.io/domain: ubuntu-vm1
    spec:
      domain:
        cpu:
          cores: 2
        resources:
          requests:
            memory: 4Gi
        devices:
          disks:
            - name: rootdisk
              disk:
                bus: virtio
            - name: cloudinit
              disk:
                bus: virtio
          interfaces:
            - name: default
              masquerade: {}
      networks:
        - name: default
          pod: {}
      volumes:
        - name: rootdisk
          dataVolume:
            name: ubuntu-vm1-dv
        - name: cloudinit
          cloudInitNoCloud:
            secretRef:
              name: ubuntu-vm1-cloudinit
EOF

Then force a Fleet re-sync so Git remains the source of truth.

KubeVirt install Job shows i/o timeout

Transient dial tcp 10.43.0.1:443: i/o timeout errors during the installer Job are normal under load. The Job has restartPolicy: OnFailure and will retry. Check final status:

kubectl get jobs -n kube-system | grep kubevirt
# Should eventually show 1/1

Zededa UI tags not syncing to Fleet cluster labels

ZedCloud tags are ZedCloud-only — they do not propagate to Rancher cluster labels or Fleet bundle selectors. Use Fleet workspace membership (cluster groups) for coarse targeting, not clusterSelector based on ZedCloud tags.

—

Cluster Group → Fleet Workspace Mapping

Zededa concept Fleet concept
Cluster group Fleet workspace
Cluster tags Not synced — ZedCloud only
Fleet GitRepo Scoped to a workspace

To target a specific cluster group, scope the Fleet GitRepo to the matching workspace in Rancher.

—

Quick Reference — Key IPs

Resource IP / Address
MetalLB pool 192.168.0.200 – 192.168.0.225
Grafana LB 192.168.0.205
network-dashboard LB 192.168.0.207
UniFi controller 192.168.0.1
NFS server 192.168.0.101
PostgreSQL 192.168.0.2:5432
Ubuntu image HTTP 192.168.0.101:1080
zks/zks-demo-cluster-deploy.txt · Last modified: by mc