====== ZEDEDA Application Purge & Update ====== ===== Overview ===== //Purge & Update// is the ZEDEDA operation that resets an edge application instance's **mutable state** back to a clean slate while applying whatever configuration change triggered it. It is driven by a **counter** on the app instance object (``PurgeCmd`` in the device config; exposed in Terraform as a ``purge { counter = N }`` block). Incrementing the counter is the trigger — zedagent/zedmanager on the device sees the counter change and executes the purge sequence. **Key point:** Purge & Update wipes out any data on the node that lives on a mutable/purgeable volume. This is what distinguishes it from a plain refresh or restart. ===== What It Does (Step by Step) ===== On every purge or force update, ZEDEDA Cloud performs, in order: - Fetches the updated details from the edge application (if any) - Regenerates user data (certificate regeneration, if configured) - Issues the purge operation for the purgeable drives (volumes) On the device side: - The app instance is stopped. - **Mutable/purgeable volumes** are deleted and recreated from the base image — the app comes back in its pristine, just-deployed state. - **Persistent volumes** are left alone (a volume created without "Purge" checked is exempt). At the drive level in Terraform this is the ``ignorepurge`` flag on a ``drives`` block. - Updated config is applied — new image version, drive definitions, cloud-init/custom-config, vTPM setting, etc. - The instance restarts. **Cloud-init caveat:** cloud-init is only reapplied on a purge (or reactivate) **if the version string changed** (set in the meta-data ``instance-id`` field). If you edit variable group *values* after deployment without bumping the version, those edits will **not** be picked up by a purge — you'd need to delete and redeploy the instance instead. **Volume versioning:** volume instance names get a version suffix (e.g. ``_0``). Purge & Update increments this automatically. Normally the older version is deleted; if a snapshot was taken, the older version is preserved instead. A snapshot rollback deletes the newer volume version and reactivates the older one. ===== What Triggers It / What It's Used For ===== The GUI shows "Purge & Update" specifically when a change is destructive to mutable state: ^ Change ^ Action Triggered ^ Data Loss? ^ | Editing custom configuration | Purge & Update | Yes | | Editing Drives | Purge & Update | Yes | | Disabling vTPM on a VM App (Marketplace) then propagating | Purge & Update | Yes | | Editing Network ACL | Update/Refresh | No | | Editing CPU/Memory | Update/Refresh | No | | Editing VNC configuration | Update/Refresh | No | Also used deliberately in **Edge Sync** (airgapped) workflows: upgrade the application template in ZEDEDA Cloud, select affected instances, trigger Purge & Update — this updates the Application Instances Manifest in the device config. ===== When NOT To Use It ===== * Any instance with meaningful data on a **mutable (non-persistent) volume** you want to keep — purge deletes that content by design. * Instead of purge, **snapshot first** if you want rollback capability — but each instance can only hold **one snapshot at a time**, and **each snapshot can only be used for rollback once**. * If the change only touches CPU/memory, network ACLs, or VNC config — use plain refresh/restart, not purge. * Watch storage headroom: you need **at least 20% free space** on the persist partition to create a snapshot. EVE-OS 10.8.0+ required for snapshot-during-purge; 14.5.0-LTS+ for the standalone Create Snapshot action. * Purge will **not** pick up variable-group value edits by itself — don't rely on it for that. ===== How To Enable / Trigger It ===== ==== GUI ==== - Edge App Instances → select instance(s) → ellipsis (**...**) → **Purge & Update** - Optionally check **"Create a snapshot on purge & update"** - Confirm. The app stops briefly, ZEDEDA Cloud captures config/volume data (if snapshotting), applies the update, then restarts. ==== ZCLI ==== zcli edge-app-instance refresh --purge [--api-version=] Plain ``refresh`` (no ``--purge``) updates without wiping data. ==== REST API ==== PUT /v1/apps/instances/id/{id}/refresh/purge # refresh + purge (destructive) PUT /v1/apps/instances/id/{id}/refresh # refresh only (non-destructive) ==== Terraform ==== The ``zedcloud_application_instance`` resource exposes dedicated counter blocks: purge { counter = 0 # bump this to trigger a purge on next terraform apply } refresh { counter = 0 # bump this for a non-destructive refresh/restart } restart { counter = 0 # plain restart, no config re-application } Terraform diffs the counter against current state — incrementing it (e.g. ``0`` → ``1`` → ``2``) is what fires the operation on ``apply``. Practical pattern: drive the counter from a variable (e.g. ``var.purge_generation``) so the trigger is explicit and auditable rather than incidental. **Note:** EVE-OS purge/teardown is async — same caveat class as the NI Reconciler race condition. If a purge needs to complete before something downstream depends on it, pair the counter bump with explicit ordering (``depends_on``) or a ``time_sleep``, the same pattern used for the NI teardown workaround. ===== How To Prevent / Exempt Data From Being Purged ===== * **Mark volumes as persistent**: when adding a volume instance, leave "Purge" **unchecked** — this makes it persistent and exempt from Purge & Update. * **Per-drive exemption**: set ``ignorepurge = true`` on the specific ``drives`` block in config/Terraform — that drive is skipped during a purge even though the volume itself is mutable. * There is no "cancel" once a purge is triggered (counter-driven, one-way). Recovery options after the fact are a snapshot rollback (if one exists) or redeploying from a known-good state. ---- ====== Migrating a Critical Data-Plane VM to a New Version ====== For anything you'd call "important" — don't use Purge & Update as the primary migration tool. It's destructive-by-design with only a single-use snapshot as a safety net. Use a **blue/green** pattern instead: run old and new side by side, validate, then cut over. ===== Option A — Blue/Green (Preferred) ===== - **Deploy the new version as a brand-new app instance** (separate name) from the updated image/bundle. Don't touch the existing instance. - **Plan networking up front** — this is the part that differs by attachment type: * **Switch/local network instance**: attach the new instance to the same network instance with its own IP. Both run simultaneously with no conflict. * **Direct-attach (SR-IOV/passthrough)**: you generally can't have both instances holding the same physical interface at once. Stand up the new instance on a different adapter/VF (or a non-live mirror network instance), validate there, and only move the physical interface binding during the actual cutover — a short, controlled edit, not a purge. - **Validate the new instance** independently (test traffic, mirroring, whatever your data-plane validation process is). - **Cut over** — repoint whatever makes it "live" (interface reassignment, routing, LB target, DNS, VRRP/keepalive priority, etc.). - **Leave the old instance deactivated (not deleted)** for a bake/soak window — instant rollback by reactivating and flipping traffic back. - **Delete the old instance only once confident.** **Result:** real rollback (a live, working instance), zero data-plane downtime risk from the migration itself. ===== Option B — Snapshot + Purge & Update (Only If Blue/Green Isn't Feasible) ===== Use only when there's a hard constraint — e.g. a single physical interface with no spare capacity for a parallel instance. - Take a snapshot of the current instance (Storage tab → Create Snapshot, or via the "Create a snapshot on purge & update" checkbox). - Trigger Purge & Update to move to the new version in place. - If it fails → roll back to the snapshot. - **This causes an outage** during the purge/restart, and depends entirely on that one snapshot working correctly. ===== The Retry Loop (Repeating Snapshot → Purge → Rollback) ===== Yes — you can repeat this cycle as many times as needed. The "one snapshot, one rollback use" limit applies to that **specific snapshot object**, not to your ability to take new snapshots. - Instance running on old version (stable, post-rollback if applicable). - Fix whatever caused the prior failure. - Trigger Purge & Update again, checking **"Create a snapshot on purge & update"** — generates a fresh snapshot of the current (old-version) state. - New version comes up — validate it. - **If it fails again** → roll back to this new snapshot → back to old version → repeat. - **If it succeeds** → bake/soak, then delete the now-unneeded snapshot (frees storage — up to 100% of volume size for purgeable volumes; deletion of files happens on next deactivation, so restart to force cleanup immediately if needed). **Things to watch across repeated attempts:** * **Storage headroom**: need at least 20% free on the persist partition for each new snapshot. If a prior attempt's snapshot files haven't been cleaned up yet (deletion is async, tied to next deactivation), verify free space before the next attempt — restart to force cleanup if tight. * **This is still a disruptive in-place cycle** — each attempt is a real stop/restart of the data-plane VM. If you expect multiple failed attempts, that favors testing the new version out-of-band first (throwaway instance, or the blue/green approach above) so the "important" instance only goes through this cycle once, with high confidence. ---- //Source: ZEDEDA Help Center (Application Snapshot Overview, Manage an Edge Application Instance, Manage Application Snapshots, Use the ZEDEDA ZCLI, Custom Configuration Edge Application) and the `zededa/zedcloud` Terraform provider documentation.//