====== ZEDEDA Application Purge & Update ======
===== Overview =====
//Purge & Update// is the ZEDEDA operation that resets an edge application instance's **mutable state** back to a clean slate while applying whatever configuration change triggered it. It is driven by a **counter** on the app instance object (``PurgeCmd`` in the device config; exposed in Terraform as a ``purge { counter = N }`` block). Incrementing the counter is the trigger — zedagent/zedmanager on the device sees the counter change and executes the purge sequence.
**Key point:** Purge & Update wipes out any data on the node that lives on a mutable/purgeable volume. This is what distinguishes it from a plain refresh or restart.
===== What It Does (Step by Step) =====
On every purge or force update, ZEDEDA Cloud performs, in order:
- Fetches the updated details from the edge application (if any)
- Regenerates user data (certificate regeneration, if configured)
- Issues the purge operation for the purgeable drives (volumes)
On the device side:
- The app instance is stopped.
- **Mutable/purgeable volumes** are deleted and recreated from the base image — the app comes back in its pristine, just-deployed state.
- **Persistent volumes** are left alone (a volume created without "Purge" checked is exempt). At the drive level in Terraform this is the ``ignorepurge`` flag on a ``drives`` block.
- Updated config is applied — new image version, drive definitions, cloud-init/custom-config, vTPM setting, etc.
- The instance restarts.
**Cloud-init caveat:** cloud-init is only reapplied on a purge (or reactivate) **if the version string changed** (set in the meta-data ``instance-id`` field). If you edit variable group *values* after deployment without bumping the version, those edits will **not** be picked up by a purge — you'd need to delete and redeploy the instance instead.
**Volume versioning:** volume instance names get a version suffix (e.g. ``_0``). Purge & Update increments this automatically. Normally the older version is deleted; if a snapshot was taken, the older version is preserved instead. A snapshot rollback deletes the newer volume version and reactivates the older one.
===== What Triggers It / What It's Used For =====
The GUI shows "Purge & Update" specifically when a change is destructive to mutable state:
^ Change ^ Action Triggered ^ Data Loss? ^
| Editing custom configuration | Purge & Update | Yes |
| Editing Drives | Purge & Update | Yes |
| Disabling vTPM on a VM App (Marketplace) then propagating | Purge & Update | Yes |
| Editing Network ACL | Update/Refresh | No |
| Editing CPU/Memory | Update/Refresh | No |
| Editing VNC configuration | Update/Refresh | No |
Also used deliberately in **Edge Sync** (airgapped) workflows: upgrade the application template in ZEDEDA Cloud, select affected instances, trigger Purge & Update — this updates the Application Instances Manifest in the device config.
===== When NOT To Use It =====
* Any instance with meaningful data on a **mutable (non-persistent) volume** you want to keep — purge deletes that content by design.
* Instead of purge, **snapshot first** if you want rollback capability — but each instance can only hold **one snapshot at a time**, and **each snapshot can only be used for rollback once**.
* If the change only touches CPU/memory, network ACLs, or VNC config — use plain refresh/restart, not purge.
* Watch storage headroom: you need **at least 20% free space** on the persist partition to create a snapshot. EVE-OS 10.8.0+ required for snapshot-during-purge; 14.5.0-LTS+ for the standalone Create Snapshot action.
* Purge will **not** pick up variable-group value edits by itself — don't rely on it for that.
===== How To Enable / Trigger It =====
==== GUI ====
- Edge App Instances → select instance(s) → ellipsis (**...**) → **Purge & Update**
- Optionally check **"Create a snapshot on purge & update"**
- Confirm. The app stops briefly, ZEDEDA Cloud captures config/volume data (if snapshotting), applies the update, then restarts.
==== ZCLI ====
zcli edge-app-instance refresh --purge [--api-version=]
Plain ``refresh`` (no ``--purge``) updates without wiping data.
==== REST API ====
PUT /v1/apps/instances/id/{id}/refresh/purge # refresh + purge (destructive)
PUT /v1/apps/instances/id/{id}/refresh # refresh only (non-destructive)
==== Terraform ====
The ``zedcloud_application_instance`` resource exposes dedicated counter blocks:
purge {
counter = 0 # bump this to trigger a purge on next terraform apply
}
refresh {
counter = 0 # bump this for a non-destructive refresh/restart
}
restart {
counter = 0 # plain restart, no config re-application
}
Terraform diffs the counter against current state — incrementing it (e.g. ``0`` → ``1`` → ``2``) is what fires the operation on ``apply``. Practical pattern: drive the counter from a variable (e.g. ``var.purge_generation``) so the trigger is explicit and auditable rather than incidental.
**Note:** EVE-OS purge/teardown is async — same caveat class as the NI Reconciler race condition. If a purge needs to complete before something downstream depends on it, pair the counter bump with explicit ordering (``depends_on``) or a ``time_sleep``, the same pattern used for the NI teardown workaround.
===== How To Prevent / Exempt Data From Being Purged =====
* **Mark volumes as persistent**: when adding a volume instance, leave "Purge" **unchecked** — this makes it persistent and exempt from Purge & Update.
* **Per-drive exemption**: set ``ignorepurge = true`` on the specific ``drives`` block in config/Terraform — that drive is skipped during a purge even though the volume itself is mutable.
* There is no "cancel" once a purge is triggered (counter-driven, one-way). Recovery options after the fact are a snapshot rollback (if one exists) or redeploying from a known-good state.
----
====== Migrating a Critical Data-Plane VM to a New Version ======
For anything you'd call "important" — don't use Purge & Update as the primary migration tool. It's destructive-by-design with only a single-use snapshot as a safety net. Use a **blue/green** pattern instead: run old and new side by side, validate, then cut over.
===== Option A — Blue/Green (Preferred) =====
- **Deploy the new version as a brand-new app instance** (separate name) from the updated image/bundle. Don't touch the existing instance.
- **Plan networking up front** — this is the part that differs by attachment type:
* **Switch/local network instance**: attach the new instance to the same network instance with its own IP. Both run simultaneously with no conflict.
* **Direct-attach (SR-IOV/passthrough)**: you generally can't have both instances holding the same physical interface at once. Stand up the new instance on a different adapter/VF (or a non-live mirror network instance), validate there, and only move the physical interface binding during the actual cutover — a short, controlled edit, not a purge.
- **Validate the new instance** independently (test traffic, mirroring, whatever your data-plane validation process is).
- **Cut over** — repoint whatever makes it "live" (interface reassignment, routing, LB target, DNS, VRRP/keepalive priority, etc.).
- **Leave the old instance deactivated (not deleted)** for a bake/soak window — instant rollback by reactivating and flipping traffic back.
- **Delete the old instance only once confident.**
**Result:** real rollback (a live, working instance), zero data-plane downtime risk from the migration itself.
===== Option B — Snapshot + Purge & Update (Only If Blue/Green Isn't Feasible) =====
Use only when there's a hard constraint — e.g. a single physical interface with no spare capacity for a parallel instance.
- Take a snapshot of the current instance (Storage tab → Create Snapshot, or via the "Create a snapshot on purge & update" checkbox).
- Trigger Purge & Update to move to the new version in place.
- If it fails → roll back to the snapshot.
- **This causes an outage** during the purge/restart, and depends entirely on that one snapshot working correctly.
===== The Retry Loop (Repeating Snapshot → Purge → Rollback) =====
Yes — you can repeat this cycle as many times as needed. The "one snapshot, one rollback use" limit applies to that **specific snapshot object**, not to your ability to take new snapshots.
- Instance running on old version (stable, post-rollback if applicable).
- Fix whatever caused the prior failure.
- Trigger Purge & Update again, checking **"Create a snapshot on purge & update"** — generates a fresh snapshot of the current (old-version) state.
- New version comes up — validate it.
- **If it fails again** → roll back to this new snapshot → back to old version → repeat.
- **If it succeeds** → bake/soak, then delete the now-unneeded snapshot (frees storage — up to 100% of volume size for purgeable volumes; deletion of files happens on next deactivation, so restart to force cleanup immediately if needed).
**Things to watch across repeated attempts:**
* **Storage headroom**: need at least 20% free on the persist partition for each new snapshot. If a prior attempt's snapshot files haven't been cleaned up yet (deletion is async, tied to next deactivation), verify free space before the next attempt — restart to force cleanup if tight.
* **This is still a disruptive in-place cycle** — each attempt is a real stop/restart of the data-plane VM. If you expect multiple failed attempts, that favors testing the new version out-of-band first (throwaway instance, or the blue/green approach above) so the "important" instance only goes through this cycle once, with high confidence.
----
//Source: ZEDEDA Help Center (Application Snapshot Overview, Manage an Edge Application Instance, Manage Application Snapshots, Use the ZEDEDA ZCLI, Custom Configuration Edge Application) and the `zededa/zedcloud` Terraform provider documentation.//