Table of Contents

Upgrading EVE-OS on an Edge Node with Terraform

This page shows how to move an edge node from one EVE-OS version to another using the zededa/zedcloud Terraform provider.

There is no need to reinstall the node, use a USB stick, or visit the site. An EVE-OS upgrade is a file download plus a reboot, both driven from the controller.

Time needed: roughly 10 minutes of editing, plus 5-20 minutes while the node downloads and reboots.


How it works

Every edge node has two root partitions, IMGA and IMGB. One is active, the other is spare.

  1. You tell the controller which version the node should run.
  2. The node downloads that version into whichever partition is not currently in use.
  3. The node reboots into that partition.
  4. If the new version does not check in with the controller, the node falls back to the

old partition automatically.

Because of that last step, an upgrade that goes wrong generally costs a reboot rather than a site visit.

Two pieces of Terraform are involved:

Piece Resource Purpose
Image record, in the image repository zedcloud_image with image_type = “IMAGE_TYPE_EVE” Tells the controller where the rootfs file lives and what its checksum is
Node setting a base_image block inside your zedcloud_edgenode Tells the node to run that image

The order matters: the image record comes first, the base_image reference second.


The image must already exist in the image repository

base_image does not accept an arbitrary version string. It is a reference to an image record that already exists in your enterprise's image repository. A node cannot be told to run a version the repository has never heard of, so the image record has to be created and applied before any base_image block points at it.

Nothing populates the repository on your behalf. There is no hidden catalogue of EVE releases, so every version you intend to deploy needs its own zedcloud_image resource. That is what Steps 2 to 4 are for; Step 5 only works once Step 4 has been applied.

To list the EVE images your repository currently holds:

TOK=$(grep -E '^zedcloud_token' secret.auto.tfvars | sed -E 's/^[^=]*= *"?([^"]*)"?.*/\1/')
CTRL=https://zedcontrol.gmwtus.zededa.net
 
curl -s -H "Authorization: Bearer $TOK" \
  "$CTRL/api/v1/apps/images?imageType=IMAGE_TYPE_EVE&next.pageSize=200" \
| python3 -c "
import json,sys
rows = json.load(sys.stdin).get('list',[])
print('%d EVE image(s) in the repository' % len(rows))
for i in rows:
    print(' ', i['imageArch'], i['imageStatus'], i['name'],
          '| sha:', (i.get('imageSha256') or '(NONE)')[:16])
"

An empty list is normal for an enterprise that has not deployed an EVE image before. Whatever you put in base_image.image_name has to appear in this list, with a SHA-256, before the node can act on it.

Two things can go wrong, and they look quite different:

Situation Result
The named version is not in the repository at all The controller cannot resolve the name to an image ID, and the apply fails with an error
The record is in the repository but has no SHA-256 The apply succeeds, and the node then reports a failure on its own

The second case shows up on the node like this:

IMGA  active  16.0.2-lts-kvm-amd64  DEVICE_SW_STATUS_FAILED
      swError: doUpdateContentTree(21f00882-...) name 16.0.2-lts-kvm-amd64:
               no content sha256 defined

The node is unharmed and still running the old version, but the upgrade never started. This is why Step 7 checks the node rather than relying on Terraform's output.


Step 1: Record what is running now

Useful to have for comparison afterwards, and if you ever want to go back.

# Pull your API token out of your tfvars file
TOK=$(grep -E '^zedcloud_token' secret.auto.tfvars | sed -E 's/^[^=]*= *"?([^"]*)"?.*/\1/')
 
# Adjust to your controller and node name
CTRL=https://zedcontrol.gmwtus.zededa.net
NODE=TF-DEMO-AI-PX-1
 
curl -s -H "Authorization: Bearer $TOK" "$CTRL/api/v1/devices/name/$NODE/status" \
| python3 -c "
import json,sys
for s in json.load(sys.stdin).get('swInfo',[]):
    print(s['partitionLabel'], s['partitionState'], repr(s['shortVersion']), s['swStatus'])
"

A settled node shows one partition active with a version and the other unused:

IMGA active '16.0.2-lts-kvm-amd64' DEVICE_SW_STATUS_UPDATED
IMGB unused '' DEVICE_SW_STATUS_UNSPECIFIED

The string in quotes is the exact version name. Keep a copy.


Step 2: Put the rootfs file where the nodes can reach it

The *.rootfs.img file needs to be on a datastore. The node downloads it itself - the controller does not push it out - so it needs to be reachable from the edge node rather than from your workstation.

HTTP, HTTPS, AWS S3, Azure Blob and SFTP all work. A plain HTTP file server is the simplest option. In this project that is the datastore TF-DEMO-AI-ATL-DS at http://192.168.0.101:1080.

Once the file is in place, confirm it:

curl -I "http://192.168.0.101:1080/0.0.0-evetest_download_burst_dead_mgmt_ports-59dbfd78-kvm-amd64.rootfs.img"

Look for HTTP/1.1 200 OK and a plausible Content-Length - a rootfs is a few hundred MB. Note the Content-Length, it is used in Step 4.

HTTP/1.1 200 OK
Content-Length: 284921856

If the datastore is on a private LAN, the nodes need to be on that LAN or routed to it. For nodes spread across several sites, S3 or similar tends to be easier.


Step 3: Get the SHA-256 checksum

The node verifies the checksum before installing, so it needs to be right.

Whoever built the image will usually supply the hash. Checking it against the file takes a couple of seconds and rules out a truncated or corrupted upload:

curl -s "http://192.168.0.101:1080/0.0.0-evetest_download_burst_dead_mgmt_ports-59dbfd78-kvm-amd64.rootfs.img" \
  | shasum -a 256
0c7e0d6f33355db40ee24a6c480168efc73f5250a4f5ab1a9ce0b02648dce431  -

It should match the hash you were given exactly, and be 64 characters long. If it does not match, the file on the datastore is not the file you think it is - worth resolving before going further.


Step 4: Add the image to the repository

This is the step that puts the version into the image repository, so that Step 5 has something to reference.

Add a zedcloud_image resource alongside your existing image resources. In this project those live in 1-Infra-ai.tf.

This is a new resource; your existing ones stay as they are.

resource "zedcloud_image" "eve_test_download_burst_kvm_amd64" {
  datastore_id        = zedcloud_datastore.demo_atl_ds.id
  image_type          = "IMAGE_TYPE_EVE"
  image_arch          = "AMD64"
  image_format        = "RAW"
  name                = "0.0.0-evetest_download_burst_dead_mgmt_ports-59dbfd78-kvm-amd64"
  title               = "0.0.0-evetest_download_burst_dead_mgmt_ports-59dbfd78-kvm-amd64"
  image_rel_url       = "0.0.0-evetest_download_burst_dead_mgmt_ports-59dbfd78-kvm-amd64.rootfs.img"
  image_sha256        = "0c7e0d6f33355db40ee24a6c480168efc73f5250a4f5ab1a9ce0b02648dce431"
  image_size_bytes    = 284921856
  project_access_list = []
}
Field Value
datastore_id Reference to the datastore from Step 2
image_type IMAGE_TYPE_EVE for a base OS, as opposed to IMAGE_TYPE_APPLICATION
image_arch AMD64 for Intel/AMD, ARM64 for Jetson and other ARM boards
image_format RAW for an EVE rootfs
name The EVE version string, without the .rootfs.img suffix
image_rel_url The filename relative to the datastore root, with the suffix
image_sha256 The 64-character hash from Step 3
image_size_bytes The Content-Length from Step 2

image_sha256 is worth filling in before you apply. Terraform will create the record without it, and nodes pointed at that record then fail with no content sha256 defined.


Step 5: Point the node at the image

This step only works for an image that is in the repository. If you skipped Step 4, or its record has no SHA-256, come back once that is sorted - the repository listing command above is the quickest way to confirm.

Add a base_image block inside the zedcloud_edgenode resource for the node you are upgrading. In this project the nodes are in 3-EdgeNodes-ai.tf.

Another addition - the rest of the node resource is unchanged.

resource "zedcloud_edgenode" "demo_en_px" {
  # ... existing settings unchanged ...

  base_image {
    image_name = zedcloud_image.eve_test_download_burst_kvm_amd64.name
    version    = "0.0.0-evetest_download_burst_dead_mgmt_ports-59dbfd78-kvm-amd64"
    activate   = true
  }
}
Field Value
image_name Reference to your Step 4 resource via .name. Referencing rather than retyping gives Terraform the dependency, so it creates the image before the node needs it.
version The same version string
activate true starts the upgrade. false stages the setting without starting it.

Naming a new image here is the whole upgrade - there is no separate upgrade command or force flag. EVE stages the new version to the spare partition and reboots into it.

A node resource using for_each needs only one base_image block; it applies to every instance.

Then check the syntax:

terraform fmt
terraform validate

Step 6: Apply

6a. Preview

terraform plan

You are looking for the new image being created and the node updated in place:

# zedcloud_image.eve_test_download_burst_kvm_amd64 will be created
# zedcloud_edgenode.demo_en_px["EVE-Node-1"] will be updated in-place
    ~ base_image {
        ~ image_name = "16.0.2-lts-kvm-amd64" -> "0.0.0-evetest_...-kvm-amd64"

Anything listed as being destroyed is worth reading before you continue. On an established codebase a plain terraform apply also picks up unrelated drift accumulated since the last run.

6b. Apply the upgrade

-target limits the apply to the image and the node:

terraform apply \
  -target=zedcloud_image.eve_test_download_burst_kvm_amd64 \
  -target='zedcloud_edgenode.demo_en_px'

Expect 1 to add and one change per node covered by the resource, with nothing destroyed.

Two reasons to scope it this way:

longer has a dependency telling it to delete that record after repointing the node,

  so it may attempt to delete an image that is still in use.

6c. Catch up afterwards

Once the upgrade has landed, a normal apply brings everything else into line, including removing any EVE image records you dropped from the configuration:

terraform plan
terraform apply

Step 7: Watch the node

Apply complete! means the controller accepted the setting. Whether the node succeeded is a separate question, and the node is the place to look.

Re-run the Step 1 query every minute or two:

TOK=$(grep -E '^zedcloud_token' secret.auto.tfvars | sed -E 's/^[^=]*= *"?([^"]*)"?.*/\1/')
CTRL=https://zedcontrol.gmwtus.zededa.net
 
for NODE in TF-DEMO-AI-PX-1 TF-DEMO-AI-PX-2 TF-DEMO-AI-PX-3; do
  echo "=== $NODE ==="
  curl -s -H "Authorization: Bearer $TOK" "$CTRL/api/v1/devices/name/$NODE/status" \
  | python3 -c "
import json,sys
for s in json.load(sys.stdin).get('swInfo',[]):
    err = (s.get('swError') or {}).get('description','')
    print(' ', s['partitionLabel'], s['partitionState'], repr(s['shortVersion']),
          s['swStatus'], 'dl=%s%%' % s['downloadProgress'], err)
"
done

The sequence to expect:

  1. Downloading. The spare partition shows the new version, with dl= climbing to 100.
  2. Rebooting. The node goes offline for a few minutes.
  3. Testing. The spare partition becomes active, with partitionState of testing.
  4. Settled. State reaches active / DEVICE_SW_STATUS_UPDATED, and the previously

active partition now reads unused. The two partitions have swapped roles.


If it does not go to plan

The node reports DEVICE_SW_STATUS_FAILED

The swError text is usually specific about the cause.

Error text Cause Fix
no content sha256 defined image_sha256 is empty on the image record Fill it in (Step 3) and apply again
Checksum or verification mismatch The hash does not match the file Re-hash the file and correct the record
Download or 404 errors image_rel_url is wrong, or the node cannot reach the datastore Check the path, and the node's route to the datastore

Once the record is corrected, the node can be asked to try again:

# Uses the node's UUID rather than its name
curl -s -X PUT -H "Authorization: Bearer $TOK" \
  "$CTRL/api/v1/devices/id/<NODE-UUID>/baseos/upgrade/retry"

Bumping base_os_retry_counter on the zedcloud_edgenode resource also triggers a retry through Terraform.

The node came back on the old version

This is the automatic fallback: the new version booted but did not check in with the controller, so EVE returned to the previous partition. The node is in a safe state. The node's logs or EdgeView will show why the new version could not reach the controller.

Going back deliberately

Point base_image at the previous version and apply again - which is what the version string from Step 1 is for.

The same repository rule applies in this direction: the older version needs to be in the image repository too. If its record was deleted after the upgrade, it has to be recreated (Steps 2 to 4) before a node can be sent back to it. The automatic fallback described above does not depend on this, since that rootfs is already on the node's other partition - but a deliberate, controller-driven rollback does.


NVIDIA Jetson and Orin

in the name. A generic ARM64 rootfs will not boot them.

version needs a newer BSP than the board currently has, that calls for the installer

  and physical access rather than this procedure, so the release notes are worth reading
  first.

Quick reference

Task Command
See current version curl …/devices/name/$NODE/status, then read swInfo
List EVE images in the repository curl …/apps/images?imageType=IMAGE_TYPE_EVE
Confirm the file is reachable curl -I http://datastore/path/file.rootfs.img
Get the checksum curl -s http://.../file.rootfs.img \| shasum -a 256
Check syntax terraform fmt && terraform validate
Preview terraform plan
Apply the upgrade terraform apply -target=IMAGE -target=NODE
Catch up afterwards terraform plan, then terraform apply

Worth keeping in mind:

zedcloud_image record first, then reference it.