User Tools

Site Tools


eve-kvm:upgrade-eve-terraform

This is an old revision of the document!


Upgrading EVE-OS on an Edge Node with Terraform

This page shows how to move an edge node from one EVE-OS version to another using the zededa/zedcloud Terraform provider.

You do not need to reinstall the node, touch a USB stick, or visit the site. An EVE-OS upgrade is a file download plus a reboot, and both are driven from the controller.

Time needed: about 10 minutes of typing, plus 5-20 minutes of waiting while the node downloads and reboots.


How it works, in plain words

Every edge node has two root partitions, called IMGA and IMGB. Only one is active at a time. The other is spare.

  1. You tell the controller “this node should run version X”.
  2. The node downloads version X into whichever partition is not currently in use.
  3. The node reboots into that partition.
  4. If the new version does not check in with the controller, the node automatically falls back to the old partition.

That last point is why this is safe: a bad upgrade costs you a reboot, not a site visit.

To make this happen in Terraform you need two things:

Thing Terraform What it is
An image record zedcloud_image with image_type = “IMAGE_TYPE_EVE” A pointer telling the controller where the rootfs file lives and what its checksum is
A node setting a base_image block inside your zedcloud_edgenode Says “this node should run that image”

The image record is the part people forget. Keep reading.


!! Read this before you start !!

An EVE version name is not magic text. You cannot simply type 16.0.2-lts-kvm-amd64 into base_image and expect it to work. That name has to already exist as an image record in your enterprise, and that record has to be complete.

There is no built-in catalogue of EVE releases that fills this in for you. If you have never created an EVE image record before, you have zero of them, no matter how many versions ZEDEDA has published.

There are two ways this bites you, and the second one is nasty:

What you did What happens How obvious is it
Named a version with no image record at all The controller cannot resolve the name to an image ID, so the apply fails Obvious. Terraform shows an error.
Named a version whose image record exists but is missing its SHA-256 terraform apply says Apply complete! and the node then fails on its own Not obvious at all. Terraform is green and the node is broken.

The second case looks like this on the node:

IMGA  active  16.0.2-lts-kvm-amd64  DEVICE_SW_STATUS_FAILED
      swError: doUpdateContentTree(21f00882-...) name 16.0.2-lts-kvm-amd64:
               no content sha256 defined

The node is fine, it is still running the old version, but the upgrade never started.

The lesson: a green Terraform apply does not mean the node upgraded. Always do Step 7.


Step 1: Write down what is running now

Do this first so you can tell later whether anything changed.

# Pull your API token out of your tfvars file
TOK=$(grep -E '^zedcloud_token' secret.auto.tfvars | sed -E 's/^[^=]*= *"?([^"]*)"?.*/\1/')
 
# Change this to your controller and your node name
CTRL=https://zedcontrol.gmwtus.zededa.net
NODE=TF-DEMO-AI-PX-1
 
curl -s -H "Authorization: Bearer $TOK" "$CTRL/api/v1/devices/name/$NODE/status" \
| python3 -c "
import json,sys
for s in json.load(sys.stdin).get('swInfo',[]):
    print(s['partitionLabel'], s['partitionState'], repr(s['shortVersion']), s['swStatus'])
"

Healthy output looks like this. One partition active with a version, the other unused:

IMGA active '16.0.2-lts-kvm-amd64' DEVICE_SW_STATUS_UPDATED
IMGB unused '' DEVICE_SW_STATUS_UNSPECIFIED

The string in quotes is the exact version name. Copy it somewhere. You will want it if you ever need to go back.


Step 2: Put the rootfs file where the nodes can reach it

You need the *.rootfs.img file on a datastore. The node downloads it itself - the controller does not push it. So “reachable” means reachable from the edge node, not from your laptop.

Any of these work: HTTP, HTTPS, AWS S3, Azure Blob, SFTP.

A plain HTTP file server is the easiest. In this project that is the datastore TF-DEMO-AI-ATL-DS, which is http://192.168.0.101:1080.

Drop the file there, then prove it is really there:

curl -I "http://192.168.0.101:1080/0.0.0-evetest_download_burst_dead_mgmt_ports-59dbfd78-kvm-amd64.rootfs.img"

You want HTTP/1.1 200 OK and a sensible Content-Length (a rootfs is a few hundred MB). Write the Content-Length number down, you will use it in Step 4.

HTTP/1.1 200 OK
Content-Length: 284921856

If your datastore is on a private LAN, remember the nodes must be on that LAN or routed to it. An S3 bucket is the safer choice for nodes spread across sites.


Step 3: Get the SHA-256 checksum

The node verifies the checksum before it installs anything. Get this wrong and the node refuses the file.

Whoever built the image should give you the hash. Verify it yourself anyway - it takes seconds and it is the single most common cause of a failed upgrade:

curl -s "http://192.168.0.101:1080/0.0.0-evetest_download_burst_dead_mgmt_ports-59dbfd78-kvm-amd64.rootfs.img" \
  | shasum -a 256

Output:

0c7e0d6f33355db40ee24a6c480168efc73f5250a4f5ab1a9ce0b02648dce431  -

That must match what you were given, character for character. It must be 64 characters long. If it does not match, stop - you have the wrong file or a corrupted upload.


Step 4: Add the image record to your Terraform

Add a new zedcloud_image resource. Put it next to your other image resources. In this project that is 1-Infra-ai.tf.

This is an addition. Do not edit or delete your existing image resources.

resource "zedcloud_image" "eve_test_download_burst_kvm_amd64" {
  datastore_id        = zedcloud_datastore.demo_atl_ds.id
  image_type          = "IMAGE_TYPE_EVE"
  image_arch          = "AMD64"
  image_format        = "RAW"
  name                = "0.0.0-evetest_download_burst_dead_mgmt_ports-59dbfd78-kvm-amd64"
  title               = "0.0.0-evetest_download_burst_dead_mgmt_ports-59dbfd78-kvm-amd64"
  image_rel_url       = "0.0.0-evetest_download_burst_dead_mgmt_ports-59dbfd78-kvm-amd64.rootfs.img"
  image_sha256        = "0c7e0d6f33355db40ee24a6c480168efc73f5250a4f5ab1a9ce0b02648dce431"
  image_size_bytes    = 284921856
  project_access_list = []
}

What each field means:

Field What to put
datastore_id Reference to the datastore from Step 2
image_type Always IMAGE_TYPE_EVE for a base OS. Not IMAGE_TYPE_APPLICATION.
image_arch AMD64 for Intel/AMD boxes, ARM64 for Jetson and other ARM
image_format Always RAW for an EVE rootfs
name The EVE version string, without the .rootfs.img suffix
image_rel_url The filename, relative to the datastore root. With the suffix.
image_sha256 The 64-character hash from Step 3
image_size_bytes The Content-Length from Step 2

Do not leave image_sha256 empty as a placeholder. Terraform will happily create the record and every node pointed at it will fail with no content sha256 defined.


Step 5: Point the node at the image

Add a base_image block inside the zedcloud_edgenode resource you want to upgrade. In this project that is 3-EdgeNodes-ai.tf.

Again, an addition. Leave the rest of the node resource alone.

resource "zedcloud_edgenode" "demo_en_px" {
  # ... all your existing settings stay exactly as they are ...

  base_image {
    image_name = zedcloud_image.eve_test_download_burst_kvm_amd64.name
    version    = "0.0.0-evetest_download_burst_dead_mgmt_ports-59dbfd78-kvm-amd64"
    activate   = true
  }
}
Field What to put
image_name Reference your Step 4 resource with .name. Do not retype the string - let Terraform link them so it creates the image before it uses it.
version The same version string
activate true means go now. false stages the setting without starting the upgrade.

Naming a new image here is the upgrade. There is no separate “upgrade” command and no force flag to set. EVE stages the new version to the spare partition and reboots into it.

If the node uses for_each, one base_image block covers every instance.

Now check your syntax:

terraform fmt
terraform validate

Step 6: Apply, in two commands

6a. Preview

terraform plan

Read the output. You are looking for your new image being created and your node being updated in place:

# zedcloud_image.eve_test_download_burst_kvm_amd64 will be created
# zedcloud_edgenode.demo_en_px["EVE-Node-1"] will be updated in-place
    ~ base_image {
        ~ image_name = "16.0.2-lts-kvm-amd64" -> "0.0.0-evetest_...-kvm-amd64"

If you see anything being destroyed that you did not expect, stop and read it. This matters a lot on an existing codebase: a plain terraform apply also picks up every unrelated drift that has built up since the last run.

6b. Apply only what you mean to

Use -target so you change the image and the node and nothing else:

terraform apply \
  -target=zedcloud_image.eve_test_download_burst_kvm_amd64 \
  -target='zedcloud_edgenode.demo_en_px'

Expected: 1 to add, 3 to change (the 3 being however many nodes the resource covers), and 0 to destroy.

Why -target and not a plain apply:

  • It cannot trip over unrelated drift elsewhere in your code.
  • It cannot delete an old image record that a node is still using. If you removed an

older EVE image from your config in the same edit, Terraform has no way to know it must

  delete that //after// repointing the node, and may try to delete an in-use image
  mid-apply.

6c. Tidy up later

Once the upgrade has landed and you are happy, run a normal apply to catch up on everything else, including deleting any EVE image records you removed from the config:

terraform plan     # read it properly first
terraform apply

Step 7: Watch the upgrade. Do not skip this.

Terraform saying Apply complete! only means the controller accepted the setting. It says nothing about whether the node succeeded.

Run the Step 1 command again, every minute or two:

TOK=$(grep -E '^zedcloud_token' secret.auto.tfvars | sed -E 's/^[^=]*= *"?([^"]*)"?.*/\1/')
CTRL=https://zedcontrol.gmwtus.zededa.net
 
for NODE in TF-DEMO-AI-PX-1 TF-DEMO-AI-PX-2 TF-DEMO-AI-PX-3; do
  echo "=== $NODE ==="
  curl -s -H "Authorization: Bearer $TOK" "$CTRL/api/v1/devices/name/$NODE/status" \
  | python3 -c "
import json,sys
for s in json.load(sys.stdin).get('swInfo',[]):
    err = (s.get('swError') or {}).get('description','')
    print(' ', s['partitionLabel'], s['partitionState'], repr(s['shortVersion']),
          s['swStatus'], 'dl=%s%%' % s['downloadProgress'], err)
"
done

What you should see, in order:

  1. Downloading. The spare partition shows the new version and dl= climbing from 0 to 100.
  2. Rebooting. The node drops offline for a few minutes. This is normal.
  3. Testing. The spare partition becomes active with partitionState of testing.
  4. Done. State settles to active / DEVICE_SW_STATUS_UPDATED, and the partition that

used to be active now reads unused. The two partitions have swapped roles.


If it goes wrong

The node says DEVICE_SW_STATUS_FAILED

Read the swError text. It is usually specific.

Error text Cause Fix
no content sha256 defined image_sha256 is empty on the image record Fill it in (Step 3), apply again
Checksum / verification mismatch Hash does not match the actual file Re-hash the file, correct the record
Download or 404 errors Node cannot reach the datastore, or image_rel_url is wrong Check the path and the node's network route to the datastore

After fixing the record, ask the node to try again:

# Needs the node's UUID, not its name
curl -s -X PUT -H "Authorization: Bearer $TOK" \
  "$CTRL/api/v1/devices/id/<NODE-UUID>/baseos/upgrade/retry"

You can also bump base_os_retry_counter on the zedcloud_edgenode resource to force a retry through Terraform.

The node came back on the old version

That is the automatic fallback doing its job. The new version booted but did not check in, so EVE rolled back. The node is safe. Look at the node's logs or EdgeView for why the new version could not reach the controller.

I want to go back on purpose

Point base_image at the old version and apply again. This is why you wrote the old version string down in Step 1. Note that the old version needs a complete image record too - the same rules apply in both directions.


Special case: NVIDIA Jetson / Orin

  • Use image_arch = “ARM64”.
  • You need the Jetson build of EVE, with nvidia-jp6 or similar in the name. A generic

ARM64 rootfs will not boot a Jetson board.

  • A rootfs upgrade cannot cross a Jetson bootloader/BSP boundary. If the target version

needs a newer BSP than the board has, you need the installer and physical access, not

  this procedure. Check the release notes before you start.

Quick reference

Step Command
See current version curl …/devices/name/$NODE/status then read swInfo
Confirm file is reachable curl -I http://datastore/path/file.rootfs.img
Get checksum curl -s http://.../file.rootfs.img \| shasum -a 256
Check syntax terraform fmt && terraform validate
Preview terraform plan
Apply the upgrade terraform apply -target=IMAGE -target=NODE
Tidy up afterwards terraform plan then terraform apply

The five rules:

  1. The version name must exist as a complete zedcloud_image record before a node can use it.
  2. image_sha256 must be correct and must never be left blank.
  3. The node, not the controller, downloads the file. It must be reachable from the node.
  4. Use -target so an upgrade cannot drag unrelated drift along with it.
  5. A green Terraform apply is not a successful upgrade. Check the node.
eve-kvm/upgrade-eve-terraform.1790624036.txt.gz · Last modified: by mc