====== Upgrading EVE-OS on an Edge Node with Terraform ======
This page shows how to move an edge node from one EVE-OS version to another using the
''zededa/zedcloud'' Terraform provider.
There is no need to reinstall the node, use a USB stick, or visit the site. An EVE-OS
upgrade is a file download plus a reboot, both driven from the controller.
**Time needed:** roughly 10 minutes of editing, plus 5-20 minutes while the node downloads
and reboots.
----
===== How it works =====
Every edge node has two root partitions, ''IMGA'' and ''IMGB''. One is active, the other is
spare.
- You tell the controller which version the node should run.
- The node downloads that version into whichever partition is //not// currently in use.
- The node reboots into that partition.
- If the new version does not check in with the controller, the node falls back to the
old partition automatically.
Because of that last step, an upgrade that goes wrong generally costs a reboot rather than
a site visit.
Two pieces of Terraform are involved:
^ Piece ^ Resource ^ Purpose ^
| Image record, in the image repository | ''zedcloud_image'' with ''image_type = "IMAGE_TYPE_EVE"'' | Tells the controller where the rootfs file lives and what its checksum is |
| Node setting | a ''base_image'' block inside your ''zedcloud_edgenode'' | Tells the node to run that image |
The order matters: the image record comes first, the ''base_image'' reference second.
----
===== The image must already exist in the image repository =====
**''base_image'' does not accept an arbitrary version string.** It is a reference to an
image record that already exists in your enterprise's image repository. A node cannot be
told to run a version the repository has never heard of, so the image record has to be
created and applied //before// any ''base_image'' block points at it.
Nothing populates the repository on your behalf. There is no hidden catalogue of EVE
releases, so every version you intend to deploy needs its own ''zedcloud_image'' resource.
That is what Steps 2 to 4 are for; Step 5 only works once Step 4 has been applied.
To list the EVE images your repository currently holds:
TOK=$(grep -E '^zedcloud_token' secret.auto.tfvars | sed -E 's/^[^=]*= *"?([^"]*)"?.*/\1/')
CTRL=https://zedcontrol.gmwtus.zededa.net
curl -s -H "Authorization: Bearer $TOK" \
"$CTRL/api/v1/apps/images?imageType=IMAGE_TYPE_EVE&next.pageSize=200" \
| python3 -c "
import json,sys
rows = json.load(sys.stdin).get('list',[])
print('%d EVE image(s) in the repository' % len(rows))
for i in rows:
print(' ', i['imageArch'], i['imageStatus'], i['name'],
'| sha:', (i.get('imageSha256') or '(NONE)')[:16])
"
An empty list is normal for an enterprise that has not deployed an EVE image before.
Whatever you put in ''base_image.image_name'' has to appear in this list, with a SHA-256,
before the node can act on it.
Two things can go wrong, and they look quite different:
^ Situation ^ Result ^
| The named version is not in the repository at all | The controller cannot resolve the name to an image ID, and the apply fails with an error |
| The record is in the repository but has no SHA-256 | The apply succeeds, and the node then reports a failure on its own |
The second case shows up on the node like this:
IMGA active 16.0.2-lts-kvm-amd64 DEVICE_SW_STATUS_FAILED
swError: doUpdateContentTree(21f00882-...) name 16.0.2-lts-kvm-amd64:
no content sha256 defined
The node is unharmed and still running the old version, but the upgrade never started.
This is why Step 7 checks the node rather than relying on Terraform's output.
----
===== Step 1: Record what is running now =====
Useful to have for comparison afterwards, and if you ever want to go back.
# Pull your API token out of your tfvars file
TOK=$(grep -E '^zedcloud_token' secret.auto.tfvars | sed -E 's/^[^=]*= *"?([^"]*)"?.*/\1/')
# Adjust to your controller and node name
CTRL=https://zedcontrol.gmwtus.zededa.net
NODE=TF-DEMO-AI-PX-1
curl -s -H "Authorization: Bearer $TOK" "$CTRL/api/v1/devices/name/$NODE/status" \
| python3 -c "
import json,sys
for s in json.load(sys.stdin).get('swInfo',[]):
print(s['partitionLabel'], s['partitionState'], repr(s['shortVersion']), s['swStatus'])
"
A settled node shows one partition ''active'' with a version and the other ''unused'':
IMGA active '16.0.2-lts-kvm-amd64' DEVICE_SW_STATUS_UPDATED
IMGB unused '' DEVICE_SW_STATUS_UNSPECIFIED
The string in quotes is the exact version name. Keep a copy.
----
===== Step 2: Put the rootfs file where the nodes can reach it =====
The ''*.rootfs.img'' file needs to be on a datastore. The node downloads it itself - the
controller does not push it out - so it needs to be reachable from the edge node rather
than from your workstation.
HTTP, HTTPS, AWS S3, Azure Blob and SFTP all work. A plain HTTP file server is the simplest
option. In this project that is the datastore ''TF-DEMO-AI-ATL-DS'' at
''http://192.168.0.101:1080''.
Once the file is in place, confirm it:
curl -I "http://192.168.0.101:1080/0.0.0-evetest_download_burst_dead_mgmt_ports-59dbfd78-kvm-amd64.rootfs.img"
Look for ''HTTP/1.1 200 OK'' and a plausible ''Content-Length'' - a rootfs is a few hundred
MB. Note the ''Content-Length'', it is used in Step 4.
HTTP/1.1 200 OK
Content-Length: 284921856
//If the datastore is on a private LAN, the nodes need to be on that LAN or routed to it.
For nodes spread across several sites, S3 or similar tends to be easier.//
----
===== Step 3: Get the SHA-256 checksum =====
The node verifies the checksum before installing, so it needs to be right.
Whoever built the image will usually supply the hash. Checking it against the file takes a
couple of seconds and rules out a truncated or corrupted upload:
curl -s "http://192.168.0.101:1080/0.0.0-evetest_download_burst_dead_mgmt_ports-59dbfd78-kvm-amd64.rootfs.img" \
| shasum -a 256
0c7e0d6f33355db40ee24a6c480168efc73f5250a4f5ab1a9ce0b02648dce431 -
It should match the hash you were given exactly, and be 64 characters long. If it does not
match, the file on the datastore is not the file you think it is - worth resolving before
going further.
----
===== Step 4: Add the image to the repository =====
This is the step that puts the version into the image repository, so that Step 5 has
something to reference.
Add a ''zedcloud_image'' resource alongside your existing image resources. In this project
those live in ''1-Infra-ai.tf''.
This is a new resource; your existing ones stay as they are.
resource "zedcloud_image" "eve_test_download_burst_kvm_amd64" {
datastore_id = zedcloud_datastore.demo_atl_ds.id
image_type = "IMAGE_TYPE_EVE"
image_arch = "AMD64"
image_format = "RAW"
name = "0.0.0-evetest_download_burst_dead_mgmt_ports-59dbfd78-kvm-amd64"
title = "0.0.0-evetest_download_burst_dead_mgmt_ports-59dbfd78-kvm-amd64"
image_rel_url = "0.0.0-evetest_download_burst_dead_mgmt_ports-59dbfd78-kvm-amd64.rootfs.img"
image_sha256 = "0c7e0d6f33355db40ee24a6c480168efc73f5250a4f5ab1a9ce0b02648dce431"
image_size_bytes = 284921856
project_access_list = []
}
^ Field ^ Value ^
| ''datastore_id'' | Reference to the datastore from Step 2 |
| ''image_type'' | ''IMAGE_TYPE_EVE'' for a base OS, as opposed to ''IMAGE_TYPE_APPLICATION'' |
| ''image_arch'' | ''AMD64'' for Intel/AMD, ''ARM64'' for Jetson and other ARM boards |
| ''image_format'' | ''RAW'' for an EVE rootfs |
| ''name'' | The EVE version string, without the ''.rootfs.img'' suffix |
| ''image_rel_url'' | The filename relative to the datastore root, with the suffix |
| ''image_sha256'' | The 64-character hash from Step 3 |
| ''image_size_bytes'' | The ''Content-Length'' from Step 2 |
''image_sha256'' is worth filling in before you apply. Terraform will create the record
without it, and nodes pointed at that record then fail with ''no content sha256 defined''.
----
===== Step 5: Point the node at the image =====
This step only works for an image that is in the repository. If you skipped Step 4, or its
record has no SHA-256, come back once that is sorted - the repository listing command above
is the quickest way to confirm.
Add a ''base_image'' block inside the ''zedcloud_edgenode'' resource for the node you are
upgrading. In this project the nodes are in ''3-EdgeNodes-ai.tf''.
Another addition - the rest of the node resource is unchanged.
resource "zedcloud_edgenode" "demo_en_px" {
# ... existing settings unchanged ...
base_image {
image_name = zedcloud_image.eve_test_download_burst_kvm_amd64.name
version = "0.0.0-evetest_download_burst_dead_mgmt_ports-59dbfd78-kvm-amd64"
activate = true
}
}
^ Field ^ Value ^
| ''image_name'' | Reference to your Step 4 resource via ''.name''. Referencing rather than retyping gives Terraform the dependency, so it creates the image before the node needs it. |
| ''version'' | The same version string |
| ''activate'' | ''true'' starts the upgrade. ''false'' stages the setting without starting it. |
Naming a new image here is the whole upgrade - there is no separate upgrade command or
force flag. EVE stages the new version to the spare partition and reboots into it.
//A node resource using ''for_each'' needs only one ''base_image'' block; it applies to every
instance.//
Then check the syntax:
terraform fmt
terraform validate
----
===== Step 6: Apply =====
==== 6a. Preview ====
terraform plan
You are looking for the new image being created and the node updated in place:
# zedcloud_image.eve_test_download_burst_kvm_amd64 will be created
# zedcloud_edgenode.demo_en_px["EVE-Node-1"] will be updated in-place
~ base_image {
~ image_name = "16.0.2-lts-kvm-amd64" -> "0.0.0-evetest_...-kvm-amd64"
Anything listed as being destroyed is worth reading before you continue. On an established
codebase a plain ''terraform apply'' also picks up unrelated drift accumulated since the
last run.
==== 6b. Apply the upgrade ====
''-target'' limits the apply to the image and the node:
terraform apply \
-target=zedcloud_image.eve_test_download_burst_kvm_amd64 \
-target='zedcloud_edgenode.demo_en_px'
Expect ''1 to add'' and one change per node covered by the resource, with nothing destroyed.
Two reasons to scope it this way:
* Unrelated drift elsewhere in the configuration stays out of the picture.
* If the same edit removed an older EVE image record from the configuration, Terraform no
longer has a dependency telling it to delete that record //after// repointing the node,
so it may attempt to delete an image that is still in use.
==== 6c. Catch up afterwards ====
Once the upgrade has landed, a normal apply brings everything else into line, including
removing any EVE image records you dropped from the configuration:
terraform plan
terraform apply
----
===== Step 7: Watch the node =====
''Apply complete!'' means the controller accepted the setting. Whether the node succeeded
is a separate question, and the node is the place to look.
Re-run the Step 1 query every minute or two:
TOK=$(grep -E '^zedcloud_token' secret.auto.tfvars | sed -E 's/^[^=]*= *"?([^"]*)"?.*/\1/')
CTRL=https://zedcontrol.gmwtus.zededa.net
for NODE in TF-DEMO-AI-PX-1 TF-DEMO-AI-PX-2 TF-DEMO-AI-PX-3; do
echo "=== $NODE ==="
curl -s -H "Authorization: Bearer $TOK" "$CTRL/api/v1/devices/name/$NODE/status" \
| python3 -c "
import json,sys
for s in json.load(sys.stdin).get('swInfo',[]):
err = (s.get('swError') or {}).get('description','')
print(' ', s['partitionLabel'], s['partitionState'], repr(s['shortVersion']),
s['swStatus'], 'dl=%s%%' % s['downloadProgress'], err)
"
done
The sequence to expect:
- **Downloading.** The spare partition shows the new version, with ''dl='' climbing to 100.
- **Rebooting.** The node goes offline for a few minutes.
- **Testing.** The spare partition becomes ''active'', with ''partitionState'' of ''testing''.
- **Settled.** State reaches ''active'' / ''DEVICE_SW_STATUS_UPDATED'', and the previously
active partition now reads ''unused''. The two partitions have swapped roles.
----
===== If it does not go to plan =====
==== The node reports DEVICE_SW_STATUS_FAILED ====
The ''swError'' text is usually specific about the cause.
^ Error text ^ Cause ^ Fix ^
| ''no content sha256 defined'' | ''image_sha256'' is empty on the image record | Fill it in (Step 3) and apply again |
| Checksum or verification mismatch | The hash does not match the file | Re-hash the file and correct the record |
| Download or 404 errors | ''image_rel_url'' is wrong, or the node cannot reach the datastore | Check the path, and the node's route to the datastore |
Once the record is corrected, the node can be asked to try again:
# Uses the node's UUID rather than its name
curl -s -X PUT -H "Authorization: Bearer $TOK" \
"$CTRL/api/v1/devices/id//baseos/upgrade/retry"
Bumping ''base_os_retry_counter'' on the ''zedcloud_edgenode'' resource also triggers a retry
through Terraform.
==== The node came back on the old version ====
This is the automatic fallback: the new version booted but did not check in with the
controller, so EVE returned to the previous partition. The node is in a safe state. The
node's logs or EdgeView will show why the new version could not reach the controller.
==== Going back deliberately ====
Point ''base_image'' at the previous version and apply again - which is what the version
string from Step 1 is for.
The same repository rule applies in this direction: the older version needs to be in the
image repository too. If its record was deleted after the upgrade, it has to be recreated
(Steps 2 to 4) before a node can be sent back to it. The automatic fallback described above
does not depend on this, since that rootfs is already on the node's other partition - but a
deliberate, controller-driven rollback does.
----
===== NVIDIA Jetson and Orin =====
* Set ''image_arch = "ARM64"''.
* Jetson boards need a Jetson build of EVE, identifiable by something like ''nvidia-jp6''
in the name. A generic ARM64 rootfs will not boot them.
* A rootfs upgrade cannot cross a Jetson bootloader/BSP boundary. Where the target
version needs a newer BSP than the board currently has, that calls for the installer
and physical access rather than this procedure, so the release notes are worth reading
first.
----
===== Quick reference =====
^ Task ^ Command ^
| See current version | ''curl .../devices/name/$NODE/status'', then read ''swInfo'' |
| List EVE images in the repository | ''curl .../apps/images?imageType=IMAGE_TYPE_EVE'' |
| Confirm the file is reachable | ''curl -I http://datastore/path/file.rootfs.img'' |
| Get the checksum | ''curl -s http://.../file.rootfs.img \| shasum -a 256'' |
| Check syntax | ''terraform fmt && terraform validate'' |
| Preview | ''terraform plan'' |
| Apply the upgrade | ''terraform apply -target=IMAGE -target=NODE'' |
| Catch up afterwards | ''terraform plan'', then ''terraform apply'' |
Worth keeping in mind:
* ''base_image'' can only name an image that is already in the image repository. Create the
''zedcloud_image'' record first, then reference it.
* ''image_sha256'' needs to be present and correct, or the node rejects the image.
* The node downloads the file itself, so the datastore has to be reachable from the node.
* ''-target'' keeps an upgrade separate from unrelated drift.
* Terraform's output covers the controller; the node's status covers the upgrade.