eve-kvm:upgrade-eve-terraform
Differences
This shows you the differences between two versions of the page.
| eve-kvm:upgrade-eve-terraform [2026/09/28 19:33] – created mc | eve-kvm:upgrade-eve-terraform [2026/09/28 19:52] (current) – mc | ||
|---|---|---|---|
| Line 4: | Line 4: | ||
| '' | '' | ||
| - | You do **not** | + | There is no need to reinstall the node, use a USB stick, or visit the site. An EVE-OS |
| - | upgrade is a file download plus a reboot, | + | upgrade is a file download plus a reboot, both driven from the controller. |
| - | **Time needed: | + | **Time needed: |
| - | downloads | + | and reboots. |
| ---- | ---- | ||
| - | ===== How it works, in plain words ===== | + | ===== How it works ===== |
| - | Every edge node has **two** root partitions, | + | Every edge node has two root partitions, '' |
| - | active | + | spare. |
| - | - You tell the controller | + | - You tell the controller |
| - | - The node downloads version | + | - The node downloads |
| - The node reboots into that partition. | - The node reboots into that partition. | ||
| - | - If the new version does not check in with the controller, the node **automatically | + | - If the new version does not check in with the controller, the node falls back to the |
| - | | + | |
| - | That last point is why this is safe: a bad upgrade costs you a reboot, not a site visit. | + | Because of that last step, an upgrade |
| + | a site visit. | ||
| - | To make this happen in Terraform | + | Two pieces of Terraform |
| - | ^ Thing ^ Terraform | + | ^ Piece ^ Resource |
| - | | An **image | + | | Image record, in the image repository |
| - | | A **node | + | | Node setting | a '' |
| - | The image record | + | The order matters: the image record |
| ---- | ---- | ||
| - | ===== !! Read this before you start !! ===== | + | ===== The image must already exist in the image repository |
| - | **An EVE version name is not magic text.** You cannot simply type '' | + | **'' |
| - | into '' | + | image record |
| - | record in your enterprise, | + | told to run a version the repository has never heard of, so the image record has to be |
| + | created and applied | ||
| - | There is no built-in | + | Nothing populates the repository on your behalf. |
| - | never created an EVE image record before, you have zero of them, no matter how many | + | releases, so every version |
| - | versions ZEDEDA | + | That is what Steps 2 to 4 are for; Step 5 only works once Step 4 has been applied. |
| - | There are two ways this bites you, and the second one is nasty: | + | To list the EVE images your repository currently holds: |
| - | ^ What you did ^ What happens ^ How obvious is it ^ | + | <code bash> |
| - | | Named a version with **no image record at all** | The controller cannot resolve the name to an image ID, so the apply fails | Obvious. Terraform shows an error. | | + | TOK=$(grep -E '^zedcloud_token' |
| - | | Named a version whose image record exists but is **missing its SHA-256** | '' | + | CTRL=https:// |
| - | The second case looks like this on the node: | + | curl -s -H " |
| + | " | ||
| + | | python3 -c " | ||
| + | import json,sys | ||
| + | rows = json.load(sys.stdin).get(' | ||
| + | print(' | ||
| + | for i in rows: | ||
| + | print(' | ||
| + | '| sha:', (i.get(' | ||
| + | " | ||
| + | </ | ||
| + | |||
| + | An empty list is normal for an enterprise that has not deployed an EVE image before. | ||
| + | Whatever you put in '' | ||
| + | before the node can act on it. | ||
| + | |||
| + | Two things can go wrong, and they look quite different: | ||
| + | |||
| + | ^ Situation ^ Result ^ | ||
| + | | The named version is not in the repository at all | The controller cannot resolve the name to an image ID, and the apply fails with an error | | ||
| + | | The record is in the repository but has no SHA-256 | The apply succeeds, and the node then reports a failure on its own | | ||
| + | |||
| + | The second case shows up on the node like this: | ||
| < | < | ||
| Line 59: | Line 83: | ||
| </ | </ | ||
| - | The node is fine, it is still running the old version, but the upgrade never started. | + | The node is unharmed and still running the old version, but the upgrade never started. |
| - | + | This is why Step 7 checks | |
| - | **The lesson:** a green Terraform apply does not mean the node upgraded. Always do Step 7. | + | |
| ---- | ---- | ||
| - | ===== Step 1: Write down what is running now ===== | + | ===== Step 1: Record |
| - | Do this first so you can tell later whether anything changed. | + | Useful to have for comparison afterwards, and if you ever want to go back. |
| <code bash> | <code bash> | ||
| Line 73: | Line 96: | ||
| TOK=$(grep -E ' | TOK=$(grep -E ' | ||
| - | # Change this to your controller and your node name | + | # Adjust |
| CTRL=https:// | CTRL=https:// | ||
| NODE=TF-DEMO-AI-PX-1 | NODE=TF-DEMO-AI-PX-1 | ||
| Line 85: | Line 108: | ||
| </ | </ | ||
| - | Healthy output looks like this. One partition '' | + | A settled node shows one partition '' |
| < | < | ||
| Line 92: | Line 115: | ||
| </ | </ | ||
| - | The string in quotes is the **exact** version name. Copy it somewhere. You will want it if | + | The string in quotes is the exact version name. Keep a copy. |
| - | you ever need to go back. | + | |
| ---- | ---- | ||
| Line 99: | Line 121: | ||
| ===== Step 2: Put the rootfs file where the nodes can reach it ===== | ===== Step 2: Put the rootfs file where the nodes can reach it ===== | ||
| - | You need the '' | + | The '' |
| - | controller does not push it. So "reachable" means reachable //from the edge node//, not | + | controller does not push it out - so it needs to be reachable from the edge node rather |
| - | from your laptop. | + | than from your workstation. |
| - | Any of these work: HTTP, HTTPS, AWS S3, Azure Blob, SFTP. | + | HTTP, HTTPS, AWS S3, Azure Blob and SFTP all work. A plain HTTP file server is the simplest |
| + | option. In this project that is the datastore '' | ||
| + | '' | ||
| - | A plain HTTP file server | + | Once the file is in place, confirm |
| - | '' | + | |
| - | + | ||
| - | Drop the file there, then prove it is really there: | + | |
| <code bash> | <code bash> | ||
| Line 114: | Line 135: | ||
| </ | </ | ||
| - | You want '' | + | Look for '' |
| - | MB). Write the '' | + | MB. Note the '' |
| < | < | ||
| Line 122: | Line 143: | ||
| </ | </ | ||
| - | //If your datastore is on a private LAN, remember | + | //If the datastore is on a private LAN, the nodes need to be on that LAN or routed to it. |
| - | to it. An S3 bucket is the safer choice for nodes spread across sites.// | + | For nodes spread across |
| ---- | ---- | ||
| Line 129: | Line 150: | ||
| ===== Step 3: Get the SHA-256 checksum ===== | ===== Step 3: Get the SHA-256 checksum ===== | ||
| - | The node verifies the checksum before it installs anything. Get this wrong and the node | + | The node verifies the checksum before |
| - | refuses the file. | + | |
| - | Whoever built the image should give you the hash. **Verify it yourself anyway** - it takes | + | Whoever built the image will usually supply |
| - | seconds and it is the single most common cause of a failed upgrade: | + | couple of seconds and rules out a truncated or corrupted upload: |
| <code bash> | <code bash> | ||
| Line 139: | Line 159: | ||
| | shasum -a 256 | | shasum -a 256 | ||
| </ | </ | ||
| - | |||
| - | Output: | ||
| < | < | ||
| Line 146: | Line 164: | ||
| </ | </ | ||
| - | That must match what you were given, | + | It should |
| - | long. If it does not match, | + | match, the file on the datastore is not the file you think it is - worth resolving before |
| + | going further. | ||
| ---- | ---- | ||
| - | ===== Step 4: Add the image record | + | ===== Step 4: Add the image to the repository |
| - | Add a new '' | + | This is the step that puts the version into the image repository, so that Step 5 has |
| - | project | + | something to reference. |
| - | This is an **addition**. Do not edit or delete | + | Add a '' |
| + | those live in '' | ||
| + | |||
| + | This is a new resource; your existing ones stay as they are. | ||
| < | < | ||
| Line 173: | Line 195: | ||
| </ | </ | ||
| - | What each field means: | + | ^ Field ^ Value ^ |
| - | + | ||
| - | ^ Field ^ What to put ^ | + | |
| | '' | | '' | ||
| - | | '' | + | | '' |
| - | | '' | + | | '' |
| - | | '' | + | | '' |
| - | | '' | + | | '' |
| - | | '' | + | | '' |
| | '' | | '' | ||
| | '' | | '' | ||
| - | **Do not leave '' | + | '' |
| - | record and every node pointed at it will fail with '' | + | without it, and nodes pointed at that record then fail with '' |
| ---- | ---- | ||
| Line 192: | Line 212: | ||
| ===== Step 5: Point the node at the image ===== | ===== Step 5: Point the node at the image ===== | ||
| - | Add a '' | + | This step only works for an image that is in the repository. If you skipped Step 4, or its |
| - | In this project | + | record has no SHA-256, come back once that is sorted |
| + | is the quickest way to confirm. | ||
| - | Again, an **addition**. Leave the rest of the node resource | + | Add a '' |
| + | upgrading. In this project the nodes are in '' | ||
| + | |||
| + | Another addition - the rest of the node resource | ||
| < | < | ||
| resource " | resource " | ||
| - | # ... all your existing settings | + | # ... existing settings |
| base_image { | base_image { | ||
| Line 209: | Line 233: | ||
| </ | </ | ||
| - | ^ Field ^ What to put ^ | + | ^ Field ^ Value ^ |
| - | | '' | + | | '' |
| | '' | | '' | ||
| - | | '' | + | | '' |
| - | Naming a new image here **is** the upgrade. There is no separate | + | Naming a new image here is the whole upgrade |
| - | force flag to set. EVE stages the new version to the spare partition and reboots into it. | + | force flag. EVE stages the new version to the spare partition and reboots into it. |
| - | //If the node uses '' | + | //A node resource using '' |
| + | instance.// | ||
| - | Now check your syntax: | + | Then check the syntax: |
| <code bash> | <code bash> | ||
| Line 228: | Line 253: | ||
| ---- | ---- | ||
| - | ===== Step 6: Apply, in two commands | + | ===== Step 6: Apply ===== |
| ==== 6a. Preview ==== | ==== 6a. Preview ==== | ||
| Line 236: | Line 261: | ||
| </ | </ | ||
| - | Read the output. | + | You are looking for the new image being created and the node updated in place: |
| - | **updated in place**: | + | |
| < | < | ||
| Line 246: | Line 270: | ||
| </ | </ | ||
| - | **If you see anything | + | Anything listed as being destroyed |
| - | This matters a lot on an existing | + | codebase a plain '' |
| - | unrelated drift that has built up since the last run. | + | last run. |
| - | ==== 6b. Apply only what you mean to ==== | + | ==== 6b. Apply the upgrade |
| - | Use '' | + | '' |
| <code bash> | <code bash> | ||
| Line 260: | Line 284: | ||
| </ | </ | ||
| - | Expected: | + | Expect |
| - | and **0 to destroy**. | + | |
| - | Why '' | + | Two reasons to scope it this way: |
| - | * It cannot trip over unrelated | + | * Unrelated |
| - | * It cannot delete an old image record that a node is still using. | + | * If the same edit removed an older EVE image record |
| - | | + | |
| - | delete that //after// repointing the node, and may try to delete an in-use image | + | so it may attempt |
| - | mid-apply. | + | |
| - | ==== 6c. Tidy up later ==== | + | ==== 6c. Catch up afterwards |
| - | Once the upgrade has landed | + | Once the upgrade has landed, a normal apply brings |
| - | everything else, including | + | removing |
| <code bash> | <code bash> | ||
| - | terraform plan # read it properly first | + | terraform plan |
| terraform apply | terraform apply | ||
| </ | </ | ||
| Line 283: | Line 305: | ||
| ---- | ---- | ||
| - | ===== Step 7: Watch the upgrade. Do not skip this. ===== | + | ===== Step 7: Watch the node ===== |
| - | Terraform saying | + | '' |
| - | says nothing about whether | + | is a separate question, and the node is the place to look. |
| - | Run the Step 1 command again, | + | Re-run |
| <code bash> | <code bash> | ||
| Line 307: | Line 329: | ||
| </ | </ | ||
| - | What you should see, in order: | + | The sequence to expect: |
| - | - **Downloading.** The spare partition shows the new version | + | - **Downloading.** The spare partition shows the new version, with '' |
| - | - **Rebooting.** The node drops offline for a few minutes. This is normal. | + | - **Rebooting.** The node goes offline for a few minutes. |
| - | - **Testing.** The spare partition becomes '' | + | - **Testing.** The spare partition becomes '' |
| - | - **Done.** State settles to '' | + | - **Settled.** State reaches |
| - | | + | active |
| ---- | ---- | ||
| - | ===== If it goes wrong ===== | + | ===== If it does not go to plan ===== |
| - | ==== The node says DEVICE_SW_STATUS_FAILED ==== | + | ==== The node reports |
| - | Read the '' | + | The '' |
| ^ Error text ^ Cause ^ Fix ^ | ^ Error text ^ Cause ^ Fix ^ | ||
| - | | '' | + | | '' |
| - | | Checksum | + | | Checksum |
| - | | Download or 404 errors | Node cannot reach the datastore, or '' | + | | Download or 404 errors | '' |
| - | After fixing | + | Once the record |
| <code bash> | <code bash> | ||
| - | # Needs the node's UUID, not its name | + | # Uses the node's UUID rather than its name |
| curl -s -X PUT -H " | curl -s -X PUT -H " | ||
| " | " | ||
| </ | </ | ||
| - | You can also bump '' | + | Bumping |
| - | retry through Terraform. | + | through Terraform. |
| ==== The node came back on the old version ==== | ==== The node came back on the old version ==== | ||
| - | That is the automatic fallback | + | This is the automatic fallback: the new version booted but did not check in with the |
| - | so EVE rolled back. The node is safe. Look at the node's logs or EdgeView | + | controller, |
| - | version could not reach the controller. | + | node's logs or EdgeView |
| + | |||
| + | ==== Going back deliberately ==== | ||
| - | ==== I want to go back on purpose ==== | + | Point '' |
| + | string from Step 1 is for. | ||
| - | Point '' | + | The same repository rule applies in this direction: |
| - | version string down in Step 1. Note that the old version needs a complete image record | + | image repository too. If its record was deleted after the upgrade, it has to be recreated |
| - | too - the same rules apply in both directions. | + | (Steps 2 to 4) before a node can be sent back to it. The automatic fallback described above |
| + | does not depend on this, since that rootfs is already on the node's other partition - but a | ||
| + | deliberate, controller-driven rollback does. | ||
| ---- | ---- | ||
| - | ===== Special case: NVIDIA Jetson | + | ===== NVIDIA Jetson |
| - | * Use '' | + | * Set '' |
| - | * You need the Jetson build of EVE, with '' | + | * Jetson boards |
| - | | + | |
| - | * A rootfs upgrade cannot cross a Jetson bootloader/ | + | * A rootfs upgrade cannot cross a Jetson bootloader/ |
| - | needs a newer BSP than the board has, you need the installer and physical access, not | + | |
| - | | + | |
| + | first. | ||
| ---- | ---- | ||
| Line 366: | Line 394: | ||
| ===== Quick reference ===== | ===== Quick reference ===== | ||
| - | ^ Step ^ Command ^ | + | ^ Task ^ Command ^ |
| - | | See current version | '' | + | | See current version | '' |
| - | | Confirm file is reachable | '' | + | | List EVE images in the repository | '' |
| - | | Get checksum | '' | + | | Confirm |
| + | | Get the checksum | '' | ||
| | Check syntax | '' | | Check syntax | '' | ||
| | Preview | '' | | Preview | '' | ||
| | Apply the upgrade | '' | | Apply the upgrade | '' | ||
| - | | Tidy up afterwards | '' | + | | Catch up afterwards | '' |
| - | **The five rules:** | + | Worth keeping in mind: |
| - | | + | |
| - | | + | |
| - | | + | |
| - | | + | |
| - | | + | |
| + | | ||
eve-kvm/upgrade-eve-terraform.txt · Last modified: by mc
