User Tools

Site Tools


eve-kvm:core-isolation

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Next revision
Previous revision
eve-kvm:core-isolation [2026/09/18 16:34] – created mceve-kvm:core-isolation [2026/09/19 17:35] (current) – mc
Line 1: Line 1:
 +====== EVE CPU Core Isolation — Step-by-Step Runbook ======
 +
 +Follow the steps in order. Every step has a command, the output you should see, and what to do if you do not see it. Do not skip **Step 5** (the BEFORE baseline) — it is what makes the AFTER checks meaningful.
 +
 +**Scope.** The ''core-pinning'' feature branch (eve [[https://github.com/lf-edge/eve/pull/6335|PR #6335]], eve-api [[https://github.com/lf-edge/eve-api/pull/155|#155]], adam [[https://github.com/lf-edge/adam/pull/158|#158]], NFR-172). Verified end to end on EVE ''0.0.0-core-pinning-gmwtus-551edfcb-kvm-amd64'', node = Intel Core i3-13100TE, on 2026-09-18. Every transcript below is real output from that run except Step 17, which is marked.
 +
 +**Time required.** About 20 minutes, including one reboot.
 +
 +**Just here to run the demo?** Use the four-act script immediately below. The numbered steps in Parts 1-4 are the full procedure behind it.
 +
 +----
 +
 +===== Quick demo script — 1, 2, 3, 4 =====
 +
 +Four commands-and-a-sentence. Run them in order, with the pinned VM **not yet deployed**.
 +
 +Set this up first so every command is a single keystroke away:
 +
 +<code bash>
 +alias iso='echo "isolated: [$(cat /sys/devices/system/cpu/isolated)]"'
 +alias pools='cat /run/domainmgr/CPUPoolStatus/*.json'
 +alias plan='cat /run/domainmgr/cpuplan.json'
 +alias ticks='awk "/^cpu[0-9] /{printf \"%s=%s  \", \$1, \$2+\$4}" /proc/stat; echo'
 +</code>
 +
 +==== 1. BEFORE — a plain node: no isolation, no assignment ====
 +
 +<code bash>
 +iso
 +pools
 +ticks
 +</code>
 +
 +<code>
 +isolated: []
 +{"Pools":[{"Kind":1,"CPUs":[0,1,2,3,4,5,6,7],"FreeCPUs":[1,2,3,4,5,6,7],"TotalCores":4,"FreeWholeCores":3},
 +          {"Kind":2,"CPUs":null},{"Kind":3,"CPUs":null}]}
 +cpu0=73851  cpu1=586  cpu2=496  cpu3=390  cpu4=1120  cpu5=980  cpu6=842  cpu7=771
 +</code>
 +
 +**Say:** "The kernel isolates nothing. There is one pool — all eight threads, shared. Every CPU is doing work. Any workload can be scheduled anywhere, and so can the kernel's own housekeeping."
 +
 +//This is the state before any configuration. If your node is already configured and you want to rehearse the whole arc, revert it with:// ''eve config mount /tmp/cfg && rm /tmp/cfg/grub.cfg && eve config unmount && reboot'' //— or just show a saved capture of this output.//
 +
 +==== 2. AFTER the config — the isolated pool exists, and nobody may use it ====
 +
 +Enable per Part 2 (property, ''grub.cfg'', reboot), then:
 +
 +<code bash>
 +iso
 +pools
 +ticks
 +</code>
 +
 +<code>
 +isolated: [4-7]
 +{"Pools":[{"Kind":1,"CPUs":[0,1,2,3,4,5,6,7],"FreeCPUs":[2,3],"AllocatedThreads":6,"FreeWholeCores":1},
 +          {"Kind":2,"CPUs":null},
 +          {"Kind":3,"CPUs":[4,5,6,7],"FreeCPUs":[4,5,6,7],"TotalCores":2,"FreeWholeCores":2}]}
 +cpu0=9525  cpu1=376  cpu2=496  cpu3=390  cpu4=0  cpu5=0  cpu6=0  cpu7=0
 +</code>
 +
 +**Say:** "Now there are three pools. Four CPUs are isolated. Look at the housekeeping pool — of eight threads, six are gone: two reserved for EVE, four withheld for isolation. Only one whole core is left for ordinary work. And cores 4 through 7 have executed **zero** cycles since boot — not the scheduler, not a workload, nothing."
 +
 +**Point at:** ''FreeWholeCores'' dropping 3 → 1, and the four zeros.
 +
 +==== 3. DEPLOY — a pinned VM takes an isolated core ====
 +
 +Deploy a **2-vCPU** app with CPU pinning enabled, then:
 +
 +<code bash>
 +plan
 +pools
 +</code>
 +
 +<code>
 +{ "display_name": "TF-STND-VM-1", "mode": "whole-core-smt", "vcpus": 2,
 +  "status": "success", "host_cpus": [4, 5] }
 +
 +{"Pools":[{"Kind":1,"CPUs":[0,1,2,3,6,7],"FreeCPUs":[2,3],"FreeWholeCores":1},
 +          {"Kind":2,"CPUs":[4,5],"AllocatedThreads":2},
 +          {"Kind":3,"CPUs":[4,5,6,7],"FreeCPUs":[6,7],"FreeWholeCores":1}]}
 +</code>
 +
 +**Say:** "Two vCPUs, one whole physical core, taken from the isolated set — cores 4 and 5. The isolated pool drops from two free cores to one. The app asked for nothing but 'CPU pinning'; the node-level switch did the rest."
 +
 +==== 4. PROVE IT — the kernel agrees, down to the thread ====
 +
 +<code bash>
 +PID=$(pgrep -f qemu-system | head -1)
 +for t in /proc/$PID/task/*; do
 +  printf '%-18s %s
 +' "$(cat $t/comm)" "$(awk '/Cpus_allowed_list/{print $2}' $t/status)"
 +done | sort -u
 +cat /sys/fs/cgroup/cpuset/eve-user-apps/*.1/cpuset.cpus
 +ticks
 +</code>
 +
 +<code>
 +qemu-system-x86    4         <-- vCPU 0
 +qemu-system-x86    5         <-- vCPU 1
 +qemu-system-x86    4-5       <-- emulator / IO threads
 +vhost-7981         4-5
 +
 +4-5
 +
 +cpu0=9525  cpu1=376  cpu2=496  cpu3=390  cpu4=1990  cpu5=813  cpu6=0  cpu7=0
 +</code>
 +
 +**Say:** "One vCPU thread per hardware thread, pinned 1:1. Everything else the VM needs is confined to the same core. The cgroup agrees. And the second isolated core is **still at zero** — it is reserved and untouchable, waiting for the next workload that asks."
 +
 +**The closing line:** cores 6 and 7 are idle and cannot be used by anything that did not ask for isolation. On this node, a workload that needs guaranteed CPU gets it — not by priority, not by best effort, but because nothing else is allowed to run there.
 +
 +----
 +
 +===== Part 0 — Understand what you are doing =====
 +
 +There are **two independent switches**. Neither works without the other.
 +
 +^ # ^ Switch ^ Where it is set ^ What it does ^ Needs reboot? ^
 +| 1 | ''isolcpus'' kernel argument | ''/config/grub.cfg'' on the node | creates the isolated CPU set | **yes** — kernel command line |
 +| 2 | ''cpu.pinning.use.isolated'' | node-level property in the controller — **Terraform or REST only, not in the UI** | lets pinned workloads use that set | no, but see Step 9 |
 +
 +A third thing is often confused with these:
 +
 +  * **"CPU pinning" on the application** (''pin_cpu'') is a //per-app// setting. It gives the workload whole physical cores out of the **dedicated** pool. It is **not** the isolation switch. If you only set this, your VM gets pinned — just not onto the isolated cores.
 +
 +Target layout on a 4-core / 8-thread node:
 +
 +^ Role ^ Physical core ^ CPUs ^
 +| EVE housekeeping | core 0 | 0, 1 |
 +| Dedicated pool (ordinary pinned workloads) | core 1 | 2, 3 |
 +| **Isolated pool** | cores 2-3 | **4, 5, 6, 7** |
 +
 +----
 +
 +===== Part 1 — BEFORE: prerequisites and baseline =====
 +
 +Run every command in this part **on the node**, over SSH or the console, before changing anything.
 +
 +==== Step 1 — Confirm the CPU can do this ====
 +
 +<code bash>
 +lscpu | grep -E 'Model name|Thread\(s\) per core|Core\(s\) per socket|NUMA node\(s\)'
 +</code>
 +
 +Expected:
 +
 +<code>
 +Model name:            13th Gen Intel(R) Core(TM) i3-13100TE
 +Thread(s) per core:    2
 +Core(s) per socket:    4
 +NUMA node(s):          1
 +</code>
 +
 +**PASS if ''Thread(s) per core'' is 2.** If it is 1, stop: either SMT is disabled in the BIOS (enable it and reboot) or the CPU has no SMT and cannot run this feature at all. There is no fallback — pinning fails closed with ''cpu.topology.unsupported''.
 +
 +You also need **at least 3 physical cores**: one for EVE, one for the dedicated pool, one to isolate.
 +
 +^ CPU ^ Topology ^ Usable ^
 +| Core i3-10100 / 10105 / 10300 | 4C/8T | yes |
 +| Core i3-12100 / 12300, i3-13100, i3-14100 | 4 P-cores, 0 E-cores, 8T | yes — ideal |
 +| Core i5-12400 | 6 P-cores, 0 E-cores, 12T | yes — more headroom |
 +| Core i3-8100 / 8300, i3-9100 / 9300 | 4C/4T, no HT | **no** |
 +| Core i3-N300 / i3-N305 (Alder Lake-N) | 8 E-cores, no HT | **no** |
 +| Core Ultra 200S / Meteor Lake / Lunar Lake | SMT removed | **no** |
 +
 +Avoid Intel hybrid parts with E-cores (i5/i7 12th gen and up): the E-cores have no SMT and are silently skipped for pinning, so they add nothing and confuse the results.
 +
 +==== Step 2 — Find out which CPUs are SMT siblings ====
 +
 +This decides the numbers you will type in Step 8. **Do not assume them.**
 +
 +<code bash>
 cat /sys/devices/system/cpu/cpu[0-9]*/topology/thread_siblings_list | sort -u cat /sys/devices/system/cpu/cpu[0-9]*/topology/thread_siblings_list | sort -u
 +</code>
 +
 +Expected on the reference node:
 +
 +<code>
 +0-1
 +2-3
 +4-5
 +6-7
 +</code>
 +
 +^ If you see ^ Enumeration ^ Use the values in ^
 +| ''0-1  2-3  4-5  6-7'' | adjacent (Intel 12th gen and newer) | **Step 8a** |
 +| ''0,4  1,5  2,6  3,7'' | split (older Intel) | **Step 8b** |
 +
 +==== Step 3 — Confirm nothing is isolated yet ====
 +
 +<code bash>
 +cat /sys/devices/system/cpu/isolated
 +</code>
 +
 +Expected: **empty output.** That is the starting state. If it already lists CPUs, someone has been here before — read the existing ''/config/grub.cfg'' before you overwrite it.
 +
 +==== Step 4 — Confirm the property is currently off ====
 +
 +<code bash>
 +grep -o 'cpu.pinning.use.isolated[^}]*}' /run/zedagent/ConfigItemValueMap/global.json
 +</code>
 +
 +Expected:
 +
 +<code>
 +cpu.pinning.use.isolated":{"Key":"cpu.pinning.use.isolated","ItemType":2,"IntValue":0,"StrValue":"","BoolValue":false,"TriStateValue":0}
 +</code>
 +
 +''BoolValue'' is ''false''. If the key is missing entirely, this EVE build does not have the feature — check ''eve version''.
 +
 +==== Step 5 — Capture the BEFORE baseline ====
 +
 +Paste this whole block. Save the output somewhere; you will compare against it in Step 12.
 +
 +<code bash>
 +echo "=== cmdline ==="
 +tr ' ' '\n' < /proc/cmdline | grep -E 'isolcpus|nohz_full|rcu_nocbs|irqaffinity|eve_max_vcpus'
 +echo "=== kernel isolated set ==="
 +echo "[$(cat /sys/devices/system/cpu/isolated)]"
 +echo "=== CPU pools ==="
 +cat /run/domainmgr/CPUPoolStatus/*.json; echo
 +echo "=== placement plan ==="
 +cat /run/domainmgr/cpuplan.json 2>/dev/null
 +echo "=== eve cpuset ==="
 +cat /sys/fs/cgroup/cpuset/eve/cpuset.cpus
 +echo "=== per-CPU busy ticks ==="
 +awk '/^cpu[0-9] /{print $1, $2+$4}' /proc/stat
 +</code>
 +
 +Expected BEFORE (no isolation, nothing pinned):
 +
 +<code>
 +=== cmdline ===
 +eve_max_vcpus=1
 +=== kernel isolated set ===
 +[]
 +=== CPU pools ===
 +{"Pools":[
 + {"Kind":1,"CPUs":[0,1,2,3,4,5,6,7],"FreeCPUs":[1,2,3,4,5,6,7],"TotalThreads":8,"AllocatedThreads":1,"FreeThreads":7,"TotalCores":4,"FreeWholeCores":3},
 + {"Kind":2,"CPUs":null,"TotalThreads":0,"FreeWholeCores":0},
 + {"Kind":3,"CPUs":null,"TotalThreads":0,"FreeWholeCores":0}
 +]}
 +=== eve cpuset ===
 +0
 +</code>
 +
 +Pool ''Kind'' values, needed for every later step: **1 = housekeeping, 2 = dedicated, 3 = isolated.**
 +
 +What this baseline says: all 8 threads are in housekeeping, only CPU 0 is taken (by EVE), and the dedicated and isolated pools do not exist.
 +
 +----
 +
 +===== Part 2 — ENABLE =====
 +
 +Do these three steps **in this order**. Doing the property last costs you a second reboot.
 +
 +==== Step 6 — Set the controller property ====
 +
 +**This property is not exposed in the ZedControl UI.** Do not go looking for it there — the UI only offers properties it knows about, and this one is new in the feature branch. It can only be set via **Terraform** or the **REST API**.
 +
 +**Terraform** — a ''config_item'' block on the **''zedcloud_edgenode''** resource. Not on the application, not on the app instance.
 +
 +<code>
 +resource "zedcloud_edgenode" "my_node" {
 +  # ... existing fields ...
 +
 +  config_item {
 +    key          = "cpu.pinning.use.isolated"
 +    string_value = "true"
 +  }
 +}
 +</code>
 +
 +Then ''terraform apply''.
 +
 +**Use ''string_value = "true"''.** This is the single most common way to get "I set it and nothing happened", and the reason is that two different schemas are involved:
 +
 +^ Hop ^ Schema ^ Fields ^
 +| you → controller | ''EDConfigItem'' | ''key'', ''valueType'', ''stringValue'', ''boolValue'', ''floatValue'', ''uint32Value'', ''uint64Value'' |
 +| controller → device | eve-api ''ConfigItem'' | ''key'', ''value'' — **both strings** |
 +
 +The controller has to collapse the typed fields into a single string ''value'' for the device. An entry carrying only ''boolValue: true'' with no ''valueType'' can arrive at the device as an empty ''value'', which pillar cannot parse, so the item keeps its default — ''false''. ''stringValue'' passes straight through. The provider also exposes an optional ''value_type'' attribute; it is not needed when you use ''string_value''.
 +
 +**REST** — ''PUT /api/v1/devices/id/{id}'' (operation ''EdgeNodeConfiguration_UpdateEdgeNode''). The property goes in the device object's ''configItem'' array, whose elements are ''EDConfigItem'':
 +
 +<code javascript>
 +{ "key": "cpu.pinning.use.isolated", "stringValue": "true" }
 +</code>
 +
 +**This is a full-object PUT, not a patch.** The body carries the whole edge node — ''name'', ''title'', ''modelId'', ''projectId'', ''interfaces'', ''adminState'', every existing ''configItem'' and about forty more fields. Send a partial body and you erase what you left out. It also requires ''revision.curr'', the current database version of the record; a stale value is rejected with //409 Version mismatch//.
 +
 +So the REST procedure is read-modify-write:
 +
 +  - ''GET /api/v1/devices/id/{id}''
 +  - append ''{ "key": "cpu.pinning.use.isolated", "stringValue": "true" }'' to the returned ''configItem'' array, keeping every existing entry
 +  - ''PUT'' the entire modified object back, including ''revision''
 +
 +This is why Terraform is the easier route: the provider does that read-modify-write for you. Schema verified against ''swagger/zedge_node_service.swagger.json'' (ZEDEDA Edge Node Service v1.0) in the provider source; confirm against your own controller's ''/api/v1/docs/'' if it runs a different version.
 +
 +**Do not expect to find the key itself in the API documentation.** In the spec ''key'' is a plain ''string'' with no enum and no validation — ''configItem'' is a free-form key/value list that the controller stores and forwards to the device verbatim. The only component that knows ''cpu.pinning.use.isolated'' is a real property is EVE itself (pillar's ''types/global.go''). That is why the property works through a controller that has never heard of it, and equally why the UI does not offer it: the UI renders a curated list of known properties, the API accepts any string.
 +
 +Whichever route you use, **Step 7 is what confirms it** — do not assume it landed.
 +
 +==== Step 7 — Confirm the property reached the node ====
 +
 +Do not go further until this passes. Allow a minute for the config poll.
 +
 +<code bash>
 +grep -o 'cpu.pinning.use.isolated[^}]*}' /run/zedagent/ConfigItemValueMap/global.json
 +logread | grep 'use.isolated'
 +</code>
 +
 +Expected:
 +
 +<code>
 +"BoolValue":true
 +domainmgr: CPU placement: cpu.pinning.use.isolated is now true; workloads placed from now on may use the kernel-isolated CPUs
 +</code>
 +
 +**If ''BoolValue'' is still ''false'':** you used ''bool_value'' instead of ''string_value'' (go back to Step 6), or the controller rejected the key — see Troubleshooting.
 +
 +==== Step 8 — Write /config/grub.cfg ====
 +
 +''/config/grub.cfg'' does **not** exist on a fresh install. The image ships only ''grub.cfg.tmpl'', and GRUB reads ''grub.cfg'' exactly. **Create the file. Do not rename the template.**
 +
 +=== Step 8a — adjacent siblings (0-1, 2-3, 4-5, 6-7) ===
 +
 +<code bash>
 +eve config mount /tmp/cfg
 +cat > /tmp/cfg/grub.cfg <<'EOF'
 +set_getty
 +set_global hv_eve_cpu_settings "eve_max_vcpus=2"
 +set_global dom0_extra_args "$dom0_extra_args isolcpus=managed_irq,domain,4,5,6,7 rcu_nocbs=4,5,6,7 nohz_full=4,5,6,7 irqaffinity=0,1"
 +EOF
 +cat /tmp/cfg/grub.cfg
 +sync
 +eve config unmount
 +</code>
 +
 +=== Step 8b — split siblings (0,4  1,5  2,6  3,7) ===
 +
 +<code bash>
 +eve config mount /tmp/cfg
 +cat > /tmp/cfg/grub.cfg <<'EOF'
 +set_getty
 +set_global dom0_extra_args "$dom0_extra_args isolcpus=managed_irq,domain,2,3,6,7 rcu_nocbs=2,3,6,7 nohz_full=2,3,6,7 irqaffinity=0,4"
 +EOF
 +cat /tmp/cfg/grub.cfg
 +sync
 +eve config unmount
 +</code>
 +
 +There is **no ''eve_max_vcpus'' line in 8b** on purpose. Reservation is by lowest CPU id and a core is dropped if any sibling is reserved, so ''eve_max_vcpus=2'' with split enumeration would reserve CPUs 0 and 1 — two halves of two different cores — and cost you a second core for nothing.
 +
 +**Check the ''cat'' output before moving on.** The third line must still contain the literal text ''$dom0_extra_args''. If that text is missing, your shell expanded it and GRUB will lose everything it had accumulated — rewrite it, making sure the heredoc marker is quoted as ''<<'EOF''' and not ''<<EOF''.
 +
 +Rules that matter here:
 +
 +  * ''irqaffinity'' lists the **housekeeping** CPUs, never the isolated ones. Pointing it at the isolated set steers interrupts onto the cores you are shielding.
 +  * ''set_getty'' keeps a console shell. It is the only thing in the shipped template, so if you do not include it you lose it.
 +  * Do **not** use GRUB's ''set_isolcpus'' helper (menu entry //isolate CPU0 (only for PREEMPT_RT)//). It emits ''isolcpus=inverse,0'', which isolates CPU 0's own sibling and produces a core that serves nobody.
 +
 +==== Step 9 — Reboot ====
 +
 +<code bash>
 +reboot
 +</code>
 +
 +Allow about 3 minutes. The reboot is required for two reasons: ''isolcpus'' is a kernel command-line argument, and a reboot is also what lets **already-running** workloads be re-placed (see Troubleshooting — a restart is not enough).
 +
 +Expect the SSH host key to change: EVE regenerates ''/etc/ssh/ssh_host_*'' on every boot. Clear the old entry with ''ssh-keygen -R <node-ip>''.
 +
 +----
 +
 +===== Part 3 — AFTER: validate =====
 +
 +Run these in order. Stop at the first failure.
 +
 +==== Step 10 — The kernel accepted the arguments ====
 +
 +<code bash>
 +tr ' ' '\n' < /proc/cmdline | grep -E 'isolcpus|nohz_full|rcu_nocbs|irqaffinity|eve_max_vcpus'
 +</code>
 +
 +<code>
 +eve_max_vcpus=2
 +isolcpus=managed_irq,domain,4,5,6,7
 +rcu_nocbs=4,5,6,7
 +nohz_full=4,5,6,7
 +irqaffinity=0,1
 +</code>
 +
 +**If nothing appears:** GRUB never read your file. Confirm it is named ''grub.cfg'' (not ''grub.cfg.tmpl'') on the CONFIG partition.
 +
 +==== Step 11 — The kernel is actually isolating ====
 +
 +<code bash>
 +cat /sys/devices/system/cpu/isolated
 +</code>
 +
 +<code>
 +4-7
 +</code>
 +
 +**This is the make-or-break check.** An empty result means there is no isolated pool and nothing after this point can work. Do not continue — fix the command line first.
 +
 +==== Step 12 — domainmgr saw the topology and the isolated set ====
 +
 +<code bash>
 +logread | grep -E 'CPU topology|Kernel isolates'
 +</code>
 +
 +<code>
 +msg":"CPU topology: 4 physical cores"
 +msg":"Kernel isolates CPUs [4 5 6 7]; dedicated workloads are placed there first and everything else is kept off them"
 +</code>
 +
 +''CPU topology: N physical cores'' **must** appear. If topology discovery failed, the model is synthetic and whole-core placement is refused outright rather than performed against a fabricated topology.
 +
 +Ignore the wording of the second line — it is printed whenever the isolated set is non-empty, whether or not the switch is on. Step 7 is what tells you the switch is on.
 +
 +==== Step 13 — The pools changed ====
 +
 +<code bash>
 +cat /run/domainmgr/CPUPoolStatus/*.json
 +</code>
 +
 +<code javascript>
 +{"Pools":[
 + {"Kind":1,"CPUs":[0,1,2,3,4,5,6,7],"FreeCPUs":[2,3],"TotalThreads":8,"AllocatedThreads":6,"FreeThreads":2,"TotalCores":4,"FreeWholeCores":1},
 + {"Kind":2,"CPUs":null,"TotalThreads":0,"FreeWholeCores":0},
 + {"Kind":3,"CPUs":[4,5,6,7],"FreeCPUs":[4,5,6,7],"TotalThreads":4,"AllocatedThreads":0,"FreeThreads":4,"TotalCores":2,"FreeWholeCores":2}
 +]}
 +</code>
 +
 +==== Step 14 — BEFORE vs AFTER at a glance ====
 +
 +^ Check ^ BEFORE ^ AFTER ^
 +| ''/sys/devices/system/cpu/isolated'' | empty | ''4-7'' |
 +| ''eve_max_vcpus'' | 1 | 2 |
 +| Isolated pool (Kind 3) | does not exist | ''[4,5,6,7]'', 2 whole cores |
 +| Housekeeping free CPUs (Kind 1) | ''[1,2,3,4,5,6,7]'' | ''[2,3]'' |
 +| Housekeeping free whole cores | 3 | 1 |
 +
 +The housekeeping line is the one to point at. Of 8 threads, 6 are now allocated: CPUs 0-1 reserved for EVE and CPUs 4-7 withheld for isolation. Only core 1 is left for an ordinary workload. That is the feature working: housekeeping freeness answers //"will an ordinary workload fit?"//, and isolated CPUs cannot serve that.
 +
 +==== Step 15 — Nothing is running on the isolated cores ====
 +
 +<code bash>
 +awk '/^cpu[4-7] /{print $1, $2+$4}' /proc/stat
 +</code>
 +
 +<code>
 +cpu4 0
 +cpu5 0
 +cpu6 0
 +cpu7 0
 +</code>
 +
 +Zero busy ticks since boot. Keep this number — it is the strongest single line in the demo.
 +
 +----
 +
 +===== Part 4 — Demo =====
 +
 +==== Step 16 — A pinned workload lands on the isolated cores ====
 +
 +Deploy a **2-vCPU** app with CPU pinning enabled (even vCPU counts only — an odd count fails closed with ''cpu.policy.odd_vcpu'').
 +
 +<code bash>
 +cat /run/domainmgr/cpuplan.json
 +cat /run/domainmgr/CPUPoolStatus/*.json
 +</code>
 +
 +<code javascript>
 +{ "display_name": "TF-STND-VM-1", "mode": "whole-core-smt", "vcpus": 2,
 +  "status": "success", "host_cpus": [4, 5] }
 +
 +{"Pools":[
 + {"Kind":1,"CPUs":[0,1,2,3,6,7],"FreeCPUs":[2,3],"AllocatedThreads":4,"FreeWholeCores":1},
 + {"Kind":2,"CPUs":[4,5],"FreeCPUs":null,"AllocatedThreads":2,"FreeWholeCores":0},
 + {"Kind":3,"CPUs":[4,5,6,7],"FreeCPUs":[6,7],"AllocatedThreads":2,"FreeWholeCores":1}
 +]}
 +</code>
 +
 +Read it as: the isolated pool went from 2 free whole cores to 1, the dedicated pool //is// ''[4,5]'' now (the isolated pool overlaps the other two rather than partitioning with them), and core 1 has been handed back to housekeeping.
 +
 +Worth saying out loud during a demo: the app asked only for ''pin_cpu''. It arrives with ''FullPCPUsOnly:false'' and ''ThreadsPerCore:0'' — the new API fields unset — and still gets whole-core placement on kernel-isolated cores. That is the point of the node-level switch: a controller that cannot express the new policy still gets the behaviour.
 +
 +==== Step 17 — Capacity fails closed instead of spilling ====
 +
 +//Not yet captured on hardware; this is the expected behaviour.//
 +
 +With the switch **on**, a second 2-vCPU pinned app is promoted onto the remaining isolated core (''[6,7]''), and the **third** fails — while core 1 (''[2,3]'') sits completely free in housekeeping. A node with a visibly idle physical core refusing the workload is the most persuasive part of the demo, and it is by design: a promoted workload may only be served from the isolated set.
 +
 +With the switch **off**, it is the //second// app that fails, for the same reason.
 +
 +The error names how many cores were withheld for isolation, so the operator is not left comparing "insufficient" against CPUs that look idle.
 +
 +==== Step 18 — Prove it at the thread level ====
 +
 +The pool report is EVE's own bookkeeping. This shows the kernel agrees.
 +
 +<code bash>
 +ls /run/hypervisor/kvm/                        # gives <app-uuid>.<gen>.<inst>
 +PID=$(pgrep -f qemu-system | head -1)
 +for t in /proc/$PID/task/*; do
 +  printf '%-18s %s\n' "$(cat $t/comm)" "$(awk '/Cpus_allowed_list/{print $2}' $t/status)"
 +done | sort -u
 +find /sys/fs/cgroup/cpuset/eve-user-apps -maxdepth 2 -name cpuset.cpus | while read f; do echo "$f = $(cat $f)"; done
 +awk '/^cpu[0-9] /{print $1, $2+$4}' /proc/stat
 +</code>
 +
 +<code>
 +qemu-system-x86    4          <-- vCPU 0
 +qemu-system-x86    5          <-- vCPU 1
 +qemu-system-x86    4-5        <-- emulator / IO threads
 +vhost-7981         4-5
 +iou-wrk-7981       4-5
 +
 +/sys/fs/cgroup/cpuset/eve-user-apps/faa814f0-….2.1/cpuset.cpus = 4-5
 +
 +cpu4 1990   cpu5 813   cpu6 0   cpu7 0
 +</code>
 +
 +The two vCPU threads are pinned **1:1** to CPUs 4 and 5; everything else is confined to ''4-5''; the other isolated core is still at zero. Inside the guest, ''lscpu'' shows ''Thread(s) per core: 2'' — EVE truthfully advertising one physical core rather than two sockets.
 +
 +==== Step 19 — One-shot cross-check ====
 +
 +The branch ships a script that cross-checks the allocator's record against real per-thread affinity masks, host SMT siblings and the guest's view, and exits non-zero on any mismatch. It also catches an ''isolcpus'' range that is not sibling-complete.
 +
 +<code bash>
 +GUEST_PASS=<password> ./cpu-assignment-report.sh <node-ip> <app-ip>
 +</code>
 +
 +Override its lab defaults (''NODE_IP'', ''SSH_KEY=~/.ssh/ztest_key'', ''GUEST_USER=pocuser'') and make sure ''debug.enable.ssh'' is set on the node.
 +
 +----
 +
 +===== Part 5 — Known limitations =====
 +
 +==== 5.1 nohz_full and rcu_nocbs do nothing on the shipped kernel ====
 +
 +<code bash>
 +dmesg | grep -iE 'nohz|Unknown kernel command line'
 +(zcat /proc/config.gz || cat /proc/config) | grep -E 'CONFIG_NO_HZ_FULL|CONFIG_RCU_NOCB_CPU|CONFIG_CPU_ISOLATION'
 +</code>
 +
 +<code>
 +Housekeeping: nohz unsupported. Build with CONFIG_NO_HZ_FULL
 +Unknown kernel command line parameters "... rcu_nocbs=4,5,6,7 nohz_full=4,5,6,7", will be passed to user space
 +# CONFIG_NO_HZ_FULL is not set
 +CONFIG_CPU_ISOLATION=y
 +</code>
 +
 +^ Argument ^ Effect ^
 +| ''isolcpus'' | **works** |
 +| ''irqaffinity'' | **works** |
 +| ''nohz_full'' | **no effect** — kernel not built with it |
 +| ''rcu_nocbs'' | **no effect** — kernel not built with it |
 +
 +So what you can demonstrate today is **scheduler isolation**: the load balancer is kept off the isolated cores and nothing else runs there. You cannot yet claim timer-tick shedding or RCU callback offload. That needs ''CONFIG_NO_HZ_FULL'' and ''CONFIG_RCU_NOCB_CPU'' in eve-kernel. Leave the arguments in place — they are harmless and become effective as soon as the kernel supports them.
 +
 +==== 5.2 EVE's own cpuset stays on CPU 0 ====
 +
 +''/sys/fs/cgroup/cpuset/eve/cpuset.cpus'' remains ''0'' even with ''eve_max_vcpus=2'', because that cpuset is driven by ''dom0_max_vcpus'' (still 1); ''eve_max_vcpus'' only sets the ''eve/services'' child, which cannot exceed its parent. No workload capacity is lost — core 0 is dropped from placement either way — but CPU 1 sits reserved and unused. To give EVE its whole core, add ''set_global hv_dom0_cpu_settings "dom0_max_vcpus=2 dom0_vcpus_pin"''.
 +
 +==== 5.3 The vault reports a PCR mismatch on the first boot ====
 +
 +Changing the kernel command line changes PCRs 8, 9 and 14, so ''eve diag'' shows ''vault: ENABLED unlock:controller-key mismatchPCRs:[8 9 14]'' on the boot right after the change. It re-seals itself and returns to ''unlock:tpm-local-sealed'' on the following boot. Expected, not a fault.
 +
 +----
 +
 +===== Part 6 — Troubleshooting =====
 +
 +^ Symptom ^ Cause ^ Fix ^
 +| Cannot find the property in the controller UI | it is not exposed there | set it via Terraform or REST (Step 6) |
 +| Property set in Terraform, node still shows ''BoolValue:false'' | ''bool_value'' serializes empty | use ''string_value = "true"'' in the ''config_item'' block |
 +| Property set on the application instead of the node | ''cpu.pinning.use.isolated'' is node-level only | move it to a ''config_item'' on ''zedcloud_edgenode'' |
 +| ''/sys/devices/system/cpu/isolated'' empty after reboot | GRUB did not read the file, or the kernel rejected the argument | check ''/proc/cmdline''; confirm the file is ''grub.cfg'', not ''grub.cfg.tmpl'' |
 +| ''$dom0_extra_args'' missing from ''/proc/cmdline'' | unquoted heredoc expanded it away | rewrite ''grub.cfg'' using ''<<'EOF''' |
 +| **Pinned VM stays on the dedicated cores after enabling the switch** | a running workload holds its allocation; ''doActivate'' returns early on ''len(status.VmConfig.CPUs) > 0'' before the promotion is applied, and ''releaseCPUs'' only runs on the delete/deactivate path | **reboot the node** (''DomainStatus'' lives in ''/run'') or **deactivate and reactivate** the app instance. An app restart or a boot retry is **not** enough |
 +| VM halted with ''BootFailed: true'', retries onto the same CPUs | someone killed the qemu process instead of stopping the workload properly | reboot, or deactivate/reactivate from the controller |
 +| ''cpu.topology.unsupported'' | no core has two hardware threads | enable SMT in the BIOS; some CPUs have none |
 +| ''cpu.policy.odd_vcpu'' | odd vCPU count on a whole-core request | use an even count |
 +| Placement refused citing a synthetic topology | sysfs topology discovery failed | check ''/sys/devices/system/cpu/cpu*/topology/'' is populated |
 +| Isolated cores free but nothing can use them | expected with the switch off — they are withheld from every workload that did not ask for isolation | set the property (Step 6), or use ''isolation_tier=hard'' |
 +| SSH host key changed after reboot | EVE regenerates ''/etc/ssh/ssh_host_*'' every boot | ''ssh-keygen -R <node-ip>'' |
 +| Boot loop, ''fatal: agent zedbox[N]: zboot curpart: err exit status 64'' | two disks carry EVE's static PARTUUIDs — install media left in, or a second EVE install | remove the media, or wipe the second disk's partition table |
 +
 +----
 +
 +===== Part 7 — Reference =====
 +
 +  * ''docs/cpu-affinity-design.md'' in the feature branch — isolation tiers, placement vocabulary, the EVE↔controller interface. Large parts are marked //not yet implemented//; it describes the intended end state, not this build.
 +  * ''docs/CONFIG-PROPERTIES.md'' — all configuration properties.
 +  * Pool kinds and error codes: ''pkg/pillar/types/cpuplacement.go''.
 +  * Placement logic: ''pkg/pillar/cpuallocator/placement.go''.
 +  * The allocation-reuse shortcut discussed in Troubleshooting: ''pkg/pillar/cmd/domainmgr/domainmgr.go'' (''doActivate''), and ''releaseCPUs'' in the same file.
 +
eve-kvm/core-isolation.1789749252.txt.gz · Last modified: by mc