User Tools

Site Tools


eve-kvm:core-isolation

This is an old revision of the document!


EVE CPU Core Isolation — Step-by-Step Runbook

Follow the steps in order. Every step has a command, the output you should see, and what to do if you do not see it. Do not skip Step 5 (the BEFORE baseline) — it is what makes the AFTER checks meaningful.

Scope. The core-pinning feature branch (eve PR #6335, eve-api #155, adam #158, NFR-172). Verified end to end on EVE 0.0.0-core-pinning-gmwtus-551edfcb-kvm-amd64, node = Intel Core i3-13100TE, on 2026-09-18. Every transcript below is real output from that run except Step 17, which is marked.

Time required. About 20 minutes, including one reboot.


Part 0 — Understand what you are doing

There are two independent switches. Neither works without the other.

# Switch Where it is set What it does Needs reboot?
1 isolcpus kernel argument /config/grub.cfg on the node creates the isolated CPU set yes — kernel command line
2 cpu.pinning.use.isolated node-level property in the controller — Terraform or REST only, not in the UI lets pinned workloads use that set no, but see Step 9

A third thing is often confused with these:

  • “CPU pinning” on the application (pin_cpu) is a per-app setting. It gives the workload whole physical cores out of the dedicated pool. It is not the isolation switch. If you only set this, your VM gets pinned — just not onto the isolated cores.

Target layout on a 4-core / 8-thread node:

Role Physical core CPUs
EVE housekeeping core 0 0, 1
Dedicated pool (ordinary pinned workloads) core 1 2, 3
Isolated pool cores 2-3 4, 5, 6, 7

Part 1 — BEFORE: prerequisites and baseline

Run every command in this part on the node, over SSH or the console, before changing anything.

Step 1 — Confirm the CPU can do this

lscpu | grep -E 'Model name|Thread\(s\) per core|Core\(s\) per socket|NUMA node\(s\)'

Expected:

Model name:            13th Gen Intel(R) Core(TM) i3-13100TE
Thread(s) per core:    2
Core(s) per socket:    4
NUMA node(s):          1

PASS if Thread(s) per core is 2. If it is 1, stop: either SMT is disabled in the BIOS (enable it and reboot) or the CPU has no SMT and cannot run this feature at all. There is no fallback — pinning fails closed with cpu.topology.unsupported.

You also need at least 3 physical cores: one for EVE, one for the dedicated pool, one to isolate.

CPU Topology Usable
Core i3-10100 / 10105 / 10300 4C/8T yes
Core i3-12100 / 12300, i3-13100, i3-14100 4 P-cores, 0 E-cores, 8T yes — ideal
Core i5-12400 6 P-cores, 0 E-cores, 12T yes — more headroom
Core i3-8100 / 8300, i3-9100 / 9300 4C/4T, no HT no
Core i3-N300 / i3-N305 (Alder Lake-N) 8 E-cores, no HT no
Core Ultra 200S / Meteor Lake / Lunar Lake SMT removed no

Avoid Intel hybrid parts with E-cores (i5/i7 12th gen and up): the E-cores have no SMT and are silently skipped for pinning, so they add nothing and confuse the results.

Step 2 — Find out which CPUs are SMT siblings

This decides the numbers you will type in Step 8. Do not assume them.

cat /sys/devices/system/cpu/cpu[0-9]*/topology/thread_siblings_list | sort -u

Expected on the reference node:

0-1
2-3
4-5
6-7
If you see Enumeration Use the values in
0-1 2-3 4-5 6-7 adjacent (Intel 12th gen and newer) Step 8a
0,4 1,5 2,6 3,7 split (older Intel) Step 8b

Step 3 — Confirm nothing is isolated yet

cat /sys/devices/system/cpu/isolated

Expected: empty output. That is the starting state. If it already lists CPUs, someone has been here before — read the existing /config/grub.cfg before you overwrite it.

Step 4 — Confirm the property is currently off

grep -o 'cpu.pinning.use.isolated[^}]*}' /run/zedagent/ConfigItemValueMap/global.json

Expected:

cpu.pinning.use.isolated":{"Key":"cpu.pinning.use.isolated","ItemType":2,"IntValue":0,"StrValue":"","BoolValue":false,"TriStateValue":0}

BoolValue is false. If the key is missing entirely, this EVE build does not have the feature — check eve version.

Step 5 — Capture the BEFORE baseline

Paste this whole block. Save the output somewhere; you will compare against it in Step 12.

echo "=== cmdline ==="
tr ' ' '\n' < /proc/cmdline | grep -E 'isolcpus|nohz_full|rcu_nocbs|irqaffinity|eve_max_vcpus'
echo "=== kernel isolated set ==="
echo "[$(cat /sys/devices/system/cpu/isolated)]"
echo "=== CPU pools ==="
cat /run/domainmgr/CPUPoolStatus/*.json; echo
echo "=== placement plan ==="
cat /run/domainmgr/cpuplan.json 2>/dev/null
echo "=== eve cpuset ==="
cat /sys/fs/cgroup/cpuset/eve/cpuset.cpus
echo "=== per-CPU busy ticks ==="
awk '/^cpu[0-9] /{print $1, $2+$4}' /proc/stat

Expected BEFORE (no isolation, nothing pinned):

=== cmdline ===
eve_max_vcpus=1
=== kernel isolated set ===
[]
=== CPU pools ===
{"Pools":[
 {"Kind":1,"CPUs":[0,1,2,3,4,5,6,7],"FreeCPUs":[1,2,3,4,5,6,7],"TotalThreads":8,"AllocatedThreads":1,"FreeThreads":7,"TotalCores":4,"FreeWholeCores":3},
 {"Kind":2,"CPUs":null,"TotalThreads":0,"FreeWholeCores":0},
 {"Kind":3,"CPUs":null,"TotalThreads":0,"FreeWholeCores":0}
]}
=== eve cpuset ===
0

Pool Kind values, needed for every later step: 1 = housekeeping, 2 = dedicated, 3 = isolated.

What this baseline says: all 8 threads are in housekeeping, only CPU 0 is taken (by EVE), and the dedicated and isolated pools do not exist.


Part 2 — ENABLE

Do these three steps in this order. Doing the property last costs you a second reboot.

Step 6 — Set the controller property

This property is not exposed in the ZedControl UI. Do not go looking for it there — the UI only offers properties it knows about, and this one is new in the feature branch. It can only be set via Terraform or the REST API.

Terraform — a config_item block on the zedcloud_edgenode resource. Not on the application, not on the app instance.

resource "zedcloud_edgenode" "my_node" {
  # ... existing fields ...

  config_item {
    key          = "cpu.pinning.use.isolated"
    string_value = "true"
  }
}

Then terraform apply.

Use string_value = “true”. This is the single most common way to get “I set it and nothing happened”, and the reason is that two different schemas are involved:

Hop Schema Fields
you → controller EDConfigItem key, valueType, stringValue, boolValue, floatValue, uint32Value, uint64Value
controller → device eve-api ConfigItem key, value — both strings

The controller has to collapse the typed fields into a single string value for the device. An entry carrying only boolValue: true with no valueType can arrive at the device as an empty value, which pillar cannot parse, so the item keeps its default — false. stringValue passes straight through. The provider also exposes an optional value_type attribute; it is not needed when you use string_value.

REST — PUT /api/v1/devices/id/{id} (operation EdgeNodeConfiguration_UpdateEdgeNode). The property goes in the device object's configItem array, whose elements are EDConfigItem:

{ "key": "cpu.pinning.use.isolated", "stringValue": "true" }

This is a full-object PUT, not a patch. The body carries the whole edge node — name, title, modelId, projectId, interfaces, adminState, every existing configItem and about forty more fields. Send a partial body and you erase what you left out. It also requires revision.curr, the current database version of the record; a stale value is rejected with 409 Version mismatch.

So the REST procedure is read-modify-write:

  1. GET /api/v1/devices/id/{id}
  2. append { “key”: “cpu.pinning.use.isolated”, “stringValue”: “true” } to the returned configItem array, keeping every existing entry
  3. PUT the entire modified object back, including revision

This is why Terraform is the easier route: the provider does that read-modify-write for you. Schema verified against swagger/zedge_node_service.swagger.json (ZEDEDA Edge Node Service v1.0) in the provider source; confirm against your own controller's /api/v1/docs/ if it runs a different version.

Do not expect to find the key itself in the API documentation. In the spec key is a plain string with no enum and no validation — configItem is a free-form key/value list that the controller stores and forwards to the device verbatim. The only component that knows cpu.pinning.use.isolated is a real property is EVE itself (pillar's types/global.go). That is why the property works through a controller that has never heard of it, and equally why the UI does not offer it: the UI renders a curated list of known properties, the API accepts any string.

Whichever route you use, Step 7 is what confirms it — do not assume it landed.

Step 7 — Confirm the property reached the node

Do not go further until this passes. Allow a minute for the config poll.

grep -o 'cpu.pinning.use.isolated[^}]*}' /run/zedagent/ConfigItemValueMap/global.json
logread | grep 'use.isolated'

Expected:

"BoolValue":true
domainmgr: CPU placement: cpu.pinning.use.isolated is now true; workloads placed from now on may use the kernel-isolated CPUs

If BoolValue is still false: you used bool_value instead of string_value (go back to Step 6), or the controller rejected the key — see Troubleshooting.

Step 8 — Write /config/grub.cfg

/config/grub.cfg does not exist on a fresh install. The image ships only grub.cfg.tmpl, and GRUB reads grub.cfg exactly. Create the file. Do not rename the template.

Step 8a — adjacent siblings (0-1, 2-3, 4-5, 6-7)

eve config mount /tmp/cfg
cat > /tmp/cfg/grub.cfg <<'EOF'
set_getty
set_global hv_eve_cpu_settings "eve_max_vcpus=2"
set_global dom0_extra_args "$dom0_extra_args isolcpus=managed_irq,domain,4,5,6,7 rcu_nocbs=4,5,6,7 nohz_full=4,5,6,7 irqaffinity=0,1"
EOF
cat /tmp/cfg/grub.cfg
sync
eve config unmount

Step 8b — split siblings (0,4 1,5 2,6 3,7)

eve config mount /tmp/cfg
cat > /tmp/cfg/grub.cfg <<'EOF'
set_getty
set_global dom0_extra_args "$dom0_extra_args isolcpus=managed_irq,domain,2,3,6,7 rcu_nocbs=2,3,6,7 nohz_full=2,3,6,7 irqaffinity=0,4"
EOF
cat /tmp/cfg/grub.cfg
sync
eve config unmount

There is no eve_max_vcpus line in 8b on purpose. Reservation is by lowest CPU id and a core is dropped if any sibling is reserved, so eve_max_vcpus=2 with split enumeration would reserve CPUs 0 and 1 — two halves of two different cores — and cost you a second core for nothing.

Check the cat output before moving on. The third line must still contain the literal text $dom0_extra_args. If that text is missing, your shell expanded it and GRUB will lose everything it had accumulated — rewrite it, making sure the heredoc marker is quoted as «'EOF' and not «EOF.

Rules that matter here:

  • irqaffinity lists the housekeeping CPUs, never the isolated ones. Pointing it at the isolated set steers interrupts onto the cores you are shielding.
  • set_getty keeps a console shell. It is the only thing in the shipped template, so if you do not include it you lose it.
  • Do not use GRUB's set_isolcpus helper (menu entry isolate CPU0 (only for PREEMPT_RT)). It emits isolcpus=inverse,0, which isolates CPU 0's own sibling and produces a core that serves nobody.

Step 9 — Reboot

reboot

Allow about 3 minutes. The reboot is required for two reasons: isolcpus is a kernel command-line argument, and a reboot is also what lets already-running workloads be re-placed (see Troubleshooting — a restart is not enough).

Expect the SSH host key to change: EVE regenerates /etc/ssh/ssh_host_* on every boot. Clear the old entry with ssh-keygen -R <node-ip>.


Part 3 — AFTER: validate

Run these in order. Stop at the first failure.

Step 10 — The kernel accepted the arguments

tr ' ' '\n' < /proc/cmdline | grep -E 'isolcpus|nohz_full|rcu_nocbs|irqaffinity|eve_max_vcpus'
eve_max_vcpus=2
isolcpus=managed_irq,domain,4,5,6,7
rcu_nocbs=4,5,6,7
nohz_full=4,5,6,7
irqaffinity=0,1

If nothing appears: GRUB never read your file. Confirm it is named grub.cfg (not grub.cfg.tmpl) on the CONFIG partition.

Step 11 — The kernel is actually isolating

cat /sys/devices/system/cpu/isolated
4-7

This is the make-or-break check. An empty result means there is no isolated pool and nothing after this point can work. Do not continue — fix the command line first.

Step 12 — domainmgr saw the topology and the isolated set

logread | grep -E 'CPU topology|Kernel isolates'
msg":"CPU topology: 4 physical cores"
msg":"Kernel isolates CPUs [4 5 6 7]; dedicated workloads are placed there first and everything else is kept off them"

CPU topology: N physical cores must appear. If topology discovery failed, the model is synthetic and whole-core placement is refused outright rather than performed against a fabricated topology.

Ignore the wording of the second line — it is printed whenever the isolated set is non-empty, whether or not the switch is on. Step 7 is what tells you the switch is on.

Step 13 — The pools changed

cat /run/domainmgr/CPUPoolStatus/*.json
{"Pools":[
 {"Kind":1,"CPUs":[0,1,2,3,4,5,6,7],"FreeCPUs":[2,3],"TotalThreads":8,"AllocatedThreads":6,"FreeThreads":2,"TotalCores":4,"FreeWholeCores":1},
 {"Kind":2,"CPUs":null,"TotalThreads":0,"FreeWholeCores":0},
 {"Kind":3,"CPUs":[4,5,6,7],"FreeCPUs":[4,5,6,7],"TotalThreads":4,"AllocatedThreads":0,"FreeThreads":4,"TotalCores":2,"FreeWholeCores":2}
]}

Step 14 — BEFORE vs AFTER at a glance

Check BEFORE AFTER
/sys/devices/system/cpu/isolated empty 4-7
eve_max_vcpus 1 2
Isolated pool (Kind 3) does not exist [4,5,6,7], 2 whole cores
Housekeeping free CPUs (Kind 1) [1,2,3,4,5,6,7] [2,3]
Housekeeping free whole cores 3 1

The housekeeping line is the one to point at. Of 8 threads, 6 are now allocated: CPUs 0-1 reserved for EVE and CPUs 4-7 withheld for isolation. Only core 1 is left for an ordinary workload. That is the feature working: housekeeping freeness answers “will an ordinary workload fit?”, and isolated CPUs cannot serve that.

Step 15 — Nothing is running on the isolated cores

awk '/^cpu[4-7] /{print $1, $2+$4}' /proc/stat
cpu4 0
cpu5 0
cpu6 0
cpu7 0

Zero busy ticks since boot. Keep this number — it is the strongest single line in the demo.


Part 4 — Demo

Step 16 — A pinned workload lands on the isolated cores

Deploy a 2-vCPU app with CPU pinning enabled (even vCPU counts only — an odd count fails closed with cpu.policy.odd_vcpu).

cat /run/domainmgr/cpuplan.json
cat /run/domainmgr/CPUPoolStatus/*.json
{ "display_name": "TF-STND-VM-1", "mode": "whole-core-smt", "vcpus": 2,
  "status": "success", "host_cpus": [4, 5] }
 
{"Pools":[
 {"Kind":1,"CPUs":[0,1,2,3,6,7],"FreeCPUs":[2,3],"AllocatedThreads":4,"FreeWholeCores":1},
 {"Kind":2,"CPUs":[4,5],"FreeCPUs":null,"AllocatedThreads":2,"FreeWholeCores":0},
 {"Kind":3,"CPUs":[4,5,6,7],"FreeCPUs":[6,7],"AllocatedThreads":2,"FreeWholeCores":1}
]}

Read it as: the isolated pool went from 2 free whole cores to 1, the dedicated pool is [4,5] now (the isolated pool overlaps the other two rather than partitioning with them), and core 1 has been handed back to housekeeping.

Worth saying out loud during a demo: the app asked only for pin_cpu. It arrives with FullPCPUsOnly:false and ThreadsPerCore:0 — the new API fields unset — and still gets whole-core placement on kernel-isolated cores. That is the point of the node-level switch: a controller that cannot express the new policy still gets the behaviour.

Step 17 — Capacity fails closed instead of spilling

Not yet captured on hardware; this is the expected behaviour.

With the switch on, a second 2-vCPU pinned app is promoted onto the remaining isolated core ([6,7]), and the third fails — while core 1 ([2,3]) sits completely free in housekeeping. A node with a visibly idle physical core refusing the workload is the most persuasive part of the demo, and it is by design: a promoted workload may only be served from the isolated set.

With the switch off, it is the second app that fails, for the same reason.

The error names how many cores were withheld for isolation, so the operator is not left comparing “insufficient” against CPUs that look idle.

Step 18 — Prove it at the thread level

The pool report is EVE's own bookkeeping. This shows the kernel agrees.

ls /run/hypervisor/kvm/                        # gives <app-uuid>.<gen>.<inst>
PID=$(pgrep -f qemu-system | head -1)
for t in /proc/$PID/task/*; do
  printf '%-18s %s\n' "$(cat $t/comm)" "$(awk '/Cpus_allowed_list/{print $2}' $t/status)"
done | sort -u
find /sys/fs/cgroup/cpuset/eve-user-apps -maxdepth 2 -name cpuset.cpus | while read f; do echo "$f = $(cat $f)"; done
awk '/^cpu[0-9] /{print $1, $2+$4}' /proc/stat
qemu-system-x86    4          <-- vCPU 0
qemu-system-x86    5          <-- vCPU 1
qemu-system-x86    4-5        <-- emulator / IO threads
vhost-7981         4-5
iou-wrk-7981       4-5

/sys/fs/cgroup/cpuset/eve-user-apps/faa814f0-….2.1/cpuset.cpus = 4-5

cpu4 1990   cpu5 813   cpu6 0   cpu7 0

The two vCPU threads are pinned 1:1 to CPUs 4 and 5; everything else is confined to 4-5; the other isolated core is still at zero. Inside the guest, lscpu shows Thread(s) per core: 2 — EVE truthfully advertising one physical core rather than two sockets.

Step 19 — One-shot cross-check

The branch ships a script that cross-checks the allocator's record against real per-thread affinity masks, host SMT siblings and the guest's view, and exits non-zero on any mismatch. It also catches an isolcpus range that is not sibling-complete.

GUEST_PASS=<password> ./cpu-assignment-report.sh <node-ip> <app-ip>

Override its lab defaults (NODE_IP, SSH_KEY=~/.ssh/ztest_key, GUEST_USER=pocuser) and make sure debug.enable.ssh is set on the node.


Part 5 — Known limitations

5.1 nohz_full and rcu_nocbs do nothing on the shipped kernel

dmesg | grep -iE 'nohz|Unknown kernel command line'
(zcat /proc/config.gz || cat /proc/config) | grep -E 'CONFIG_NO_HZ_FULL|CONFIG_RCU_NOCB_CPU|CONFIG_CPU_ISOLATION'
Housekeeping: nohz unsupported. Build with CONFIG_NO_HZ_FULL
Unknown kernel command line parameters "... rcu_nocbs=4,5,6,7 nohz_full=4,5,6,7", will be passed to user space
# CONFIG_NO_HZ_FULL is not set
CONFIG_CPU_ISOLATION=y
Argument Effect
isolcpus works
irqaffinity works
nohz_full no effect — kernel not built with it
rcu_nocbs no effect — kernel not built with it

So what you can demonstrate today is scheduler isolation: the load balancer is kept off the isolated cores and nothing else runs there. You cannot yet claim timer-tick shedding or RCU callback offload. That needs CONFIG_NO_HZ_FULL and CONFIG_RCU_NOCB_CPU in eve-kernel. Leave the arguments in place — they are harmless and become effective as soon as the kernel supports them.

5.2 EVE's own cpuset stays on CPU 0

/sys/fs/cgroup/cpuset/eve/cpuset.cpus remains 0 even with eve_max_vcpus=2, because that cpuset is driven by dom0_max_vcpus (still 1); eve_max_vcpus only sets the eve/services child, which cannot exceed its parent. No workload capacity is lost — core 0 is dropped from placement either way — but CPU 1 sits reserved and unused. To give EVE its whole core, add set_global hv_dom0_cpu_settings “dom0_max_vcpus=2 dom0_vcpus_pin”.

5.3 The vault reports a PCR mismatch on the first boot

Changing the kernel command line changes PCRs 8, 9 and 14, so eve diag shows vault: ENABLED unlock:controller-key mismatchPCRs:[8 9 14] on the boot right after the change. It re-seals itself and returns to unlock:tpm-local-sealed on the following boot. Expected, not a fault.


Part 6 — Troubleshooting

Symptom Cause Fix
Cannot find the property in the controller UI it is not exposed there set it via Terraform or REST (Step 6)
Property set in Terraform, node still shows BoolValue:false bool_value serializes empty use string_value = “true” in the config_item block
Property set on the application instead of the node cpu.pinning.use.isolated is node-level only move it to a config_item on zedcloud_edgenode
/sys/devices/system/cpu/isolated empty after reboot GRUB did not read the file, or the kernel rejected the argument check /proc/cmdline; confirm the file is grub.cfg, not grub.cfg.tmpl
$dom0_extra_args missing from /proc/cmdline unquoted heredoc expanded it away rewrite grub.cfg using «'EOF'
Pinned VM stays on the dedicated cores after enabling the switch a running workload holds its allocation; doActivate returns early on len(status.VmConfig.CPUs) > 0 before the promotion is applied, and releaseCPUs only runs on the delete/deactivate path reboot the node (DomainStatus lives in /run) or deactivate and reactivate the app instance. An app restart or a boot retry is not enough
VM halted with BootFailed: true, retries onto the same CPUs someone killed the qemu process instead of stopping the workload properly reboot, or deactivate/reactivate from the controller
cpu.topology.unsupported no core has two hardware threads enable SMT in the BIOS; some CPUs have none
cpu.policy.odd_vcpu odd vCPU count on a whole-core request use an even count
Placement refused citing a synthetic topology sysfs topology discovery failed check /sys/devices/system/cpu/cpu*/topology/ is populated
Isolated cores free but nothing can use them expected with the switch off — they are withheld from every workload that did not ask for isolation set the property (Step 6), or use isolation_tier=hard
SSH host key changed after reboot EVE regenerates /etc/ssh/ssh_host_* every boot ssh-keygen -R <node-ip>
Boot loop, fatal: agent zedbox[N]: zboot curpart: err exit status 64 two disks carry EVE's static PARTUUIDs — install media left in, or a second EVE install remove the media, or wipe the second disk's partition table

Part 7 — Reference

  • docs/cpu-affinity-design.md in the feature branch — isolation tiers, placement vocabulary, the EVE↔controller interface. Large parts are marked not yet implemented; it describes the intended end state, not this build.
  • docs/CONFIG-PROPERTIES.md — all configuration properties.
  • Pool kinds and error codes: pkg/pillar/types/cpuplacement.go.
  • Placement logic: pkg/pillar/cpuallocator/placement.go.
  • The allocation-reuse shortcut discussed in Troubleshooting: pkg/pillar/cmd/domainmgr/domainmgr.go (doActivate), and releaseCPUs in the same file.
eve-kvm/core-isolation.1789754736.txt.gz · Last modified: by mc