This is an old revision of the document!
Table of Contents
EVE CPU Core Pinning and Isolation
How to enable kernel CPU isolation on an EVE node, prove it took effect, and demo it.
Scope. Applies to the core-pinning feature branch (eve PR #6335, eve-api #155, adam #158, NFR-172). Verified on EVE 0.0.0-core-pinning-gmwtus-551edfcb-kvm-amd64.
Provenance. The transcripts in §3, §4.1 and §4.4 were captured on a live node on 2026-09-18. The outputs in §4.2 and §4.3 are expected values, not yet captured.
1. Requirements
1.1 CPU
Pinning on this branch means whole-core SMT: each vCPU gets one hardware thread of a physical core, and both siblings go to the same workload. The allocator will only use a core that has exactly two hardware threads.
- SMT / Hyper-Threading is mandatory and must be enabled in the BIOS. A node with no sibling threads fails closed with
cpu.topology.unsupported— there is no fallback to thread-granular placement. - At least 3 physical cores to be useful: one for EVE housekeeping, one for the dedicated pool, one or more to isolate.
- Even vCPU counts only. An odd count fails closed with
cpu.policy.odd_vcpu. - Homogeneous cores are strongly preferred. On Intel hybrid parts the E-cores have no SMT and are silently skipped for pinning.
| CPU | Topology | Usable |
|---|---|---|
| Core i3-10100 / 10105 / 10300 | 4C/8T | yes |
| Core i3-12100 / 12300, i3-13100, i3-14100 | 4 P-cores, 0 E-cores, 8T | yes — ideal |
| Core i5-12400 | 6 P-cores, 0 E-cores, 12T | yes — more headroom |
| Core i3-8100 / 8300, i3-9100 / 9300 | 4C/4T, no HT | no |
| Core i3-N300 / i3-N305 (Alder Lake-N) | 8 E-cores, no HT | no |
| Core Ultra 200S / Meteor Lake / Lunar Lake | SMT removed | no |
Reference node for this page: Intel Core i3-13100TE, 4 P-cores / 8 threads, no E-cores, 1 NUMA node, 12 MiB shared L3, VT-x.
1.2 Check the node before you start
lscpu | grep -E 'Model name|Thread\(s\) per core|Core\(s\) per socket|NUMA node\(s\)'
Thread(s) per core: 2 is the requirement. A 1 here means either the CPU has no SMT or it is disabled in the BIOS; both fail the same way.
2. Enable core isolation
2.1 Determine the SMT sibling enumeration
This decides your CPU indexes. Do not assume it.
cat /sys/devices/system/cpu/cpu[0-9]*/topology/thread_siblings_list | sort -u
Two possible answers on a 4C/8T part:
| Output | Meaning | Core map |
|---|---|---|
0-1 2-3 4-5 6-7 | adjacent siblings (Intel 12th gen and newer) | core0={0,1} core1={2,3} core2={4,5} core3={6,7} |
0,4 1,5 2,6 3,7 | split siblings (older Intel) | core0={0,4} core1={1,5} core2={2,6} core3={3,7} |
2.2 Decide the split
Three roles. The isolated set must be sibling-complete — a core with one thread isolated and the other in housekeeping serves neither ordinary requests nor the isolated pool, and simply disappears from usable capacity.
| Role | Cores | CPUs (adjacent enumeration) |
|---|---|---|
| EVE housekeeping | core 0 | 0,1 |
| Dedicated pool (ordinary pinned workloads) | core 1 | 2,3 |
| Isolated pool | cores 2-3 | 4,5,6,7 |
2.3 Set the controller property FIRST
In the controller, set the edge-node configuration property:
cpu.pinning.use.isolated = true
Order matters. Set this before the reboot that applies isolcpus. The switch only affects workloads placed after it is set, so setting it after the reboot silently does nothing until the next placement. Setting it on a node whose kernel isolates nothing warns rather than errors.
Without this property, kernel-isolated CPUs are reported in the pool report and usable by nobody: they are withheld from every workload that did not explicitly ask for isolation, and only isolation_tier=hard can claim them — which current controllers cannot express.
2.4 Write /config/grub.cfg
/config/grub.cfg does not exist on a fresh install. The image ships only grub.cfg.tmpl, and GRUB reads grub.cfg exactly — so create the file, do not rename the template.
eve config mount /tmp/cfg cat > /tmp/cfg/grub.cfg <<'EOF' set_getty set_global hv_eve_cpu_settings "eve_max_vcpus=2" set_global dom0_extra_args "$dom0_extra_args isolcpus=managed_irq,domain,4,5,6,7 rcu_nocbs=4,5,6,7 nohz_full=4,5,6,7 irqaffinity=0,1" EOF cat /tmp/cfg/grub.cfg sync eve config unmount
For the split enumeration use isolcpus=managed_irq,domain,2,3,6,7 rcu_nocbs=2,3,6,7 nohz_full=2,3,6,7 irqaffinity=0,4 and omit the eve_max_vcpus line: reservation is by lowest CPU id and a core is dropped if any sibling is reserved, so eve_max_vcpus=2 there would reserve CPUs 0 and 1, which belong to two different cores, and cost a second core for nothing.
Notes:
- The quoted heredoc (
«'EOF') is required. Unquoted, the shell expands$dom0_extra_argsto nothing and GRUB loses everything it had accumulated. irqaffinitynames the housekeeping CPUs, not the isolated ones. Pointing it at the isolated set steers interrupts onto the cores you are trying to shield.set_gettykeeps a console shell available; it is the only content of the shipped template, so without it you lose that.- Do not use GRUB's
set_isolcpushelper (the interactive menu entry isolate CPU0 (only for PREEMPT_RT)). It emitsisolcpus=inverse,0, which with the defaulteve_max_vcpus=1isolates CPU 0's own SMT sibling and produces a sibling-incomplete core. /config/grub.cfgis sourced after EVE's defaults and before the boot entry is built, soset_globaloverrides the shippedeve_max_vcpus=1.
2.5 Reboot
reboot
Kernel isolation is a command-line decision, so it is reboot-gated. Expect roughly 3 minutes before SSH answers again.
3. Validate
Run these in order. Stop at the first one that disagrees.
3.1 The kernel accepted the arguments
tr ' ' '\n' < /proc/cmdline | grep -E 'isolcpus|nohz_full|rcu_nocbs|irqaffinity|eve_max_vcpus'
eve_max_vcpus=2 isolcpus=managed_irq,domain,4,5,6,7 rcu_nocbs=4,5,6,7 nohz_full=4,5,6,7 irqaffinity=0,1
3.2 The kernel is actually isolating
cat /sys/devices/system/cpu/isolated
4-7
An empty result means there is no isolated pool and nothing downstream will work. Stop and fix the command line.
3.3 domainmgr saw the topology and the isolated set
logread | grep -E 'CPU topology|Kernel isolates'
msg":"CPU topology: 4 physical cores" msg":"Kernel isolates CPUs [4 5 6 7]; dedicated workloads are placed there first and everything else is kept off them"
CPU topology: N physical cores must appear. If topology discovery failed, the model is synthetic and whole-core placement is refused outright rather than performed against a fabricated topology.
Note: the second line is printed whenever the isolated set is non-empty, regardless of cpu.pinning.use.isolated. It does not prove the promotion switch is on — check the property itself (§3.5).
3.4 The pool report
cat /run/domainmgr/CPUPoolStatus/*.json
Pool Kind values: 1 = housekeeping, 2 = dedicated, 3 = isolated.
{"Pools":[ {"Kind":1,"CPUs":[0,1,2,3,4,5,6,7],"FreeCPUs":[2,3],"TotalThreads":8,"AllocatedThreads":6,"FreeThreads":2,"TotalCores":4,"FreeWholeCores":1}, {"Kind":2,"CPUs":null,"FreeCPUs":null,"TotalThreads":0,"AllocatedThreads":0,"FreeThreads":0,"TotalCores":0,"FreeWholeCores":0}, {"Kind":3,"CPUs":[4,5,6,7],"FreeCPUs":[4,5,6,7],"TotalThreads":4,"AllocatedThreads":0,"FreeThreads":4,"TotalCores":2,"FreeWholeCores":2} ]}
What to read from it:
- Kind 3 is non-empty — the isolated set is now allocatable capacity, not just a number in the device info.
- Kind 1
FreeCPUsis[2,3]only — of 8 threads, 6 are allocated: CPUs 0-1 reserved for EVE and CPUs 4-7 withheld for isolation. This is the point of the feature: housekeeping freeness answers “will an ordinary workload fit?”, which isolated CPUs cannot serve. - Kind 1
FreeWholeCoresis 1 — exactly core 1.
3.5 The promotion switch
grep -o 'cpu.pinning.use.isolated[^}]*}' /run/zedagent/ConfigItemValueMap/global.json
cpu.pinning.use.isolated":{"Key":"cpu.pinning.use.isolated","ItemType":2,"IntValue":0,"StrValue":"","BoolValue":true,"TriStateValue":0}
BoolValue must be true. This value comes from the controller only. A value written locally into /config/GlobalConfig/global.json survives just until the first config poll, because zedagent rebuilds the whole map from defaults plus the controller's items.
3.6 Derived placement plan
cat /run/domainmgr/cpuplan.json
{ "workloads": [ { "uuid": "...", "display_name": "TF-STND-VM-1", "mode": "whole-core-smt", "vcpus": 2, "status": "success", "host_cpus": [4, 5] } ] }
Diagnostic only — nothing reads it back — but it is the quickest way to see what the allocator decided, including for workloads that are configured but not running.
3.7 Known limitations on the shipped kernel
| Argument | Effect | Why |
|---|---|---|
isolcpus | works | CONFIG_CPU_ISOLATION=y |
irqaffinity | works | accepted by the kernel |
nohz_full | no effect | # CONFIG_NO_HZ_FULL is not set |
rcu_nocbs | no effect | CONFIG_RCU_NOCB_CPU not set |
dmesg | grep -iE 'nohz|Unknown kernel command line' (zcat /proc/config.gz || cat /proc/config) | grep -E 'CONFIG_NO_HZ_FULL|CONFIG_RCU_NOCB_CPU|CONFIG_CPU_ISOLATION'
Housekeeping: nohz unsupported. Build with CONFIG_NO_HZ_FULL Unknown kernel command line parameters "... rcu_nocbs=4,5,6,7 nohz_full=4,5,6,7", will be passed to user space # CONFIG_NO_HZ_FULL is not set CONFIG_CPU_ISOLATION=y
So today you get the isolcpus half of hard isolation — the scheduler's load balancer is kept off the isolated cores — but not tick shedding or RCU callback offload. Getting those requires CONFIG_NO_HZ_FULL and CONFIG_RCU_NOCB_CPU in eve-kernel. Leave the arguments in place; they are harmless and become effective as soon as the kernel supports them.
Two more expected observations:
/sys/fs/cgroup/cpuset/eve/cpuset.cpusstays0even witheve_max_vcpus=2. That cpuset is driven bydom0_max_vcpus(still 1);eve_max_vcpusonly sets theeve/serviceschild, which cannot exceed its parent. No workload capacity is lost — core 0 is dropped from placement either way — but CPU 1 sits reserved and unused. To give EVE its whole core, also setset_global hv_dom0_cpu_settings “dom0_max_vcpus=2 dom0_vcpus_pin”.- The vault reports a PCR mismatch. Changing the kernel command line changes PCRs 8, 9 and 14, so
eve diagshowsvault: ENABLED unlock:controller-key mismatchPCRs:[8 9 14]. The vault unlocks with the controller key instead of the TPM seal. Expected after anygrub.cfgchange.
4. Demo
Four steps, each one showing a distinct guarantee. Run §3 first so you are demoing from a known-good base.
4.1 An ordinary pinned workload avoids the isolated cores
With cpu.pinning.use.isolated = false, deploy a 2-vCPU app with CPU pinning enabled.
Expected: it lands on the dedicated pool — CPUs 2,3 — and never on 4-7.
cat /run/domainmgr/cpuplan.json cat /run/domainmgr/CPUPoolStatus/*.json
{ "uuid": "faa814f0-23d0-4f67-a0d0-ae5b39b3a354", "display_name": "TF-STND-VM-1", "mode": "whole-core-smt", "vcpus": 2, "status": "success", "host_cpus": [2, 3] } {"Pools":[ {"Kind":1,"CPUs":[0,1,4,5,6,7],"FreeCPUs":null,"TotalThreads":6,"AllocatedThreads":6,"FreeWholeCores":0}, {"Kind":2,"CPUs":[2,3],"FreeCPUs":null,"TotalThreads":2,"AllocatedThreads":2,"FreeWholeCores":0}, {"Kind":3,"CPUs":[4,5,6,7],"FreeCPUs":[4,5,6,7],"TotalThreads":4,"AllocatedThreads":0,"FreeWholeCores":2} ]}
Kind 2 (dedicated) becomes [2,3]; Kind 3 (isolated) stays fully free. This is the guarantee: an operator's isolated cores are not spent on a workload that never asked for them.
Note that a legacy pin_cpu app arrives with FullPCPUsOnly:false and ThreadsPerCore:0 — the new API fields unset — and is still placed as whole-core-smt. That reinterpretation is what makes the feature demonstrable from a controller that cannot send the new placement fields.
4.2 Capacity fails closed rather than spilling
Deploy a second 2-vCPU pinned app on the same node.
Expected: it fails rather than taking an isolated core, and the error states how many cores were withheld for isolation — so the operator is not left comparing “insufficient” against CPUs that look idle.
This is the demo that makes the point: the node has 4 free threads at that moment and still refuses, by design.
4.3 The node-wide switch promotes a pinned workload
Set cpu.pinning.use.isolated = true in the controller, then reboot the node.
Expected: the same 2-vCPU pinned app is now placed on 4,5 (or 6,7). A controller that can only send pin_cpu has reached the isolated cores without expressing anything new.
cat /run/domainmgr/cpuplan.json # host_cpus should now be [4,5]
Only whole-core workloads are promoted: kernel isolation means nothing on a core whose sibling the kernel still schedules freely.
4.4 Prove it at the thread level
The pool report is EVE's own bookkeeping. This shows the kernel agrees.
# find the domain ls /run/hypervisor/kvm/ # <app-uuid>.<gen>.<inst> PID=$(pgrep -f '<app-uuid>') # per-thread affinity of the VM's vCPU threads for t in /proc/$PID/task/*; do printf '%s %s\n' "$(cat $t/comm)" "$(grep Cpus_allowed_list $t/status)" done
Captured on the dedicated-pool run of §4.1 (2 vCPUs on core 1):
qemu pid=8752 qemu-system-x86 2-3 qemu-system-x86 2 qemu-system-x86 3 trace-thread 2-3 vhost-8752 2-3 iou-wrk-8752 2-3
The two vCPU threads are pinned 1:1 to host CPUs 2 and 3; every other thread is confined to 2-3. Nothing touches 4-7. The app's cgroup agrees:
cat /sys/fs/cgroup/cpuset/eve-user-apps/<app-uuid>.1.1/cpuset.cpus
2-3
Expected on the isolated run of §4.3: the same shape with 4 and 5 (or 6 and 7) in place of 2 and 3, and the emulator threads confined to the housekeeping CPUs — excluding the isolated ones. That exclusion matters: isolcpus keeps the load balancer off a core but honours an explicit affinity, so the cpuset confining non-pinned workloads has to exclude the isolated CPUs or the kernel will happily run them there.
Inside the guest, the topology EVE advertised should match:
lscpu | grep -E 'CPU\(s\)|Thread\(s\) per core'
A 2-vCPU whole-core-SMT workload sees Thread(s) per core: 2 — a truthful topology, not a flat list of 2 sockets.
4.5 One-shot cross-check
The branch ships a script that cross-checks the allocator's record against real per-thread affinity masks, host SMT siblings and the guest's view, and exits non-zero on any mismatch. It also catches an isolcpus range that is not sibling-complete.
GUEST_PASS=<password> ./cpu-assignment-report.sh <node-ip> <app-ip>
Override its lab defaults (NODE_IP, SSH_KEY=~/.ssh/ztest_key, GUEST_USER=pocuser) and set debug.enable.ssh on the node.
5. Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
/sys/devices/system/cpu/isolated empty after reboot | kernel rejected the argument, or grub.cfg not read | check /proc/cmdline; confirm the file is /config/grub.cfg, not grub.cfg.tmpl |
cpu.topology.unsupported | no core has 2 hardware threads | enable SMT in the BIOS; the CPU may have none |
cpu.policy.odd_vcpu | odd vCPU count on a whole-core request | use an even count |
| Pinned workload placed on dedicated cores, never isolated | cpu.pinning.use.isolated is false, or it was set after the reboot | set the property, then reboot |
| Isolated cores show as free but nothing can use them | expected without the property — they are withheld from every workload that did not ask for isolation | set the property, or use isolation_tier=hard |
| Placement refused with a synthetic-topology error | sysfs topology discovery failed | check /sys/devices/system/cpu/cpu*/topology/ is populated |
Boot loop, fatal: agent zedbox[N]: zboot curpart: err exit status 64 | two disks carry EVE's static PARTUUIDs — install media left in, or a second EVE install | remove the media or wipe the second disk's partition table |
$dom0_extra_args missing from /proc/cmdline | unquoted heredoc expanded it away when writing grub.cfg | rewrite with «'EOF' |
6. Reference
docs/cpu-affinity-design.mdin the feature branch — isolation tiers, placement vocabulary, the EVE↔controller interface. Large parts are marked not yet implemented; it describes the intended end state.docs/CONFIG-PROPERTIES.md— all configuration properties.- Pool kinds and error codes:
pkg/pillar/types/cpuplacement.go. - Placement logic:
pkg/pillar/cpuallocator/placement.go.
