User Tools

Site Tools


eve-kvm:core-isolation

This is an old revision of the document!


EVE CPU Core Pinning and Isolation

How to enable kernel CPU isolation on an EVE node, prove it took effect, and demo it.

Scope. Applies to the core-pinning feature branch (eve PR #6335, eve-api #155, adam #158, NFR-172). Verified on EVE 0.0.0-core-pinning-gmwtus-551edfcb-kvm-amd64.

Provenance. The transcripts in §3, §4.1 and §4.4 were captured on a live node on 2026-09-18. The outputs in §4.2 and §4.3 are expected values, not yet captured.


1. Requirements

1.1 CPU

Pinning on this branch means whole-core SMT: each vCPU gets one hardware thread of a physical core, and both siblings go to the same workload. The allocator will only use a core that has exactly two hardware threads.

  • SMT / Hyper-Threading is mandatory and must be enabled in the BIOS. A node with no sibling threads fails closed with cpu.topology.unsupported — there is no fallback to thread-granular placement.
  • At least 3 physical cores to be useful: one for EVE housekeeping, one for the dedicated pool, one or more to isolate.
  • Even vCPU counts only. An odd count fails closed with cpu.policy.odd_vcpu.
  • Homogeneous cores are strongly preferred. On Intel hybrid parts the E-cores have no SMT and are silently skipped for pinning.
CPU Topology Usable
Core i3-10100 / 10105 / 10300 4C/8T yes
Core i3-12100 / 12300, i3-13100, i3-14100 4 P-cores, 0 E-cores, 8T yes — ideal
Core i5-12400 6 P-cores, 0 E-cores, 12T yes — more headroom
Core i3-8100 / 8300, i3-9100 / 9300 4C/4T, no HT no
Core i3-N300 / i3-N305 (Alder Lake-N) 8 E-cores, no HT no
Core Ultra 200S / Meteor Lake / Lunar Lake SMT removed no

Reference node for this page: Intel Core i3-13100TE, 4 P-cores / 8 threads, no E-cores, 1 NUMA node, 12 MiB shared L3, VT-x.

1.2 Check the node before you start

lscpu | grep -E 'Model name|Thread\(s\) per core|Core\(s\) per socket|NUMA node\(s\)'

Thread(s) per core: 2 is the requirement. A 1 here means either the CPU has no SMT or it is disabled in the BIOS; both fail the same way.


2. Enable core isolation

2.1 Determine the SMT sibling enumeration

This decides your CPU indexes. Do not assume it.

cat /sys/devices/system/cpu/cpu[0-9]*/topology/thread_siblings_list | sort -u

Two possible answers on a 4C/8T part:

Output Meaning Core map
0-1 2-3 4-5 6-7 adjacent siblings (Intel 12th gen and newer) core0={0,1} core1={2,3} core2={4,5} core3={6,7}
0,4 1,5 2,6 3,7 split siblings (older Intel) core0={0,4} core1={1,5} core2={2,6} core3={3,7}

2.2 Decide the split

Three roles. The isolated set must be sibling-complete — a core with one thread isolated and the other in housekeeping serves neither ordinary requests nor the isolated pool, and simply disappears from usable capacity.

Role Cores CPUs (adjacent enumeration)
EVE housekeeping core 0 0,1
Dedicated pool (ordinary pinned workloads) core 1 2,3
Isolated pool cores 2-3 4,5,6,7

2.3 Set the controller property FIRST

In the controller, set the edge-node configuration property:

cpu.pinning.use.isolated = true

Order matters. Set this before the reboot that applies isolcpus. The switch only affects workloads placed after it is set, so setting it after the reboot silently does nothing until the next placement. Setting it on a node whose kernel isolates nothing warns rather than errors.

Without this property, kernel-isolated CPUs are reported in the pool report and usable by nobody: they are withheld from every workload that did not explicitly ask for isolation, and only isolation_tier=hard can claim them — which current controllers cannot express.

2.4 Write /config/grub.cfg

/config/grub.cfg does not exist on a fresh install. The image ships only grub.cfg.tmpl, and GRUB reads grub.cfg exactly — so create the file, do not rename the template.

eve config mount /tmp/cfg
cat > /tmp/cfg/grub.cfg <<'EOF'
set_getty
set_global hv_eve_cpu_settings "eve_max_vcpus=2"
set_global dom0_extra_args "$dom0_extra_args isolcpus=managed_irq,domain,4,5,6,7 rcu_nocbs=4,5,6,7 nohz_full=4,5,6,7 irqaffinity=0,1"
EOF
cat /tmp/cfg/grub.cfg
sync
eve config unmount

For the split enumeration use isolcpus=managed_irq,domain,2,3,6,7 rcu_nocbs=2,3,6,7 nohz_full=2,3,6,7 irqaffinity=0,4 and omit the eve_max_vcpus line: reservation is by lowest CPU id and a core is dropped if any sibling is reserved, so eve_max_vcpus=2 there would reserve CPUs 0 and 1, which belong to two different cores, and cost a second core for nothing.

Notes:

  • The quoted heredoc («'EOF') is required. Unquoted, the shell expands $dom0_extra_args to nothing and GRUB loses everything it had accumulated.
  • irqaffinity names the housekeeping CPUs, not the isolated ones. Pointing it at the isolated set steers interrupts onto the cores you are trying to shield.
  • set_getty keeps a console shell available; it is the only content of the shipped template, so without it you lose that.
  • Do not use GRUB's set_isolcpus helper (the interactive menu entry isolate CPU0 (only for PREEMPT_RT)). It emits isolcpus=inverse,0, which with the default eve_max_vcpus=1 isolates CPU 0's own SMT sibling and produces a sibling-incomplete core.
  • /config/grub.cfg is sourced after EVE's defaults and before the boot entry is built, so set_global overrides the shipped eve_max_vcpus=1.

2.5 Reboot

reboot

Kernel isolation is a command-line decision, so it is reboot-gated. Expect roughly 3 minutes before SSH answers again.


3. Validate

Run these in order. Stop at the first one that disagrees.

3.1 The kernel accepted the arguments

tr ' ' '\n' < /proc/cmdline | grep -E 'isolcpus|nohz_full|rcu_nocbs|irqaffinity|eve_max_vcpus'
eve_max_vcpus=2
isolcpus=managed_irq,domain,4,5,6,7
rcu_nocbs=4,5,6,7
nohz_full=4,5,6,7
irqaffinity=0,1

3.2 The kernel is actually isolating

cat /sys/devices/system/cpu/isolated
4-7

An empty result means there is no isolated pool and nothing downstream will work. Stop and fix the command line.

3.3 domainmgr saw the topology and the isolated set

logread | grep -E 'CPU topology|Kernel isolates'
msg":"CPU topology: 4 physical cores"
msg":"Kernel isolates CPUs [4 5 6 7]; dedicated workloads are placed there first and everything else is kept off them"

CPU topology: N physical cores must appear. If topology discovery failed, the model is synthetic and whole-core placement is refused outright rather than performed against a fabricated topology.

Note: the second line is printed whenever the isolated set is non-empty, regardless of cpu.pinning.use.isolated. It does not prove the promotion switch is on — check the property itself (§3.5).

3.4 The pool report

cat /run/domainmgr/CPUPoolStatus/*.json

Pool Kind values: 1 = housekeeping, 2 = dedicated, 3 = isolated.

{"Pools":[
 {"Kind":1,"CPUs":[0,1,2,3,4,5,6,7],"FreeCPUs":[2,3],"TotalThreads":8,"AllocatedThreads":6,"FreeThreads":2,"TotalCores":4,"FreeWholeCores":1},
 {"Kind":2,"CPUs":null,"FreeCPUs":null,"TotalThreads":0,"AllocatedThreads":0,"FreeThreads":0,"TotalCores":0,"FreeWholeCores":0},
 {"Kind":3,"CPUs":[4,5,6,7],"FreeCPUs":[4,5,6,7],"TotalThreads":4,"AllocatedThreads":0,"FreeThreads":4,"TotalCores":2,"FreeWholeCores":2}
]}

What to read from it:

  • Kind 3 is non-empty — the isolated set is now allocatable capacity, not just a number in the device info.
  • Kind 1 FreeCPUs is [2,3] only — of 8 threads, 6 are allocated: CPUs 0-1 reserved for EVE and CPUs 4-7 withheld for isolation. This is the point of the feature: housekeeping freeness answers “will an ordinary workload fit?”, which isolated CPUs cannot serve.
  • Kind 1 FreeWholeCores is 1 — exactly core 1.

3.5 The promotion switch

grep -o 'cpu.pinning.use.isolated[^}]*}' /run/zedagent/ConfigItemValueMap/global.json
cpu.pinning.use.isolated":{"Key":"cpu.pinning.use.isolated","ItemType":2,"IntValue":0,"StrValue":"","BoolValue":true,"TriStateValue":0}

BoolValue must be true. This value comes from the controller only. A value written locally into /config/GlobalConfig/global.json survives just until the first config poll, because zedagent rebuilds the whole map from defaults plus the controller's items.

3.6 Derived placement plan

cat /run/domainmgr/cpuplan.json
{
  "workloads": [
    { "uuid": "...", "display_name": "TF-STND-VM-1", "mode": "whole-core-smt",
      "vcpus": 2, "status": "success", "host_cpus": [4, 5] }
  ]
}

Diagnostic only — nothing reads it back — but it is the quickest way to see what the allocator decided, including for workloads that are configured but not running.

3.7 Known limitations on the shipped kernel

Argument Effect Why
isolcpus works CONFIG_CPU_ISOLATION=y
irqaffinity works accepted by the kernel
nohz_full no effect # CONFIG_NO_HZ_FULL is not set
rcu_nocbs no effect CONFIG_RCU_NOCB_CPU not set
dmesg | grep -iE 'nohz|Unknown kernel command line'
(zcat /proc/config.gz || cat /proc/config) | grep -E 'CONFIG_NO_HZ_FULL|CONFIG_RCU_NOCB_CPU|CONFIG_CPU_ISOLATION'
Housekeeping: nohz unsupported. Build with CONFIG_NO_HZ_FULL
Unknown kernel command line parameters "... rcu_nocbs=4,5,6,7 nohz_full=4,5,6,7", will be passed to user space
# CONFIG_NO_HZ_FULL is not set
CONFIG_CPU_ISOLATION=y

So today you get the isolcpus half of hard isolation — the scheduler's load balancer is kept off the isolated cores — but not tick shedding or RCU callback offload. Getting those requires CONFIG_NO_HZ_FULL and CONFIG_RCU_NOCB_CPU in eve-kernel. Leave the arguments in place; they are harmless and become effective as soon as the kernel supports them.

Two more expected observations:

  • /sys/fs/cgroup/cpuset/eve/cpuset.cpus stays 0 even with eve_max_vcpus=2. That cpuset is driven by dom0_max_vcpus (still 1); eve_max_vcpus only sets the eve/services child, which cannot exceed its parent. No workload capacity is lost — core 0 is dropped from placement either way — but CPU 1 sits reserved and unused. To give EVE its whole core, also set set_global hv_dom0_cpu_settings “dom0_max_vcpus=2 dom0_vcpus_pin”.
  • The vault reports a PCR mismatch. Changing the kernel command line changes PCRs 8, 9 and 14, so eve diag shows vault: ENABLED unlock:controller-key mismatchPCRs:[8 9 14]. The vault unlocks with the controller key instead of the TPM seal. Expected after any grub.cfg change.

4. Demo

Four steps, each one showing a distinct guarantee. Run §3 first so you are demoing from a known-good base.

4.1 An ordinary pinned workload avoids the isolated cores

With cpu.pinning.use.isolated = false, deploy a 2-vCPU app with CPU pinning enabled.

Expected: it lands on the dedicated pool — CPUs 2,3 — and never on 4-7.

cat /run/domainmgr/cpuplan.json
cat /run/domainmgr/CPUPoolStatus/*.json
{ "uuid": "faa814f0-23d0-4f67-a0d0-ae5b39b3a354", "display_name": "TF-STND-VM-1",
  "mode": "whole-core-smt", "vcpus": 2, "status": "success", "host_cpus": [2, 3] }
 
{"Pools":[
 {"Kind":1,"CPUs":[0,1,4,5,6,7],"FreeCPUs":null,"TotalThreads":6,"AllocatedThreads":6,"FreeWholeCores":0},
 {"Kind":2,"CPUs":[2,3],"FreeCPUs":null,"TotalThreads":2,"AllocatedThreads":2,"FreeWholeCores":0},
 {"Kind":3,"CPUs":[4,5,6,7],"FreeCPUs":[4,5,6,7],"TotalThreads":4,"AllocatedThreads":0,"FreeWholeCores":2}
]}

Kind 2 (dedicated) becomes [2,3]; Kind 3 (isolated) stays fully free. This is the guarantee: an operator's isolated cores are not spent on a workload that never asked for them.

Note that a legacy pin_cpu app arrives with FullPCPUsOnly:false and ThreadsPerCore:0 — the new API fields unset — and is still placed as whole-core-smt. That reinterpretation is what makes the feature demonstrable from a controller that cannot send the new placement fields.

4.2 Capacity fails closed rather than spilling

Deploy a second 2-vCPU pinned app on the same node.

Expected: it fails rather than taking an isolated core, and the error states how many cores were withheld for isolation — so the operator is not left comparing “insufficient” against CPUs that look idle.

This is the demo that makes the point: the node has 4 free threads at that moment and still refuses, by design.

4.3 The node-wide switch promotes a pinned workload

Set cpu.pinning.use.isolated = true in the controller, then reboot the node.

Expected: the same 2-vCPU pinned app is now placed on 4,5 (or 6,7). A controller that can only send pin_cpu has reached the isolated cores without expressing anything new.

cat /run/domainmgr/cpuplan.json     # host_cpus should now be [4,5]

Only whole-core workloads are promoted: kernel isolation means nothing on a core whose sibling the kernel still schedules freely.

4.4 Prove it at the thread level

The pool report is EVE's own bookkeeping. This shows the kernel agrees.

# find the domain
ls /run/hypervisor/kvm/                       # <app-uuid>.<gen>.<inst>
PID=$(pgrep -f '<app-uuid>')
 
# per-thread affinity of the VM's vCPU threads
for t in /proc/$PID/task/*; do
  printf '%s  %s\n' "$(cat $t/comm)" "$(grep Cpus_allowed_list $t/status)"
done

Captured on the dedicated-pool run of §4.1 (2 vCPUs on core 1):

qemu pid=8752
qemu-system-x86    2-3
qemu-system-x86    2
qemu-system-x86    3
trace-thread       2-3
vhost-8752         2-3
iou-wrk-8752       2-3

The two vCPU threads are pinned 1:1 to host CPUs 2 and 3; every other thread is confined to 2-3. Nothing touches 4-7. The app's cgroup agrees:

cat /sys/fs/cgroup/cpuset/eve-user-apps/<app-uuid>.1.1/cpuset.cpus
2-3

Expected on the isolated run of §4.3: the same shape with 4 and 5 (or 6 and 7) in place of 2 and 3, and the emulator threads confined to the housekeeping CPUs — excluding the isolated ones. That exclusion matters: isolcpus keeps the load balancer off a core but honours an explicit affinity, so the cpuset confining non-pinned workloads has to exclude the isolated CPUs or the kernel will happily run them there.

Inside the guest, the topology EVE advertised should match:

lscpu | grep -E 'CPU\(s\)|Thread\(s\) per core'

A 2-vCPU whole-core-SMT workload sees Thread(s) per core: 2 — a truthful topology, not a flat list of 2 sockets.

4.5 One-shot cross-check

The branch ships a script that cross-checks the allocator's record against real per-thread affinity masks, host SMT siblings and the guest's view, and exits non-zero on any mismatch. It also catches an isolcpus range that is not sibling-complete.

GUEST_PASS=<password> ./cpu-assignment-report.sh <node-ip> <app-ip>

Override its lab defaults (NODE_IP, SSH_KEY=~/.ssh/ztest_key, GUEST_USER=pocuser) and set debug.enable.ssh on the node.


5. Troubleshooting

Symptom Cause Fix
/sys/devices/system/cpu/isolated empty after reboot kernel rejected the argument, or grub.cfg not read check /proc/cmdline; confirm the file is /config/grub.cfg, not grub.cfg.tmpl
cpu.topology.unsupported no core has 2 hardware threads enable SMT in the BIOS; the CPU may have none
cpu.policy.odd_vcpu odd vCPU count on a whole-core request use an even count
Pinned workload placed on dedicated cores, never isolated cpu.pinning.use.isolated is false, or it was set after the reboot set the property, then reboot
Isolated cores show as free but nothing can use them expected without the property — they are withheld from every workload that did not ask for isolation set the property, or use isolation_tier=hard
Placement refused with a synthetic-topology error sysfs topology discovery failed check /sys/devices/system/cpu/cpu*/topology/ is populated
Boot loop, fatal: agent zedbox[N]: zboot curpart: err exit status 64 two disks carry EVE's static PARTUUIDs — install media left in, or a second EVE install remove the media or wipe the second disk's partition table
$dom0_extra_args missing from /proc/cmdline unquoted heredoc expanded it away when writing grub.cfg rewrite with «'EOF'

6. Reference

  • docs/cpu-affinity-design.md in the feature branch — isolation tiers, placement vocabulary, the EVE↔controller interface. Large parts are marked not yet implemented; it describes the intended end state.
  • docs/CONFIG-PROPERTIES.md — all configuration properties.
  • Pool kinds and error codes: pkg/pillar/types/cpuplacement.go.
  • Placement logic: pkg/pillar/cpuallocator/placement.go.
eve-kvm/core-isolation.1789750033.txt.gz · Last modified: by mc