SO_REUSEPORT_LB / receive-CPU affinity / commit 1

Receive-Queue Spread

commit 1: iflib optionally spreads queue CPUs across cache groups companion to commit 2, Receive-CPU Dealing status: proposed; core_spread not separately exercised on HW 2026-09-14, validated (deal) 2026-09-26

Two questions from the dealing design. With only eight queues, is locality improved per CCD? Not on this machine: all eight queues sit on CCD0, so the deal can upgrade the sixteen CCD0 workers from same-L3 to exact or same-core, but the sixteen CCD1 workers stay cross-CCD, exactly as under the plain hash. Can the queues be spread across CCDs? Yes, and it is a placement change in iflib, not a setting: x86 offers one thread per core as an interrupt target and iflib hands queue i the i-th of those in numeric order, so eight queues always take cores 0-7. Once four queues sit on each CCD, the deal keeps every worker inside its own L3.

The change is one tunable, dev.<drv>.<unit>.iflib.core_spread, and one function, cpuid_advance(), which every queue placement goes through. Nothing in the dealing changes; the placement is what gives it something to deal from on both CCDs.

Validated indirectly on the two-socket machine, 2026-09-26. On the two-socket test box the deal was exercised with receive queues already on both sockets - by using two NICs, one bound to each NUMA domain (aq0 on socket 0, aq2 on socket 1), which is the placement core_spread would produce on a single multi-queue card. With queues on both sockets the deal kept each socket's connections local and won 11.3% cyc/req over a plain hash. The core_spread path itself (spreading one NIC's queues across cache groups) is built in but was not separately exercised here; the per-domain NIC placement made it unnecessary for this run. On the 5950X, where all eight queues of one card land on CCD0, it is core_spread that the deal would need.

Where eight queues land today

CCD0: L3 group, CPUs 0-15, cores 0-7 core 0 0q0 1 core 1 2q1 3 core 2 4q2 5 core 3 6q3 7 core 4 8q4 9 core 5 10q5 11 core 6 12q6 13 core 7 14q7 15 CCD1: L3 group, CPUs 16-31, cores 8-15 core 8 16intr 17 core 9 18intr 19 core 10 20intr 21 core 11 22intr 23 core 12 24intr 25 core 13 26intr 27 core 14 28intr 29 core 15 30intr 31 Interrupt CPU set: thread 0 of every core, {0, 2, ..., 30}, sixteen CPUs. Hyperthreads are left out unless machdep.hyperthreading_intr_allowed=1 (mp_x86.c:1199). iflib walks that set in numeric order (cpuid_advance, iflib.c:5116), so queue i takes core i: q0-q7 fill cores 0-7 and every queue sits on CCD0. The RX task and its IRQ are bound at attach through taskqgroup_attach_cpu(); nothing moves them afterwards, and cpuset -x on the IRQ alone does not move the processing. queue bound here eligible, no queue not an interrupt target (second thread)
The placement, not the NIC, decides which CCD receives. The card has eight queues and the machine sixteen eligible cores; the numeric walk fills the first eight and stops. Every connection this host accepts on aq0 is therefore processed on CCD0, whatever the dealing does afterwards.

What the deal fixes, and what only placement can

worker on CPU 0 15 | 16 31 Hash (today's default) every worker fed, no steering 16 same L3 (random core), 16 other CCD Exact claims (current branch) mode 2 or explicit tags 8 exact, 24 never selected Deal, queues on CCD0 set 3, as placed today 8 exact, 8 same core, 16 other CCD (one steady producer each) Deal + spread, 4 queues per CCD this set 8 exact, 8 same core, 16 same L3, none on the other CCD exact same core same L3, other core other CCD never selected Measured on this host: other CCD costs +19% cycles per request; same L3 on another core is about even with exact but takes a reschedule IPI per wakeup; exact and same core wake locally. Rows 1 and 3 have the same CCD split. The deal cannot move the sixteen CCD1 workers onto a receive CPU that does not exist there; only row 4's placement does.
The honest accounting. Compared with the hash, the deal on today's placement buys exact or same-core service for the CCD0 half and one steady producer for the CCD1 half, but the CCD1 half is still paying the cross-CCD penalty on every connection. Spreading the queues is what turns that half into same-L3 service, which is the regime where the August measurement showed the double-digit win.

The spread order

Numeric order (today): cpuid_advance steps through the interrupt set by CPU id 0 2 4 6 8 10 12 14 16 18 20 22 24 26 28 30 CCD0CCD1 positions 0-7 = q0-q7 = cores 0-7, all on CCD0; the shared cursor leaves the next device at position 8, CPU 16 Spread order (core_spread=1): interleave the children of the widest split of the local set 0 16 2 18 4 20 6 22 8 24 10 26 12 28 14 30 CCD0 and CCD1 alternate positions 0-7 = q0-q7 = cores 0-3 of CCD0 and cores 8-11 of CCD1; the next device starts at position 8, CPU 8 The cursor shared between devices advances by positions, as it does today. Devices with a different order (spread flag or local domain) keep their own cursor. Single-L3 parts: the split's children are the cores and the order equals today's. Multi-socket parts: the local domain is spread first, then the remote CPUs follow in numeric order.
One permutation, computed once at attach. Take the lowest scheduler group that contains the device's local share of the interrupt set, interleave its children round-robin with numeric order inside each child, and append anything outside the local domain. Sixteen queues would cover both CCDs under either order; eight queues cover both only under this one.

After the spread: every worker inside its own L3

CCD0: queues q0, q2, q4, q6 core 0 0q0 1sibling core 1 2q2 3sibling core 2 4q4 5sibling core 3 6q6 7sibling core 4 8from 0 9from 0 core 5 10from 2 11from 2 core 6 12from 4 13from 4 core 7 14from 6 15from 6 CCD1: queues q1, q3, q5, q7 core 8 16q1 17sibling core 9 18q3 19sibling core 10 20q5 21sibling core 11 22q7 23sibling core 12 24from 16 25from 16 core 13 26from 18 27from 18 core 14 28from 20 29from 20 core 15 30from 22 31from 22 receive CPU, exact worker (8) same core (8) same L3, dealt in member order (16) other CCD: none Each receive CPU supplies 64 of the 512 slots (equal counts): 16 to itself, 16 to its sibling, then 32 to the first two same-L3 workers still owed, in member order, which is one core's pair. The level order guarantees it: exact for all eight receive CPUs first, then siblings, then the same L3; nobody reaches ring 3, so CCD1 is never crossed. With the queues on one CCD the same loop reached ring 3 for sixteen workers. The placement changed, not the deal.
The deal is unchanged; its input is. Four receive CPUs per CCD means the same-L3 ring already holds enough workers to absorb every remaining slot, so the other-L3 ring is never consulted. Each of the sixteen non-adjacent workers still has one producer, now on its own CCD.

Where the knob sits in iflib

attach iflib_register() bus_get_cpus(INTR_CPUS) ifc_cpus: one thread per core build ifc_cpu_order[] numeric, or spread if core_spread get_ctx_core_offset() cursor keyed (set, spread, domain) get_cpuid_for_queue(qid) per queue, unchanged cpuid_advance(): order[(pos(cpu) + n) % count] was: the n-th numeric successor inside ifc_cpus taskqgroup_attach_cpu() binds the RX task and its IRQ rxq<i>.cpu sysctl read-only report of the result Building the spread order S = ifc_cpus, narrowed to bus_get_cpus(LOCAL_CPUS) when that leaves anything G = lowest cpu_group covering S (descend from cpu_top while one child still covers S) order = round-robin over G's children, numeric inside each; then ifc_cpus minus S, numeric no children, or a kernel without SMP: numeric order, unchanged
Two boxes change, everything else is read through them. The order is built once after the interrupt set is known, and the walk that every placement uses reads it. The shared cursor that lets several devices start where the previous one stopped still counts positions, so it keeps working; it only needs to know that a device with a different order is a different sequence. Teal marks what the commit touches.

How the rule lands on other machines

machine and queuesnumeric order (today)spread order
this host, 5950X, aq with 8 queuescores 0-7, all CCD0cores 0-3 and 8-11, four per CCD
this host, a 16-queue NICcores 0-15, both CCDsthe same sixteen cores in a different order; nothing to gain
single-L3 desktop (one cache group)consecutive coresthe children of the split are the cores: identical to today
EPYC, one NUMA node, 8 CCDs, 8 queuesall eight on CCD0one per CCD
two sockets, NIC local to socket 0, 8 queuessocket 0's first eight coressocket 0's CCDs interleaved; socket 1 only after socket 0 is exhausted
bhyve guest, two sockets configuredsocket 0 firstinterleaved across whatever the guest's top-level split is; check kern.sched.topology_spec
kernel without SMPCPU 0CPU 0

Validation in the guest

runexpect
guest 16 vCPUs as cpu_sockets=2 cpu_cores=4 cpu_threads=2, aq passthrough, override_nrxqs=4, core_spread=0rxq0-3.cpu all in socket 0; with 16 pinned listeners the eight socket-1 listeners' handshake histograms are entirely cross-package
same guest, core_spread=1rxq*.cpu lands 2+2; every listener's histogram sits inside its own package; shares still about 1/16 each
host, after the next planned kernel install, dev.aq.0.iflib.core_spread=1rxq*.cpu = 0, 16, 2, 18, 4, 20, 6, 22; the 32-worker perf harness is where the machine-wide win should appear
iflib is compiled into the kernel, so the host only sees the spread after a kernel install; the guest is where the placement and the dealing are proven together.

What the drawings commit to

  • Default off; the numeric order is byte-for-byte today's behaviour, and on a single-L3 machine the spread order equals it.
  • Spread never leaves the local NUMA domain while it has CPUs to give; remote CPUs are appended, never interleaved.
  • Only the walk changes. core_offset, separate_txrx, use_logical_cores and the shared cursor keep their semantics in positions of the order.
  • Devices with different orders never share a cursor; devices with the same order hand off exactly as today.
  • The deal is untouched. With four receive CPUs per CCD its own level order stops at the same-L3 ring, so no connection crosses a CCD.