SO_REUSEPORT_LB / receive-CPU affinity / commit 1
Two questions from the dealing design. With only eight queues, is locality improved per CCD? Not on this machine: all eight queues sit on CCD0, so the deal can upgrade the sixteen CCD0 workers from same-L3 to exact or same-core, but the sixteen CCD1 workers stay cross-CCD, exactly as under the plain hash. Can the queues be spread across CCDs? Yes, and it is a placement change in iflib, not a setting: x86 offers one thread per core as an interrupt target and iflib hands queue i the i-th of those in numeric order, so eight queues always take cores 0-7. Once four queues sit on each CCD, the deal keeps every worker inside its own L3.
The change is one tunable, dev.<drv>.<unit>.iflib.core_spread, and one function, cpuid_advance(), which every queue placement goes through. Nothing in the dealing changes; the placement is what gives it something to deal from on both CCDs.
Validated indirectly on the two-socket machine, 2026-09-26. On the two-socket test box the deal was exercised with receive queues already on both sockets - by using two NICs, one bound to each NUMA domain (aq0 on socket 0, aq2 on socket 1), which is the placement core_spread would produce on a single multi-queue card. With queues on both sockets the deal kept each socket's connections local and won 11.3% cyc/req over a plain hash. The core_spread path itself (spreading one NIC's queues across cache groups) is built in but was not separately exercised here; the per-domain NIC placement made it unnecessary for this run. On the 5950X, where all eight queues of one card land on CCD0, it is core_spread that the deal would need.
| machine and queues | numeric order (today) | spread order |
|---|---|---|
| this host, 5950X, aq with 8 queues | cores 0-7, all CCD0 | cores 0-3 and 8-11, four per CCD |
| this host, a 16-queue NIC | cores 0-15, both CCDs | the same sixteen cores in a different order; nothing to gain |
| single-L3 desktop (one cache group) | consecutive cores | the children of the split are the cores: identical to today |
| EPYC, one NUMA node, 8 CCDs, 8 queues | all eight on CCD0 | one per CCD |
| two sockets, NIC local to socket 0, 8 queues | socket 0's first eight cores | socket 0's CCDs interleaved; socket 1 only after socket 0 is exhausted |
| bhyve guest, two sockets configured | socket 0 first | interleaved across whatever the guest's top-level split is; check kern.sched.topology_spec |
| kernel without SMP | CPU 0 | CPU 0 |
| run | expect |
|---|---|
guest 16 vCPUs as cpu_sockets=2 cpu_cores=4 cpu_threads=2, aq passthrough, override_nrxqs=4, core_spread=0 | rxq0-3.cpu all in socket 0; with 16 pinned listeners the eight socket-1 listeners' handshake histograms are entirely cross-package |
same guest, core_spread=1 | rxq*.cpu lands 2+2; every listener's histogram sits inside its own package; shares still about 1/16 each |
host, after the next planned kernel install, dev.aq.0.iflib.core_spread=1 | rxq*.cpu = 0, 16, 2, 18, 4, 20, 6, 22; the 32-worker perf harness is where the machine-wide win should appear |
core_offset, separate_txrx, use_logical_cores and the shared cursor keep their semantics in positions of the order.