SO_REUSEPORT_LB / receive-CPU affinity / overview
A connection's locality is decided twice. The NIC and iflib decide, at attach, which CPU every packet of a flow will be processed on; the socket layer decides, once at the SYN, which worker gets the connection. The exact-claim design only ever touched the second decision and so could not use more workers than receive queues. This page draws both decisions and what we learned by measuring them: where packets land is a placement policy, where connections go is a dealing policy, and the measured cost that matters is crossing an L3 group.
Three commits follow from that. iflib gains an optional spread of queue CPUs across cache groups. The load-balancing group gains a dealt map that serves every worker from the nearest receiving CPU with an equal share. accept(2) learns both which CPUs receive and where each worker runs, so unmodified servers get all of it.
Measured 2026. On this host (Ryzen 9 5950X, two CCDs on one package) the expensive hop is cross-CCD: it costs +19% cycles per request, and the exact-claim predecessor gained up to -15% with workers spanning both CCDs (affinity vs hash). On the two-socket Xeon E5-2470 v2 (two sockets over QPI) the ring deal itself was run: with receive spanning both sockets and one worker per receive CPU across both, it cut cross-CPU wakeups from ~292k to ~15k and cycles per request by 11.3% against a plain hash - the same win as exact-claim while starving no worker and needing no application change. Same-socket, the automatic accept(2) learning co-located ~96% of connections. Caveat: supply is learned in two ten-second windows, so runs shorter than that under-converge. Ring detail: the ring model.
| commit | touches | delivers | proven by |
|---|---|---|---|
1. iflib: optionally spread queue CPUs across cache groups | sys/net/iflib.c, iflib.4: an order array built at attach, cpuid_advance() walking it, the cursor key, tunable core_spread | queues on every L3 group of the local domain; default off, numeric order unchanged | guest with two sockets and 4 queues: rxq*.cpu lands 2+2; host after the next kernel install: 0, 16, 2, 18, 4, 20, 6, 22 |
2. netinet: add SO_REUSEPORT_LB_CPU, dealing members by receive CPU | in_pcb_var.h, in_pcb.h, in_pcb.c, tcp_syncache.c, option plumbing, getsockopt.2, ATF tests | the dealt map, the selection, explicit tags standing in as receive CPUs until something is learned | loopback ATF: steering, duplicates, untagged share, partial claims; all fail on the pristine kernel |
3. netinet: learn SO_REUSEPORT_LB receive and worker CPUs at accept(2) | protosw.h, uipc_syscalls.c, tcp_usrreq.c, in_pcb.c, tcp.4, ATF tests | counts, windows, capture, the 0/1 sysctl; unmodified servers get everything | aq passthrough: 8 exact, 4 queues on SMT pairs, 6 on 4, 1 queue; sysctl 0 control |
| situation | what happens |
|---|---|
| no member is tagged | no map is built; selection and cost are identical to today |
| the NIC reports no hash (bhyve e1000, vtnet), workers pinned and captured | nothing is learned, no map: the hash distributes, nobody starves |
| explicit tags, nothing learned (loopback tests, UDP) | the tagged CPUs stand in with equal supply: exact for them, other members dealt from them |
| something learned, an explicit tag names a CPU that never receives | that member is dealt from the nearest receiving CPU; a claim cannot starve anyone |
| a packet arrives on a CPU with no slots (loopback, a stray netisr CPU) | hash over all members |
| two members tag the same CPU | both filled from it when its supply covers them; otherwise the later one takes its quota from the next ring |
| one receive queue, thirty-two workers | all thirty-two are dealt from that CPU: the same distribution as the hash |
| queues or NICs receive at different rates | supply follows the counts; a trickle CPU supplies a trickle and its workers get the rest elsewhere |
| all queues on one L3, workers on every CPU | the far workers are dealt from the nearest L3 that receives; only placement (commit 1) can do better |
| a member joins, leaves or is retagged; the map allocation fails | remap in the same critical section, old map freed after an epoch; on failure the map is NULL, the hash takes over, the next trigger retries |
| an unbound netisr thread carries the traffic (deferred dispatch, options RSS without the RSS-kernel dispatch fix) | counts follow the thread; after it migrates, flows on the new CPU hash until it is counted, then at most two windows of skew; the precondition of a per-flow-stable CPU is unchanged |
| single-L3 machine, or a kernel without SMP | the spread order equals the numeric order; the rings degrade to exact / everyone |
| two sockets, NIC local to one | the spread interleaves the local socket's cache groups and never crosses to the other while local CPUs remain |