FreeBSD networking / receive path
The hardware works hard to put a flow on one CPU and keep it there: the NIC hashes the four-tuple to a receive queue, and iflib bound that queue to a CPU at attach, so every packet of the flow is processed on the same core. Then, at the last step, SO_REUSEPORT_LB chooses which listening socket receives the connection by a separate software hash that knows nothing about the queue - so the worker that wakes may sit on a different CPU, and the locality the hardware set up is thrown away. This page draws that path, the dispatch fork that decides whether the stack even runs on the receive CPU, and where the receive-CPU affinity work reconnects the two.
SO_REUSEPORT_LB's software hash steers to a listener and thus a worker. They are independent, so the worker that wakes is on the receive CPU only by chance. That gap is what the receive-CPU affinity work closes.options RSS the IP handlers register NETISR_DISPATCH_HYBRID, which defers to a single unbound netisr worker whose CPU floats - and the receive CPU the socket layer would key on is no longer stable. The affinity work sets those handlers back to NETISR_DISPATCH_DEFAULT so options RSS honours the direct policy and inbound stays on c. That is the precondition the measurements ran under.The socket-layer half is drawn end to end in End-to-End Receive Locality, with the deal itself in Receive-CPU Dealing and The Ring Model. The per-NIC RSS controls are in Per-NIC RSS Controls. Measured on a two-socket Ivy Bridge machine (Xeon E5-2470 v2), reconnecting listener selection to the receive CPU cut cycles per request 11-12% when workers span both sockets, and on aq(4) the key and table get/set are proven on both A1 and A2.