FreeBSD networking / receive path

From Packet to Processor

how iflib and RSS steer inbound traffic and where the socket layer drops the locality primer for the receive-CPU affinity series

The hardware works hard to put a flow on one CPU and keep it there: the NIC hashes the four-tuple to a receive queue, and iflib bound that queue to a CPU at attach, so every packet of the flow is processed on the same core. Then, at the last step, SO_REUSEPORT_LB chooses which listening socket receives the connection by a separate software hash that knows nothing about the queue - so the worker that wakes may sit on a different CPU, and the locality the hardware set up is thrown away. This page draws that path, the dispatch fork that decides whether the stack even runs on the receive CPU, and where the receive-CPU affinity work reconnects the two.

One flow, wire to worker, and where locality is lost

NIC: hardware RSS Toeplitz hash H over the 4-tuple, key K; queue q = table[H mod 64] hardware iflib: queue q bound to CPU c chosen once, at attach; every packet of the flow lands on c CPU c IP + TCP input direct dispatch: the stack runs inline on the RX CPU CPU c locality dropped here SO_REUSEPORT_LB: choose a listener a SEPARATE 4-tuple software hash, mod the member count hash, not c worker wakes and runs on whatever CPU it was pinned to - decided by the hash, not the queue CPU w != c The first three steps keep the flow on CPU c. The last two choose the worker by a hash that never sees c, so w is unrelated to c: a cross-CPU wakeup, and on a multi-socket machine often a cross-socket one. Everything RSS set up is spent at the socket layer.
Two hashes, only one of them aware of the machine. The NIC's Toeplitz hash steers to a queue and thus a CPU; SO_REUSEPORT_LB's software hash steers to a listener and thus a worker. They are independent, so the worker that wakes is on the receive CPU only by chance. That gap is what the receive-CPU affinity work closes.

The dispatch fork: does the stack run on the receive CPU at all?

GENERIC (and the affinity work under options RSS) RX ithread on CPU c MSI-X interrupt for queue q IP + TCP inline on CPU c direct dispatch: no hand-off inbound stays on c options RSS, stock RX ithread on CPU c MSI-X interrupt for queue q IP forces hybrid dispatch defers to one unbound netisr worker CPU not fixed: the workstream ID does not pin it
Affinity only means something if inbound runs on a consistent per-flow CPU. Stock GENERIC dispatches IP directly, so the stack runs on the RX ithread's CPU. Under options RSS the IP handlers register NETISR_DISPATCH_HYBRID, which defers to a single unbound netisr worker whose CPU floats - and the receive CPU the socket layer would key on is no longer stable. The affinity work sets those handlers back to NETISR_DISPATCH_DEFAULT so options RSS honours the direct policy and inbound stays on c. That is the precondition the measurements ran under.

What the current work adds

before: independent hash SO_REUSEPORT_LB: 4-tuple hash mod member count, ignores c worker on CPU w != c after: dealt by receive CPU deal picks the member on CPU c learned at accept(2) or set with the option worker on CPU w = c (or nearest) per-NIC RSS controls (rss-ioctl-api), the layer this builds on ifconfig reads and programs the card's hash key K and indirection table through iflib; aq(4) is the reference driver. caveat: setting the per-NIC indirection table does not update the kernel's software RSS map - the two are separate tables.
The deal reconnects the two hashes. Instead of a hash that ignores the receive CPU, the group is dealt by it - each receiving CPU's connections go to the nearest listener - so the worker runs where the flow was processed. The per-NIC RSS controls are the neighbouring piece: they let an operator read and set the card's key and table, which is what decides the queue in the very first step of the pipeline above.

Where this goes

The socket-layer half is drawn end to end in End-to-End Receive Locality, with the deal itself in Receive-CPU Dealing and The Ring Model. The per-NIC RSS controls are in Per-NIC RSS Controls. Measured on a two-socket Ivy Bridge machine (Xeon E5-2470 v2), reconnecting listener selection to the receive CPU cut cycles per request 11-12% when workers span both sockets, and on aq(4) the key and table get/set are proven on both A1 and A2.