SO_REUSEPORT_LB / receive-CPU affinity / design history
Exact Claim to Ring Deal
consolidates the earlier exact-claim designs (superseded)proposed design: the ring dealvalidated 2026-09-26 on a two-socket Xeon E5-2470 v2 and a Ryzen 9 5950X
Three designs in turn tried to give a SO_REUSEPORT_LB group the receive locality that RSS sets up in hardware. This page keeps the parts of the first two that explain why the third exists, and retires their separate pages. The exact-claim design added a per-connection receive-CPU match - an exact claim on one CPU - and could not use more workers than receive queues. The first-accept design kept the exact claim but learned the CPU at accept(2), so unmodified servers benefited; it still starved any worker off a receive queue. The ring deal, the design we'd propose, prototyped and measured, replaces the exact claim with a nearest-CPU deal: every worker is owed an equal share, spent on the CPUs nearest it, so a machine uses every worker without giving up locality.
The figures below are the ones worth keeping from the first two: why the exact claim starved workers, the selection schemes it forced a choice between, and the duplicate-claim problem it could never solve cleanly. The last section is where the ring deal leaves all of it, and what it measured.
Three designs, and what each could not do
The deal is the two earlier ideas kept and the one flaw dropped. Learning the CPU at accept, so an unmodified server benefits, was the good idea of the middle design and survives unchanged. The exact claim - one member per CPU, everyone else starved - was the flaw of both, and the deal replaces it.
Why the exact claim starved workers
Panel B is the case that sank the exact claim. The recommended nginx configuration on this host gives thirty-two pinned workers against eight receive queues, so a rule that tags any pinned acceptor hands every connection to the eight workers on receive CPUs and leaves the rest idle. The first-accept design answered it with a coverage gate - fall back to the hash until every member is decided - which recovers today's distribution but wins nothing for the twenty-four. The ring deal is the answer that actually feeds them; see the last section.
The selection schemes the exact claim forced a choice between
Row 2 is the flaw in the shortcut, row 3 is the fix for the scan - and both are still only an exact claim. Indexing the member array by CPU assumes bind order equals CPU order and that receive CPUs are dense from zero; on this rig neither holds, and odd-numbered members are never chosen. The per-CPU index table fixes that in one load. But even the fixed version answers only "which single member sits on this CPU" - the question the ring deal discards in favour of "which members are nearest, and what does each one owe."
Duplicate claims: the problem the deal dissolves
The exact claim could not even share one CPU among several workers. A first-match scan hands all of CPU 3's flows to one member and starves the rest; the best the exact claim managed was Linux's start-at-hash walk, which shares but not equally. The ring deal removes the question: several claimants on one CPU are just several members owed a share, dealt equally at ring 0.
Where the ring deal leaves all of it
The nearest-CPU deal answers every limitation above at once. Each receiving CPU is given a share of the group's connections and spends it on the members nearest it - itself, then its SMT sibling, then its cache group, then outward - and every member is owed an equal number of slots, so no worker starves however few CPUs receive. The coverage-gate case (panel B) is resolved by feeding the far workers from the nearest receiving CPU instead of idling them. The selection schemes collapse to a single map read of two loads. The duplicate-claim walk is gone: several claimants on one CPU are just members owed a share, dealt at ring 0.
Measured 2026-09-26. On the two-socket Xeon E5-2470 v2 (two sockets over QPI), with receive spanning both sockets and sixteen workers one per receive CPU, the deal cut cycles per request 11.3% against a plain hash while using every worker - matching the exact claim's 12.1% (which had to starve half the workers to get it) with no application change. On one socket the automatic accept(2) learning co-located ~96% of connections. On the 5950X (two CCDs) the cross-cache-group penalty the deal exists to avoid is +19%. The mechanics of the deal are drawn in the ring model; the whole path from wire to worker is in end-to-end receive locality.