SO_REUSEPORT_LB / receive-CPU affinity / commit 2 detail
A receiving CPU spends its connections on the workers nearest to it, and "nearest" is not a number this design invents. It is read from the tree the scheduler already built to describe the machine. Ring 0 is the CPU itself; each ring after that is one step up that tree; and one final ring catches whatever is left so the deal always closes.
On this host that comes to four rings that mean something - the CPU, its SMT sibling, the rest of its L3 group, the rest of the package - plus the catch-all. On a different machine it is a different number: the rings are however deep that machine's tree is, and the code asks rather than assumes.
cg_parent to the root, and the n-th group it passes is ring n. A machine that reports more levels gets more rings; one that reports none gets two.| machine | chain from a leaf | rings that mean something | what the deal does |
|---|---|---|---|
| this host, 5950X | SMT pair, L3 group, root | 0 exact, 1 sibling, 2 same L3, 3 package | the worked example above |
| 2x Xeon E5-2470 v2 (two sockets) | SMT pair, socket L3, root | 0 exact, 1 sibling, 2 same socket, 3 other socket (QPI) | validated 2026-09-26: same four rings, ring 3 crosses QPI; two-DUT deal -11.3% cyc/req vs hash, no worker starved |
| one L3, no SMT | core, root | 0 exact, 1 same L3 | exact first, then anyone in the group |
flat topology (smp_topo_none) | root only | 0 exact, 1 everyone | degrades to exact-or-share, which is the honest answer for a machine that reports no structure |
| two sockets, 8 CCDs | SMT pair, L3, socket, root | 0 to 4, five rings | fills its own CCD, then its own socket, then crosses |
| chains of unequal depth | varies by CPU | as many as that CPU has | a CPU whose chain is shorter than the current level sits that pass out and waits for the catch-all, rather than reaching further than its neighbours |
| kernel without SMP | none | 0 exact, 1 everyone | one CPU, so the question does not arise |
kern.smp.topology override | whatever was asked for | follows it | nearness follows the administrator's description, right or wrong |
in_pcblbgroup_rings() counts the chain for each receiving CPU and the deal runs one pass per level plus the catch-all, so a machine the scheduler describes in more detail gets a finer-grained deal without any change here.Widening rings over a machine's own topology are established practice in schedulers and memory allocators, where they place a thread or a page near the work that needs it. Equal-share slot tables are established practice in load balancers, where they divide traffic evenly and survive a membership change without reshuffling everything.
Neither has been brought to the socket layer's choice of which listener receives a connection. That choice has until now been an exact CPU match or a hash: the first starves every worker not sitting on a receive queue, the second throws locality away. Putting the two together there is new. The rings decide who is nearest, the slot table guarantees the share, and between them a machine can use every worker it has without giving up the locality the hardware set up.
This host (Ryzen 9 5950X, two CCDs on one package - the topology the drawings above use as their example), 2026-08. Here ring 3 is the cross-CCD hop over Infinity Fabric, and its cost is measured, not assumed: against a same-CCD wakeup it is +19% cycles per request. The exact-claim predecessor that the ring deal generalizes was run across three worker layouts - workers all on the receiving CCD gained nothing (the wakeup already shares L3), workers spanning both CCDs gained -15% (affinity vs hash), and forcing every worker to the far CCD cost -19% (affinity vs anti, the pure ring-3 penalty). The ring deal itself was not exercised here; it needs the two-socket receive topology below. These are the numbers that make ring 3 the one expensive step in the drawings.
Two-socket 2x Xeon E5-2470 v2 (Ivy Bridge-EP), two sockets over QPI, 2026-09-26. The ring chain from any leaf CPU is the SMT pair, then the socket's L3 group, then the root, so the deal has the same four rings drawn above, with ring 3 the cross-socket step. Booted the affinity kernel with direct netisr dispatch; cycles per request from a process-scoped PMC over the worker threads only.
The deal delivers when receive spans both sockets. One receive-queue set on each socket, one worker per receive CPU across both (16 workers): plain hash scatters about half the connections to the far socket, while the ring deal keeps each socket's connections on its own socket and still uses every worker. Cross-CPU wakeups fell from ~292k to ~15k and cycles per request dropped 11.3% against the hash - matching the exact-claim design's 12.1% while starving no worker and needing no application change (the learning is on by default; the workers only pin themselves).
Same-socket, the mechanism co-locates ~96% (workers one per receive CPU on one socket): cross-CPU wakeups ~266k to ~8k, learned automatically at accept(2). Cycles per request are flat there, because same-socket wakeups already share the L3 - the ring-3 hop is the only expensive step, exactly as drawn.
One caveat the runs surfaced: supply is learned in two ten-second windows, so a benchmark shorter than that under-converges and the deal stays partly on the hash. Sustained traffic converges it; short bursty runs show variance. Worth weighing where the window length meets real workloads.