SO_REUSEPORT_LB / receive-CPU affinity / commit 2 detail

The Ring Model

detail for commit 2, in_pcblbgroup_remap() hardware-validated on a two-socket Ivy Bridge machine (Xeon E5-2470 v2) 2026-09-14, validated 2026-09-26

A receiving CPU spends its connections on the workers nearest to it, and "nearest" is not a number this design invents. It is read from the tree the scheduler already built to describe the machine. Ring 0 is the CPU itself; each ring after that is one step up that tree; and one final ring catches whatever is left so the deal always closes.

On this host that comes to four rings that mean something - the CPU, its SMT sibling, the rest of its L3 group, the rest of the package - plus the catch-all. On a different machine it is a different number: the rings are however deep that machine's tree is, and the code asks rather than assumes.

Where a ring comes from

The chain the code walks, for CPU 2 smp_topo_find(cpu_top, 2) the leaf group holding CPU 2 L3 group CPUs 0-15 root, the package CPUs 0-31 NULL chain ends What each step becomes ring 0tag == cpu ring 1leaf mask, {2,3} ring 2L3 mask, {0-15} ring 3root mask, {0-31} ring 4anyone still owed 4 rings + catch-all The two functions that read it in_pcblbgroup_rings(cpu) how many groups the chain has: 3 here, so the catch-all is level 4 in_pcblbgroup_ring(cpu, n) the n-th group's cg_mask, or NULL when the chain is shorter than n Each arrow is one cg_parent step. A member is in ring n of cpu when its tag is set in that mask. Ring 0 is tested against the CPU id itself rather than a mask, because a member that named this very CPU has to be settled before anything sharing its core can take the slots. Both are #ifdef SMP: cpu_top and smp_topo_find exist only there. Without SMP the chain is empty, and only ring 0 and the catch-all remain.
The rings are the scheduler's group chain, numbered. There is no table of cache levels in this code and no assumption about how many there are. It walks from the leaf group containing the CPU up through cg_parent to the root, and the n-th group it passes is ring n. A machine that reports more levels gets more rings; one that reports none gets two.

The four rings from CPU 2, drawn to what they cost

CPU 2 ring 0 ring 1 ring 2 ring 3 catch-all ring 0 the CPU itself members whose tag is 2, on this host 1 CPU wakes locally, the cheapest there is ring 1 its SMT sibling the leaf group {2, 3}, 1 further CPU shares L1 and L2, about the same cost ring 2 the rest of its L3 group CCD0, {0-15}, 14 further CPUs shares L3; costs one reschedule IPI per wakeup ring 3 the rest of the package CCD1, {16-31}, 16 further CPUs crosses the fabric: +19% cycles per request catch-all anyone still owed slots includes members that named no CPU at all exists so the deal always closes, not for locality Ring 3's mask is the root, so it already covers every tagged member; on a machine with more levels between L3 and the package there would be more rings between these two.
Each ring is strictly larger than the last, and the cost steps once. Rings 0 and 1 are free relative to each other. Ring 2 keeps the data in the same L3 but takes an inter-processor interrupt to wake the worker. Ring 3 is the only expensive step on this machine, which is why the placement change exists: with queues on both CCDs, no receive CPU ever has to reach ring 3.

A deal, traced ring by ring

One L3 group, 4 cores x 2 threads. Receive CPUs 0 and 2. Six workers, tagged 0 to 5. 96 slots: 48 per receive CPU, 16 owed to each worker. supply left deficit left, worker 0 to 5 what this pass did start 48 48 16 16 16 16 16 16 nothing dealt yet level 0 exact 32 32 016016 16 16 cpu0 to worker 0 (16), cpu2 to worker 2 (16) level 1 sibling 16 16 000016 16 cpu0 to worker 1 (16), cpu2 to worker 3 (16) level 2 same L3 0 0 000000 cpu0 to worker 4 (16), cpu2 to worker 5 (16) level 3 catch-all 0 0 0 0 0 0 0 0 nothing to do: both sides already zero Every worker ends with 16 of the 96 slots, one sixth of the connections, and no worker was served from further away than it had to be. This machine has two rings in its chain (leaf and root), so the catch-all is level 3. The L3 group is the root here, which is why level 2 reads "same L3". The loop that produced it for level in 0 .. nlevels: outermost, so no ring is entered before every CPU has finished the one before for each receiving CPU with supply: in CPU id order for each member still owed, in tag order: give it min(deficit, supply)
Levels outermost is the whole of the policy. Each pass lets every receiving CPU settle the members at one distance before any CPU is allowed to look further. Within a pass the receiving CPUs are taken in CPU order and the members in tag order, so the result depends on the machine and the tags, never on the order in which listeners happened to join the group.

Why the loops are nested that way round

Levels outermost, as the code walks them level 0: cpu0 takes worker 0 cpu2 takes worker 2 level 1: cpu0 takes worker 1 cpu2 takes worker 3 level 2: cpu0 takes worker 4 cpu2 takes worker 5 every worker that named a receiving CPU is served by it CPUs outermost, the mistake it avoids cpu0: worker 0, then worker 1, then worker 2 and worker 3 at ring 2, using up its slots cpu2: worker 2 already satisfied, takes 4 and 5 instead worker 2 named cpu2 and is served by cpu0 anyway Both orders deal every slot and give every worker an equal share. The difference is only locality, and it is not small: in the right-hand order a worker that named the very CPU its connections arrive on can be served from another core entirely, which is the property the whole design exists to provide. The same argument applies at every ring, not just the first: a CPU must not reach into ring n+1 while any CPU still has ring n business, or it can take a worker that a nearer CPU was about to claim. Cost of the ordering: the level loop runs nlevels + 1 times over the receiving CPUs, so the whole deal is about (rings + 2) x CPUs x members cheap set tests, typically 5 x 8 x 32, once per membership or tag change and never on the packet path.
The nesting is what makes "nearest" true rather than approximate. Reversing the loops would still balance the load and still fill every slot, so a test that only checked shares would pass either way. That is why the invariant asserted alongside the balance one is that an exact match is never displaced.

Machines where the rings are not four

machinechain from a leafrings that mean somethingwhat the deal does
this host, 5950XSMT pair, L3 group, root0 exact, 1 sibling, 2 same L3, 3 packagethe worked example above
2x Xeon E5-2470 v2 (two sockets)SMT pair, socket L3, root0 exact, 1 sibling, 2 same socket, 3 other socket (QPI)validated 2026-09-26: same four rings, ring 3 crosses QPI; two-DUT deal -11.3% cyc/req vs hash, no worker starved
one L3, no SMTcore, root0 exact, 1 same L3exact first, then anyone in the group
flat topology (smp_topo_none)root only0 exact, 1 everyonedegrades to exact-or-share, which is the honest answer for a machine that reports no structure
two sockets, 8 CCDsSMT pair, L3, socket, root0 to 4, five ringsfills its own CCD, then its own socket, then crosses
chains of unequal depthvaries by CPUas many as that CPU hasa CPU whose chain is shorter than the current level sits that pass out and waits for the catch-all, rather than reaching further than its neighbours
kernel without SMPnone0 exact, 1 everyoneone CPU, so the question does not arise
kern.smp.topology overridewhatever was asked forfollows itnearness follows the administrator's description, right or wrong
The count of rings is never written down in the code. in_pcblbgroup_rings() counts the chain for each receiving CPU and the deal runs one pass per level plus the catch-all, so a machine the scheduler describes in more detail gets a finer-grained deal without any change here.

What is new here

Widening rings over a machine's own topology are established practice in schedulers and memory allocators, where they place a thread or a page near the work that needs it. Equal-share slot tables are established practice in load balancers, where they divide traffic evenly and survive a membership change without reshuffling everything.

Neither has been brought to the socket layer's choice of which listener receives a connection. That choice has until now been an exact CPU match or a hash: the first starves every worker not sitting on a receive queue, the second throws locality away. Putting the two together there is new. The rings decide who is nearest, the slot table guarantees the share, and between them a machine can use every worker it has without giving up the locality the hardware set up.

Measured on hardware

This host (Ryzen 9 5950X, two CCDs on one package - the topology the drawings above use as their example), 2026-08. Here ring 3 is the cross-CCD hop over Infinity Fabric, and its cost is measured, not assumed: against a same-CCD wakeup it is +19% cycles per request. The exact-claim predecessor that the ring deal generalizes was run across three worker layouts - workers all on the receiving CCD gained nothing (the wakeup already shares L3), workers spanning both CCDs gained -15% (affinity vs hash), and forcing every worker to the far CCD cost -19% (affinity vs anti, the pure ring-3 penalty). The ring deal itself was not exercised here; it needs the two-socket receive topology below. These are the numbers that make ring 3 the one expensive step in the drawings.

Two-socket 2x Xeon E5-2470 v2 (Ivy Bridge-EP), two sockets over QPI, 2026-09-26. The ring chain from any leaf CPU is the SMT pair, then the socket's L3 group, then the root, so the deal has the same four rings drawn above, with ring 3 the cross-socket step. Booted the affinity kernel with direct netisr dispatch; cycles per request from a process-scoped PMC over the worker threads only.

The deal delivers when receive spans both sockets. One receive-queue set on each socket, one worker per receive CPU across both (16 workers): plain hash scatters about half the connections to the far socket, while the ring deal keeps each socket's connections on its own socket and still uses every worker. Cross-CPU wakeups fell from ~292k to ~15k and cycles per request dropped 11.3% against the hash - matching the exact-claim design's 12.1% while starving no worker and needing no application change (the learning is on by default; the workers only pin themselves).

Same-socket, the mechanism co-locates ~96% (workers one per receive CPU on one socket): cross-CPU wakeups ~266k to ~8k, learned automatically at accept(2). Cycles per request are flat there, because same-socket wakeups already share the L3 - the ring-3 hop is the only expensive step, exactly as drawn.

One caveat the runs surfaced: supply is learned in two ten-second windows, so a benchmark shorter than that under-converges and the deal stays partly on the hash. Sustained traffic converges it; short bursty runs show variance. Worth weighing where the window length meets real workloads.

What the drawings commit to

  • The ring count is read, not assumed. It is the length of the scheduler's group chain for each receiving CPU, plus one catch-all.
  • Ring 0 is tested against the CPU id, every other ring against a group mask, so a member that named this exact CPU is settled before anything sharing its core.
  • Levels are the outer loop, so no CPU reaches into ring n+1 while any CPU still has ring n business.
  • The catch-all admits every member still owed slots, including those that named no CPU, which is why supply and deficit always both reach zero.
  • A CPU whose chain is shorter than the current level sits that pass out rather than jumping ahead of its neighbours.
  • Nothing here runs on the packet path: the whole deal is a few thousand cheap set tests, once per membership or tag change.