SO_REUSEPORT_LB / receive-CPU affinity / commit 2

Receive-CPU Dealing

commit 2: the nearest-CPU deal, in_pcblbgroup_remap() replaces the earlier exact-claim work status: proposed; prototyped and measured on a two-socket Xeon E5-2470 v2; not yet submitted 2026-09-14, validated 2026-09-26

The two earlier designs drew a tag as an exact claim on one CPU. That works when there is one worker per receive queue and fails the moment there are more workers than queues: either the group never leaves the hash, or the few workers sitting on receive CPUs take everything. This commit draws the replacement. One rule: each receiving CPU deals its share of connections to the group members nearest to it, and every member ends up with an equal share.

Three things exist per group. A tag says where a worker runs (captured at its first accept, or set with the option). Receive counts say which CPUs complete hardware-hashed handshakes and how many, in two ten-second windows. The map is the dealt result: for every CPU, a range of slots holding member pointers, rebuilt off the packet path and read with two loads.

Hardware-validated 2026-09-26 on a two-socket machine (2x Xeon E5-2470 v2, two sockets over QPI). With receive spanning both sockets and one worker per receive CPU across both, the dealt map cut cross-CPU wakeups from ~292k to ~15k and cycles per request by 11.3% against a plain hash - matching the exact-claim design while starving no worker and needing no application change. On one socket the automatic accept(2) learning co-located ~96% of connections (~266k to ~8k), which also confirms direct netisr dispatch places the handshake on the receive CPU. On the 5950X (two CCDs) the earlier exact-claim runs set the ring-3 cost: +19% cross-CCD, -15% for workers spanning both CCDs. Caveat: supply is learned in two ten-second windows, so runs shorter than that under-converge; sustained traffic converges it.

Thirty-two pinned workers, eight receive queues

A. Exact claims (the current branch, mode 2 or explicit tags) receive queues q0-q7 bound to CPUs 0,2,...,14 by iflib; one pinned worker per CPU q0q1q2q3q4q5q6q7 CCD0 CCD1 0 2 4 6 8 10 12 14 1 3 5 7 9 11 13 15 16 18 20 22 24 26 28 30 17 19 21 23 25 27 29 31 fed: worker on a receive CPU (8) never selected (24) share per worker, CPU 0 to 31 1/8 each 0 8 workers do all the work; 24 pinned workers never see a connection. B. Dealt by nearest receive CPU (this design) same machine, same queues, no application change q0q1q2q3q4q5q6q7 CCD0 CCD1 0 2 4 6 8 10 12 14 1 3 5 7 9 11 13 15 16from 0 18from 2 20from 4 22from 6 24from 8 26from 10 28from 12 30from 14 17from 0 19from 2 21from 4 23from 6 25from 8 27from 10 29from 12 31from 14 exact (8) same core (8) other L3, 2 per receive CPU (16) share per worker, CPU 0 to 31 1/32 each 32 workers each take 1/32; 16 are exact or on the same core, 16 sit on the other CCD.
The starvation the current branch cannot escape, and the deal that replaces it. Left: a tag is an exclusive claim, so only the eight workers whose CPU owns a receive queue are ever selected. Right: the same eight receive CPUs deal their connections outward by distance. The eight exact matches keep what they had, each SMT sibling is served by its own core, and the sixteen workers on the other CCD are dealt two per receive CPU. Every worker gets 1/32. The cross-CCD workers were going to be cross-CCD under any scheme, since no receive queue lives there; now they at least have one steady producer each.

Two levers, and which one this is

Lever 1, config only: bound per-CPU netisr threads queue the packet to CPU t = flowid mod nthreads, IPI, run the stack there cost: a queue hop on most packets; needs net.isr tunables (and the RSS-kernel dispatch fix) every packet NIC ring q Toeplitz hash of the 4-tuple ithread on CPU c iflib binds ring q to CPU c IP and TCP input on c direct dispatch, inline group lookup on c one member is chosen worker on CPU w accept(2), read, write once, at the SYN Lever 2, this design: deal the connection at the group lookup map[c] holds the members nearest to c; the tuple hash picks among them cost: none per packet; one remap when membership, tags or counts change Lever 1 already exists (net.isr.maxthreads=-1, bindthreads=1, dispatch=hybrid) and makes every CPU a receive CPU; the same map then holds only exact matches.
The packet's CPU is decided by the NIC and the driver; the socket layer can only decide the worker. Moving packets to more CPUs is lever 1, which FreeBSD already has in the netisr tunables, at a hop per packet. Lever 2 accepts where the packet runs and chooses the nearest worker once per connection. The two compose: with bound netisr threads on every CPU, the counts show every CPU receiving and the deal degenerates to exact matches.

What one group holds after the change

Members: il_inp[] with tags A inp_lb_cpu = 0 B inp_lb_cpu = 1 C inp_lb_cpu = 2 D inp_lb_cpu = 9 E untagged (NONE or HASH) a tag says where the worker runs; captured at first accept(2) from a pinned thread, or set with SO_REUSEPORT_LB_CPU Receive counts: il_rxcnt[2][cpu] two 10 s windows, bumped at accept(2) from the child's handshake CPU cpu 0123456789 prev 60 0 60 0 30 0 0 0 0 0 cur 40 0 40 0 20 0 0 0 0 0 sum 10010050 T = 16 slots x 5 members = 80 slots. Supply in proportion to 100:100:50, rounded cumulatively: cpu0 32, cpu2 32, cpu4 16. Every member is owed 16. Nothing counted yet? Explicitly tagged CPUs stand in with a count of one each. A group with no tagged member builds no map and keeps today's cost. Map: il_map points at an immutable struct inpcblbmap c=0c=1c=2c=3c=4c=5c=6c=7c=8c=9 lm_cpu[c] off 0n 32 none off 32n 32 none off 64n 16 none none none none none lm_slot[] A x16 exact B x16 same core C x16 exact D x16 other L3 E x16 untagged 01632486480 cpu0 range = A,B | cpu2 range = C,D (D's own CPU 9 receives nothing, so it is dealt from the nearest that does) | cpu4 range = E | cpu1, 3, 5-9: hash over all five
Tags, counts, map. The tag is a uint16 in the inpcb hole, as before. The counts live in the group block, sized at allocation like the old per-CPU index was, and are bumped without a lock. The map is a separate immutable block: a packed offset and count per CPU, then the slot array of member pointers. It is rewritten under the bucket lock and published with a pointer swap; the old one is freed after a net epoch. Readers never validate anything, because nothing they read can change under them.

Distance rings from one receive CPU

package (root cpu_group) L3 group, CCD0 core 0 0 1 core 1 2 3 core 2 4 5 core 3 6 7 core 4 8 9 core 5 10 11 core 6 12 13 core 7 14 15 L3 group, CCD1 core 8 16 17 core 9 18 19 core 10 20 21 core 11 22 23 core 12 24 25 core 13 26 27 core 14 28 29 core 15 30 31 ring 0: CPU 2 itself (exact) ring 1: its SMT sibling ring 2: the rest of its L3 ring 3: the other L3 ring 4: any member with deficit for level in 0..4: for each receiving CPU with supply left: hand slots to ring(level) members that still have deficit even chunk = min(deficit, supply / candidates), then one slot at a time until supply or deficit is gone rings come from smp_topo_find(cpu_top, c) walked upward; a flat topology degrades to exact / everyone; ring 4 also holds untagged members and guarantees the deal closes
Nearness is the scheduler's own tree. From receive CPU 2 the rings are: itself, its sibling, the rest of CCD0, CCD1, then anyone still owed slots. The level loop is outermost, so a farther CPU can never take a member that a nearer CPU is about to claim: exact matches are settled for every receive CPU before any sibling is dealt, siblings before the L3, and so on. The last ring admits every member with deficit, which is why supply and deficit always both reach zero.

The deal as a ledger: twelve workers on eight queues

12 pinned workers on CPUs 0-11; receive queues on CPUs 0,2,...,14 with equal counts. T = 16 x 12 = 192 slots: each receive CPU supplies 24, each worker is owed 16. cpu 0cpu 1cpu 2cpu 3cpu 4cpu 5cpu 6cpu 7cpu 8cpu 9cpu 10cpu 11 w0w1w2w3w4w5w6w7w8w9w10w11 supply rx cpu 0rx cpu 2rx cpu 4rx cpu 6rx cpu 8rx cpu 10rx cpu 12rx cpu 14 2424242424242424 16 8 16 8 16 8 16 8 16 8 16 8 4 4 4 4 4 4 4 4 4 4 4 4 owed 16, got: 161616161616161616161616 level 0, exact: 16 level 1, same core: the 8 left over level 2, same L3: 24 / 6 candidates = 4 each, from the two receive CPUs with no worker of their own
A ratio that is not an integer still comes out exact. Under a plain partition, twelve on eight would give four workers twice the load of the other eight. Here every receive CPU has 24 slots to give and every worker is owed 16, so the six exact matches take 16, their siblings take the remaining 8, and the two receive CPUs with no worker of their own fill the siblings' last 8 at the L3 level, 4 each. Rows sum to the supply, columns to the quota, and unequal receive rates simply change the row supplies.

Selecting a member: two loads, one fallback, no validation

packet on CPU c tuple hash h already computed map = il_map atomic_load_acq_ptr map == NULL? no e = map->lm_cpu[c] load 1 count(e) == 0? yes yes no member = il_inp[h % il_inpcnt] today's path, unchanged; never NULL for a non-empty group member = map->lm_slot[off(e) + (h * count(e) >> 32)] load 2; the pointer is the answer, no tag check caller: inp_smr_lock(member) same window as today's il_inp[] read freed? INP_LOOKUP_AGAIN -> in_pcblookup_with_lock() takes the bucket lock and sees the fresh map A group that never tagged anyone pays one load for the NULL map. A steered lookup pays two more and a multiply; there is no division and no per-slot validation, because the map is immutable.
The lookup got simpler than the previous set's. The old index had to validate the slot it found against the member's tag because the member array is compacted in place on removal. The map is never modified after publication, so a reader either sees the old map or the new one, whole. A member that left the group after the map was read is caught where it was always caught: the caller fails to lock it and retries the lookup under the bucket lock.

Learning at accept(2): where the counts and the tags come from

on receive CPU c, in the input path SYN on CPU c NIC-hashed, M_HASHTYPE RSS syncache entry no socket yet ACK on CPU c same flow, same ring syncache_socket() stamps the child inp_lb_cpu = c (the observable), and RXHASH if M_HASHTYPE_ISRSS(m) and rcvif is not IFF_LOOPBACK queued on the listener on worker CPU w, in the accept(2) syscall, no socket lock held worker on CPU w accept(2) dequeues the child pr_accepted(head, child) reads the child's CPU and RXHASH bit RXHASH and count[win][c] > 0: count[win][c]++ the steady state: one atomic add, no lock, no remap RXHASH and count[win][c] == 0: lock, count, remap a CPU seen for the first time in this window; at most N times per window window older than 10 s: roll prev = cur, cur = 0, remap a CPU with zero counts in both windows drops out of the map listener untagged and the thread pinned to one CPU: tag w, remap the capture from the previous set, minus its coverage gate remap = rebuild the map under the lbgroup bucket lock, publish it with a pointer swap, free the old one after an epoch. Also on every member join, leave and setsockopt, as before. The listener's write lock is only taken on those triggers; SYN processing contends with it otherwise. Two windows of counts prev (10 s) cur (10 s) rollrollrollnow cpu 2 seen in both: keeps its supply cpu 9 seen in neither: dropped at a roll Which arrivals are learned: NIC RSS hash, rcvif aq0: yes. lo0: the sender's software hash rides on the mbuf, rcvif is lo0: no. inp_flowtype is not consulted: the syncache fills it by software when no hash arrived (tcp_syncache.c:894), so it is set on every child.
Nothing new runs in the packet path except one stamp. The syncache already recorded the handshake CPU in the previous set; it now also records whether the segment was hardware-hashed on a real interface. Everything else happens in accept(2), which the worker was going to call anyway: the count for the handshake CPU goes up without a lock, and the bucket lock is taken only on the four triggers. The first draft of this design read the child's inp_flowtype instead of a stamped bit; the review showed that field is set on every child and that loopback SYNs carry the client's own software hash, which would have made loopback tests nondeterministic and would have let a wandering netisr thread be learned as a receive CPU.

Cases the rule decides without special code

situationwhat happens
no member is taggedno map is built; selection and cost are identical to today
the NIC reports no hash (bhyve e1000, vtnet), workers are pinned and capturednothing is learned, no map: the hash distributes, nobody starves
explicit tags, nothing learned (loopback tests, UDP)the tagged CPUs stand in with equal supply: exact for them, other members dealt from them
something learned, an explicit tag names a CPU that never receivesthat member is dealt from the nearest receiving CPU; a claim cannot starve anyone
a packet arrives on a CPU with no slots (loopback, a stray netisr CPU)hash over all members
two members tag the same CPUthey split that CPU's slots evenly, then fill their quota from the next ring
one receive queue, thirty-two workersall thirty-two are dealt from that CPU: the same distribution as the hash
queues or NICs receive at different ratessupply follows the counts; a trickle CPU supplies a trickle and its workers get the rest elsewhere
a member joins, leaves or is retaggedremap in the same critical section; the old map is freed after an epoch
the map allocation fails (M_NOWAIT)the map is set to NULL, the hash takes over, the next trigger retries
an unbound netisr thread carries the traffic (deferred dispatch, options RSS without the RSS-kernel dispatch fix)counts follow the thread; after it migrates, flows on the new CPU hash until it is counted, then at most two windows of skew; the precondition of a per-flow-stable CPU is unchanged from the earlier designs

What changed from the exact-claim design

The exact-claim design this replaces - the per-connection exact match and the first-accept work - and why it was superseded are consolidated in Exact Claim to Ring Deal. In short:

  • The per-CPU index became a dealt map. One member per CPU with a shared bit is replaced by a slot range per CPU. Duplicates, fan-out and fan-in are all the same mechanism.
  • The coverage gate is gone. It existed to avoid starving workers on non-receive CPUs; the deal gives them their share instead, so a captured tag can steer as soon as it exists.
  • The seen-CPU set became counts. A cpuset that only grew is replaced by two windows of per-CPU counts, which also weight the supply by rate and age out dead CPUs.
  • The learn hook moved out of tcp_input. One bit stamped in the syncache replaces the packet-path call; the accept path does the rest.
  • Sysctl mode 2 is gone. The sysctl is on or off; explicit tags work either way.
  • Kept: the option and its plumbing, the uint16 tag and flags in the inpcb hole, the group back-pointer in the spare slot, the handshake-CPU observable on accepted sockets, the pr_accepted hook with its CURVNET_SET, and the sticky ANY.

What the drawings commit to

  • Selection never returns NULL for a non-empty group, on either path.
  • Every member is owed exactly 16 slots and gets them; supply equals demand by construction and the last ring closes the deal.
  • An exact match is never displaced by a farther receive CPU, because levels are the outer loop.
  • Supply is proportional to observed handshakes, rounded so the supplies sum to the total; nothing is learned from loopback or from the syncache's software fallback.
  • No new code in the packet path beyond one stamp in the syncache; the bucket lock is taken only on membership, tag, first-count and roll events.
  • Failure degrades to today's hash: no tags, no counts, a failed allocation, or a CPU without slots all land there.