SO_REUSEPORT_LB / receive-CPU affinity / overview

End-to-End Receive Locality

three commits: queue spread, dealing, learning at accept(2) the overview for the series status: proposed; prototyped and measured on a two-socket Xeon E5-2470 v2; not yet submitted 2026-09-14, validated 2026-09-26

A connection's locality is decided twice. The NIC and iflib decide, at attach, which CPU every packet of a flow will be processed on; the socket layer decides, once at the SYN, which worker gets the connection. The exact-claim design only ever touched the second decision and so could not use more workers than receive queues. This page draws both decisions and what we learned by measuring them: where packets land is a placement policy, where connections go is a dealing policy, and the measured cost that matters is crossing an L3 group.

Three commits follow from that. iflib gains an optional spread of queue CPUs across cache groups. The load-balancing group gains a dealt map that serves every worker from the nearest receiving CPU with an equal share. accept(2) learns both which CPUs receive and where each worker runs, so unmodified servers get all of it.

Measured 2026. On this host (Ryzen 9 5950X, two CCDs on one package) the expensive hop is cross-CCD: it costs +19% cycles per request, and the exact-claim predecessor gained up to -15% with workers spanning both CCDs (affinity vs hash). On the two-socket Xeon E5-2470 v2 (two sockets over QPI) the ring deal itself was run: with receive spanning both sockets and one worker per receive CPU across both, it cut cross-CPU wakeups from ~292k to ~15k and cycles per request by 11.3% against a plain hash - the same win as exact-claim while starving no worker and needing no application change. Same-socket, the automatic accept(2) learning co-located ~96% of connections. Caveat: supply is learned in two ten-second windows, so runs shorter than that under-converge. Ring detail: the ring model.

One flow, from the wire to the worker

NIC: Toeplitz hash 4-tuple, key, indirection table picks queue q iflib: queue q on CPU c i-th of the interrupt set commit 1: spread order ithread and stack on c direct dispatch, inline IP, TCP, syncache group lookup on c map[c] then hash within commit 2: the dealt map worker on CPU w accept(2): count[c]++, tag w commit 3: learning remap under the bucket lock only on triggers per packet, in the NIC once, at attach per packet once, at the SYN per accept, no lock on triggers only alternative, config only: bound netisr threads re-queue each packet to CPU t = flowid mod n; a hop per packet Two decisions fix a connection's locality: the placement of queue q (box 2, at attach) and the choice of worker (box 4, at the SYN). Commit 1 changes the first, commits 2 and 3 the second. Nothing new runs per packet; the only added stamp is in the syncache when the handshake completes. The netisr detour is the existing way to make every CPU a receiving CPU; it pays a queue hop on most packets and needs tunables, so it stays documented rather than default. changed by this work existing per-packet alternative
Every hash and decision between the wire and the worker. The NIC's hash and iflib's placement fix the processing CPU before any packet arrives; the group lookup then has one job, choosing the worker. The old exact-claim index answered that with "the worker on this very CPU or the hash". The dealt map answers it with "the nearest worker still owed connections", and accept(2) supplies the two facts that need: which CPUs receive, and where each worker runs.

The cost ladder, measured on this host

Where the worker sits relative to the receive CPU cycles per request, relative to exact (process-scoped PMC, 102400 requests, Ryzen 9 5950X) exact, same CPU 1.00, local wakeup same core, SMT sibling about 1.00, shares L1 and L2 (inferred from the topology, not measured separately) same L3, other core about 1.00, plus one reschedule IPI per wakeup other CCD 1.19, about 1.8 us more per request cross-CPU wakeups per 102400 requests: 89457 when the worker is on another core, 9 when it is exact; the IPI lands in the interrupt thread, not in the worker's cycles What steering was worth, by configuration measured change in cycles per request, dealing versus tuple hash, same pinning in both 8 workers on one CCD, 8 queues on it 0.0%: the hash already kept them on the receiving L3 16 workers on both CCDs, queues on one -15.0%half the connections were cross-CCD under the hash 32 threads on both CCDs, netisr on all -9.2%bound netisr threads everywhere, the config-only lever worker forced to the far CCD (worst case) +15 to +19%every connection cross-CCD win = (share of connections the hash would serve cross-CCD) x 19%. Inside one L3 the difference is an IPI; crossing to the other CCD is the only expensive step. So the design is worth exactly as much as it keeps connections on the L3 that received them, which is why placement comes first. Sources: _ATTIC/rss-ioctl-api/RUNBOOK-perf.md, runs of 2026-08-20 (configs #1 to #3 and the anti mode).
What we learned about receive locality, in one figure. The August harness measured the worker-side cost of every distance between the receive CPU and the worker. Inside an L3 the difference is an IPI; across CCDs it is a fifth of the request. Steering therefore pays in proportion to the connections it keeps on the receiving CCD, and on this host that fraction is set by where the eight queues sit, not by any selection policy.

This host: where the queues land, today and spread

Today: numeric walk of the interrupt set {0, 2, ..., 30} CCD0: L3 group, CPUs 0-15 core 0 0q0 1 core 1 2q1 3 core 2 4q2 5 core 3 6q3 7 core 4 8q4 9 core 5 10q5 11 core 6 12q6 13 core 7 14q7 15 CCD1: L3 group, CPUs 16-31, eligible but unused core 8 16 17 core 9 18 19 core 10 20 21 core 11 22 23 core 12 24 25 core 13 26 27 core 14 28 29 core 15 30 31 cpuid_advance() takes the i-th CPU of the set by id, so q0-q7 fill cores 0-7. core_offset would shift the block, not split it. The RX task and IRQ are bound here at attach and never move. Spread (core_spread=1): order 0, 16, 2, 18, 4, 20, 6, 22, ... and the deal on top CCD0: queues q0, q2, q4, q6 core 0 0q0 1sibling core 1 2q2 3sibling core 2 4q4 5sibling core 3 6q6 7sibling core 4 8from 0 9from 0 core 5 10from 2 11from 2 core 6 12from 4 13from 4 core 7 14from 6 15from 6 CCD1: queues q1, q3, q5, q7 core 8 16q1 17sibling core 9 18q3 19sibling core 10 20q5 21sibling core 11 22q7 23sibling core 12 24from 16 25from 16 core 13 26from 18 27from 18 core 14 28from 20 29from 20 core 15 30from 22 31from 22 receive CPU, exact worker same core same L3, dealt in member order not an interrupt target Each receive CPU supplies 64 of the 512 slots: 16 to itself, 16 to its sibling, 32 to the first two same-L3 workers still owed. No connection crosses a CCD. Same rule, different input.
Placement is the half that was missing. The top half is the host as it runs today: eight queues, one CCD. The bottom half is the same host after commit 1 with commits 2 and 3 dealing on top: every receive CPU serves itself, its sibling and one further core's pair inside its own L3. The dealing code is identical in both halves; only the set of receiving CPUs differs.

Thirty-two workers, four ways

worker on CPU 0 15 | 16 31 Hash (today's default) every worker fed, no steering 16 same L3 (random core), 16 other CCD Exact claims (current branch) mode 2 or explicit tags 8 exact, 24 never selected Deal, queues on CCD0 commits 2 and 3 on today's placement 8 exact, 8 same core, 16 other CCD (one steady producer each) Deal + spread, 4 queues per CCD all three commits 8 exact, 8 same core, 16 same L3, none on the other CCD exact same core same L3, other core other CCD never selected Rows 1 and 3 share the CCD split, because both accept where the packets land. Row 2 is the starvation that prompted the redesign. Row 4 is the target state. Expected effect of row 4 over row 1 on this host: the 19% cross-CCD penalty removed from half the connections, the regime that measured -15% in August.
The four outcomes side by side. Fairness (every worker fed) comes from the hash and from the deal alike; exact claims lose it. Locality inside the receiving CCD comes from the deal. Locality for the other CCD comes only from placement. The design needs all three commits to reach row 4 and degrades gracefully to row 3 without the first.

Distance rings and the dealing loop

package (root cpu_group) L3 group, CCD0 core 0 0 1 core 1 2 3 core 2 4 5 core 3 6 7 core 4 8 9 core 5 10 11 core 6 12 13 core 7 14 15 L3 group, CCD1 core 8 16 17 core 9 18 19 core 10 20 21 core 11 22 23 core 12 24 25 core 13 26 27 core 14 28 29 core 15 30 31 ring 0: CPU 2 itself (exact) ring 1: its SMT sibling ring 2: the rest of its L3 ring 3: the other L3 ring 4: any member with deficit for level in 0..4: for each receiving CPU with supply left: fill ring(level) members that still have deficit, in member order, each up to its deficit, until the supply is gone; one loop, no chunk arithmetic supply per receiving CPU is proportional to its handshake count; every member is owed 16 slots; ring 4 also holds untagged members and guarantees the deal closes
Nearness is the scheduler's own tree, and the loop is one line of policy. Levels are the outer loop, so exact matches are settled for every receive CPU before any sibling is dealt, siblings before the L3, and so on. With queues spread, every receive CPU's supply is exhausted by ring 2; with queues on one CCD, ring 3 absorbs the other CCD's workers. Either way the last ring admits every member with deficit, which is why supply and demand both reach zero.

What one group holds

Members: il_inp[] with tags A inp_lb_cpu = 0 B inp_lb_cpu = 1 C inp_lb_cpu = 2 D inp_lb_cpu = 9 E untagged (NONE or HASH) a tag says where the worker runs; captured at first accept(2) from a pinned thread, or set with SO_REUSEPORT_LB_CPU Receive counts: il_rxcnt[2][cpu] two 10 s windows, bumped at accept(2) from the child's handshake CPU cpu 0123456789 prev 60 0 60 0 30 0 0 0 0 0 cur 40 0 40 0 20 0 0 0 0 0 sum 10010050 T = 16 slots x 5 members = 80 slots. Supply in proportion to 100:100:50, rounded cumulatively: cpu0 32, cpu2 32, cpu4 16. Every member is owed 16. Nothing counted yet? Explicitly tagged CPUs stand in with a count of one each. A group with no tagged member builds no map and keeps today's cost. Map: il_map points at an immutable struct inpcblbmap c=0c=1c=2c=3c=4c=5c=6c=7c=8c=9 lm_cpu[c] off 0n 32 none off 32n 32 none off 64n 16 none none none none none lm_slot[] A x16 exact B x16 same core C x16 exact D x16 other L3 E x16 untagged 01632486480 cpu0 range = A,B | cpu2 range = C,D (D's own CPU 9 receives nothing, so it is dealt from the nearest that does) | cpu4 range = E | cpu1, 3, 5-9: hash over all five
Tags, counts, map. The tag is a uint16 in the inpcb hole. The counts live in the group block and are bumped without a lock. The map is a separate immutable block: a packed offset and count per CPU, then the slot array of member pointers, rewritten under the bucket lock, published with a pointer swap and freed after a net epoch. Readers never validate anything, because nothing they read can change under them.

Selecting a member: two loads, one fallback, no validation

packet on CPU c tuple hash h already computed map = il_map atomic_load_acq_ptr map == NULL? no e = map->lm_cpu[c] load 1 count(e) == 0? yes yes no member = il_inp[h % il_inpcnt] today's path, unchanged; never NULL for a non-empty group member = map->lm_slot[off(e) + (h * count(e) >> 32)] load 2; the pointer is the answer, no tag check caller: inp_smr_lock(member) same window as today's il_inp[] read freed? INP_LOOKUP_AGAIN -> in_pcblookup_with_lock() takes the bucket lock and sees the fresh map A group that never tagged anyone pays one load for the NULL map. A steered lookup pays two more and a multiply; there is no division and no per-slot validation, because the map is immutable.
The lookup got simpler than the exact-claim version. The old index had to validate the slot it found against the member's tag because the member array is compacted in place on removal. The map is never modified after publication, so a reader either sees the old map or the new one, whole. A member that left the group after the map was read is caught where it was always caught: the caller fails to lock it and retries under the bucket lock.

Learning at accept(2): where the counts and the tags come from

on receive CPU c, in the input path SYN on CPU c NIC-hashed, M_HASHTYPE RSS syncache entry no socket yet ACK on CPU c same flow, same ring syncache_socket() stamps the child inp_lb_cpu = c (the observable), and RXHASH if M_HASHTYPE_ISRSS(m) and rcvif is not IFF_LOOPBACK queued on the listener on worker CPU w, in the accept(2) syscall, no socket lock held worker on CPU w accept(2) dequeues the child pr_accepted(head, child) reads the child's CPU and RXHASH bit RXHASH and count[win][c] > 0: count[win][c]++ the steady state: one atomic add, no lock, no remap RXHASH and count[win][c] == 0: lock, count, remap a CPU seen for the first time in this window; at most N times per window window older than 10 s: roll prev = cur, cur = 0, remap a CPU with zero counts in both windows drops out of the map listener untagged and the thread pinned to one CPU: tag w, remap the capture from the previous set, minus its coverage gate remap = rebuild the map under the lbgroup bucket lock, publish it with a pointer swap, free the old one after an epoch. Also on every member join, leave and setsockopt, as before. The listener's write lock is only taken on those triggers; SYN processing contends with it otherwise. Two windows of counts prev (10 s) cur (10 s) rollrollrollnow cpu 2 seen in both: keeps its supply cpu 9 seen in neither: dropped at a roll Which arrivals are learned: NIC RSS hash, rcvif aq0: yes. lo0: the sender's software hash rides on the mbuf, rcvif is lo0: no. inp_flowtype is not consulted: the syncache fills it by software when no hash arrived (tcp_syncache.c:894), so it is set on every child.
Nothing new runs in the packet path except one stamp. The syncache already recorded the handshake CPU; it now also records whether the segment was hardware-hashed on a real interface. Everything else happens in accept(2), which the worker was going to call anyway. The first draft read the child's flow type instead; the review showed that field is set on every child and that loopback SYNs carry the client's own software hash, which would have made loopback tests nondeterministic and let a wandering netisr thread be learned as a receive CPU.

The three commits

committouchesdeliversproven by
1. iflib: optionally spread queue CPUs across cache groupssys/net/iflib.c, iflib.4: an order array built at attach, cpuid_advance() walking it, the cursor key, tunable core_spreadqueues on every L3 group of the local domain; default off, numeric order unchangedguest with two sockets and 4 queues: rxq*.cpu lands 2+2; host after the next kernel install: 0, 16, 2, 18, 4, 20, 6, 22
2. netinet: add SO_REUSEPORT_LB_CPU, dealing members by receive CPUin_pcb_var.h, in_pcb.h, in_pcb.c, tcp_syncache.c, option plumbing, getsockopt.2, ATF teststhe dealt map, the selection, explicit tags standing in as receive CPUs until something is learnedloopback ATF: steering, duplicates, untagged share, partial claims; all fail on the pristine kernel
3. netinet: learn SO_REUSEPORT_LB receive and worker CPUs at accept(2)protosw.h, uipc_syscalls.c, tcp_usrreq.c, in_pcb.c, tcp.4, ATF testscounts, windows, capture, the 0/1 sysctl; unmodified servers get everythingaq passthrough: 8 exact, 4 queues on SMT pairs, 6 on 4, 1 queue; sysctl 0 control

Cases the rules decide without special code

situationwhat happens
no member is taggedno map is built; selection and cost are identical to today
the NIC reports no hash (bhyve e1000, vtnet), workers pinned and capturednothing is learned, no map: the hash distributes, nobody starves
explicit tags, nothing learned (loopback tests, UDP)the tagged CPUs stand in with equal supply: exact for them, other members dealt from them
something learned, an explicit tag names a CPU that never receivesthat member is dealt from the nearest receiving CPU; a claim cannot starve anyone
a packet arrives on a CPU with no slots (loopback, a stray netisr CPU)hash over all members
two members tag the same CPUboth filled from it when its supply covers them; otherwise the later one takes its quota from the next ring
one receive queue, thirty-two workersall thirty-two are dealt from that CPU: the same distribution as the hash
queues or NICs receive at different ratessupply follows the counts; a trickle CPU supplies a trickle and its workers get the rest elsewhere
all queues on one L3, workers on every CPUthe far workers are dealt from the nearest L3 that receives; only placement (commit 1) can do better
a member joins, leaves or is retagged; the map allocation failsremap in the same critical section, old map freed after an epoch; on failure the map is NULL, the hash takes over, the next trigger retries
an unbound netisr thread carries the traffic (deferred dispatch, options RSS without the RSS-kernel dispatch fix)counts follow the thread; after it migrates, flows on the new CPU hash until it is counted, then at most two windows of skew; the precondition of a per-flow-stable CPU is unchanged
single-L3 machine, or a kernel without SMPthe spread order equals the numeric order; the rings degrade to exact / everyone
two sockets, NIC local to onethe spread interleaves the local socket's cache groups and never crosses to the other while local CPUs remain

What the drawings commit to

  • Two decisions, two policies. Placement fixes the receiving CPU at attach; dealing fixes the worker at the SYN. Neither runs per packet.
  • Selection never returns NULL for a non-empty group, on either path, and costs one load when the group has no tags.
  • Every member is owed 16 slots and gets them; supply is proportional to observed handshakes, each CPU fills its ring in member order, and the last ring closes the deal.
  • An exact match is never displaced by a farther receive CPU; siblings come before the L3, the L3 before the other CCD.
  • Nothing is learned from loopback or from the syncache's software fallback; only a hardware-hashed handshake on a real interface counts.
  • Spread is off by default and never leaves the local NUMA domain while it has CPUs to give; on a single-L3 machine it changes nothing.
  • Failure degrades to today's hash: no tags, no counts, a failed allocation, or a CPU without slots all land there.
  • The expected win is bounded by the cost ladder: at most the cross-CCD share of connections times 19%, which on this host requires the spread to be nonzero.