commit 2: the nearest-CPU deal, in_pcblbgroup_remap()replaces the earlier exact-claim workstatus: proposed; prototyped and measured on a two-socket Xeon E5-2470 v2; not yet submitted2026-09-14, validated 2026-09-26
The two earlier designs drew a tag as an exact claim on one CPU. That works when there is one worker per receive queue and fails the moment there are more workers than queues: either the group never leaves the hash, or the few workers sitting on receive CPUs take everything. This commit draws the replacement. One rule: each receiving CPU deals its share of connections to the group members nearest to it, and every member ends up with an equal share.
Three things exist per group. A tag says where a worker runs (captured at its first accept, or set with the option). Receive counts say which CPUs complete hardware-hashed handshakes and how many, in two ten-second windows. The map is the dealt result: for every CPU, a range of slots holding member pointers, rebuilt off the packet path and read with two loads.
Hardware-validated 2026-09-26 on a two-socket machine (2x Xeon E5-2470 v2, two sockets over QPI). With receive spanning both sockets and one worker per receive CPU across both, the dealt map cut cross-CPU wakeups from ~292k to ~15k and cycles per request by 11.3% against a plain hash - matching the exact-claim design while starving no worker and needing no application change. On one socket the automatic accept(2) learning co-located ~96% of connections (~266k to ~8k), which also confirms direct netisr dispatch places the handshake on the receive CPU. On the 5950X (two CCDs) the earlier exact-claim runs set the ring-3 cost: +19% cross-CCD, -15% for workers spanning both CCDs. Caveat: supply is learned in two ten-second windows, so runs shorter than that under-converge; sustained traffic converges it.
Thirty-two pinned workers, eight receive queues
The starvation the current branch cannot escape, and the deal that replaces it. Left: a tag is an exclusive claim, so only the eight workers whose CPU owns a receive queue are ever selected. Right: the same eight receive CPUs deal their connections outward by distance. The eight exact matches keep what they had, each SMT sibling is served by its own core, and the sixteen workers on the other CCD are dealt two per receive CPU. Every worker gets 1/32. The cross-CCD workers were going to be cross-CCD under any scheme, since no receive queue lives there; now they at least have one steady producer each.
Two levers, and which one this is
The packet's CPU is decided by the NIC and the driver; the socket layer can only decide the worker. Moving packets to more CPUs is lever 1, which FreeBSD already has in the netisr tunables, at a hop per packet. Lever 2 accepts where the packet runs and chooses the nearest worker once per connection. The two compose: with bound netisr threads on every CPU, the counts show every CPU receiving and the deal degenerates to exact matches.
What one group holds after the change
Tags, counts, map. The tag is a uint16 in the inpcb hole, as before. The counts live in the group block, sized at allocation like the old per-CPU index was, and are bumped without a lock. The map is a separate immutable block: a packed offset and count per CPU, then the slot array of member pointers. It is rewritten under the bucket lock and published with a pointer swap; the old one is freed after a net epoch. Readers never validate anything, because nothing they read can change under them.
Distance rings from one receive CPU
Nearness is the scheduler's own tree. From receive CPU 2 the rings are: itself, its sibling, the rest of CCD0, CCD1, then anyone still owed slots. The level loop is outermost, so a farther CPU can never take a member that a nearer CPU is about to claim: exact matches are settled for every receive CPU before any sibling is dealt, siblings before the L3, and so on. The last ring admits every member with deficit, which is why supply and deficit always both reach zero.
The deal as a ledger: twelve workers on eight queues
A ratio that is not an integer still comes out exact. Under a plain partition, twelve on eight would give four workers twice the load of the other eight. Here every receive CPU has 24 slots to give and every worker is owed 16, so the six exact matches take 16, their siblings take the remaining 8, and the two receive CPUs with no worker of their own fill the siblings' last 8 at the L3 level, 4 each. Rows sum to the supply, columns to the quota, and unequal receive rates simply change the row supplies.
Selecting a member: two loads, one fallback, no validation
The lookup got simpler than the previous set's. The old index had to validate the slot it found against the member's tag because the member array is compacted in place on removal. The map is never modified after publication, so a reader either sees the old map or the new one, whole. A member that left the group after the map was read is caught where it was always caught: the caller fails to lock it and retries the lookup under the bucket lock.
Learning at accept(2): where the counts and the tags come from
Nothing new runs in the packet path except one stamp. The syncache already recorded the handshake CPU in the previous set; it now also records whether the segment was hardware-hashed on a real interface. Everything else happens in accept(2), which the worker was going to call anyway: the count for the handshake CPU goes up without a lock, and the bucket lock is taken only on the four triggers. The first draft of this design read the child's inp_flowtype instead of a stamped bit; the review showed that field is set on every child and that loopback SYNs carry the client's own software hash, which would have made loopback tests nondeterministic and would have let a wandering netisr thread be learned as a receive CPU.
Cases the rule decides without special code
situation
what happens
no member is tagged
no map is built; selection and cost are identical to today
the NIC reports no hash (bhyve e1000, vtnet), workers are pinned and captured
nothing is learned, no map: the hash distributes, nobody starves
the tagged CPUs stand in with equal supply: exact for them, other members dealt from them
something learned, an explicit tag names a CPU that never receives
that member is dealt from the nearest receiving CPU; a claim cannot starve anyone
a packet arrives on a CPU with no slots (loopback, a stray netisr CPU)
hash over all members
two members tag the same CPU
they split that CPU's slots evenly, then fill their quota from the next ring
one receive queue, thirty-two workers
all thirty-two are dealt from that CPU: the same distribution as the hash
queues or NICs receive at different rates
supply follows the counts; a trickle CPU supplies a trickle and its workers get the rest elsewhere
a member joins, leaves or is retagged
remap in the same critical section; the old map is freed after an epoch
the map allocation fails (M_NOWAIT)
the map is set to NULL, the hash takes over, the next trigger retries
an unbound netisr thread carries the traffic (deferred dispatch, options RSS without the RSS-kernel dispatch fix)
counts follow the thread; after it migrates, flows on the new CPU hash until it is counted, then at most two windows of skew; the precondition of a per-flow-stable CPU is unchanged from the earlier designs
What changed from the exact-claim design
The exact-claim design this replaces - the per-connection exact match and the first-accept work - and why it was superseded are consolidated in Exact Claim to Ring Deal. In short:
The per-CPU index became a dealt map. One member per CPU with a shared bit is replaced by a slot range per CPU. Duplicates, fan-out and fan-in are all the same mechanism.
The coverage gate is gone. It existed to avoid starving workers on non-receive CPUs; the deal gives them their share instead, so a captured tag can steer as soon as it exists.
The seen-CPU set became counts. A cpuset that only grew is replaced by two windows of per-CPU counts, which also weight the supply by rate and age out dead CPUs.
The learn hook moved out of tcp_input. One bit stamped in the syncache replaces the packet-path call; the accept path does the rest.
Sysctl mode 2 is gone. The sysctl is on or off; explicit tags work either way.
Kept: the option and its plumbing, the uint16 tag and flags in the inpcb hole, the group back-pointer in the spare slot, the handshake-CPU observable on accepted sockets, the pr_accepted hook with its CURVNET_SET, and the sticky ANY.
What the drawings commit to
Selection never returns NULL for a non-empty group, on either path.
Every member is owed exactly 16 slots and gets them; supply equals demand by construction and the last ring closes the deal.
An exact match is never displaced by a farther receive CPU, because levels are the outer loop.
Supply is proportional to observed handshakes, rounded so the supplies sum to the total; nothing is learned from loopback or from the syncache's software fallback.
No new code in the packet path beyond one stamp in the syncache; the bucket lock is taken only on membership, tag, first-count and roll events.
Failure degrades to today's hash: no tags, no counts, a failed allocation, or a CPU without slots all land there.