FreeBSD / networking / flow steering

The Flow-Steering Landscape

reference / whole-territory survey FreeBSD main: sys/net, sys/netinet, sys/dev 2026-09-26

Receive-side scaling and hardware flow steering exist all over the FreeBSD tree, but nowhere as one programmable thing. Every piece an operator would want - the hash key, the indirection table, an n-tuple filter, a listener that follows the queue its packets landed on - lives behind a different interface, and most of those interfaces are read-only, dead, or private to one driver.

This is the map of the whole territory: which layer owns what, which direction it moves, and where a generic API could slot in. Two companion pages go narrow instead - one walks a single mechanism, one documents a single get-and-set API. This one draws the gaps between them.

Where steering can and cannot be programmed

write-capable read-only / observe dead / missing vendor-private the generic datapath userland: app accept loop, ifconfig(8) socket layer IP_RECVRSSBUCKETID observes; SO_REUSEPORT_LB jhashes neither can ask for a queue generic ifnet ioctls SIOCGIFRSSKEY / SIOCGIFRSSHASH - GET only no set path, no table ioctl, no consumers iflib: no ifdi_ RSS method unknown ioctl -> ether_ioctl() -> EINVAL x write path dies here, never reaches silicon boot-time, kernel-owned net.inet.rss (rss_config.c) key CTLFLAG_RD (write -> EINVAL) bits RDTUN, bucket_mapping RD, basecpu const under options RSS the kernel owns key + LUT non-iflib GET: mlx4en, mlx5en, mana, hn, qlnx programs key + LUT at attach the write-capable escape hatches per-driver sysctls hn, ixl, bnxt, ena, cxgbe five incompatible namespaces; root writes vendor filter engines cxgbe t4_filter (cxgbetool); mlx5 fs_core cxgbe user-programmable; mlx5 kTLS-only IF_SND_TAG_TYPE_TLS_RX the only generic per-flow RX hook lagg/vlan proxy, refcount - welded to kTLS kTLS RX only NIC RSS engine Toeplitz key + indirection table + N rx queues ithreads pinned to CPUs at attach
The generic path down the middle is read-only or dead; every route that can actually write sits off to the side. The socket layer can read the bucket a packet landed in but cannot ask for one; the generic SIOCGIFRSS* ioctls are GET-only with no in-tree consumer; and iflib - which backs nearly every modern NIC - has no RSS method, so a write falls through to ether_ioctl() and returns EINVAL. Real programmability lives in the boot-time config the kernel owns, five per-driver sysctls, two vendor-private filter engines, and one send-tag type meant for kTLS. None of them share an interface.

Every layer, direction, and reach

layer / APIdirectionwho implements itreachable from userland?notes
net.inet.rss (rss_config.c)read-onlykernel core, at bootyes, sysctl readkey CTLFLAG_RD (write returns EINVAL), bits RDTUN, bucket_mapping RD, basecpu const; kernel owns key and LUT under options RSS
SIOCGIFRSSKEY / SIOCGIFRSSHASHread-onlymlx4en, mlx5en, mana, hn, qlnxyes, but no consumerno-privilege GET block in ifhwioctl(); RSS_KEYLEN 128 vs RSS_KEYSIZE 40 for one concept; only in-tree caller is hn mirroring its own VF
indirection-table ioctldead / missing-nothe one thing an operator wants to change had no generic representation in the baseline tree
iflib (ifdi_if.m)deadbacks ixgbe, ice, ixl, em, igc, bnxt, aq, vmx, enic, axgbe...nono ifdi_ RSS method; unknown ioctl falls to ether_ioctl() -> EINVAL; ifdi_i2c_req is the thin-method precedent it never got
per-driver sysctlswritehn, ixl, bnxt, ena, cxgbeyes, rootfive incompatible namespaces; hn writable only #ifndef RSS; cxgbe rss_base/rss_size read-only
IP_FLOWID / IP_RECVFLOWID / IP_RECVRSSBUCKETIDobservenetinet socket options (90 / 93 / 94)yesread the flowid and bucket of a received datagram; nothing lets an app request a bucket
SO_REUSEPORT_LBread-onlyin_pcb lbgroup lookupyesINP_PCBLBGROUP_PKTHASH jhashes the 3-tuple and mods by member count; NUMA-domain aware but ignores m_pkthdr.flowid; replaces the removed IP_RSS_LISTEN_BUCKET (slot reserved)
cxgbe t4_filtervendor-private (write)cxgbeyes, via cxgbetool / /dev/t4nexCHELSIO_T4_{GET,SET,DEL}_FILTER, filter mode/mask; the only user-programmable hardware filter API in the tree
mlx5 fs_corevendor-privatemlx5nofull flow-steering core plus a per-inpcb entry point (mlx5e_accel_fs_add_inpcb); reachable only from kTLS RX
IF_SND_TAG_TYPE_TLS_RXwrite (kTLS-only)snd_tag framework (type 4)only via kTLSthe only generic per-flow RX steering hook; already has lagg/vlan proxying, a refcounted lifecycle, and an inpcb in the alloc params - just hardcoded to kTLS
Direction is the whole story: almost everything an application can reach is read-only or observe, everything that can write is either private to one driver or welded to one consumer, and the one central place a generic set would belong - iflib - is dead. The removed IP_RSS_LISTEN_BUCKET and PCBGROUP KPI mean FreeBSD once had an application-facing "bind this listener to bucket N" and now has nothing.

The question any set-side API has to answer first

who owns the key and the indirection table? options RSS on the kernel owns both rss_m2cpuid() and the netisr path assume hardware matches rss_table[] a write behind the stack breaks the bucket -> CPU contract userland gets nothing options RSS off the driver owns both nothing in the stack cares one driver's private sysctls are the only way to change anything no cross-driver interface the missing third state kernel and userland agree no path exists today hn just refuses writes when RSS is defined - crude but safe proposal: get reports owner kernel | driver | user; set transitions it or refuses There is no third state today where userland and kernel agree on ownership, and getting that contract right is the prerequisite to any generic set path: the get side has to name the owner before the set side can be allowed to change anything.
The get path has to define the contract before the set path can exist. With options RSS the kernel's affinity plumbing depends on the hardware table matching rss_table[]; without it, the driver is free and the stack is blind. hn's answer - refuse writes when RSS is compiled in - is defensible but binary. A real answer reports a three-state owner on the get and either fails the set or transitions ownership explicitly.

Where a generic API could slot in, ranked

insertion pointshapereusesnote
1. bridge the read ioctls into iflibthin ifdi_ methods behind the existing GET ioctlsifdi_i2c_req, ifdi_get_downreasoncloses the hole in the middle; a generic ioctl backed by a per-driver method is already the established pattern
2. add the indirection tablea table structure plus get/set ioctls, with an owner fieldthe ifnet ioctl idiomgives the one control an operator actually wants a generic representation; owner field carries the three-state answer
3. make SO_REUSEPORT_LB RSS-awareselect by m_pkthdr.flowid; let a listener declare its bucket(s)the lbgroup NUMA-affinity machineryhighest application value; the modern replacement for IP_RSS_LISTEN_BUCKET, what a per-core accept loop needs, no PCBGROUPS revival
4. a generic n-tuple flow-rule APInetlink family (preferred) or an ifflowrule ioctlcxgbe, mlx5, ice backends need only adaptersthe real gap versus Linux ethtool -N; netlink is the modern, dump-friendly, extensible surface
5. a flow-steering send-tag typea new snd_tag type carrying a match spec and target queuethe TLS_RX lifecycle, lagg/vlan proxying, driver plumbinggives the kernel (NVMe-oF, iSCSI, zero-copy RX) programmable steering without designing a userland API first; composes with 4
6. runtime-writable rss_table[]an EVENTHANDLER so drivers reprogram hardware on changethe rss_config.c reader/writer pathsenables actual rebalancing; needs synchronization against the datapath readers; do after 1-2, since the get path defines the contract
The cheap wins reuse a precedent that already exists; the valuable ones are new plumbing. Bridging the existing GET ioctls through iflib copies the ifdi_i2c_req pattern verbatim. Teaching SO_REUSEPORT_LB to follow m_pkthdr.flowid is the single change with the best value-to-risk ratio, because the NUMA-aware group machinery is already there and just does not know about RSS. A generic n-tuple API and an in-kernel steering send-tag are the two that close the gap with what other stacks already expose.

What reading this map leaves you with

  • RSS in FreeBSD is boot-time, kernel-internal, and read-only. Every piece to make it programmable exists somewhere; the pieces were never joined.
  • The generic ifnet ioctls are near-dead. GET-only, implemented by a handful of non-iflib drivers, with zero userland consumers and no indirection-table command.
  • iflib is the hole in the middle. It backs nearly every modern NIC and has no RSS method, so the generic path returns EINVAL exactly where a generic set would belong.
  • All real write capability is fragmented. Five per-driver sysctl namespaces, two vendor-private filter engines, and one kTLS-welded send-tag hook, none sharing an API.
  • The socket layer can watch but not steer. It reads the bucket a packet landed in and picks listeners with a software hash that ignores the hardware's own result.
  • Ownership is the gating question. Kernel-owns or driver-owns, with no third state where userland and kernel agree - and that contract has to be settled on the get side before any set can be trusted.