FreeBSD / networking / flow steering
Receive-side scaling and hardware flow steering exist all over the FreeBSD tree, but nowhere as one programmable thing. Every piece an operator would want - the hash key, the indirection table, an n-tuple filter, a listener that follows the queue its packets landed on - lives behind a different interface, and most of those interfaces are read-only, dead, or private to one driver.
This is the map of the whole territory: which layer owns what, which direction it moves, and where a generic API could slot in. Two companion pages go narrow instead - one walks a single mechanism, one documents a single get-and-set API. This one draws the gaps between them.
SIOCGIFRSS* ioctls are GET-only with no in-tree consumer; and iflib - which backs nearly every modern NIC - has no RSS method, so a write falls through to ether_ioctl() and returns EINVAL. Real programmability lives in the boot-time config the kernel owns, five per-driver sysctls, two vendor-private filter engines, and one send-tag type meant for kTLS. None of them share an interface.| layer / API | direction | who implements it | reachable from userland? | notes |
|---|---|---|---|---|
net.inet.rss (rss_config.c) | read-only | kernel core, at boot | yes, sysctl read | key CTLFLAG_RD (write returns EINVAL), bits RDTUN, bucket_mapping RD, basecpu const; kernel owns key and LUT under options RSS |
SIOCGIFRSSKEY / SIOCGIFRSSHASH | read-only | mlx4en, mlx5en, mana, hn, qlnx | yes, but no consumer | no-privilege GET block in ifhwioctl(); RSS_KEYLEN 128 vs RSS_KEYSIZE 40 for one concept; only in-tree caller is hn mirroring its own VF |
| indirection-table ioctl | dead / missing | - | no | the one thing an operator wants to change had no generic representation in the baseline tree |
iflib (ifdi_if.m) | dead | backs ixgbe, ice, ixl, em, igc, bnxt, aq, vmx, enic, axgbe... | no | no ifdi_ RSS method; unknown ioctl falls to ether_ioctl() -> EINVAL; ifdi_i2c_req is the thin-method precedent it never got |
| per-driver sysctls | write | hn, ixl, bnxt, ena, cxgbe | yes, root | five incompatible namespaces; hn writable only #ifndef RSS; cxgbe rss_base/rss_size read-only |
IP_FLOWID / IP_RECVFLOWID / IP_RECVRSSBUCKETID | observe | netinet socket options (90 / 93 / 94) | yes | read the flowid and bucket of a received datagram; nothing lets an app request a bucket |
SO_REUSEPORT_LB | read-only | in_pcb lbgroup lookup | yes | INP_PCBLBGROUP_PKTHASH jhashes the 3-tuple and mods by member count; NUMA-domain aware but ignores m_pkthdr.flowid; replaces the removed IP_RSS_LISTEN_BUCKET (slot reserved) |
cxgbe t4_filter | vendor-private (write) | cxgbe | yes, via cxgbetool / /dev/t4nex | CHELSIO_T4_{GET,SET,DEL}_FILTER, filter mode/mask; the only user-programmable hardware filter API in the tree |
mlx5 fs_core | vendor-private | mlx5 | no | full flow-steering core plus a per-inpcb entry point (mlx5e_accel_fs_add_inpcb); reachable only from kTLS RX |
IF_SND_TAG_TYPE_TLS_RX | write (kTLS-only) | snd_tag framework (type 4) | only via kTLS | the only generic per-flow RX steering hook; already has lagg/vlan proxying, a refcounted lifecycle, and an inpcb in the alloc params - just hardcoded to kTLS |
IP_RSS_LISTEN_BUCKET and PCBGROUP KPI mean FreeBSD once had an application-facing "bind this listener to bucket N" and now has nothing.options RSS the kernel's affinity plumbing depends on the hardware table matching rss_table[]; without it, the driver is free and the stack is blind. hn's answer - refuse writes when RSS is compiled in - is defensible but binary. A real answer reports a three-state owner on the get and either fails the set or transitions ownership explicitly.| insertion point | shape | reuses | note |
|---|---|---|---|
| 1. bridge the read ioctls into iflib | thin ifdi_ methods behind the existing GET ioctls | ifdi_i2c_req, ifdi_get_downreason | closes the hole in the middle; a generic ioctl backed by a per-driver method is already the established pattern |
| 2. add the indirection table | a table structure plus get/set ioctls, with an owner field | the ifnet ioctl idiom | gives the one control an operator actually wants a generic representation; owner field carries the three-state answer |
| 3. make SO_REUSEPORT_LB RSS-aware | select by m_pkthdr.flowid; let a listener declare its bucket(s) | the lbgroup NUMA-affinity machinery | highest application value; the modern replacement for IP_RSS_LISTEN_BUCKET, what a per-core accept loop needs, no PCBGROUPS revival |
| 4. a generic n-tuple flow-rule API | netlink family (preferred) or an ifflowrule ioctl | cxgbe, mlx5, ice backends need only adapters | the real gap versus Linux ethtool -N; netlink is the modern, dump-friendly, extensible surface |
| 5. a flow-steering send-tag type | a new snd_tag type carrying a match spec and target queue | the TLS_RX lifecycle, lagg/vlan proxying, driver plumbing | gives the kernel (NVMe-oF, iSCSI, zero-copy RX) programmable steering without designing a userland API first; composes with 4 |
| 6. runtime-writable rss_table[] | an EVENTHANDLER so drivers reprogram hardware on change | the rss_config.c reader/writer paths | enables actual rebalancing; needs synchronization against the datapath readers; do after 1-2, since the get path defines the contract |
ifdi_i2c_req pattern verbatim. Teaching SO_REUSEPORT_LB to follow m_pkthdr.flowid is the single change with the best value-to-risk ratio, because the NUMA-aware group machinery is already there and just does not know about RSS. A generic n-tuple API and an in-kernel steering send-tag are the two that close the gap with what other stacks already expose.EINVAL exactly where a generic set would belong.