module/zfs/zio.c

ZIO batch arrival race

On FreeBSD/arm64, the thread that drops a batch's last hold can read the list of arrived members before its own hold drop has finished. It then walks a list that is missing a member whose hold it just counted: that member's completion never runs, its parent never completes, and the txg sync thread waits in zio_wait() forever.

Symptom
The txg sync thread parks in zio_wait() under dsl_pool_sync(). Writers queue behind the txg, none of them can be killed, and a reboot hangs in sync(2).
Affected
FreeBSD on arm64, where it was observed. FreeBSD's riscv64 atomics are just as unordered but untested. Linux and FreeBSD/amd64 are not affected.
Needs
A mirror, RAIDZ or dRAID vdev whose leaves bypass the I/O scheduler. That is the default for SSD and NVMe leaves, and what scheduler=off forces.
Introduced
The "ZIO: Batch vdev children completions" change (openzfs/zfs#18921). On OpenZFS master since 2026-09-05 and FreeBSD main since 2026-09-14; in no release.
Fix
An acquire barrier before the list is walked, and a release barrier between each member's push and its hold drop.

How a batch works

A mirror, RAIDZ or dRAID parent that issues several leaf children opens a batch for them. Each child that comes back from the block layer pushes itself onto zb_arrived and drops one hold, with no lock held. The parent keeps one extra hold until it has created every child.

typedef struct zio_batch {
	zio_t		*zb_arrived;	/* lock-free LIFO of arrived members */
	uint64_t	zb_holds;	/* members not yet arrived + creator */
} zio_batch_t;

Whoever drops the last hold runs the batch: it walks zb_arrived once, runs every member's completion, then frees the batch. That walk is the only time anyone looks at the list, so a member that is not on it at that moment is never run.

What goes wrong

On arm64, FreeBSD's atomic_dec_64_nv() is a bare LSE LDADD, with no acquire, no release and no barrier around it. An LDADD can take a long time to complete, and nothing stops a plain load that comes after it in the code from completing first. In zio_batch_rele() and zio_batch_leave(), the load that comes after it is the read of zb_arrived.

parent CPU 0 zio_batch_rele() child A CPU 1 zb_arrived zb_holds drop hold LDADD in flight completes: last hold walk list read list list as read: B push A drop hold A's completion never runs B A, B 2 1 0 reads B A pushes and drops its hold inside this window time ->
Schematic, not to scale; the window is nanoseconds wide. B arrived earlier. The parent drops its hold and then reads the list, in that order in the code, but the read completes first and sees only B. Child A pushes and drops its hold while the parent's LDADD is still in flight, so the LDADD finds zb_holds at 1, takes it to 0, and the parent walks the list it had already read. A is never run.

Nothing here needs A's own two writes to be reordered: A pushes, then drops its hold, and both land in that order. The reader looked at the list before it looked at the counter, and on the hardware tested that is where every observed miss came from.

Which paths are exposed

The last hold can be dropped in three places, and every batch is walked in zio_batch_run(). What differs is what sits between the drop and the walk.

who drops the last hold where the list is read an arriving child zio_batch_arrive() the parent zio_batch_rele() a member leaving zio_batch_leave() taskq worker zio_batch_execute() dispatch, return queue lock orders the read same CPU, straight to the read same CPU, straight to the read zio_batch_run() membar_consumer(); added by the fix walk zb_arrived
Every batch is walked in zio_batch_run(). On the taskq path the worker reads the list only after taking the queue lock, which orders the read. On the two inline paths the thread that dropped the last hold goes straight to the read on the same CPU, with nothing in between in the shipped code. One acquire at the top of zio_batch_run() covers all three.

What the fix orders

@@ zio_batch_run(zio_batch_t *zb) 	zio_t *list = NULL, *zio, *next; +	/* Pairs with zio_batch_arrive(). */+	membar_consumer();+ 	for (zio = zb->zb_arrived; zio != NULL; zio = next) {@@ zio_batch_arrive(zio_t *zio) 	} while (atomic_cas_ptr(&zb->zb_arrived, head, zio) != head); +	/* Publish the arrival before dropping the hold that runs the batch. */+	membar_producer();+ 	if (atomic_dec_64_nv(&zb->zb_holds) == 0) {

The acquire goes where the list is read. At the top of zio_batch_run(), one barrier covers both inline paths. On the taskq path the queue lock already provides it, so there it is redundant. The thread that hands a batch to the taskq never reads the list and needs nothing of its own.

The release covers the other half. The architecture also lets another CPU see a member's hold drop before its push, since the two are unordered atomics to different addresses. That was not observed on the Cortex-A76 tested here, where the acquire alone removed every miss. But nothing rules it out on other arm64 cores, and no lock could make up for it: the member that pushed takes no part in any handoff and has usually returned by the time anyone walks the list. The release makes each push visible before the hold drop that follows it.

On x86 both barriers compile to compiler barriers and nothing else. On arm64 each is one DMB: one per arriving member and one per batch walked.

What it looks like

  1. The missed member never leaves ZIO_VDEV_IO_DONE. Nothing will run its completion.
  2. Its parent's count of outstanding children never reaches zero, so the parent never completes.
  3. When that parent belongs to the txg being synced, the sync thread waits in zio_wait() forever.
  4. Every writer then waits behind the txg in cv_wait(), which no signal interrupts, so nothing can be killed.
  5. reboot(8) calls sync(2), which waits on every pool, so the machine won't reboot cleanly either, even when the stuck pool isn't the root pool.
sync thread
    _cv_timedwait_sbt
    zio_wait+0x63c
    dsl_pool_sync+0x171
    spa_sync+0x7fe
    txg_sync_thread+0x4b9

a writer
    _cv_wait
    txg_wait_synced_flags
    dmu_tx_wait
    dmu_tx_assign
    zfs_write

Stacks from the reproduction described under Evidence, where one missed member was forced on purpose.

Why only FreeBSD on arm64

Platformatomic_dec_64_nv() becomesOrders the read after it
Linux, any architectureatomic64_dec_return(), fully ordered by definitionYes
FreeBSD, amd64LOCK XADD, a full barrierYes
FreeBSD, arm64LDADD (LSE) or an LDXR/STXR loop, with no acquire and no DMBNo

The Linux SPL maps the illumos atomic names onto Linux primitives that are fully ordered by definition. The FreeBSD SPL maps them onto the unsuffixed atomic(9) primitives, which promise no ordering, and on arm64 emit none. The fix asks for the ordering explicitly, so it no longer depends on which SPL is underneath.

Evidence

A litmus test runs a member's two operations on one core and the last-hold drop and list read on another, using the instructions FreeBSD emits, and counts the times the second core took the last hold but did not see the member.

initially   zb_arrived = NULL     zb_holds = 2

member, CPU 1                     last hold, CPU 2
  CAS   zb_arrived: NULL -> A       r1 = LDADD zb_holds, -1
  LDADD zb_holds, -1                r2 = LDR   zb_arrived

missed:  r1 == 1 (took the last hold)  and  r2 == NULL (no A)
BarriersLast-hold observationsMissed a member
None, as shipped8,647,3453,162,461
Release only4,147,73894
Acquire only8,592,9080
Both, the fix4,168,3350

Cortex-A76 r4p1 with LSE atomics, FreeBSD 16.0-CURRENT. Each row is ten runs of 2,000,000 iterations: five timings, each with a single CAS and with zio_batch_arrive()'s CAS loop. The miss rate swings from almost never to nearly always with timing, and the barriers change the timing too, which is why the observation counts differ between rows. The same test on amd64 never missed in about six million observations.

The acquire alone removes every miss; the release alone does not. That is why the reader side is the one drawn here, and why the fix carries both: one for what was observed, one for what the architecture still allows.

The consequence was reproduced separately, on amd64, by forcing exactly this outcome once: a batch run that sees only its own member. The pool stopped syncing with the stacks above, and the same kernel without the forced miss synced normally. Hitting it unforced takes a member finishing on another CPU while the last-hold drop is in flight, a window of nanoseconds, which fits a hang that appears after minutes or hours of load rather than at once.