module/zfs/zio.c
ZIO batch arrival race
On FreeBSD/arm64, the thread that drops a batch's last hold can read the list of arrived members before its own hold drop has finished. It then walks a list that is missing a member whose hold it just counted: that member's completion never runs, its parent never completes, and the txg sync thread waits in zio_wait() forever.
- Symptom
- The txg sync thread parks in
zio_wait()underdsl_pool_sync(). Writers queue behind the txg, none of them can be killed, and a reboot hangs insync(2). - Affected
- FreeBSD on arm64, where it was observed. FreeBSD's riscv64 atomics are just as unordered but untested. Linux and FreeBSD/amd64 are not affected.
- Needs
- A mirror, RAIDZ or dRAID vdev whose leaves bypass the I/O scheduler. That is the default for SSD and NVMe leaves, and what
scheduler=offforces. - Introduced
- The "ZIO: Batch vdev children completions" change (openzfs/zfs#18921). On OpenZFS master since 2026-09-05 and FreeBSD main since 2026-09-14; in no release.
- Fix
- An acquire barrier before the list is walked, and a release barrier between each member's push and its hold drop.
How a batch works
A mirror, RAIDZ or dRAID parent that issues several leaf children opens a batch for them. Each child that comes back from the block layer pushes itself onto zb_arrived and drops one hold, with no lock held. The parent keeps one extra hold until it has created every child.
typedef struct zio_batch {
zio_t *zb_arrived; /* lock-free LIFO of arrived members */
uint64_t zb_holds; /* members not yet arrived + creator */
} zio_batch_t;
Whoever drops the last hold runs the batch: it walks zb_arrived once, runs every member's completion, then frees the batch. That walk is the only time anyone looks at the list, so a member that is not on it at that moment is never run.
What goes wrong
On arm64, FreeBSD's atomic_dec_64_nv() is a bare LSE LDADD, with no acquire, no release and no barrier around it. An LDADD can take a long time to complete, and nothing stops a plain load that comes after it in the code from completing first. In zio_batch_rele() and zio_batch_leave(), the load that comes after it is the read of zb_arrived.
LDADD is still in flight, so the LDADD finds zb_holds at 1, takes it to 0, and the parent walks the list it had already read. A is never run.Nothing here needs A's own two writes to be reordered: A pushes, then drops its hold, and both land in that order. The reader looked at the list before it looked at the counter, and on the hardware tested that is where every observed miss came from.
Which paths are exposed
The last hold can be dropped in three places, and every batch is walked in zio_batch_run(). What differs is what sits between the drop and the walk.
zio_batch_run(). On the taskq path the worker reads the list only after taking the queue lock, which orders the read. On the two inline paths the thread that dropped the last hold goes straight to the read on the same CPU, with nothing in between in the shipped code. One acquire at the top of zio_batch_run() covers all three.- A child arriving from the block layer that drops the last hold hands the batch to a taskq and returns without reading the list. The worker reads it only after taking the queue's lock, and the lock orders that read.
- The parent, in
zio_batch_rele()once it has created every child, and a member leaving the batch, inzio_batch_leave(), go straight from the drop to the walk on the same CPU. In the shipped code nothing sits between them. That is the shape in the timing diagram above, and it is the one that fails.
What the fix orders
@@ zio_batch_run(zio_batch_t *zb) zio_t *list = NULL, *zio, *next; + /* Pairs with zio_batch_arrive(). */+ membar_consumer();+ for (zio = zb->zb_arrived; zio != NULL; zio = next) {@@ zio_batch_arrive(zio_t *zio) } while (atomic_cas_ptr(&zb->zb_arrived, head, zio) != head); + /* Publish the arrival before dropping the hold that runs the batch. */+ membar_producer();+ if (atomic_dec_64_nv(&zb->zb_holds) == 0) {
The acquire goes where the list is read. At the top of zio_batch_run(), one barrier covers both inline paths. On the taskq path the queue lock already provides it, so there it is redundant. The thread that hands a batch to the taskq never reads the list and needs nothing of its own.
The release covers the other half. The architecture also lets another CPU see a member's hold drop before its push, since the two are unordered atomics to different addresses. That was not observed on the Cortex-A76 tested here, where the acquire alone removed every miss. But nothing rules it out on other arm64 cores, and no lock could make up for it: the member that pushed takes no part in any handoff and has usually returned by the time anyone walks the list. The release makes each push visible before the hold drop that follows it.
On x86 both barriers compile to compiler barriers and nothing else. On arm64 each is one DMB: one per arriving member and one per batch walked.
What it looks like
- The missed member never leaves
ZIO_VDEV_IO_DONE. Nothing will run its completion. - Its parent's count of outstanding children never reaches zero, so the parent never completes.
- When that parent belongs to the txg being synced, the sync thread waits in
zio_wait()forever. - Every writer then waits behind the txg in
cv_wait(), which no signal interrupts, so nothing can be killed. reboot(8)callssync(2), which waits on every pool, so the machine won't reboot cleanly either, even when the stuck pool isn't the root pool.
sync thread
_cv_timedwait_sbt
zio_wait+0x63c
dsl_pool_sync+0x171
spa_sync+0x7fe
txg_sync_thread+0x4b9
a writer
_cv_wait
txg_wait_synced_flags
dmu_tx_wait
dmu_tx_assign
zfs_write
Stacks from the reproduction described under Evidence, where one missed member was forced on purpose.
Why only FreeBSD on arm64
| Platform | atomic_dec_64_nv() becomes | Orders the read after it |
|---|---|---|
| Linux, any architecture | atomic64_dec_return(), fully ordered by definition | Yes |
| FreeBSD, amd64 | LOCK XADD, a full barrier | Yes |
| FreeBSD, arm64 | LDADD (LSE) or an LDXR/STXR loop, with no acquire and no DMB | No |
The Linux SPL maps the illumos atomic names onto Linux primitives that are fully ordered by definition. The FreeBSD SPL maps them onto the unsuffixed atomic(9) primitives, which promise no ordering, and on arm64 emit none. The fix asks for the ordering explicitly, so it no longer depends on which SPL is underneath.
Evidence
A litmus test runs a member's two operations on one core and the last-hold drop and list read on another, using the instructions FreeBSD emits, and counts the times the second core took the last hold but did not see the member.
initially zb_arrived = NULL zb_holds = 2 member, CPU 1 last hold, CPU 2 CAS zb_arrived: NULL -> A r1 = LDADD zb_holds, -1 LDADD zb_holds, -1 r2 = LDR zb_arrived missed: r1 == 1 (took the last hold) and r2 == NULL (no A)
| Barriers | Last-hold observations | Missed a member |
|---|---|---|
| None, as shipped | 8,647,345 | 3,162,461 |
| Release only | 4,147,738 | 94 |
| Acquire only | 8,592,908 | 0 |
| Both, the fix | 4,168,335 | 0 |
Cortex-A76 r4p1 with LSE atomics, FreeBSD 16.0-CURRENT. Each row is ten runs of 2,000,000 iterations: five timings, each with a single CAS and with zio_batch_arrive()'s CAS loop. The miss rate swings from almost never to nearly always with timing, and the barriers change the timing too, which is why the observation counts differ between rows. The same test on amd64 never missed in about six million observations.
The acquire alone removes every miss; the release alone does not. That is why the reader side is the one drawn here, and why the fix carries both: one for what was observed, one for what the architecture still allows.
The consequence was reproduced separately, on amd64, by forcing exactly this outcome once: a batch run that sees only its own member. The pool stopped syncing with the stacks above, and the same kernel without the forced miss synced normally. Hitting it unforced takes a member finishing on another CPU while the last-hold drop is in flight, a window of nanoseconds, which fits a hang that appears after minutes or hours of load rather than at once.