x86-64, arm64, Linux, FreeBSD
Memory ordering on x86-64 and arm64
Each CPU runs its own code in order as far as it can tell. Ordering is about what the other CPUs see, and that is where the two architectures differ. This page shows how each one behaves, how code asks for order on each, and where the ZIO batch race fits.
- x86-64 keeps three of the four orderings, and every atomic read-modify-write on it is a full barrier. Code that forgets to ask for ordering usually works anyway.
- arm64 keeps none of them between different addresses, and an atomic carries only the ordering its suffix asks for. Code that forgets to ask gets none.
- The batch race is the arm64 case. The reader's list read finished before its own hold drop. x86 prevents that without being asked. Linux's arm64 atomics prevent it too, because Linux asks for ordering by default.
Program order and what other CPUs see
A CPU always sees its own loads and stores take effect in program order. Other CPUs only see the results in memory, and they can see them in a different order. The ordering rules say which pairs of accesses to two different addresses the other CPUs are guaranteed to see in program order. Accesses to the same address stay in order on every architecture here.
Any two accesses in program order form one of four pairs, depending on whether each is a load or a store. Atomic read-modify-write instructions get their own row, because that is where x86 and arm64 differ most.
| Earlier, then later | Sequentially consistent | x86-64 | arm64 | riscv64 |
|---|---|---|---|---|
| load, then load | kept | kept | may reorder | may reorder |
| load, then store | kept | kept | may reorder | may reorder |
| store, then store | kept | kept | may reorder | may reorder |
| store, then load | kept | may reorder | may reorder | may reorder |
| atomic read-modify-write with no ordering suffix | kept | full barrier | may reorder | may reorder |
Sequential consistency is the reference model, in which nothing reorders. No mainstream CPU works that way, because it would stall on every store. On arm64 and riscv64 a dependency can still keep a pair in order. A load whose address comes from an earlier load always stays after it.
x86-64 queues its stores
An x86 core puts each store into a store buffer and keeps going. The buffer drains to the cache in program order, so other CPUs see its stores in order. Loads go straight to the cache and can finish while older stores are still waiting in the buffer. The one pair x86 lets move is a store followed by a load.
A LOCK-prefixed instruction waits for the store buffer to drain, and no load or store can pass it, which makes it a full barrier. Every atomic read-modify-write on x86 is locked (LOCK XADD, LOCK CMPXCHG, and XCHG, which locks implicitly), so every atomic is a full barrier whether the code asked for ordering or not. MFENCE gives the same barrier without an atomic.
arm64 lets anything finish first
An arm64 core only has to keep its own results right. Between different addresses it can let any of the four pairs finish out of order. Loads can be answered early from the cache. Stores can drain in any order. An LSE atomic such as LDADD can be sent out to be performed in the memory system, away from the core, so it can finish well after instructions that come after it.
The batch handoff is the classic message-passing test. The member publishes and then signals. The reader sees the signal and then reads what was published.
initially zb_arrived = NULL zb_holds = 2 member reader, drops the last hold CAS zb_arrived: NULL -> A publish r1 = atomic_dec(zb_holds) sees signal atomic_dec(zb_holds) signal r2 = zb_arrived reads outcome in question r1 == 0 and r2 == NULL x86-64 never arm64 allowed, and observed
On x86 both halves stay in order. The member's two writes drain in order, and the reader's atomic is a full barrier. On arm64 either half can come apart. On the Cortex-A76 that was tested it was always the reader's half, where the load of zb_arrived finished before the LDADD in front of it.
The reader's two instructions on each build
The reader drops its hold and then reads the list. The only question is whether the read can finish before the drop does.
DMB LD that the fix adds, or by the acquire half of Linux's LDADDAL.Asking for order
Code asks for order with fences, or with instructions that carry ordering themselves. Most of them block only one direction.
| To keep | x86-64 | arm64 |
|---|---|---|
| later accesses after a load (acquire) | a plain load already does | LDAR, or LDADDA / CASA / SWPA |
| earlier accesses before a store (release) | a plain store already does | STLR, or LDADDL / CASL / SWPL |
| both, around an atomic | every locked atomic, always | LDADDAL / CASAL / SWPAL |
| loads before later loads and stores | nothing extra | DMB LD, or DMB ISHLD |
| stores before later stores | nothing extra | DMB ST, or DMB ISHST |
| everything, including a store before a later load | MFENCE, or any LOCK | DMB SY, or DMB ISH |
x86 needs no extra instruction for acquire or release, but the compiler must still keep the order. That is why some of the kernel macros below emit nothing on x86 and still exist. They are compiler barriers.
Each DMB comes in a plain form and an ISH form. Up to revision M.a.a of the Arm ARM, the ISH form limits the ordering to the Inner Shareable domain, which holds all the CPUs an OS runs on, and the plain form extends it to the whole system, devices included. From revision M.b the ISH forms are deprecated and defined to behave exactly like the plain ones. For ordering between CPUs, which is all the batch handoff needs, the two definitions give the same guarantee. Linux still emits the ISH forms and FreeBSD emits the plain ones.
An atomic's acquire half applies only when its result goes into a real register. An atomic that discards its result into XZR, which disassembles as STADD for an add, gets no acquire, and a later DMB LD doesn't order it either. Linux compiles its atomics that return nothing to exactly that. The ZFS reader is unaffected because atomic_dec_64_nv() uses its result.
One source line, four builds
The OpenZFS SPL gives ZFS one set of names and maps them onto each kernel's own primitives. These are the four calls the batch handoff uses, on Linux and then on FreeBSD, with the instruction each becomes and what that instruction orders. Every instruction here was read back from a built module: Ubuntu 24.04's zfs.ko for Linux, and FreeBSD's own zfs.ko on amd64 and arm64.
| SPL call | Linux primitive | x86-64 | arm64 |
|---|---|---|---|
atomic_cas_ptr() | atomic64_cmpxchg() | LOCK CMPXCHGfull barrier | CASALfull barrier |
atomic_dec_64_nv() | atomic64_dec_return() | LOCK XADDfull barrier | LDADDALfull barrier |
membar_producer() | smp_wmb() | no instructioncompiler barrier only | DMB ISHSTstores before stores |
membar_consumer() | smp_rmb() | no instructioncompiler barrier only | DMB ISHLDloads before loads and stores |
| SPL call | FreeBSD primitive | amd64 | arm64 |
|---|---|---|---|
atomic_cas_ptr() | atomic_fcmpset_64() | LOCK CMPXCHGfull barrier | CASno ordering |
atomic_dec_64_nv() | atomic_fetchadd_64() | LOCK XADDfull barrier | LDADDno ordering |
membar_producer() | atomic_thread_fence_rel() | no instructioncompiler barrier only | DMB SYfull barrier |
membar_consumer() | atomic_thread_fence_acq() | no instructioncompiler barrier only | DMB LDloads before loads and stores |
The arm64 columns show LSE atomics, which both kernels pick at run time when the CPU has them. Without LSE, Linux emits LDXR/STLXR followed by DMB ISH and stays fully ordered, while FreeBSD emits a bare LDXR/STXR loop.
The barriers differ from build to build. On arm64, FreeBSD's membar_producer() is a full barrier and Linux's orders stores against stores and nothing more. The handoff needs only what they share. The member needs its push ordered before its hold drop, which is a store before a store. The reader needs its hold drop ordered before its list read, which is a load before a load.
The consumer barriers look different but do the same job. Linux emits DMB ISHLD and FreeBSD emits DMB LD. Current revisions of the Arm ARM define them as the same barrier. Older revisions limit ISHLD to the CPUs, which is everyone the handoff involves.
What they share is not always enough. smp_wmb() doesn't order earlier loads, so it can't stand in for a release when a reference-count drop must follow earlier reads of the object. On Linux, membar_producer() is smp_wmb().
The atomics differ more. FreeBSD asked for no ordering on amd64 and got full barriers anyway, because x86 has no other kind of atomic. The same source on arm64 got exactly what it asked for, which was no ordering at all.
How each API names it
| API | The plain name means | Stronger | Weaker |
|---|---|---|---|
Linux atomic_t | fully ordered for read-modify-writes that return a value; unordered for the ones that don't, and for plain reads and sets | smp_mb__before_atomic() and smp_mb__after_atomic() around the ones that don't | _acquire, _release and _relaxed suffixes |
C11 stdatomic.h | sequentially consistent | already the strongest | a memory_order argument |
FreeBSD atomic(9) | no ordering | _acq and _rel suffixes (none for fetchadd or swap), atomic_thread_fence_*() | already the weakest |
| illumos and the OpenZFS SPL | no ordering promised | membar_producer(), membar_consumer(), membar_sync() | already the weakest |
The OpenZFS SPL maps the illumos names onto Linux's plain names on Linux and onto FreeBSD's plain names on FreeBSD. The spelling is the same in both. On Linux it means fully ordered, and on FreeBSD it means unordered.
The batch handoff on each build
| Build | Member keeps push before drop | Reader keeps read after drop | Can miss a member |
|---|---|---|---|
| Linux x86-64 | yes LOCK | yes LOCK | no |
| FreeBSD amd64 | yes LOCK | yes LOCK | no |
| Linux arm64 | yes release half of LDADDAL | yes acquire half of LDADDAL | no |
| FreeBSD arm64, shipped | not guaranteed never observed | no observed | yes |
| FreeBSD arm64, with the fix | yes DMB SY | yes DMB LD | no |
What ordering costs
On arm64 the same ordering can be bought two ways at very different prices. This is one hold and one drop on a reference count, on a Cortex-A76, with no other CPU touching the counter.
| Ordering on the hold and the drop | ns per pair | vs none |
|---|---|---|
| none, as FreeBSD ships | 17.37 | |
| DMB fences on both | 36.07 | +108% |
| DMB fences, acquire only on the last drop | 32.73 | +88% |
LDADDAL on both | 17.37 | +0.0% |
LDADDL, acquire only on the last drop | 18.70 | +7.7% |
With two or more CPUs contending for the counter's cache line, every row lands within 2% of the others, because moving the line costs far more than any barrier. On x86 the table would have one row. Every atomic is a locked full barrier, and there is no unordered one to fall back to.
Ordering and atomicity are separate problems
Barriers control the order in which one CPU's accesses become visible. They do nothing to stop two CPUs from interleaving. The same batching change turned two plain stores in vdev_queue_io_done() into an unlocked check-then-store of two fields. Two CPUs can both pass the check and both store, on x86 as easily as on arm64, and the two fields can end up describing different I/Os. No barrier fixes that. It needs an atomic read-modify-write such as a CAS loop, or a lock.