Memory ordering on x86-64 and arm64

x86-64, arm64, Linux, FreeBSD

Memory ordering on x86-64 and arm64

Each CPU runs its own code in order as far as it can tell. Ordering is about what the other CPUs see, and that is where the two architectures differ. This page shows how each one behaves, how code asks for order on each, and where the ZIO batch race fits.

Program order and what other CPUs see

A CPU always sees its own loads and stores take effect in program order. Other CPUs only see the results in memory, and they can see them in a different order. The ordering rules say which pairs of accesses to two different addresses the other CPUs are guaranteed to see in program order. Accesses to the same address stay in order on every architecture here.

Any two accesses in program order form one of four pairs, depending on whether each is a load or a store. Atomic read-modify-write instructions get their own row, because that is where x86 and arm64 differ most.

Earlier, then laterSequentially consistentx86-64arm64riscv64
load, then loadkeptkeptmay reordermay reorder
load, then storekeptkeptmay reordermay reorder
store, then storekeptkeptmay reordermay reorder
store, then loadkeptmay reordermay reordermay reorder
atomic read-modify-write with no ordering suffixkeptfull barriermay reordermay reorder

Sequential consistency is the reference model, in which nothing reorders. No mainstream CPU works that way, because it would stall on every store. On arm64 and riscv64 a dependency can still keep a pair in order. A load whose address comes from an earlier load always stays after it.

x86-64 queues its stores

An x86 core puts each store into a store buffer and keeps going. The buffer drains to the cache in program order, so other CPUs see its stores in order. Loads go straight to the cache and can finish while older stores are still waiting in the buffer. The one pair x86 lets move is a store followed by a load.

CPU 0 program x = 1; r0 = y; store load store buffer, drains in order x = 1 drains later reads y = 0 CPU 1 program y = 1; r1 = x; store load store buffer, drains in order y = 1 drains later reads x = 0 caches and memory x = 0 y = 0 result: r0 = 0 and r1 = 0
The store-buffering test. Each CPU stores to one variable and then loads the other. Both loads run while both stores are still in the buffers, so both read 0. No interleaving of the two programs gives that result, and x86 allows it.

A LOCK-prefixed instruction waits for the store buffer to drain, and no load or store can pass it, which makes it a full barrier. Every atomic read-modify-write on x86 is locked (LOCK XADD, LOCK CMPXCHG, and XCHG, which locks implicitly), so every atomic is a full barrier whether the code asked for ordering or not. MFENCE gives the same barrier without an atomic.

arm64 lets anything finish first

An arm64 core only has to keep its own results right. Between different addresses it can let any of the four pairs finish out of order. Loads can be answered early from the cache. Stores can drain in any order. An LSE atomic such as LDADD can be sent out to be performed in the memory system, away from the core, so it can finish well after instructions that come after it.

The batch handoff is the classic message-passing test. The member publishes and then signals. The reader sees the signal and then reads what was published.

initially   zb_arrived = NULL     zb_holds = 2

member                                   reader, drops the last hold
  CAS zb_arrived: NULL -> A    publish     r1 = atomic_dec(zb_holds)    sees signal
  atomic_dec(zb_holds)         signal      r2 = zb_arrived              reads

outcome in question   r1 == 0 and r2 == NULL
  x86-64              never
  arm64               allowed, and observed

On x86 both halves stay in order. The member's two writes drain in order, and the reader's atomic is a full barrier. On arm64 either half can come apart. On the Cortex-A76 that was tested it was always the reader's half, where the load of zb_arrived finished before the LDADD in front of it.

The reader's two instructions on each build

The reader drops its hold and then reads the list. The only question is whether the read can finish before the drop does.

build and instructions time -> read after drop? x86-64, either kernel LOCK XADD, then MOV drop hold read list held by the LOCK yes FreeBSD arm64, shipped LDADD, then LDR drop hold read list finishes before the drop no FreeBSD arm64, fixed LDADD, DMB LD, LDR drop hold read list held by DMB LD yes Linux arm64 LDADDAL, then LDR drop hold read list held by the acquire half of LDADDAL yes
Schematic, not to scale. The dotted line marks when the hold drop completes. On FreeBSD/arm64 as shipped the read finishes before it. On every other build the read is held back, by the LOCK on x86, by the DMB LD that the fix adds, or by the acquire half of Linux's LDADDAL.

Asking for order

Code asks for order with fences, or with instructions that carry ordering themselves. Most of them block only one direction.

acquire LDAR, LDADDA, CASA earlier access earlier access load-acquire later access later access earlier accesses may sink below it later accesses stay below it release STLR, LDADDL, CASL earlier access earlier access store-release later access later access earlier accesses stay above it later accesses may rise above it full barrier DMB SY or ISH, MFENCE, any LOCK earlier access earlier access full fence later access later access earlier accesses stay above it later accesses stay below it
Green bars mark a direction the gate blocks; red dashed arrows mark a direction accesses may still move. An acquire keeps later accesses below it. A release keeps earlier accesses above it. A full barrier does both. In the batch handoff, the member's hold drop has to act as a release so its push stays above it, and the reader's hold drop has to act as an acquire so its list read stays below it.
To keepx86-64arm64
later accesses after a load (acquire)a plain load already doesLDAR, or LDADDA / CASA / SWPA
earlier accesses before a store (release)a plain store already doesSTLR, or LDADDL / CASL / SWPL
both, around an atomicevery locked atomic, alwaysLDADDAL / CASAL / SWPAL
loads before later loads and storesnothing extraDMB LD, or DMB ISHLD
stores before later storesnothing extraDMB ST, or DMB ISHST
everything, including a store before a later loadMFENCE, or any LOCKDMB SY, or DMB ISH

x86 needs no extra instruction for acquire or release, but the compiler must still keep the order. That is why some of the kernel macros below emit nothing on x86 and still exist. They are compiler barriers.

Each DMB comes in a plain form and an ISH form. Up to revision M.a.a of the Arm ARM, the ISH form limits the ordering to the Inner Shareable domain, which holds all the CPUs an OS runs on, and the plain form extends it to the whole system, devices included. From revision M.b the ISH forms are deprecated and defined to behave exactly like the plain ones. For ordering between CPUs, which is all the batch handoff needs, the two definitions give the same guarantee. Linux still emits the ISH forms and FreeBSD emits the plain ones.

An atomic's acquire half applies only when its result goes into a real register. An atomic that discards its result into XZR, which disassembles as STADD for an add, gets no acquire, and a later DMB LD doesn't order it either. Linux compiles its atomics that return nothing to exactly that. The ZFS reader is unaffected because atomic_dec_64_nv() uses its result.

One source line, four builds

The OpenZFS SPL gives ZFS one set of names and maps them onto each kernel's own primitives. These are the four calls the batch handoff uses, on Linux and then on FreeBSD, with the instruction each becomes and what that instruction orders. Every instruction here was read back from a built module: Ubuntu 24.04's zfs.ko for Linux, and FreeBSD's own zfs.ko on amd64 and arm64.

SPL callLinux primitivex86-64arm64
atomic_cas_ptr()atomic64_cmpxchg()
LOCK CMPXCHGfull barrier
CASALfull barrier
atomic_dec_64_nv()atomic64_dec_return()
LOCK XADDfull barrier
LDADDALfull barrier
membar_producer()smp_wmb()
no instructioncompiler barrier only
DMB ISHSTstores before stores
membar_consumer()smp_rmb()
no instructioncompiler barrier only
DMB ISHLDloads before loads and stores
SPL callFreeBSD primitiveamd64arm64
atomic_cas_ptr()atomic_fcmpset_64()
LOCK CMPXCHGfull barrier
CASno ordering
atomic_dec_64_nv()atomic_fetchadd_64()
LOCK XADDfull barrier
LDADDno ordering
membar_producer()atomic_thread_fence_rel()
no instructioncompiler barrier only
DMB SYfull barrier
membar_consumer()atomic_thread_fence_acq()
no instructioncompiler barrier only
DMB LDloads before loads and stores

The arm64 columns show LSE atomics, which both kernels pick at run time when the CPU has them. Without LSE, Linux emits LDXR/STLXR followed by DMB ISH and stays fully ordered, while FreeBSD emits a bare LDXR/STXR loop.

The barriers differ from build to build. On arm64, FreeBSD's membar_producer() is a full barrier and Linux's orders stores against stores and nothing more. The handoff needs only what they share. The member needs its push ordered before its hold drop, which is a store before a store. The reader needs its hold drop ordered before its list read, which is a load before a load.

The consumer barriers look different but do the same job. Linux emits DMB ISHLD and FreeBSD emits DMB LD. Current revisions of the Arm ARM define them as the same barrier. Older revisions limit ISHLD to the CPUs, which is everyone the handoff involves.

What they share is not always enough. smp_wmb() doesn't order earlier loads, so it can't stand in for a release when a reference-count drop must follow earlier reads of the object. On Linux, membar_producer() is smp_wmb().

The atomics differ more. FreeBSD asked for no ordering on amd64 and got full barriers anyway, because x86 has no other kind of atomic. The same source on arm64 got exactly what it asked for, which was no ordering at all.

How each API names it

APIThe plain name meansStrongerWeaker
Linux atomic_tfully ordered for read-modify-writes that return a value; unordered for the ones that don't, and for plain reads and setssmp_mb__before_atomic() and smp_mb__after_atomic() around the ones that don't_acquire, _release and _relaxed suffixes
C11 stdatomic.hsequentially consistentalready the strongesta memory_order argument
FreeBSD atomic(9)no ordering_acq and _rel suffixes (none for fetchadd or swap), atomic_thread_fence_*()already the weakest
illumos and the OpenZFS SPLno ordering promisedmembar_producer(), membar_consumer(), membar_sync()already the weakest

The OpenZFS SPL maps the illumos names onto Linux's plain names on Linux and onto FreeBSD's plain names on FreeBSD. The spelling is the same in both. On Linux it means fully ordered, and on FreeBSD it means unordered.

The batch handoff on each build

BuildMember keeps push before dropReader keeps read after dropCan miss a member
Linux x86-64yes LOCKyes LOCKno
FreeBSD amd64yes LOCKyes LOCKno
Linux arm64yes release half of LDADDALyes acquire half of LDADDALno
FreeBSD arm64, shippednot guaranteed never observedno observedyes
FreeBSD arm64, with the fixyes DMB SYyes DMB LDno

What ordering costs

On arm64 the same ordering can be bought two ways at very different prices. This is one hold and one drop on a reference count, on a Cortex-A76, with no other CPU touching the counter.

Ordering on the hold and the dropns per pairvs none
none, as FreeBSD ships17.37
DMB fences on both36.07+108%
DMB fences, acquire only on the last drop32.73+88%
LDADDAL on both17.37+0.0%
LDADDL, acquire only on the last drop18.70+7.7%

With two or more CPUs contending for the counter's cache line, every row lands within 2% of the others, because moving the line costs far more than any barrier. On x86 the table would have one row. Every atomic is a locked full barrier, and there is no unordered one to fall back to.

Ordering and atomicity are separate problems

Barriers control the order in which one CPU's accesses become visible. They do nothing to stop two CPUs from interleaving. The same batching change turned two plain stores in vdev_queue_io_done() into an unlocked check-then-store of two fields. Two CPUs can both pass the check and both store, on x86 as easily as on arm64, and the two fields can end up describing different I/Os. No barrier fixes that. It needs an atomic read-modify-write such as a CAS loop, or a lock.