metalworkingGitHub ↗
/machine/special-paths

Special Paths

Two hardware asymmetries that reshape kernels on this platform: a fast exp2 path, and float atomics that don't really exist.

The fast exp2

CUDA equivalent: the SFU's MUFU.EX2. Both platforms exponentiate in base 2 at hardware speed; what differs is how much the local kernels lean on it. Every attention implementation in this glossary computes softmax in base 2: fold log₂e into the scale factor once, then use fast::exp2 in the hot loop.

  const AccumType scale = params->scale * M_LOG2E_F;

Source: MLX steel_attention.h:166↗

This is the same math as exp with one fewer multiply per element, and it is guaranteed to take the hardware path. See online softmax for where this lands in the algorithm.

Emulated FP32 atomics

CUDA equivalent: atomicAdd(float*) has been a cheap hardware instruction since Kepler, and CUDA kernels use it casually for cross-block accumulation, gradient reduction, histogram bins. Unlearn that here. Apple hardware emulates float atomics (compare-and-swap loops), and they are slow enough to dictate architecture:

  • metal-flash-attention splits its backward pass into two kernels, dQ in one and dK/dV in the other, specifically so that no output ever needs atomic accumulation from multiple threadgroups.
  • Reduction-shaped problems prefer the two-pass pattern: partial results to a scratch buffer, then a small combine kernel. See the decode attention kernels, where llama.cpp splits the KV cache across simdgroups and merges partial softmaxes in a second kernel rather than accumulating atomically.

Integer atomics exist and are usable; it's specifically 32-bit float atomics that are a trap. If your port from CUDA contains atomicAdd on floats in a hot path, that's the first thing to redesign, usually into a split-and-combine or a fusion that makes the accumulation local to one simdgroup, where a simd_sum does it in registers.