source: PyTorch Blog: Low Precision Flash Attention 4: End-to-End Block-Scaled Attention for Blackwell

level: technical

pytorch released low precision flash attention 4, adding mxfp8 forward and backward passes for nvidia blackwell gpus. the kernel reaches 2.85 petaflops per second forward and 2 petaflops per second backward on large language model shapes. on internal shapes, it hits 2.54 petaflops forward and 1.58 backward, up to 1.6 times faster than bf16. the code is open sourced in the ads model kernel library.

the implementation fuses quantization into surrounding producers and uses a zero-gather jagged module where most activations stay in fp8. scale factors are managed in tmem, which is fully utilized, requiring overlapping scale storage with accumulator regions. online quantization of softmax output uses square 32 by 32 blocks, making the payload transpose invariant. backward pass quantizes ds once and reuses it for both dk and dq, avoiding extra conversions.

the work targets production training at meta for gem models. blackwell tensor cores give two to four times throughput for mxfp8 over bf16, but real workloads need careful scale factor handling. the team added an unroll-kv optimization to hide softmax latency at tile boundaries. fp16 dq reduction with static scaling was chosen to cut memory bandwidth, as dq values stay small in practice.

why it matters: this kernel lets ai teams train large models faster on blackwell gpus by using low precision attention without losing accuracy.


source: PyTorch Blog: Low Precision Flash Attention 4: End-to-End Block-Scaled Attention for Blackwell