block attention kernel speeds up sparse transformers on blackwell gpus
a new triton kernel for fixed-block sparse self-attention achieves up to 3.5x speedup over flash attention v2 on nvidia b200 gpus by exploiting compile-time knowledge of block-diagonal patterns.
ai