source: PyTorch Blog: Helion x 🤗 HF Kernels: Building and Shipping Out-of-the-box Performant Kernels

level: technical

meta's helion dsl and hugging face's kernels project now work together. developers can write high-level tiled kernels in helion, autotune them for specific hardware, and ship the source plus pre-tuned configs on the hugging face hub. users load kernels with a simple get_kernel call, avoiding dependency issues. the workflow uses kernel-builder to scaffold, build, and upload noarch kernels that compile on first use.

the autotuner searches tile sizes and lowering strategies like memory access patterns and loop order. a runner script collects, measures, and builds a decision tree of configs that keeps each shape within 1% of its best performance, up to eight configs. shipped attention kernels beat pytorch's sdpa on 19 of 19 tuned shapes with a 1.20 geomean speedup, and on 9 of 10 held-out shapes with 1.17. linear attention kernels beat flash-linear-attention on all seven variants with a 1.41 geomean speedup on device time.

pre-tuned configs eliminate cold-start tuning on user machines. helion kernels are plain python and noarch, so only source and config trees are shipped. the kernels project standardizes packaging for both ahead-of-time and just-in-time kernels, with prebuilt binaries for popular kernels like flash attention 3. this reduces build times and fragmentation in kernel distribution, making high-performance kernels more accessible to the machine learning community.

why it matters: data scientists can now use pre-tuned, high-performance kernels without manual tuning or dependency management, speeding up model training and inference.


source: PyTorch Blog: Helion x 🤗 HF Kernels: Building and Shipping Out-of-the-box Performant Kernels