source: Hugging Face Blog: Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI

level: technical

hugging face released @huggingface/kernels, a javascript library for loading and running optimized webgpu kernels from the hub. the initial collection includes 207 kernels covering common machine learning operations like matrix multiplication, normalization, and attention. each kernel is a versioned package with shader templates, correctness tests, and benchmark cases. the team also launched fleet, an in-browser benchmarking suite that runs kernels on user hardware and collects performance data with consent.

benchmarks on an apple m4 gpu compared the kernels against onnx runtime web. across 809 matching test cases, the new kernels were 2.57 times faster by geometric mean and 1.90 times faster at the median. some operations showed dramatic gains, such as a bilinear einsum case running over 10,000 times faster. however, the measurements excluded setup overhead like shader compilation and data transfer, and results varied by device and browser.

the kernels are published as individual repositories with manifests, metadata, and wgsl shader templates. the loader lets developers call kernels by repository id and version, keeping a stable api while implementations evolve. the collection is part of hugging face's broader kernel ecosystem, which includes cuda, rocm, and metal kernels. the team plans to connect these kernels to higher-level model tooling and expand operation coverage.

why it matters: faster browser-based ai inference could make local machine learning more practical on consumer devices without server costs.


source: Hugging Face Blog: Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI