source: PyTorch Blog: Hardware-Agnostic Models in vLLM
level: technical
vllm is adding a set of hardware-agnostic layers to preserve support for diverse hardware and models. the project is shifting to flat, hardware-specific model definitions for frontier gpus, which break fullgraph torch.compile compatibility. this change risks leaving out-of-tree accelerators, older gpus, and exotic models behind. the new layers aim to keep vllm portable without slowing down performance work on the latest hardware.
on nvidia h100 gpus, the hardware-agnostic layers achieve total token throughput within 3.4% of the native implementation, measured as a geometric mean across three recent models. the layers are designed to be compilable, extensible, isolated, and portable. they use native pytorch code or portable dsls like triton and helion. support has landed in the main branch for a limited number of layers, enabled by setting use_hw_agnostic=1 with the transformers backend.
the new layers will serve the transformers backend and provide model.py files for flat models. legacy model definitions are being removed, so models will either use flat implementations or fall back to transformers. this effort helps users who run vllm on older gpus, consumer gpus, or custom accelerators like ibm spyre. it keeps vllm useful for the broader open-source ecosystem while frontier development continues on specialized hardware.
why it matters: data scientists and ml engineers can keep using vllm on non-frontier hardware without major performance loss, avoiding a costly rewrite for each accelerator.