source: Hugging Face Blog: Accelerating vision-language models with LFM2.5-VL-DSpark
level: technical
liquid ai released an experimental dspark draft model for its vision-language model lfm2.5-vl-3b. the drafter adds a speculative decoding path that trades a small memory increase for faster inference. it works by capturing hidden states from tapped layers of the target model and drafting blocks of candidate tokens. image patches and text tokens are projected into a shared representation, so the drafter operates on identical dimensionality regardless of input modality. the inference algorithm remains unchanged from the text models.
the drafter has about 280 million parameters, adding 8.9 percent to the 3 billion parameter target model. on an m5 max with mlx, decoding speeds up 2.30 to 3.13 times by task, and end-to-end latency improves 1.56 to 2.62 times. on an h100 gpu, decoding is 2.04 to 2.66 times faster, with end-to-end gains of 1.64 to 2.27 times. however, speculative decoding only accelerates the decode stage, not vision encoding or prefill. on edge devices, prefill takes a larger share of total latency, so end-to-end gains are smaller than decode gains.
the model ships with day-one support for llama.cpp, mlx-vlm, and sglang. training used a mixture of vision-language sft data, with ablations across 3, 4, and 5 layers leading to a 4-layer attention-only drafter with block size 9. at inference, a block size of 8 or 9 is recommended depending on hardware. speculative decoding is exact: the target verifies every proposed token, so greedy output equals the target alone. the drafter is available on hugging face in safetensors and gguf formats.
why it matters: faster vision-language model inference on edge devices and gpus can reduce latency and cost for real-time applications like image captioning and visual question answering.
source: Hugging Face Blog: Accelerating vision-language models with LFM2.5-VL-DSpark