source: Hugging Face Blog: Up to 3.2x Faster Inference with LFM2.5-DSpark

level: technical

Liquid AI released DSpark draft model checkpoints for LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B. These checkpoints add a speculative decoding path that trades a small memory increase for faster decoding while keeping output identical to greedy decoding. The draft models are attention-only with 5 layers and about 300M parameters each. They were trained for 15 epochs on a mix of SFT, chat, code, and function-calling data, selecting the epoch with the highest acceptance rate.

On an H100 GPU using SGLang, speedups range from 1.29x to 3.18x across five benchmarks, with mean speedups of 2.10x to 2.67x depending on the model. On an M4 Max MacBook using llama.cpp with Metal, speedups reach up to 2.87x for the 2.6B model, but only 1.18x on average for the 8B-A1B MoE model due to current Metal backend limitations. For function calling, DSpark reduces latency by 57% on average for LFM2.5-2.6B. Acceptance rates vary by dataset, from 3.90 to 8.52 out of 10.

Speculative decoding uses a lightweight draft model to propose tokens, which the target model verifies in one forward pass, reducing memory-bound weight loading. DSpark combines a DFlash-style parallel backbone, a Markov chain head for inter-token dependency, and a confidence-scheduled verifier that prunes low-confidence suffixes. The checkpoints are available on Hugging Face in Safetensors and GGUF formats, with day-one support in llama.cpp and SGLang. This release targets both cloud and on-device inference, especially for agentic workloads.

why it matters: Faster token generation reduces latency and cost for AI applications, making large language models more practical for real-time and on-device use.


source: Hugging Face Blog: Up to 3.2x Faster Inference with LFM2.5-DSpark