Liquid AI Released Speculative Vision-Language Model
The LFM2.5-VL-3B-DSpark model enables faster vision-language inference by applying speculative decoding to multimodal tokens.
Updated on Sept. 25, 2026 in Artificial Intelligence

Live Poll
Should open-source AI models require commercial licenses for companies earning over ten million dollars?
Liquid AI has released the LFM2.5-VL-3B-DSpark, a new vision-language model that accelerates inference using speculative decoding. The released model weights are currently available on Hugging Face.
Why it matters
This release extends the application of speculative decoding to vision-language tasks by treating image and text tokens as identical tensors within hidden layers. It provides an optimized pathway for developers to increase throughput on standard hardware architectures.
The model features 280 million additional parameters and utilizes a 4-layer drafter with a block size of 9. This architecture achieves a 3.13x speedup on Apple silicon and 2.66x on NVIDIA H100 hardware compared to standard inference methods.
The players
Liquid AI
A developer of AI models focused on alternative architectures and efficient inference techniques.
Hugging Face
A central platform for hosting open-source machine learning models and datasets.
The details
The model functions via speculative decoding, where a smaller 'drafter' model predicts upcoming token sequences that the larger target model then verifies in a single pass. By reading hidden states from the target model layers, the drafter predicts tokens even for visual inputs, as the system processes image and text data as unified tensors. This approach allows the system to validate multiple token blocks simultaneously, significantly reducing the computational overhead per generation.
Timeline
September 25, 2026: Liquid AI released the LFM2.5-VL-3B-DSpark model.
The Tech Race
The release builds on the established use of speculative decoding in LLMs by adapting the process for multimodal data. It directly challenges the current latency bottlenecks faced when deploying vision-language models on local or specialized silicon.
Developers can download the model weights from Hugging Face in Safetensors and GGUF formats for immediate integration. Commercial use is free for organizations with less than $10 million in annual revenue under the LFM Open License v1.0.
The takeaway
The model demonstrates that speculative decoding can be successfully applied to multimodal tokens to reduce inference time. Researchers and developers should monitor the Hugging Face repository for future updates to the LFM model family and potential benchmarks.
Further reading
For more developments in this field, explore our latest coverage of Artificial Intelligence.
Live Poll
Should open-source AI models require commercial licenses for companies earning over ten million dollars?









