RU

Liquid AI speeds up LFM2.5-VL-3B by 2–3x with speculative decoding

Published: 2026-09-26 · Author: AI Release · @ai_release1
Liquid AI speeds up LFM2.5-VL-3B by 2–3x with speculative decoding

⚡ The gist in 5 seconds - Liquid AI has introduced LFM2.5-VL-DSpark — a 280-million-parameter drafter designed to speed up the 3-billion-parameter vision-language model LFM2.5-VL-3B via speculative decoding. - The model is already available on Hugging Face; the company claims decoding speedups of up to 3.13x on an Apple M5 Max and up to 2.66x on an Nvidia H100. - Limitation: only token generation is accelerated, while the vision encoder and prefill remain unchanged, so end-to-end latency improves less — up to 2.62x on the M5 Max. ### 🔍 What was found LFM2.5-VL-DSpark uses the hidden states of selected layers of LFM2.5-VL-3B and proposes blocks of candidate tokens. Images and text are projected into a shared representation, allowing the drafter to work with the same vector dimensionality for any modality. The architecture includes four attention-only layers, a block size of nine during training and eight or nine at inference. The drafter adds about 8.9% to the parameter count of the deployed model. The company evaluated the system on six visual workloads from the MMSpec benchmark: general and text-oriented VQA, caption generation, diagrams, complex reasoning, and multi-turn dialogue. On MLX with an M5 Max, decoding sped up by 2.30–3.13x and end-to-end latency by 1.56–2.62x. On llama.cpp with an M3 Ultra — 1.57–2.14x and 1.30–1.77x respectively. On the H100 — up to 2.66x decoding and up to 2.27x end-to-end throughput. These are Liquid AI's internal measurements, not an independent benchmark; the actual gain depends on the token acceptance rate, prompt length, image complexity, quantization, batch size, and the share of time spent processing the image. ### 💡 Why it matters Speculative decoding removes the main bottleneck of multimodal inference — the latency after the image and prompt have been processed. Since the target model verifies every proposed token, greedy output matches the original, meaning the speedup does not change the model's behavior. The model is available on Hugging Face and extends the DSpark approach to multimodal inputs without changing the algorithm. However, Amdahl's law limits the overall gain: the vision encoder and prefill are not accelerated. For image-based chats, document analysis, and diagram interpretation, time-to-first-token and total response time need to be measured separately — especially on edge devices. ### 🧩 Context The release extends the DSpark approach from text models to multimodal workloads without requiring a different speculative decoding algorithm. The strongest speedups are claimed on Apple Silicon; the gain on the H100 is smaller but still noticeable, making the solution interesting both for local inference and GPU services. That said, efficiency depends heavily on the hardware and the task, so one shouldn't expect a universal speedup.

🔗 Read on creati.ai

🤖 AI summary
Liquidspeculativevision-language
📖
Read the guide on this topic
Read →
← PreviousAppBrain: Google Play Statistics — 2,010,562 Apps, Ratings and Top ChartsNext →US Appeals Court Upholds Anthropic's Supply Chain Risk Designation

Source: creati.ai · post in Telegram