Skip to content
Models · Aug 12, 2026

LiquidAI releases LFM2.5-VL-3B, a 3.1B-parameter vision-language model optimized for edge devices

The model introduces improved screen/UI understanding, grounding, multi-image reasoning, and function calling, with benchmark results showing strong performance in its size class.

Trust79
HypeLow hype

1 source · cross-referenced

ShareXLinkedInEmail
TL;DR
  • LiquidAI released LFM2.5-VL-3B, a 3.1B-parameter vision-language model designed for on-device and edge applications.
  • It achieves up to 228 tokens/s on an M5 Max and fits in ~3 GB memory, enabling real-time on-device use.
  • Benchmark results show it leads its size class on tasks like visual question answering, document understanding, and screen/UI comprehension.
  • The model supports function calling and multi-image reasoning, with training involving 34T tokens and 4x more vision data than prior releases.

LiquidAI released LFM2.5-VL-3B, a 3.1B-parameter vision-language model optimized for edge and on-device deployment. The model introduces four key improvements over prior releases: enhanced screen and UI understanding, improved grounding and object detection with natural language queries, better multi-image reasoning, and significantly stronger function-calling capabilities in both text-only and vision-text contexts.

The model pairs a SigLIP2 400M NaFlex vision encoder with the same pre-trained backbone as LiquidAI’s LFM2.5-2.6B text model. Training involved approximately 34T tokens, with 4x more vision data than previous versions, drawn from curated and synthetic datasets covering image-caption pairs, OCR, grounding, and instruction-following tasks. To support non-Latin scripts, the tokenizer’s vocabulary was expanded to 128K tokens without retraining from scratch.

Benchmark results across 20 vision and multimodal benchmarks show LFM2.5-VL-3B leading its size class on real-world image tasks, including multilingual visual comprehension, instruction following, document understanding, and screen/UI comprehension. On text-only tasks, the model demonstrates competitive instruction-following and tool-use performance, matching or exceeding models up to twice its size in some evaluations.

LFM2.5-VL-3B achieves up to 228 tokens per second on an M5 Max and 116 tokens per second on a Ryzen AI Max+ 395, with a memory footprint of about 3 GB. On a Galaxy S25 Ultra, it reaches 20 tokens per second, enabling fully on-device operation. On GPU, the model delivers high throughput, reaching approximately 11K tokens per second at high concurrency, roughly twice the speed of larger 4B-class models.

The model ships with day-one support across multiple inference backends, including llama.cpp, MLX, vLLM, SGLang, and ONNX. LiquidAI provides installation instructions via the Transformers library, with compatibility for versions 5.10.1 and higher.

Sources
  1. 01Hugging FaceLFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge
Also on Models

Stories may contain errors. Dispatch is assembled with AI assistance and curated by human editors; despite the trust-score filter, mistakes happen. We correct publicly — every article links to its revision history. Nothing here is financial, legal, or medical advice. Verify before relying on any claim.

© 2026 Dispatch. No ads. No sponsorships. No paid placement. Reader-supported via Ko-fi.

Built by a person who cares about honest AI news.