LiquidAI releases LFM2.5-VL-3B, a 3.1B-parameter vision-language model optimized for edge devices
The model introduces improved screen/UI understanding, grounding, multi-image reasoning, and function calling, with benchmark results showing strong performance in its size class.
1 source · cross-referenced
- LiquidAI released LFM2.5-VL-3B, a 3.1B-parameter vision-language model designed for on-device and edge applications.
- It achieves up to 228 tokens/s on an M5 Max and fits in ~3 GB memory, enabling real-time on-device use.
- Benchmark results show it leads its size class on tasks like visual question answering, document understanding, and screen/UI comprehension.
- The model supports function calling and multi-image reasoning, with training involving 34T tokens and 4x more vision data than prior releases.
LiquidAI released LFM2.5-VL-3B, a 3.1B-parameter vision-language model optimized for edge and on-device deployment. The model introduces four key improvements over prior releases: enhanced screen and UI understanding, improved grounding and object detection with natural language queries, better multi-image reasoning, and significantly stronger function-calling capabilities in both text-only and vision-text contexts.
The model pairs a SigLIP2 400M NaFlex vision encoder with the same pre-trained backbone as LiquidAI’s LFM2.5-2.6B text model. Training involved approximately 34T tokens, with 4x more vision data than previous versions, drawn from curated and synthetic datasets covering image-caption pairs, OCR, grounding, and instruction-following tasks. To support non-Latin scripts, the tokenizer’s vocabulary was expanded to 128K tokens without retraining from scratch.
Benchmark results across 20 vision and multimodal benchmarks show LFM2.5-VL-3B leading its size class on real-world image tasks, including multilingual visual comprehension, instruction following, document understanding, and screen/UI comprehension. On text-only tasks, the model demonstrates competitive instruction-following and tool-use performance, matching or exceeding models up to twice its size in some evaluations.
LFM2.5-VL-3B achieves up to 228 tokens per second on an M5 Max and 116 tokens per second on a Ryzen AI Max+ 395, with a memory footprint of about 3 GB. On a Galaxy S25 Ultra, it reaches 20 tokens per second, enabling fully on-device operation. On GPU, the model delivers high throughput, reaching approximately 11K tokens per second at high concurrency, roughly twice the speed of larger 4B-class models.
The model ships with day-one support across multiple inference backends, including llama.cpp, MLX, vLLM, SGLang, and ONNX. LiquidAI provides installation instructions via the Transformers library, with compatibility for versions 5.10.1 and higher.
- Aug 12, 2026 · OpenAI — News
OpenAI’s Daybreak models now accessible via Amazon Bedrock
Trust79 - Aug 11, 2026 · Google AI — Blog
Google’s AMIE research system shows real-time clinical video consultation capabilities in simulated study
Trust79 - Aug 11, 2026 · Simon Willison — everything
Meta releases Muse Glimmer, a 30B open-weight vision model optimized for agentic tasks
Trust78