LFM2.5-VL-3B is our most capable vision-language model you can run on your own hardware. It understands documents and screens alike, grounds objects, and can call tools. It answers directly instead of reasoning, so responses stay fast in real-time and on-device apps. LFM2.5-VL-3B extends the vision-language capabilities of our previous releases with four major improvements: Screen/UI understanding: Strong understanding of digital screens across different devices. Grounding: Improved grounding and object detection with natural language queries. Multi-image input: Improved reasoning across multiple images. Function calling: Significantly stronger at function calling, in text-only and vision-text situations. How we trained our most capable vision-language model LFM2.5-VL-3B pairs a …
From the source
LFM2.5-VL-3B decodes 228 tokens/s on an M5 Max and 116 tokens/s on a Ryzen AI Max+ 395, and fits in about 3 GB of memory. It even reaches 20 tokens/s on a Galaxy S26 Ultra, so you can run it fully on-device.
huggingface.co