Small Model, Big Vision: What LFM2.5-VL-3B Means for Malaysian SMEs
You know that sinking feeling when a staff member squints at a blurry receipt, or when someone has to copy order numbers from a screenshot into your inventory system by hand? That wasted time is where your business quietly loses momentum. This week, Liquid AI released LFM2.5-VL-3B, a small vision-language model that reads screens, locates objects, and triggers tools—right on a device you could actually own. For Malaysian SME owners, this is not another abstract AI headline. It is a practical signal: the era of affordable, private, on-premise computer vision has arrived.
What Happened
LFM2.5-VL-3B is a 3.1B-parameter vision-language model designed specifically for on-device deployment. It can read digital screens across mobile, web, and desktop interfaces, parse documents and charts, ground objects to specific coordinates, and call tools from text or image input. The model is non-reasoning, which means it answers directly rather than thinking step-by-step—a deliberate design choice that keeps latency low for real-time tasks.
The performance numbers are striking. Liquid AI reports an average of 69.4 across 28 vision benchmarks, matching the much larger InternVL-3.5-4B and landing just 0.7 points behind Qwen3.5-4B. It fits in roughly 3 GB of memory and decodes 228 tokens per second on an Apple M5 Max. Beyond raw benchmarks, the model ships in four formats—native, GGUF, ONNX, and MLX—with day-one runtime support for llama.cpp, MLX, vLLM, SGLang, and ONNX. That means you are not locked into one vendor’s environment.
What changed compared to the previous release? Four areas stand out. First, screen understanding improved to 80.7 on ScreenSpot-v2 across desktop, mobile, and web. Second, function calling is now part of the vision-language line, with ToolSandbox moving from 26.4 to 59.5. Third, grounding—the ability to point at objects—jumped from 57.1 to 87.9 on RefCOCO-avg precision@1. Fourth, multi-image input improved significantly, with BLINK rising from 50.2 to 61.5. All of these are documented in the technical release notes.
Why This Matters for Malaysian SMEs
Let’s bring this down to your office floor. If you run a retail, logistics, healthcare, or accounting business, you are drowning in visual information: delivery receipts, supplier invoices, lab reports, and screenshots of bank statements. A model like LFM2.5-VL-3B can convert a photo of a delivery order into structured text, or read a chart from a government portal and file the key figures into your spreadsheet. Because it runs on-device, your customer and supplier data never has to leave your laptop or workstation. For SMEs in Malaysia, where data privacy concerns are rising, that is a meaningful advantage. The model also supports 16 languages, which is valuable in our multilingual environment of Bahasa Malaysia, Mandarin, Tamil, and English.
If you are a software vendor or an automation provider serving Malaysian SMEs, this release matters even more. Many local businesses still run on legacy accounting or inventory systems that lack modern APIs. LFM2.5-VL-3B can act as a screen agent—it can look at your screen, understand the interface, and call the right action. Its function calling capability emits structured Pythonic calls between special tokens, so it can hand off directly to your existing automation workflows. Imagine building a “co-pilot” that helps your staff navigate MyInvois, SSM, or LHDN portals without requiring a full integration project. Or consider GUI test automation: your team could record a test once, and the model could visually verify each screen state, catching UI changes immediately.
“Screen understanding is where this model makes its leap: it averages 80.7 on ScreenSpot-v2 across desktop (78.7), mobile (81.2), and web (82.2).”
The licensing also appears tailored for businesses like yours. The LFM Open License v1.0 is based on Apache-2.0 but with a tiered commercial structure: smaller companies get straightforward commercial use, while larger organisations must negotiate directly with Liquid AI. Research, education, and non-profit use carry no such limits. For a Malaysian SME under that threshold, this means you can build a commercial product or internal tool without the administrative headache of complex enterprise licensing.
The Bigger Picture
What is truly significant here is the direction of travel. A few years ago, a model that could read screens and call tools would have required a rack of GPUs and a team of engineers. Today, a 3.1B-parameter model does it in roughly 3 GB of memory. The trend is clear: intelligence is shrinking, and it is moving onto the edge. For Malaysian SMEs, this shifts the conversation from “how do we access AI?” to “how do we choose which AI to run locally?” You are no longer renting vision by the API call; you are installing it like a piece of software.
This release also highlights a larger movement toward open, portable AI. With support for GGUF and ONNX, the same model can run on a laptop today and on a different vendor’s device next year. That portability matters if you want to avoid lock-in. And because the model is non-reasoning, it is fast enough for point-of-sale checks, inventory photo scans, and other real-time tasks—speed that matters when your staff is staring at a queue of customers.
| Key Feature | Why It Matters for Your SME |
|---|---|
| 80.7 on ScreenSpot-v2 | Automates reading from web portals, desktop apps, and mobile screens—perfect for legacy systems. |
| 87.9 on RefCOCO grounding | Locates objects in photos, ideal for stock-taking, product inspection, and visual QA. |
| Function calling, ToolSandbox 59.5 | Connects visual input directly to your existing tools—no manual copy-paste. |
| Runs in ~3 GB with 4 formats | Deploys on standard laptops or edge devices; no cloud dependency. |
For your business, the question is not whether to pay attention to small vision models—it is how soon you should test one. Whether you start with a simple invoice-reading prototype or a full on-device screen agent, LFM2.5-VL-3B points to a future where the “eyes” of your operations are local, fast, and under your control. That is a future worth preparing for.
Ready to Streamline Your Operations?
Technology moves fast. Your operations should keep up. AutoRunBiz builds AI systems that run your daily workflows — from WhatsApp order capture to accounting. Book a free 15-min ops audit →
