This Tiny AI Model Runs on Your Device, Not the Cloud

This Tiny AI Model Runs on Your Device, Not the Cloud — featured image

by

Your Next Employee Might Not Need a Desk

Imagine this: a worker that never sleeps, never asks for annual leave, and processes documents or answers queries even when your shop Wi-Fi is down. That’s not a distant fantasy — it’s a small AI model called LFM2.5-2.6B, just released by Liquid AI. For Malaysian SMEs running on tight budgets and unreliable connectivity, this kind of on-device intelligence could be your most practical automation tool yet.

What Happened

Liquid AI released a new agentic model that runs entirely on a device — no cloud, no API calls, no internet connection required. The model, LFM2.5-2.6B, has 2.69 billion total parameters, a 131,072-token context window, and a 128,000-token vocabulary. It was pre-trained on roughly 34 trillion tokens, according to Marktechpost. Because inference stays local, data never leaves the device and the marginal cost of each run is near zero.

The model can plan, call tools, and work through multi-step tasks on phones, laptops, PCs, and robots. Two checkpoints are available: a base version for fine-tuning, and a post-trained version for agentic workloads. Both are open weights on Hugging Face under the lfm1.0 license, and they ship in native, GGUF, MLX, and ONNX formats with day-one support in llama.cpp, vLLM, SGLang, and LM Studio. That’s about as developer-friendly as it gets.

What’s impressive is the benchmark performance. Liquid AI compared LFM2.5-2.6B against models nearly four times its size, including Qwen3.5-9B (9.7B parameters). The small model actually beat Qwen3.5-9B on ToolSandbox (77.83 vs 76.44), Multi-IF (80.07 vs 62.55), and IFStruct (85.49 vs 78.50), as shown in the benchmark table. It only trails on BFCLv4, and for coding-heavy tasks, larger models still win — LiveCodeBenchv6 is 59.41 versus 69.86 for Qwen3.5-9B. So this is not a coding model, but it’s a powerhouse for tool use and instructions.

Why This Matters for Malaysian SMEs

Think about your daily operations. If you run a logistics company, you’re juggling driver schedules, GST filings, and customer queries. If you run a clinic, you have patient records and appointment reminders that must stay private. LFM2.5-2.6B’s on-device nature means you can deploy an agent that extracts data from invoices, triages long documents, or handles form processing — all without sending sensitive information to a third-party API. For Malaysian businesses in regulated industries like healthcare or finance, this is a game-changer for compliance, because no prompt ever leaves your computer.

The practical hardware side is equally relevant. The model decodes at 220 tokens per second on an M5 Max and under 2.5 GB of memory, according to Liquid AI’s deployment notes. That means you can run it on hardware you already own — a mid-range laptop, a desktop PC, or even a phone at 30 tokens per second. For a small team of 5 to 50 employees, you don’t need to buy expensive cloud subscriptions or sign up for enterprise AI plans. You just download the weights and start building an agent that handles your document triage, invoice extraction, or even robot command parsing if you’re in light manufacturing.

But here’s the quiet superpower: continuous background agents. Because there’s no per-token cost and no network latency, you can have the model running 24/7, monitoring a folder for incoming PDFs, extracting key fields, and updating your spreadsheet — while your staff sleeps. For an SME owner, that’s like hiring a diligent night-shift assistant for the price of electricity.

“Because inference stays local, data never leaves the device and the marginal cost of each run is near zero.”

The Bigger Picture

What Liquid AI just demonstrated is that the frontier of AI is shifting — not to bigger clouds, but to smaller, better, edge-deployed models. For Malaysian SMEs, this signals a shift away from dependence on foreign cloud providers and expensive API credits. You can now own the model, control it, and fine-tune it using tools like LoRA with TRL and Unsloth, as referenced in the report. That’s the kind of autonomy that lets a bakery in George Town or a hardware shop in Kuching build automation tailored to their own workflows, offline, secure.

The bigger implication is this: your competitive advantage will no longer come from buying expensive AI tools, but from how creatively you apply open-weight models to your specific niche. Whether it’s an on-device assistant for your service counter, offline document triage for your archive room, or a background agent that keeps your inventory ledgers updated, the barrier to entry just dropped dramatically. Liquid AI’s model supports 16 languages — more than enough to handle Bahasa Malaysia, Mandarin, Tamil, and English — and with 128K context, it can chew through an entire annual report in one go.

Of course, the model isn’t perfect. Liquid AI explicitly does not recommend it for agentic coding or knowledge-heavy tasks. But for the 80% of SME operations that involve form filling, scheduling, data extraction, and straightforward decision-making, this small model is more than enough. And because it’s open weights, you’re not locked into any vendor. You can upgrade when the next generation arrives.

Key Takeaways at a Glance

Feature LFM2.5-2.6B
Parameters 2.69B total
Context Window 131,072 tokens
Trillion Training Tokens ~34T
Decode Speed (M5 Max) 220 tok/s under 2.5GB
Phone Speed 30 tok/s
Key Benchmarks Beats 9.7B model on ToolSandbox, Multi-IF, IFStruct
Run Formats GGUF, MLX, ONNX
License lfm1.0 (open weights)

So, what’s your next move? Start exploring this model today. Download it, test it on your own documents, and see how it handles your everyday tasks. The small, on-device AI revolution has arrived — and it’s ready to work for your Malaysian SME.

Ready to Streamline Your Operations?

Technology moves fast. Your operations should keep up. AutoRunBiz builds AI systems that run your daily workflows — from WhatsApp order capture to accounting. Book a free 15-min ops audit →