Liquid AI launches compact on-device model with enhanced multi-step tool use for sensitive environments

Liquid AI has introduced LFM2.5-2.6B, a small yet powerful AI model designed for local deployment in regulated sectors, capable of complex workflows without data leaving the hardware.

Liquid AI has unveiled LFM2.5-2.6B, a compact model it says can carry out multi-step, tool-using tasks entirely on the device where it is running. The company is pitching the system as a local alternative for phones, laptops, PCs and robots, with the key promise that prompts and data do not need to leave the hardware. That places it squarely in the same on-device strategy Liquid AI has been building across its LFM2.5 line, which has already included smaller text, audio and multilingual variants aimed at edge deployment.

According to Liquid AI, the new model has 2.69 billion parameters, a 131,072-token context window and a 128,000-word vocabulary. It was trained on about 34 trillion tokens and ships in two versions: a base checkpoint for custom fine-tuning and a post-trained version aimed at agentic workloads. The company says the model’s architecture combines 22 double-gated short convolution blocks with 8 grouped-query attention blocks, and that it supports 16 languages while remaining text-only.

Liquid AI is framing the release for regulated and disconnected environments, including automotive, consumer electronics, industrial robotics, healthcare, financial services, e-commerce and defence. It says the model is best suited to agentic workflows such as tool use, data extraction, retrieval-augmented generation and long-context document processing, including offline triage of large files, form and invoice extraction and command parsing for robots. The company is less enthusiastic about using it for coding or knowledge-intensive work, where larger systems still appear to hold an edge.

The release also comes with a distribution push designed to make deployment straightforward. Liquid AI says the weights are available on Hugging Face under its lfm1.0 licence and are packaged in native, GGUF, MLX and ONNX formats, with support from llama.cpp, vLLM, SGLang and LM Studio. The company claims the model can decode at 220 tokens a second on an Apple M5 Max while using less than 2.5GB of memory, and says a single NVIDIA H100 SXM5 can serve about 1.3 billion tokens a day for teams that prefer to self-host rather than run everything locally.

Benchmark claims are central to the pitch. Liquid AI says LFM2.5-2.6B outperforms larger rivals on several instruction-following and tool-use tests, including ToolSandbox, Multi-IF and IFStruct, while remaining competitive on BFCLv4. But the company’s own figures also show that larger models still lead in some coding tasks, most notably LiveCodeBenchv6, suggesting the new release is strongest as a specialised on-device agent rather than a universal replacement for bigger cloud systems.

Disclaimer: This content is intended for informational purposes only. Readers are advised to exercise their own judgement, conduct due diligence, or consult a qualified expert before acting on any information provided.