LFM
Liquid AI's open-weight LFM2.5 model family runs native text, vision, and audio AI locally on CPU, GPU, or NPU.
If your product has to work offline or keep audio and images on-device, LFM2.5 is the most complete open-weight stack we've reviewed. You get sub-1GB options, native speech in and out via LFM2.5-Audio-1.5B, and multilingual vision at 1.6B — free until you cross $10M in revenue. The 1B numbers against Llama 3.2 1B Instruct and Gemma 3 1B IT aren't close on instruction following or tool use (86.23 vs 52.37 and 63.25 on IFEval). What you don't get is a hosted endpoint: you own quantization, hardware targets and latency tuning. If you want to rent an API, look at a cloud provider instead.
Verified 4d ago · liveness 79/100 · cite: rightaichoice.com/tools/lfm
- Developers building local copilots or productivity tools that must run offline
- Automotive and IoT teams shipping speech interfaces on vehicles, phones or embedded hardware
- Product teams with privacy or compliance rules that forbid sending audio and images to a cloud API
- Japanese-language app developers who need nuance at 1B scale
- Teams that want a managed inference endpoint instead of choosing checkpoints and quantization
- Workloads that need frontier-scale reasoning beyond the few-billion-parameter range
- Projects with no target hardware plan or NPU/CPU deployment skills in house
We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.
- Honest verdict, not marketing
- Real pros & cons from real users
- Attributed quotes with receipts
3 free scans · no card needed
Skip LFM2.5 if you want a managed inference endpoint and would rather rent per-token API calls than pick checkpoints, quantize them and tune latency against your own CPU, GPU or NPU hardware.
Your free commercial rights end the moment your company passes $10M in annual revenue, and from that point a commercial license is required to keep running the models in production.
Free until your company passes $10M in annual revenue — that undercuts per-token cloud APIs outright for sustained workloads and beats most open edge weights on licensing breadth. Above $10M you move to a negotiated enterprise commercial license, so cost becomes deployment-scale dependent. It fits startups, non-profits and researchers best; funded scale-ups should model the license transition before standardizing.
In short
LFM — Liquid AI's open-weight LFM2.5 model family runs native text, vision, and audio AI locally on CPU, GPU, or NPU. Best for Developers building local copilots or productivity tools that must run offline, Automotive and IoT teams shipping speech interfaces on vehicles, phones or embedded hardware, Product teams with privacy or compliance rules that forbid sending audio and images to a cloud API. Free to use.
What's new in LFM
Checked 4 days agoAcross the latest 5 updates: 5 launches.
LFM2.5-VL-DSpark: Accelerating vision-language models on edge and beyond
Liquid AI published LFM2.5-VL-DSpark, a variant aimed at faster vision-language inference on edge hardware and beyond.
LFM2.5-DSpark: Up to 3.2x Faster Inference from H100 to MacBook
LFM2.5-DSpark claims up to 3.2x faster inference across H100 data-center GPUs and MacBooks.
LFM2.5 Q4_0: Quantization-Aware Distillation for Edge Deployment
Liquid AI shipped LFM2.5 Q4_0, a quantization-aware distilled variant built for edge deployment.
LFM2.5-VL-3B: A Better and Faster Vision-Language Model for the Edge
LFM2.5-VL-3B released as a faster 3B vision-language model positioned for edge use.
LFM2.5-2.6B: Deploy Agents Everywhere
LFM2.5-2.6B introduced for deploying agents across device classes.
What people actually say about LFM — is it worth it?
We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.
62 mentions across 5 sources (Hacker News, YouTube, Stack Overflow, GitHub, Lemmy) · researched Aug 27, 2026.
Average across the 5 sources that answered — each source counts once, not each post.
- +Extremely fast CPU inference, 35-40 t/s on 8B-A1B model
- +Small models punch above their weight, outperforming larger ones
- +Open-weight with no copyleft, fine-tunes stay private
- +Free commercial use under $10M revenue, no per-token fees
- +Day-zero support for llama.cpp, MLX, vLLM, ONNX integrations
- −Limited third-party fine-tunes available on Hugging Face
- −Users report models can be 'situational' and not universal
- −Instruction following degrades with longer, complex instructions
- −Name collision with unrelated LFM project causes confusion
- −YouTube viewers struggled to understand real-world use cases
- • Potential requirement for a paid license once crossing $10M revenue
- • No per-token fees but deployment infrastructure, like hardware and storage, is your responsibility
Viability Score
How well maintained and how widely used is LFM? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this
Last calculated: October 2026
How we score →Key Features
- Open-weight LFM2.5 model family for on-device and edge deployment
- LFM2.5-1.2B-Instruct for instruction following and tool use
- LFM2.5-1.2B-Base pretrained checkpoint for heavy fine-tuning
- LFM2.5-1.2B-JP chat model tuned for Japanese knowledge and instructions
- LFM2.5-VL-1.6B vision-language model with multi-image understanding
- Multilingual vision prompts in Arabic, Chinese, French, German, Japanese, Korean and Spanish
- LFM2.5-Audio-1.5B with native speech and text input and output
- LFM-based audio detokenizer 8x faster than Mimi on mobile CPU
- INT4 quantization-aware training for the audio detokenizer
- LFM2.5-VL-3B faster vision-language model for the edge
- LFM2.5-2.6B for deploying agents across environments
- LFM2.5-Encoders that stay fast at long context on CPU
- LFM2.5-DSpark for up to 3.2x faster inference from H100 to MacBook
- LFM2.5-VL-DSpark for faster vision-language inference on edge hardware
- LFM2.5 Q4_0 quantization-aware distillation for edge deployment
About LFM
LFM2.5 is Liquid AI's open-weight foundation model family for on-device AI. The January 2026 release covers Base, Instruct, Japanese, vision-language, and audio-language checkpoints, built on an updated LFM2 architecture with pretraining extended from 10T to 28T tokens and a much larger reinforcement-learning post-training pipeline. Sizes run from a 230M-parameter model small enough to fit under 1GB up through 1.2B, 1.6B, 2.6B and 3B variants. LFM2.5-1.2B-Instruct posts 38.89 on GPQA, 44.35 on MMLU-Pro, 86.23 on IFEval and 49.12 on BFCLv3 — well clear of Llama 3.2 1B Instruct (16.57 / 20.80 / 52.37 / 21.44) and Gemma 3 1B IT (24.24 / 14.04 / 63.25 / 16.64) on the same tests. LFM2.5-1.2B-JP scores 50.7 on JMMLU and 58.1 on M-IFEval (ja). LFM2.5-VL-1.6B handles multi-image and multilingual prompts across Arabic, Chinese, French, German, Japanese, Korean and Spanish. LFM2.5-Audio-1.5B takes speech and text as both input and output natively rather than chaining transcription, an LLM, and TTS, and its LFM-based detokenizer runs 8x faster than the previous Mimi detokenizer on a mobile CPU. Since launch Liquid has added LFM2.5-Encoders for long context on CPU (July 2026), LFM2.5-2.6B for agent deployment (August 2026), LFM2.5-VL-3B, Q4_0 quantization-aware distillation, LFM2.5-DSpark for up to 3.2x faster inference from H100 to MacBook, and LFM2.5-VL-DSpark for edge vision. This is for developers who need private, offline inference without per-token cloud bills — you trade managed serving for hardware you control. All weights ship under the LFM Open License: free to download, run, fine-tune and use commercially until your company passes $10M in annual revenue.
Behind the Verdict
Liquid AI's pitch is narrow and it holds up: put the model where the data already is. The strongest evidence is LFM2.5-Audio-1.5B. Most 'voice AI' is three systems glued together — a transcriber, a text LLM, and a TTS voice — and each seam costs latency and discards prosody. LFM2.5-Audio takes speech and text as both input and output in one model, and its LFM-based detokenizer runs 8x faster than the previous Mimi detokenizer on a mobile CPU. It was quantization-aware trained at INT4, so you can ship at low precision without much quality loss. For automotive, IoT and mobile teams that is a materially different engineering starting point. On text, the 1.2B Instruct checkpoint is the workhorse: 86.23 IFEval, 44.35 MMLU-Pro, 49.12 BFCLv3. The BFCLv3 gap against Granite-4.0-1b (52.43) and Granite-4.0-h-1b (50.69) is worth noting if tool calling is your whole product — IBM's small models are genuinely competitive there, and Qwen3-1.7B is in the same band on several rows. Liquid's advantage is the combination of instruction following, Japanese quality, vision and audio in one coherent family rather than one runaway benchmark. Japanese deserves its own line. LFM2.5-1.2B-JP scores 50.7 JMMLU / 58.1 M-IFEval (ja) / 56.0 GSM8K (ja), ahead of TinySwallow-1.5B-Instruct (48.0 / 36.6 / 47.2) and Qwen3-1.7B in instruct mode (47.7 / 40.3 / 46.0). If you're building a Japanese-language app on the edge, that's a reason to pick this family by itself. Where it doesn't fit: there is no managed inference endpoint in this offering, so you need hardware targets and quantization skills in house. Nothing here is frontier-scale reasoning — the range tops out at a few billion parameters. And the licensing has a hard cliff: cross $10M in annual revenue and your free commercial rights end, at which point you're negotiating a commercial license with sales. For a seed-stage startup that's a non-issue. For a scaling company it's a planning constraint you should model before you build your stack on it. Recent releases — LFM2.5-Encoders for CPU-friendly long context, LFM2.5-2.6B for agents, LFM2.5-DSpark for 3.2x faster inference from H100 down to MacBook, and LFM2.5-VL-DSpark for edge vision — show the cadence is steady rather than a one-off launch.
Researching LFM? Get your full AI stack in 60 seconds.
Free, no signup — tell us your goal and get tools matched to your budget & existing stack.
Real-world workflow fit
Concrete scenarios for the personas LFM actually fits — and what changes day-one when you adopt it.
Downloads LFM2.5-1.2B-Instruct via Hugging Face, converts to GGUF, runs it through llama.cpp on a MacBook, and wires in a tool-calling loop for calendar and file actions.
Outcome: A private local assistant with millisecond latency and no per-request bill, shipped without paying a licence fee.
Deploys LFM2.5-Audio-1.5B natively for speech-in/speech-out instead of chaining transcriber, LLM and TTS, using the INT4 quantization-aware build and the 8x faster LFM detokenizer on the head unit.
Outcome: Lower end-to-end latency and no audio leaving the vehicle, with NPU tuning from AMD or Nexa AI.
Fine-tunes LFM2.5-1.2B-JP on proprietary support transcripts, then ships it on-device for a consumer app that must work without connectivity.
Outcome: A support assistant that scores 58.1 on M-IFEval (ja) at launch and keeps fine-tuned weights private under the no-copyleft license.
Use Cases
- Deploy a local copilot on a laptop or phone that keeps prompts and outputs private.
- Integrate an in-car assistant with native real-time speech input and output.
- Build a Japanese-language support chatbot using LFM2.5-1.2B-JP's 58.1 M-IFEval (ja) score.
- Run a tool-calling agent on consumer hardware with no cloud dependency.
- Create a multimodal edge app combining VL-1.6B vision with Audio-1.5B speech.
- Fine-tune the 1.2B Base checkpoint on proprietary data for a domain-specific assistant.
- Ship long-context encoding on CPU-only hardware with LFM2.5-Encoders.
Models Under the Hood
as of 2026-09-24
Limitations
- LFM2.5 is designed for on-device and edge deployment, running locally on CPU, GPU or NPU — there is no managed hosted endpoint in this offering, so you own quantization, hardware targets and latency tuning.
- Model sizes run from 230M to a few billion parameters, which puts frontier-scale reasoning out of reach.
- The free license permits commercial use, including fine-tuning, only while your company's annual revenue stays under $10 million USD; above that threshold your free commercial rights end and you need a commercial license.
- Weights are distributed through Hugging Face and LEAP rather than as a cloud API.
as of 2026-10-04
Verification history
We have re-verified LFM 9 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
Showing the 6 most recent of 9 verification passes.
Free to cite with attribution — this page re-verifies continuously.
12-month cost
Project the real annual outlay, including the implied monthly cost when only an annual tier is published.
Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.
Plans compared
For each published LFM tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.
Free (< $10M annual revenue)
$0
Ideal for
Startups, indie developers, researchers and non-profits under $10M annual revenue who want commercial on-device AI with no licence fee.
What this tier adds
Starting tier and free entry point: download, run, fine-tune and commercially ship every open LFM model on mobile, edge or on-prem hardware.
Enterprise (> $10M annual revenue)
Custom
Ideal for
Companies past $10M annual revenue running LFM models in production — the tier Liquid cites Mercedes-Benz and Shopify working under.
What this tier adds
Adds a commercial licence beyond the $10M threshold plus bespoke architecture and optimization, OEM and on-prem deployment support, and dedicated support with SLAs.
Where the pricing makes sense
The company stage and team size where LFM's pricing actually pencils out — and where peers do it cheaper.
Free until your company passes $10M in annual revenue — that undercuts per-token cloud APIs outright for sustained workloads and beats most open edge weights on licensing breadth. Above $10M you move to a negotiated enterprise commercial license, so cost becomes deployment-scale dependent. It fits startups, non-profits and researchers best; funded scale-ups should model the license transition before standardizing.
Setup time & first value
How long it actually takes to get something useful out of LFM — broken out by persona, not the marketing-page minute.
Solo developer: minutes to first local generation with a GGUF build in llama.cpp or MLX. Automotive or IoT team: days to weeks, driven by NPU toolchain work with AMD or Nexa AI and on-device power and memory constraints. Fine-tuning team: days for a first domain-tuned checkpoint from LFM2.5-1.2B-Base, longer to validate against a hardware matrix.
Switching to or from LFM
How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.
- →From Llama 3.2 1B Instruct: swap weights in llama.cpp and expect large jumps on IFEval (52.37 to 86.23) and BFCLv3 (21.44 to 49.12).
- →From Gemma 3 1B IT: move to LFM2.5-1.2B-Instruct for stronger tool use, and to LFM2.5-VL-1.6B if you need multilingual vision.
- →From a pipelined cloud voice stack: replace transcriber + LLM + TTS with the single native LFM2.5-Audio-1.5B model.
- →From LFM2-VL-1.6B: LFM2.5-VL-1.6B improves MMStar 49.87 to 50.67 and OCRBench v2 35.11 to 41.44 on the same evaluation harness.
- →From a per-token cloud API: move sustained workloads to local LFM2.5 inference once you have target hardware qualified.
- ↗To a managed cloud API: required if you need frontier-scale reasoning beyond a few billion parameters.
- ↗To Granite-4.0-1b: consider if BFCLv3 tool calling is your single deciding metric (52.43 vs 49.12).
- ↗To Qwen3-1.7B in instruct mode: a same-band alternative if you need its specific reasoning profile (34.85 GPQA).
- ↗To a licensed commercial deployment: mandatory once your company passes $10M in annual revenue.
Integrations
Resources & Guides
Tutorials & Learning
YouTube returned 6 videos for “LFM”, and we withheld 6: 6 could not be judged, because “LFM” is a single word that other videos use for other things. We are showing none, because we could not prove any of them are about LFM.
Official links
Tools that pair well with LFM
Common stack mates teams adopt alongside LFM, with the specific reason each pairing earns its keep.
Falcon LLM
Apache 2.0 open-weight model family from TII Abu Dhabi, spanning hybrid Transformer-Mamba, Arabic, reasoning, and multimodal vision models.
Ollama
Ollama runs open-weight LLMs locally or in the cloud from one command, with unlimited local inference and per-token cloud pricing.
StableLM
StableLM: Stability AI's open-weight language model suite from 2023, self-hosted and licensed for commercial or research use
Featured Head-to-Head Comparisons
Lfm vs Spider Cloud
Choose LFM if you're deploying AI on edge devices (copilots, assistants, IoT) and need private, low-latency, on-device models. Choose Spider Cloud if your AI agents need real-time web data, scraping, and browser automation. They are complementary rather than direct competitors.
Lfm vs Temporal Ai
Choose LFM if your priority is private, low-latency on-device AI with strong multimodal capabilities under 1.6B parameters. Choose Temporal AI if you need a durable execution platform to make AI agents and workflows crash-proof. They are complementary: LFM handles inference, Temporal handles orchestration.
Lfm vs Presto Voice
Presto Voice and LFM serve entirely different domains: Presto Voice is a specialized drive-thru voice automation platform for QSR chains seeking revenue lift, while LFM offers on-device AI models for developers building private, low-latency edge applications. Buyers should choose based on their need: restaurant operations vs. edge AI development. If you run a QSR chain, Presto Voice is the clear choice; if you’re an AI engineer needing local inference, LFM is the better fit.
Alternatives to LFM
View allFalcon LLM
Apache 2.0 open-weight model family from TII Abu Dhabi, spanning hybrid Transformer-Mamba, Arabic, reasoning, and multimodal vision models.
Frequently Asked Questions
Best-of guides
Used LFM? Help shape our editorial sentiment research.