Instructions to use LiquidAI/LFM2.5-Encoder-230M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use LiquidAI/LFM2.5-Encoder-230M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="LiquidAI/LFM2.5-Encoder-230M", trust_remote_code=True)# Load model directly from transformers import AutoTokenizer, AutoModelForMaskedLM tokenizer = AutoTokenizer.from_pretrained("LiquidAI/LFM2.5-Encoder-230M", trust_remote_code=True) model = AutoModelForMaskedLM.from_pretrained("LiquidAI/LFM2.5-Encoder-230M", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
LFM2.5-Encoder-230M
LFM2.5-Encoder is a family of multilingual bidirectional encoders built on the LFM2 architecture, available in two sizes:
- LFM2.5-Encoder-230M (this model) โ a lightweight encoder for tight latency and memory budgets, punching above its size class.
- LFM2.5-Encoder-350M โ a larger sibling for maximum downstream quality.
Both are masked language models with full bidirectional attention, designed to be fine-tuned into task-specific models (classification, token classification, retrieval, reranking, and semantic similarity) across 15 languages, and to run efficiently on-device.
Find more details about our encoders in our blog post.
Key highlights:
- Highly capable for its size. On par with the best similarly sized encoders and well ahead of our own retrieval siblings.
- General-purpose. 8k context, strong across NLI, paraphrase, sentiment, and multilingual tasks.
- Fast and on-device. Matches or beats ModernBERT throughput, with a long-context edge on CPU; runs in the browser on WebGPU.
๐ป Demos: We built the demos below from fine-tuned LFM2.5-Encoders. Each one runs in a CPU-only Hugging Face space:
- Zero-shot prompt routing โ define your own routing lanes as free text. The model scores the whole prompt against every lane in one pass.
- Zero-shot policy linting โ check text against your company's rules, written as free text. It scores every token against every rule in one pass.
- Spell checking โ correct misspellings token by token.
- PII detection โ spot and remove 40 kinds of personal information across 16 languages.
- Masked-diffusion text generation โ bonus: run the encoder as a chatbot that generates text by iteratively unmasking instead of left to right.
๐ Model details
| Property | LFM2.5-Encoder-230M | LFM2.5-Encoder-350M |
|---|---|---|
| Type | Bidirectional encoder (masked language model) | Bidirectional encoder (masked language model) |
| Backbone | LFM2 | LFM2 |
| Total parameters | ~229.7M | ~354.5M |
| Hidden size | 1024 | 1024 |
| Vocabulary size | 65,536 | 65,536 |
| Context length | 8,192 tokens | 8,192 tokens |
| License | LFM Open License v1.0 | LFM Open License v1.0 |
Supported languages: English, German, Spanish, French, Italian, Dutch, Polish, Portuguese, Arabic, Hindi, Japanese, Russian, Turkish, Vietnamese, Chinese (15).
Architecture. LFM2.5-Encoder is built on the LFM2 hybrid backbone, which interleaves gated short-convolution blocks with grouped-query attention. For encoder use, the causal mask is replaced with full bidirectional (non-causal) attention and the model is trained with a masked language modeling head. The encoder body is exposed as Lfm2BidirectionalModel; masked-LM loading uses Lfm2BidirectionalForMaskedLM. Both are wired through auto_map and require trust_remote_code=True.
Lfm2BidirectionalForMaskedLM(
(lfm2): Lfm2BidirectionalModel
(lm_head): Linear(in_features=1024, out_features=65536, bias=False)
)
Training. LFM2.5-Encoder-230M is adapted from the LFM2 base and trained with a masked language modeling objective on a large multilingual corpus. Pre-training uses a two-stage schedule that extends the context window to up to 8,192 tokens.
We recommend fine-tuning LFM2.5-Encoder-230M for a range of downstream tasks, such as:
- Text classification: sentiment, topic, intent/routing, moderation, and business-text linting.
- Token classification: named-entity recognition, span extraction, and sequence labeling.
- Retrieval and reranking: a backbone for dense embedding or late-interaction (ColBERT-style) retrievers.
- Semantic similarity: STS, paraphrase, and duplicate detection.
- Natural language inference and extractive QA: sentence-pair reasoning and answer-span extraction.
Its small footprint makes it especially well-suited to on-device and browser (WebGPU) deployment.
๐ How to run
Install the latest version of transformers:
pip install -U transformers
Run masked-token prediction:
from transformers import AutoModelForMaskedLM, AutoTokenizer
import torch
tok = AutoTokenizer.from_pretrained("LiquidAI/LFM2.5-Encoder-230M", trust_remote_code=True)
mlm = AutoModelForMaskedLM.from_pretrained("LiquidAI/LFM2.5-Encoder-230M", trust_remote_code=True)
text = f"The capital of France is {tok.mask_token}."
enc = tok(text, return_tensors="pt")
with torch.no_grad():
logits = mlm(**enc).logits
pos = (enc["input_ids"][0] == tok.mask_token_id).nonzero()[0].item()
print([tok.decode([t]).strip() for t in logits[0, pos].topk(5).indices.tolist()])
# -> ['Paris', 'Strasbourg', 'Paris', 'Lyon', 'Versailles']
For downstream tasks, load the encoder body and attach your own head (classification, token classification, regression, retrieval):
from transformers import AutoModel
body = AutoModel.from_pretrained("LiquidAI/LFM2.5-Encoder-230M", trust_remote_code=True)
If your GPU supports it, we recommend using LFM2.5-Encoder-230M with Flash Attention 2 to reach the highest efficiency. To do so, install Flash Attention as follows, then use the model as normal:
pip install flash-attn
๐ Performance
For each benchmark task, we run a full supervised fine-tune and report that fine-tuned model's score.
The results below span 14 models across 17 tasks from GLUE, SuperGLUE, and multilingual classification tasks.
The full evaluation harness is open-sourced in the
eurobert-repro repository.
17-task results (avg@5 fresh seeds ยฑ std)
| Rank | Model | Params | 17-task mean | ยฑ std |
|---|---|---|---|---|
| 1 | XLM-R XL (3.5B) | 3.5B | 83.06 | ยฑ1.16 |
| 2 | ModernBERT-large (395M) | 395M | 81.68 | ยฑ2.49 |
| 3 | XLM-R large (560M) | 560M | 81.34 | ยฑ1.66 |
| 4 | LFM2.5-Encoder-350M (ours) | 350M | 81.02 | ยฑ1.00 |
| 5 | mDeBERTa-v3 (280M) | 280M | 80.37 | ยฑ1.06 |
| 6 | LFM2.5-Encoder-230M (ours) | 230M | 79.29 | ยฑ1.02 |
| 7 | ModernBERT-base (149M) | 149M | 78.19 | ยฑ1.39 |
| 8 | XLM-R base (280M) | 280M | 77.46 | ยฑ1.63 |
| 9 | EuroBERT-210M | 210M | 76.87 | ยฑ2.00 |
| 10 | mGTE-MLM (305M) | 305M | 76.53 | ยฑ1.85 |
| 11 | LFM2.5-ColBERT-350M | 350M | 76.18 | ยฑ1.25 |
| 12 | EuroBERT-610M | 610M | 75.87 | ยฑ2.03 |
| 13 | LFM2.5-Embedding-350M | 350M | 75.68 | ยฑ0.83 |
| 14 | EuroBERT-2.1B | 2.1B | 72.19 | ยฑ5.59 |
Click to expand per-task results โ all 17 tasks (avg@5 fresh seeds ยฑ std)
### Per-task results โ all 17 tasks (avg@5 fresh seeds ยฑ std)| Model | XNLI | PAWS-X | Amazon | MASSIVE | SeaHorse | CoLA* | SST-2* | MRPC* | STS-B* | QQP* | MNLI* | QNLI* | RTE* | BoolQ* | CB* | WiC* | WSC* | ALL |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| XLM-R XL (3.5B) | 87.12ยฑ0.45 | 93.30ยฑ0.41 | 62.29ยฑ0.05 | 88.21ยฑ0.32 | 59.51ยฑ3.52 | 84.58ยฑ1.18 | 95.69ยฑ0.30 | 87.65ยฑ1.69 | 90.20ยฑ0.93 | 91.72ยฑ0.07 | 90.09ยฑ0.11 | 94.47ยฑ0.28 | 82.38ยฑ3.51 | 83.70ยฑ0.24 | 89.88ยฑ3.72 | 66.55ยฑ1.37 | 64.62ยฑ1.58 | 83.06 |
| ModernBERT-large (395M) | 81.76ยฑ0.38 | 92.46ยฑ0.18 | 60.42ยฑ0.15 | 85.65ยฑ0.94 | 40.20ยฑ17.18 | 83.37ยฑ0.20 | 96.10ยฑ0.53 | 88.14ยฑ1.79 | 92.16ยฑ0.24 | 91.81ยฑ0.13 | 90.65ยฑ0.18 | 94.36ยฑ0.10 | 81.59ยฑ4.92 | 81.68ยฑ2.34 | 88.21ยฑ3.24 | 70.16ยฑ2.57 | 69.81ยฑ7.18 | 81.68 |
| XLM-R large (560M) | 84.69ยฑ0.59 | 93.23ยฑ0.78 | 61.58ยฑ0.13 | 88.50ยฑ0.17 | 56.12ยฑ2.31 | 83.34ยฑ1.75 | 93.83ยฑ1.04 | 88.77ยฑ2.15 | 91.35ยฑ0.23 | 90.48ยฑ0.23 | 88.29ยฑ0.07 | 93.08ยฑ0.23 | 80.79ยฑ3.14 | 80.54ยฑ0.79 | 78.21ยฑ8.69 | 66.24ยฑ5.41 | 63.65ยฑ0.43 | 81.34 |
| LFM2.5-Encoder-350M (ours) | 79.82ยฑ0.29 | 91.53ยฑ0.73 | 60.57ยฑ0.09 | 85.70ยฑ0.13 | 54.96ยฑ0.41 | 84.43ยฑ0.81 | 95.11ยฑ0.26 | 87.21ยฑ1.86 | 91.59ยฑ0.05 | 92.08ยฑ0.10 | 89.03ยฑ0.17 | 93.97ยฑ0.23 | 75.23ยฑ3.84 | 81.52ยฑ0.69 | 83.21ยฑ2.04 | 69.66ยฑ2.23 | 61.73ยฑ3.15 | 81.02 |
| mDeBERTa-v3 (280M) | 83.01ยฑ0.47 | 92.59ยฑ0.39 | 60.62ยฑ0.31 | 87.64ยฑ0.51 | 54.97ยฑ2.04 | 83.91ยฑ1.03 | 92.41ยฑ0.96 | 85.39ยฑ2.59 | 89.87ยฑ0.22 | 90.23ยฑ0.15 | 86.30ยฑ0.18 | 91.99ยฑ0.35 | 69.75ยฑ2.03 | 78.29ยฑ1.39 | 88.21ยฑ2.40 | 67.71ยฑ3.05 | 63.46ยฑ0.00 | 80.37 |
| LFM2.5-Encoder-230M (ours) | 77.63ยฑ0.31 | 90.86ยฑ0.24 | 59.97ยฑ0.21 | 85.52ยฑ0.69 | 54.61ยฑ0.62 | 81.42ยฑ1.56 | 94.08ยฑ0.37 | 80.20ยฑ2.66 | 90.99ยฑ0.12 | 91.71ยฑ0.07 | 87.98ยฑ0.22 | 92.96ยฑ0.37 | 67.29ยฑ1.97 | 76.54ยฑ1.01 | 83.21ยฑ4.48 | 70.31ยฑ1.34 | 62.69ยฑ1.05 | 79.29 |
| ModernBERT-base (149M) | 76.64ยฑ0.31 | 92.17ยฑ0.15 | 58.98ยฑ0.11 | 85.32ยฑ0.21 | 45.19ยฑ1.72 | 83.07ยฑ1.93 | 94.79ยฑ0.52 | 84.46ยฑ2.70 | 90.75ยฑ0.14 | 91.23ยฑ0.11 | 88.68ยฑ0.17 | 93.04ยฑ0.36 | 58.70ยฑ1.94 | 74.78ยฑ4.10 | 81.07ยฑ6.75 | 66.90ยฑ2.39 | 63.46ยฑ0.00 | 78.19 |
| XLM-R base (280M) | 78.20ยฑ0.80 | 91.36ยฑ0.41 | 60.01ยฑ0.12 | 87.47ยฑ0.49 | 51.08ยฑ3.75 | 81.17ยฑ1.07 | 91.97ยฑ0.18 | 86.47ยฑ0.76 | 88.27ยฑ0.36 | 89.21ยฑ0.04 | 83.07ยฑ0.21 | 90.17ยฑ0.32 | 62.60ยฑ7.03 | 71.43ยฑ1.64 | 79.64ยฑ7.53 | 61.25ยฑ3.00 | 63.46ยฑ0.00 | 77.46 |
| EuroBERT-210M | 80.83ยฑ0.35 | 91.94ยฑ0.31 | 59.94ยฑ0.16 | 86.36ยฑ0.67 | 45.16ยฑ16.60 | 72.75ยฑ1.30 | 90.64ยฑ0.92 | 80.74ยฑ2.99 | 89.29ยฑ0.23 | 90.75ยฑ0.09 | 85.63ยฑ0.28 | 91.49ยฑ0.27 | 54.95ยฑ2.56 | 71.68ยฑ2.35 | 86.79ยฑ2.40 | 64.64ยฑ2.01 | 63.27ยฑ0.43 | 76.87 |
| mGTE-MLM (305M) | 80.32ยฑ0.20 | 91.73ยฑ0.26 | 60.26ยฑ0.10 | 87.79ยฑ0.20 | 51.58ยฑ1.31 | 75.44ยฑ4.66 | 91.19ยฑ1.00 | 86.32ยฑ1.48 | 87.77ยฑ0.64 | 89.82ยฑ0.09 | 84.14ยฑ0.15 | 90.94ยฑ0.40 | 58.34ยฑ3.09 | 69.32ยฑ3.72 | 73.21ยฑ6.80 | 59.34ยฑ7.41 | 63.46ยฑ0.00 | 76.53 |
| LFM2.5-ColBERT-350M | 78.77ยฑ0.47 | 89.74ยฑ0.44 | 59.92ยฑ0.13 | 86.65ยฑ0.17 | 47.95ยฑ1.13 | 71.06ยฑ1.06 | 90.94ยฑ0.78 | 73.43ยฑ7.22 | 89.38ยฑ0.28 | 91.11ยฑ0.14 | 84.70ยฑ0.23 | 90.66ยฑ0.18 | 59.13ยฑ2.30 | 74.25ยฑ1.61 | 81.79ยฑ2.93 | 62.04ยฑ2.12 | 63.46ยฑ0.00 | 76.18 |
| EuroBERT-610M | 84.61ยฑ0.34 | 91.84ยฑ0.94 | 60.64ยฑ0.08 | 86.03ยฑ0.99 | 12.91ยฑ8.05 | 70.60ยฑ2.12 | 92.52ยฑ0.66 | 85.20ยฑ1.34 | 89.82ยฑ0.22 | 91.13ยฑ0.09 | 87.95ยฑ0.19 | 92.57ยฑ0.36 | 59.28ยฑ7.93 | 76.86ยฑ1.55 | 85.71ยฑ4.37 | 58.71ยฑ5.30 | 63.46ยฑ0.00 | 75.87 |
| LFM2.5-Embedding-350M | 78.59ยฑ0.11 | 89.13ยฑ0.63 | 60.47ยฑ0.13 | 87.03ยฑ0.21 | 50.19ยฑ0.90 | 72.54ยฑ0.68 | 91.70ยฑ0.69 | 77.45ยฑ1.31 | 89.38ยฑ0.09 | 91.14ยฑ0.13 | 84.70ยฑ0.12 | 90.62ยฑ0.50 | 55.38ยฑ1.74 | 70.17ยฑ1.99 | 71.43ยฑ3.57 | 63.10ยฑ1.28 | 63.46ยฑ0.00 | 75.68 |
| EuroBERT-2.1B | 70.52ยฑ14.22 | 92.34ยฑ0.19 | 60.45ยฑ0.70 | 85.40ยฑ1.36 | 6.84ยฑ6.81 | 68.99ยฑ0.67 | 92.50ยฑ1.03 | 82.94ยฑ3.10 | 66.44ยฑ32.36 | 91.03ยฑ0.29 | 81.56ยฑ16.62 | 93.56ยฑ0.25 | 53.29ยฑ0.79 | 77.23ยฑ6.77 | 82.86ยฑ5.14 | 57.90ยฑ4.75 | 63.46ยฑ0.00 | 72.19 |
* = dev split (GLUE/SuperGLUE test labels hidden). The 5 multilingual columns are labeled test.
SeaHorse & STS-B are Spearmanร100. All other tasks are accuracy.
Inference speed
The LFM2 backbone was built for fast inference, and the encoders inherit it. While ModernBERT-base is faster at short sequences in Apple GPU inputs, LFM2.5-Encoders overtake it as inputs grow. At long input sequences of 8k on CPU, the encoders run 3.3ร faster than ModernBERT-base.
๐ง Fine-tuning
LFM2.5-Encoder-230M follows standard BERT-style fine-tuning. Attach a task head to the encoder body and train end-to-end. Suggested starting points (tune per task):
| Hyperparameter | Suggested range |
|---|---|
| Learning rate | 1e-5 โ 5e-5 |
| Warmup ratio | 0.1 |
| Weight decay | 0.1 |
| Epochs | 3 โ 20 (early stopping, patience 3) |
| Precision | bf16 autocast (fp32 master weights) |
๐ฌ Contact
- Got questions or want to connect? Join our Discord community
- If you are interested in custom solutions with edge deployment, please contact our sales team.
Citation
@article{liquidAI2026Encoders,
author = {Liquid AI},
title = {LFM2.5-Encoders: Fast at Long Context, Even on CPU},
journal = {Liquid AI Blog},
year = {2026},
note = {www.liquid.ai/blog/lfm2-5-encoders},
}
- Downloads last month
- 8,295
Model tree for LiquidAI/LFM2.5-Encoder-230M
Base model
LiquidAI/LFM2.5-230M-Base


