Liquid AI
Try LFM โ€ข Docs โ€ข LEAP โ€ข Discord

LFM2.5-Encoder-230M

LFM2.5-Encoder is a family of multilingual bidirectional encoders built on the LFM2 architecture, available in two sizes:

  • LFM2.5-Encoder-230M (this model) โ€” a lightweight encoder for tight latency and memory budgets, punching above its size class.
  • LFM2.5-Encoder-350M โ€” a larger sibling for maximum downstream quality.

Both are masked language models with full bidirectional attention, designed to be fine-tuned into task-specific models (classification, token classification, retrieval, reranking, and semantic similarity) across 15 languages, and to run efficiently on-device.

Find more details about our encoders in our blog post.

Key highlights:

  • Highly capable for its size. On par with the best similarly sized encoders and well ahead of our own retrieval siblings.
  • General-purpose. 8k context, strong across NLI, paraphrase, sentiment, and multilingual tasks.
  • Fast and on-device. Matches or beats ModernBERT throughput, with a long-context edge on CPU; runs in the browser on WebGPU.

๐Ÿ’ป Demos: We built the demos below from fine-tuned LFM2.5-Encoders. Each one runs in a CPU-only Hugging Face space:

  • Zero-shot prompt routing โ€” define your own routing lanes as free text. The model scores the whole prompt against every lane in one pass.
  • Zero-shot policy linting โ€” check text against your company's rules, written as free text. It scores every token against every rule in one pass.
  • Spell checking โ€” correct misspellings token by token.
  • PII detection โ€” spot and remove 40 kinds of personal information across 16 languages.
  • Masked-diffusion text generation โ€” bonus: run the encoder as a chatbot that generates text by iteratively unmasking instead of left to right.

๐Ÿ“„ Model details

Property LFM2.5-Encoder-230M LFM2.5-Encoder-350M
Type Bidirectional encoder (masked language model) Bidirectional encoder (masked language model)
Backbone LFM2 LFM2
Total parameters ~229.7M ~354.5M
Hidden size 1024 1024
Vocabulary size 65,536 65,536
Context length 8,192 tokens 8,192 tokens
License LFM Open License v1.0 LFM Open License v1.0

Supported languages: English, German, Spanish, French, Italian, Dutch, Polish, Portuguese, Arabic, Hindi, Japanese, Russian, Turkish, Vietnamese, Chinese (15).

Architecture. LFM2.5-Encoder is built on the LFM2 hybrid backbone, which interleaves gated short-convolution blocks with grouped-query attention. For encoder use, the causal mask is replaced with full bidirectional (non-causal) attention and the model is trained with a masked language modeling head. The encoder body is exposed as Lfm2BidirectionalModel; masked-LM loading uses Lfm2BidirectionalForMaskedLM. Both are wired through auto_map and require trust_remote_code=True.

Lfm2BidirectionalForMaskedLM(
  (lfm2): Lfm2BidirectionalModel
  (lm_head): Linear(in_features=1024, out_features=65536, bias=False)
)

Training. LFM2.5-Encoder-230M is adapted from the LFM2 base and trained with a masked language modeling objective on a large multilingual corpus. Pre-training uses a two-stage schedule that extends the context window to up to 8,192 tokens.

We recommend fine-tuning LFM2.5-Encoder-230M for a range of downstream tasks, such as:

  • Text classification: sentiment, topic, intent/routing, moderation, and business-text linting.
  • Token classification: named-entity recognition, span extraction, and sequence labeling.
  • Retrieval and reranking: a backbone for dense embedding or late-interaction (ColBERT-style) retrievers.
  • Semantic similarity: STS, paraphrase, and duplicate detection.
  • Natural language inference and extractive QA: sentence-pair reasoning and answer-span extraction.

Its small footprint makes it especially well-suited to on-device and browser (WebGPU) deployment.

๐Ÿƒ How to run

Install the latest version of transformers:

pip install -U transformers

Run masked-token prediction:

from transformers import AutoModelForMaskedLM, AutoTokenizer
import torch

tok = AutoTokenizer.from_pretrained("LiquidAI/LFM2.5-Encoder-230M", trust_remote_code=True)
mlm = AutoModelForMaskedLM.from_pretrained("LiquidAI/LFM2.5-Encoder-230M", trust_remote_code=True)

text = f"The capital of France is {tok.mask_token}."
enc = tok(text, return_tensors="pt")
with torch.no_grad():
    logits = mlm(**enc).logits
pos = (enc["input_ids"][0] == tok.mask_token_id).nonzero()[0].item()
print([tok.decode([t]).strip() for t in logits[0, pos].topk(5).indices.tolist()])
# -> ['Paris', 'Strasbourg', 'Paris', 'Lyon', 'Versailles']

For downstream tasks, load the encoder body and attach your own head (classification, token classification, regression, retrieval):

from transformers import AutoModel
body = AutoModel.from_pretrained("LiquidAI/LFM2.5-Encoder-230M", trust_remote_code=True)

If your GPU supports it, we recommend using LFM2.5-Encoder-230M with Flash Attention 2 to reach the highest efficiency. To do so, install Flash Attention as follows, then use the model as normal:

pip install flash-attn

๐Ÿ“Š Performance

For each benchmark task, we run a full supervised fine-tune and report that fine-tuned model's score. The results below span 14 models across 17 tasks from GLUE, SuperGLUE, and multilingual classification tasks. The full evaluation harness is open-sourced in the eurobert-repro repository.

benchmark_ranking

17-task results (avg@5 fresh seeds ยฑ std)

Rank Model Params 17-task mean ยฑ std
1 XLM-R XL (3.5B) 3.5B 83.06 ยฑ1.16
2 ModernBERT-large (395M) 395M 81.68 ยฑ2.49
3 XLM-R large (560M) 560M 81.34 ยฑ1.66
4 LFM2.5-Encoder-350M (ours) 350M 81.02 ยฑ1.00
5 mDeBERTa-v3 (280M) 280M 80.37 ยฑ1.06
6 LFM2.5-Encoder-230M (ours) 230M 79.29 ยฑ1.02
7 ModernBERT-base (149M) 149M 78.19 ยฑ1.39
8 XLM-R base (280M) 280M 77.46 ยฑ1.63
9 EuroBERT-210M 210M 76.87 ยฑ2.00
10 mGTE-MLM (305M) 305M 76.53 ยฑ1.85
11 LFM2.5-ColBERT-350M 350M 76.18 ยฑ1.25
12 EuroBERT-610M 610M 75.87 ยฑ2.03
13 LFM2.5-Embedding-350M 350M 75.68 ยฑ0.83
14 EuroBERT-2.1B 2.1B 72.19 ยฑ5.59
Click to expand per-task results โ€” all 17 tasks (avg@5 fresh seeds ยฑ std) ### Per-task results โ€” all 17 tasks (avg@5 fresh seeds ยฑ std)
Model XNLI PAWS-X Amazon MASSIVE SeaHorse CoLA* SST-2* MRPC* STS-B* QQP* MNLI* QNLI* RTE* BoolQ* CB* WiC* WSC* ALL
XLM-R XL (3.5B) 87.12ยฑ0.45 93.30ยฑ0.41 62.29ยฑ0.05 88.21ยฑ0.32 59.51ยฑ3.52 84.58ยฑ1.18 95.69ยฑ0.30 87.65ยฑ1.69 90.20ยฑ0.93 91.72ยฑ0.07 90.09ยฑ0.11 94.47ยฑ0.28 82.38ยฑ3.51 83.70ยฑ0.24 89.88ยฑ3.72 66.55ยฑ1.37 64.62ยฑ1.58 83.06
ModernBERT-large (395M) 81.76ยฑ0.38 92.46ยฑ0.18 60.42ยฑ0.15 85.65ยฑ0.94 40.20ยฑ17.18 83.37ยฑ0.20 96.10ยฑ0.53 88.14ยฑ1.79 92.16ยฑ0.24 91.81ยฑ0.13 90.65ยฑ0.18 94.36ยฑ0.10 81.59ยฑ4.92 81.68ยฑ2.34 88.21ยฑ3.24 70.16ยฑ2.57 69.81ยฑ7.18 81.68
XLM-R large (560M) 84.69ยฑ0.59 93.23ยฑ0.78 61.58ยฑ0.13 88.50ยฑ0.17 56.12ยฑ2.31 83.34ยฑ1.75 93.83ยฑ1.04 88.77ยฑ2.15 91.35ยฑ0.23 90.48ยฑ0.23 88.29ยฑ0.07 93.08ยฑ0.23 80.79ยฑ3.14 80.54ยฑ0.79 78.21ยฑ8.69 66.24ยฑ5.41 63.65ยฑ0.43 81.34
LFM2.5-Encoder-350M (ours) 79.82ยฑ0.29 91.53ยฑ0.73 60.57ยฑ0.09 85.70ยฑ0.13 54.96ยฑ0.41 84.43ยฑ0.81 95.11ยฑ0.26 87.21ยฑ1.86 91.59ยฑ0.05 92.08ยฑ0.10 89.03ยฑ0.17 93.97ยฑ0.23 75.23ยฑ3.84 81.52ยฑ0.69 83.21ยฑ2.04 69.66ยฑ2.23 61.73ยฑ3.15 81.02
mDeBERTa-v3 (280M) 83.01ยฑ0.47 92.59ยฑ0.39 60.62ยฑ0.31 87.64ยฑ0.51 54.97ยฑ2.04 83.91ยฑ1.03 92.41ยฑ0.96 85.39ยฑ2.59 89.87ยฑ0.22 90.23ยฑ0.15 86.30ยฑ0.18 91.99ยฑ0.35 69.75ยฑ2.03 78.29ยฑ1.39 88.21ยฑ2.40 67.71ยฑ3.05 63.46ยฑ0.00 80.37
LFM2.5-Encoder-230M (ours) 77.63ยฑ0.31 90.86ยฑ0.24 59.97ยฑ0.21 85.52ยฑ0.69 54.61ยฑ0.62 81.42ยฑ1.56 94.08ยฑ0.37 80.20ยฑ2.66 90.99ยฑ0.12 91.71ยฑ0.07 87.98ยฑ0.22 92.96ยฑ0.37 67.29ยฑ1.97 76.54ยฑ1.01 83.21ยฑ4.48 70.31ยฑ1.34 62.69ยฑ1.05 79.29
ModernBERT-base (149M) 76.64ยฑ0.31 92.17ยฑ0.15 58.98ยฑ0.11 85.32ยฑ0.21 45.19ยฑ1.72 83.07ยฑ1.93 94.79ยฑ0.52 84.46ยฑ2.70 90.75ยฑ0.14 91.23ยฑ0.11 88.68ยฑ0.17 93.04ยฑ0.36 58.70ยฑ1.94 74.78ยฑ4.10 81.07ยฑ6.75 66.90ยฑ2.39 63.46ยฑ0.00 78.19
XLM-R base (280M) 78.20ยฑ0.80 91.36ยฑ0.41 60.01ยฑ0.12 87.47ยฑ0.49 51.08ยฑ3.75 81.17ยฑ1.07 91.97ยฑ0.18 86.47ยฑ0.76 88.27ยฑ0.36 89.21ยฑ0.04 83.07ยฑ0.21 90.17ยฑ0.32 62.60ยฑ7.03 71.43ยฑ1.64 79.64ยฑ7.53 61.25ยฑ3.00 63.46ยฑ0.00 77.46
EuroBERT-210M 80.83ยฑ0.35 91.94ยฑ0.31 59.94ยฑ0.16 86.36ยฑ0.67 45.16ยฑ16.60 72.75ยฑ1.30 90.64ยฑ0.92 80.74ยฑ2.99 89.29ยฑ0.23 90.75ยฑ0.09 85.63ยฑ0.28 91.49ยฑ0.27 54.95ยฑ2.56 71.68ยฑ2.35 86.79ยฑ2.40 64.64ยฑ2.01 63.27ยฑ0.43 76.87
mGTE-MLM (305M) 80.32ยฑ0.20 91.73ยฑ0.26 60.26ยฑ0.10 87.79ยฑ0.20 51.58ยฑ1.31 75.44ยฑ4.66 91.19ยฑ1.00 86.32ยฑ1.48 87.77ยฑ0.64 89.82ยฑ0.09 84.14ยฑ0.15 90.94ยฑ0.40 58.34ยฑ3.09 69.32ยฑ3.72 73.21ยฑ6.80 59.34ยฑ7.41 63.46ยฑ0.00 76.53
LFM2.5-ColBERT-350M 78.77ยฑ0.47 89.74ยฑ0.44 59.92ยฑ0.13 86.65ยฑ0.17 47.95ยฑ1.13 71.06ยฑ1.06 90.94ยฑ0.78 73.43ยฑ7.22 89.38ยฑ0.28 91.11ยฑ0.14 84.70ยฑ0.23 90.66ยฑ0.18 59.13ยฑ2.30 74.25ยฑ1.61 81.79ยฑ2.93 62.04ยฑ2.12 63.46ยฑ0.00 76.18
EuroBERT-610M 84.61ยฑ0.34 91.84ยฑ0.94 60.64ยฑ0.08 86.03ยฑ0.99 12.91ยฑ8.05 70.60ยฑ2.12 92.52ยฑ0.66 85.20ยฑ1.34 89.82ยฑ0.22 91.13ยฑ0.09 87.95ยฑ0.19 92.57ยฑ0.36 59.28ยฑ7.93 76.86ยฑ1.55 85.71ยฑ4.37 58.71ยฑ5.30 63.46ยฑ0.00 75.87
LFM2.5-Embedding-350M 78.59ยฑ0.11 89.13ยฑ0.63 60.47ยฑ0.13 87.03ยฑ0.21 50.19ยฑ0.90 72.54ยฑ0.68 91.70ยฑ0.69 77.45ยฑ1.31 89.38ยฑ0.09 91.14ยฑ0.13 84.70ยฑ0.12 90.62ยฑ0.50 55.38ยฑ1.74 70.17ยฑ1.99 71.43ยฑ3.57 63.10ยฑ1.28 63.46ยฑ0.00 75.68
EuroBERT-2.1B 70.52ยฑ14.22 92.34ยฑ0.19 60.45ยฑ0.70 85.40ยฑ1.36 6.84ยฑ6.81 68.99ยฑ0.67 92.50ยฑ1.03 82.94ยฑ3.10 66.44ยฑ32.36 91.03ยฑ0.29 81.56ยฑ16.62 93.56ยฑ0.25 53.29ยฑ0.79 77.23ยฑ6.77 82.86ยฑ5.14 57.90ยฑ4.75 63.46ยฑ0.00 72.19

* = dev split (GLUE/SuperGLUE test labels hidden). The 5 multilingual columns are labeled test. SeaHorse & STS-B are Spearmanร—100. All other tasks are accuracy.

Inference speed

The LFM2 backbone was built for fast inference, and the encoders inherit it. While ModernBERT-base is faster at short sequences in Apple GPU inputs, LFM2.5-Encoders overtake it as inputs grow. At long input sequences of 8k on CPU, the encoders run 3.3ร— faster than ModernBERT-base.

inference_perf_cpu

inference_perf_gpu

๐Ÿ”ง Fine-tuning

LFM2.5-Encoder-230M follows standard BERT-style fine-tuning. Attach a task head to the encoder body and train end-to-end. Suggested starting points (tune per task):

Hyperparameter Suggested range
Learning rate 1e-5 โ€“ 5e-5
Warmup ratio 0.1
Weight decay 0.1
Epochs 3 โ€“ 20 (early stopping, patience 3)
Precision bf16 autocast (fp32 master weights)

๐Ÿ“ฌ Contact

Citation

@article{liquidAI2026Encoders,
  author = {Liquid AI},
  title = {LFM2.5-Encoders: Fast at Long Context, Even on CPU},
  journal = {Liquid AI Blog},
  year = {2026},
  note = {www.liquid.ai/blog/lfm2-5-encoders},
}
Downloads last month
8,295
Safetensors
Model size
0.2B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for LiquidAI/LFM2.5-Encoder-230M

Finetuned
(11)
this model
Finetunes
6 models
Quantizations
4 models

Space using LiquidAI/LFM2.5-Encoder-230M 1

Collection including LiquidAI/LFM2.5-Encoder-230M

Article mentioning LiquidAI/LFM2.5-Encoder-230M