JevEmbed now supports CLM-v0.1-8B for Choice, Score, and Noul decisions.
🔄 In a Training Loop
Xinping Zhao
Yuki131
AI & ML interests
LLMs, RAG, Embedding, Reranker——A Pokémon Trainer on a journey to become a Pokémon Master.
Recent Activity
liked a model about 9 hours ago
HIT-TMG/JevEmbed-Qwen3-Embedding-4B updated a collection about 14 hours ago
Lychee-JevEmbed updated a model about 14 hours ago
HIT-TMG/JevEmbed-Qwen3-Embedding-4BOrganizations
replied to their post 4 days ago
replied to their post 4 days ago
JevEmbed-Data is available for fine-tuning, with 1.67 million labeled Choice, Score, and Noul questions.
Post
2777
Meet JevEmbed: an open-source framework for embedding-based decisions
Turn embeddings into decisions. Choose, score, and judge with your choice of embedding model.
We’ve open-sourced JevEmbed, a Python framework for three structured decision tasks:
🎯 Choice: select from a set of candidates
📊 Score: rate against ordered criteria
✅ Noul: judge whether a statement or question holds
🔧 JevEmbed currently includes configurations for KaLM, Qwen3, and E5 embedding models. You can use it through a Python API, CLI, or optional HTTP server. It also supports local LoRA fine-tuning, so you can adapt an embedding model to your own decision tasks and load the resulting adapter for local inference.
Fine-tuning results
📈 We trained KaLM-Embedding-V2.5 and Qwen3-Embedding-0.6B on the 79,116-example training split of Open-Jev’s release-v2-redistributable subset. We then evaluated them on 3,495 hard-label questions from the same subset’s held-out validation split.
ZefanCai/Open-Jev
KaLM-Embedding-V2.5: 30.24% base accuracy → 76.68% after LoRA fine-tuning
Qwen3-Embedding-0.6B: 30.73% base accuracy → 84.06% after LoRA fine-tuning
KaLM-Embedding/KaLM-embedding-multilingual-mini-instruct-v2.5
Qwen/Qwen3-Embedding-0.6B
These results are specific to that validation split. Performance on other tasks and datasets may differ.
JevEmbed also supports Choice tasks with more than 255 candidates, making it useful for classification and routing problems with large candidate sets.
Explore the framework, open an issue, or tell us what decision task you would try it on:
🔗 https://github.com/HITsz-TMG/JevEmbed
#Embeddings #LoRA #SentenceTransformers #OpenSource #JevEmbed
Turn embeddings into decisions. Choose, score, and judge with your choice of embedding model.
We’ve open-sourced JevEmbed, a Python framework for three structured decision tasks:
🎯 Choice: select from a set of candidates
📊 Score: rate against ordered criteria
✅ Noul: judge whether a statement or question holds
🔧 JevEmbed currently includes configurations for KaLM, Qwen3, and E5 embedding models. You can use it through a Python API, CLI, or optional HTTP server. It also supports local LoRA fine-tuning, so you can adapt an embedding model to your own decision tasks and load the resulting adapter for local inference.
Fine-tuning results
📈 We trained KaLM-Embedding-V2.5 and Qwen3-Embedding-0.6B on the 79,116-example training split of Open-Jev’s release-v2-redistributable subset. We then evaluated them on 3,495 hard-label questions from the same subset’s held-out validation split.
ZefanCai/Open-Jev
KaLM-Embedding-V2.5: 30.24% base accuracy → 76.68% after LoRA fine-tuning
Qwen3-Embedding-0.6B: 30.73% base accuracy → 84.06% after LoRA fine-tuning
KaLM-Embedding/KaLM-embedding-multilingual-mini-instruct-v2.5
Qwen/Qwen3-Embedding-0.6B
These results are specific to that validation split. Performance on other tasks and datasets may differ.
JevEmbed also supports Choice tasks with more than 255 candidates, making it useful for classification and routing problems with large candidate sets.
Explore the framework, open an issue, or tell us what decision task you would try it on:
🔗 https://github.com/HITsz-TMG/JevEmbed
#Embeddings #LoRA #SentenceTransformers #OpenSource #JevEmbed
posted an update 5 days ago
Post
2777
Meet JevEmbed: an open-source framework for embedding-based decisions
Turn embeddings into decisions. Choose, score, and judge with your choice of embedding model.
We’ve open-sourced JevEmbed, a Python framework for three structured decision tasks:
🎯 Choice: select from a set of candidates
📊 Score: rate against ordered criteria
✅ Noul: judge whether a statement or question holds
🔧 JevEmbed currently includes configurations for KaLM, Qwen3, and E5 embedding models. You can use it through a Python API, CLI, or optional HTTP server. It also supports local LoRA fine-tuning, so you can adapt an embedding model to your own decision tasks and load the resulting adapter for local inference.
Fine-tuning results
📈 We trained KaLM-Embedding-V2.5 and Qwen3-Embedding-0.6B on the 79,116-example training split of Open-Jev’s release-v2-redistributable subset. We then evaluated them on 3,495 hard-label questions from the same subset’s held-out validation split.
ZefanCai/Open-Jev
KaLM-Embedding-V2.5: 30.24% base accuracy → 76.68% after LoRA fine-tuning
Qwen3-Embedding-0.6B: 30.73% base accuracy → 84.06% after LoRA fine-tuning
KaLM-Embedding/KaLM-embedding-multilingual-mini-instruct-v2.5
Qwen/Qwen3-Embedding-0.6B
These results are specific to that validation split. Performance on other tasks and datasets may differ.
JevEmbed also supports Choice tasks with more than 255 candidates, making it useful for classification and routing problems with large candidate sets.
Explore the framework, open an issue, or tell us what decision task you would try it on:
🔗 https://github.com/HITsz-TMG/JevEmbed
#Embeddings #LoRA #SentenceTransformers #OpenSource #JevEmbed
Turn embeddings into decisions. Choose, score, and judge with your choice of embedding model.
We’ve open-sourced JevEmbed, a Python framework for three structured decision tasks:
🎯 Choice: select from a set of candidates
📊 Score: rate against ordered criteria
✅ Noul: judge whether a statement or question holds
🔧 JevEmbed currently includes configurations for KaLM, Qwen3, and E5 embedding models. You can use it through a Python API, CLI, or optional HTTP server. It also supports local LoRA fine-tuning, so you can adapt an embedding model to your own decision tasks and load the resulting adapter for local inference.
Fine-tuning results
📈 We trained KaLM-Embedding-V2.5 and Qwen3-Embedding-0.6B on the 79,116-example training split of Open-Jev’s release-v2-redistributable subset. We then evaluated them on 3,495 hard-label questions from the same subset’s held-out validation split.
ZefanCai/Open-Jev
KaLM-Embedding-V2.5: 30.24% base accuracy → 76.68% after LoRA fine-tuning
Qwen3-Embedding-0.6B: 30.73% base accuracy → 84.06% after LoRA fine-tuning
KaLM-Embedding/KaLM-embedding-multilingual-mini-instruct-v2.5
Qwen/Qwen3-Embedding-0.6B
These results are specific to that validation split. Performance on other tasks and datasets may differ.
JevEmbed also supports Choice tasks with more than 255 candidates, making it useful for classification and routing problems with large candidate sets.
Explore the framework, open an issue, or tell us what decision task you would try it on:
🔗 https://github.com/HITsz-TMG/JevEmbed
#Embeddings #LoRA #SentenceTransformers #OpenSource #JevEmbed
Post
3618
Meet KaLM-Jev — your local, Jev-style judgment engine, available in Nano, Small, and Large.
Building an agent or automation workflow? Sometimes all you need is a choice, a score, or a signal that a condition holds.
Built on KaLM-Reranker-R2, KaLM-Jev turns these decisions into structured outputs through three primitives:
🔀 Choice — select among candidates, with a probability distribution.
📊 Score — return a continuous score over your defined levels.
🔍 Noul — evaluate conditions independently, so multiple conditions can hold at once.
Think support-ticket routing, bug severity scoring, human-escalation detection, or candidate tool selection for agents.
🖥️ Run locally with downloaded weights
📦 Choose from Nano / Small / Large
🔌 Integrate through HTTP or Python
⚡ Reuse cached candidate/rule representations to reduce repeated encoding
🧪 Explore included examples, bilingual semantic smoke tests, and recorded GPU validation results
No answer-text generation:
KaLM-Jev is an independent implementation based on KaLM-Reranker, not an official TypeSafe project or a guarantee of full Jev compatibility. Scores are uncalibrated; validate thresholds on your own tasks.
Code & quickstart:
https://github.com/KaLM-Embedding/KaLM-Jev
Yuki131/KaLM-Jev
We’d love to hear what you’d build with it. Try it out, share feedback, or open an issue! 🤗
#Jev #Reranker #Agents #LocalAI #OpenSource
Building an agent or automation workflow? Sometimes all you need is a choice, a score, or a signal that a condition holds.
Built on KaLM-Reranker-R2, KaLM-Jev turns these decisions into structured outputs through three primitives:
🔀 Choice — select among candidates, with a probability distribution.
📊 Score — return a continuous score over your defined levels.
🔍 Noul — evaluate conditions independently, so multiple conditions can hold at once.
Think support-ticket routing, bug severity scoring, human-escalation detection, or candidate tool selection for agents.
🖥️ Run locally with downloaded weights
📦 Choose from Nano / Small / Large
🔌 Integrate through HTTP or Python
⚡ Reuse cached candidate/rule representations to reduce repeated encoding
🧪 Explore included examples, bilingual semantic smoke tests, and recorded GPU validation results
No answer-text generation:
output_tokens = 0. Inference still runs to compute the judgments.KaLM-Jev is an independent implementation based on KaLM-Reranker, not an official TypeSafe project or a guarantee of full Jev compatibility. Scores are uncalibrated; validate thresholds on your own tasks.
Code & quickstart:
https://github.com/KaLM-Embedding/KaLM-Jev
Yuki131/KaLM-Jev
We’d love to hear what you’d build with it. Try it out, share feedback, or open an issue! 🤗
#Jev #Reranker #Agents #LocalAI #OpenSource
posted an update 8 days ago
Post
3618
Meet KaLM-Jev — your local, Jev-style judgment engine, available in Nano, Small, and Large.
Building an agent or automation workflow? Sometimes all you need is a choice, a score, or a signal that a condition holds.
Built on KaLM-Reranker-R2, KaLM-Jev turns these decisions into structured outputs through three primitives:
🔀 Choice — select among candidates, with a probability distribution.
📊 Score — return a continuous score over your defined levels.
🔍 Noul — evaluate conditions independently, so multiple conditions can hold at once.
Think support-ticket routing, bug severity scoring, human-escalation detection, or candidate tool selection for agents.
🖥️ Run locally with downloaded weights
📦 Choose from Nano / Small / Large
🔌 Integrate through HTTP or Python
⚡ Reuse cached candidate/rule representations to reduce repeated encoding
🧪 Explore included examples, bilingual semantic smoke tests, and recorded GPU validation results
No answer-text generation:
KaLM-Jev is an independent implementation based on KaLM-Reranker, not an official TypeSafe project or a guarantee of full Jev compatibility. Scores are uncalibrated; validate thresholds on your own tasks.
Code & quickstart:
https://github.com/KaLM-Embedding/KaLM-Jev
Yuki131/KaLM-Jev
We’d love to hear what you’d build with it. Try it out, share feedback, or open an issue! 🤗
#Jev #Reranker #Agents #LocalAI #OpenSource
Building an agent or automation workflow? Sometimes all you need is a choice, a score, or a signal that a condition holds.
Built on KaLM-Reranker-R2, KaLM-Jev turns these decisions into structured outputs through three primitives:
🔀 Choice — select among candidates, with a probability distribution.
📊 Score — return a continuous score over your defined levels.
🔍 Noul — evaluate conditions independently, so multiple conditions can hold at once.
Think support-ticket routing, bug severity scoring, human-escalation detection, or candidate tool selection for agents.
🖥️ Run locally with downloaded weights
📦 Choose from Nano / Small / Large
🔌 Integrate through HTTP or Python
⚡ Reuse cached candidate/rule representations to reduce repeated encoding
🧪 Explore included examples, bilingual semantic smoke tests, and recorded GPU validation results
No answer-text generation:
output_tokens = 0. Inference still runs to compute the judgments.KaLM-Jev is an independent implementation based on KaLM-Reranker, not an official TypeSafe project or a guarantee of full Jev compatibility. Scores are uncalibrated; validate thresholds on your own tasks.
Code & quickstart:
https://github.com/KaLM-Embedding/KaLM-Jev
Yuki131/KaLM-Jev
We’d love to hear what you’d build with it. Try it out, share feedback, or open an issue! 🤗
#Jev #Reranker #Agents #LocalAI #OpenSource
replied to their post about 2 months ago
Great point—recall@20 after the 32× stage would clearly show what is lost during early filtering. We’ll include it alongside final nDCG and latency. 😏
Post
2264
Test-Time Scaling for Rerankers?
Can rerankers scale at test time—not by generating longer reasoning traces, but by selectively using richer document representations?
KaLM-Reranker-V1 supports Matryoshka compression from 1× to 32×, which suggests a progressive multi-fidelity pipeline:
- Embedding retrieval → Top-100
- KaLM-Reranker @ 32× compression → Top-20
- The same reranker @ 2× compression → final ranking
The intuition is simple: cheaply screen many candidates, then allocate higher-fidelity cross-attention only to the most promising ones.
For 100@32× → 20@2×, the passage-token interaction budget is roughly 31.8% of directly running 100@2×, before fixed model overheads. The key question is whether it can retain nearly the same ranking quality.
We’re considering evaluating nDCG–latency Pareto curves.
Would you consider this a useful form of test-time scaling for retrieval?
KaLM-Embedding/KaLM-Reranker-V1-Nano
KaLM-Embedding/KaLM-Reranker-V1-Small
KaLM-Embedding/KaLM-Reranker-V1-Large
KaLM-Reranker-V1: Fast but Not Late Interaction for Compressed Document Reranking (2606.22807)
https://huggingface.co/collections/KaLM-Embedding/lychee-kalm-reranker
KaLM-Embedding
Can rerankers scale at test time—not by generating longer reasoning traces, but by selectively using richer document representations?
KaLM-Reranker-V1 supports Matryoshka compression from 1× to 32×, which suggests a progressive multi-fidelity pipeline:
- Embedding retrieval → Top-100
- KaLM-Reranker @ 32× compression → Top-20
- The same reranker @ 2× compression → final ranking
The intuition is simple: cheaply screen many candidates, then allocate higher-fidelity cross-attention only to the most promising ones.
For 100@32× → 20@2×, the passage-token interaction budget is roughly 31.8% of directly running 100@2×, before fixed model overheads. The key question is whether it can retain nearly the same ranking quality.
We’re considering evaluating nDCG–latency Pareto curves.
Would you consider this a useful form of test-time scaling for retrieval?
KaLM-Embedding/KaLM-Reranker-V1-Nano
KaLM-Embedding/KaLM-Reranker-V1-Small
KaLM-Embedding/KaLM-Reranker-V1-Large
KaLM-Reranker-V1: Fast but Not Late Interaction for Compressed Document Reranking (2606.22807)
https://huggingface.co/collections/KaLM-Embedding/lychee-kalm-reranker
posted an update about 2 months ago
Post
2264
Test-Time Scaling for Rerankers?
Can rerankers scale at test time—not by generating longer reasoning traces, but by selectively using richer document representations?
KaLM-Reranker-V1 supports Matryoshka compression from 1× to 32×, which suggests a progressive multi-fidelity pipeline:
- Embedding retrieval → Top-100
- KaLM-Reranker @ 32× compression → Top-20
- The same reranker @ 2× compression → final ranking
The intuition is simple: cheaply screen many candidates, then allocate higher-fidelity cross-attention only to the most promising ones.
For 100@32× → 20@2×, the passage-token interaction budget is roughly 31.8% of directly running 100@2×, before fixed model overheads. The key question is whether it can retain nearly the same ranking quality.
We’re considering evaluating nDCG–latency Pareto curves.
Would you consider this a useful form of test-time scaling for retrieval?
KaLM-Embedding/KaLM-Reranker-V1-Nano
KaLM-Embedding/KaLM-Reranker-V1-Small
KaLM-Embedding/KaLM-Reranker-V1-Large
KaLM-Reranker-V1: Fast but Not Late Interaction for Compressed Document Reranking (2606.22807)
https://huggingface.co/collections/KaLM-Embedding/lychee-kalm-reranker
KaLM-Embedding
Can rerankers scale at test time—not by generating longer reasoning traces, but by selectively using richer document representations?
KaLM-Reranker-V1 supports Matryoshka compression from 1× to 32×, which suggests a progressive multi-fidelity pipeline:
- Embedding retrieval → Top-100
- KaLM-Reranker @ 32× compression → Top-20
- The same reranker @ 2× compression → final ranking
The intuition is simple: cheaply screen many candidates, then allocate higher-fidelity cross-attention only to the most promising ones.
For 100@32× → 20@2×, the passage-token interaction budget is roughly 31.8% of directly running 100@2×, before fixed model overheads. The key question is whether it can retain nearly the same ranking quality.
We’re considering evaluating nDCG–latency Pareto curves.
Would you consider this a useful form of test-time scaling for retrieval?
KaLM-Embedding/KaLM-Reranker-V1-Nano
KaLM-Embedding/KaLM-Reranker-V1-Small
KaLM-Embedding/KaLM-Reranker-V1-Large
KaLM-Reranker-V1: Fast but Not Late Interaction for Compressed Document Reranking (2606.22807)
https://huggingface.co/collections/KaLM-Embedding/lychee-kalm-reranker
reacted to Banaxi-Tech's post with 👀 3 months ago
Post
10768
A new model is coming!
Its going to take a long time on my 5070 Ti so expect a release in ~1 month.
We think this model is going to be SOTA For its size.
Our Mini Version will be 25M Parameters and Pro with 140M.
The Pro version has a 3072 Context Window (Extensible to up to 6K with RoPE) And the Mini version has a context window of 4096 (Up to 8K with RoPE)
Meanwhile we are currently working on a Instruct Version of our BananaMind 1.5 Base.
The training will start this weekend
We are very exited to release it when its done!
Its going to take a long time on my 5070 Ti so expect a release in ~1 month.
We think this model is going to be SOTA For its size.
Our Mini Version will be 25M Parameters and Pro with 140M.
The Pro version has a 3072 Context Window (Extensible to up to 6K with RoPE) And the Mini version has a context window of 4096 (Up to 8K with RoPE)
Meanwhile we are currently working on a Instruct Version of our BananaMind 1.5 Base.
The training will start this weekend
We are very exited to release it when its done!
reacted to tomaarsen's post with 🔥 3 months ago
Post
1719
🤗 Announcing the Ettin Reranker family: six new state-of-the-art CrossEncoder rerankers for search from 17M to 1B parameters, plus the full training data and the ~150-line recipe. Built on the Ettin ModernBERT encoders, Apache 2.0. Details:
All six were trained with the same single-stage pointwise MSE distillation recipe, with mixedbread-ai/mxbai-rerank-large-v2 (1.54B) as the teacher. Only the learning rate and per-device batch size change between sizes. The 1B student matches the teacher within 0.0001 NDCG@10 on MTEB(eng, v2) Retrieval, the 150M is the strongest reranker I tested in the under-600M range, and the 17M beats the 33M ms-marco-MiniLM-L12-v2 by +0.051 NDCG@10 at roughly half the parameter count.
Speed matters as much as quality for a reranker, since it determines whether the model fits the latency budget between retrieval and showing results. Our 17M is the fastest reranker in the whole comparison at 7517 pairs/sec on an H100. Our 150M runs 2.3x faster than the two other 150M ModernBERT-base rerankers (gte-reranker-modernbert-base and granite-embedding-reranker-english-r2) because the modular Transformer module propagates unpadded inputs through every layer rather than just the FA2 attention kernel. And our 1B is 2.4x faster than its 1.5B teacher while matching it on quality.
I bootstrapped the training recipe with the new train-sentence-transformers Agent Skill shipped in Sentence Transformers v5.5.0. Install it with
I wrote a blog post walking through usage, results across six embedder pairings, the speed story, and the complete training script. Check it out, or just point your Agent to the URL:
https://huggingface.co/blog/ettin-reranker
Collection: https://huggingface.co/collections/cross-encoder/ettin-rerankers
All six were trained with the same single-stage pointwise MSE distillation recipe, with mixedbread-ai/mxbai-rerank-large-v2 (1.54B) as the teacher. Only the learning rate and per-device batch size change between sizes. The 1B student matches the teacher within 0.0001 NDCG@10 on MTEB(eng, v2) Retrieval, the 150M is the strongest reranker I tested in the under-600M range, and the 17M beats the 33M ms-marco-MiniLM-L12-v2 by +0.051 NDCG@10 at roughly half the parameter count.
Speed matters as much as quality for a reranker, since it determines whether the model fits the latency budget between retrieval and showing results. Our 17M is the fastest reranker in the whole comparison at 7517 pairs/sec on an H100. Our 150M runs 2.3x faster than the two other 150M ModernBERT-base rerankers (gte-reranker-modernbert-base and granite-embedding-reranker-english-r2) because the modular Transformer module propagates unpadded inputs through every layer rather than just the FA2 attention kernel. And our 1B is 2.4x faster than its 1.5B teacher while matching it on quality.
I bootstrapped the training recipe with the new train-sentence-transformers Agent Skill shipped in Sentence Transformers v5.5.0. Install it with
hf skills add train-sentence-transformers --claude and ask Claude Code (or Codex / Cursor / Gemini CLI) to fine-tune a SentenceTransformer, CrossEncoder, or SparseEncoder model on your data.I wrote a blog post walking through usage, results across six embedder pairings, the speed story, and the complete training script. Check it out, or just point your Agent to the URL:
https://huggingface.co/blog/ettin-reranker
Collection: https://huggingface.co/collections/cross-encoder/ettin-rerankers
reacted to mlabonne's post with 🚀🔥🔥 12 months ago
Post
8539
LiquidAI/LFM2-8B-A1B just dropped!
8.3B params with only 1.5B active/token 🚀
> Quality ≈ 3–4B dense, yet faster than Qwen3-1.7B
> MoE designed to run on phones/laptops (llama.cpp / vLLM)
> Pre-trained on 12T tokens → strong math/code/IF
8.3B params with only 1.5B active/token 🚀
> Quality ≈ 3–4B dense, yet faster than Qwen3-1.7B
> MoE designed to run on phones/laptops (llama.cpp / vLLM)
> Pre-trained on 12T tokens → strong math/code/IF