Learning Multimodal Embeddings with Evidence-Aligned Readout
Abstract
Multimodal large language models can expose task-relevant evidence through generation, but producing useful evidence does not by itself determine how it enters a retrieval embedding. We study whether the semantic organization of that evidence can also specify where representations are read. To address this question, we introduce EviAlign, which couples Semantic Evidence Generation with Boundary Readout in a shared multimodal large language model. It organizes evidence into five semantic units, reads the contextualized state at each unit boundary, and aggregates these states into a single normalized embedding. Generation and contrastive retrieval objectives jointly train this shared structure. With the same trailing readout, semantic evidence and free-form CoT yield nearly identical retrieval performance, suggesting that evidence organization alone does not explain the full gain. A controlled 2times3 study compares consistent and permuted evidence organization across three readout strategies, using training targets with matched evidence spans. With five readout states and the same mean pooling, the advantage of consistent semantic organization grows from 0.65 points at length-based training positions to 2.39 at evidence boundaries, yielding a 1.74-point co-design interaction. Across 12 MMEB retrieval tasks, EviAlign achieves 76.9 average Recall@1 with 500K training pairs while retaining single-vector indexing and scoring.
Community
Excited to share our work EviAlign: Learning Multimodal Embeddings with Evidence-Aligned Readout!
Can the semantic structure of generated evidence guide how multimodal representations are extracted?
We introduce EviAlign, a framework that jointly designs semantic evidence generation and boundary-aware embedding readout within a unified MLLM.
🔹 Evidence-Aligned Readout: Extract representations at semantic evidence boundaries rather than relying on a single final token.
🔹 Generation–Retrieval Co-training: Jointly optimize evidence generation and contrastive retrieval.
🔹 Strong Retrieval Performance: Achieve 76.9 average Recall@1 across 12 MMEB tasks using 500K training pairs, while maintaining efficient single-vector retrieval.
Our findings highlight the importance of co-designing evidence organization and representation readout for multimodal retrieval.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Reason What Matters: Retrieval-Grounded Reasoning for Universal Multimodal Embeddings (2026)
- GLaQ: Grounding Latent Queries in Visual Evidence for Multimodal Reasoning (2026)
- MCD: Causal Distillation of Multimodal In-Context Learning in Large Vision-Language Models (2026)
- Is Reasoning Always Useful? Rethinking Reasoning Utility in Universal Multimodal Embeddings (2026)
- Retrieval Grounding Latent Reasoning for Dense Retrieval (2026)
- Latent-Aligned Reasoning for Multimodal Recommendation (2026)
- Skip the Talk, Re-Focus on Vision: Latent Reasoning for Reasoning Segmentation in Multimodal Large Language Models (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.33659 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper