用强化学习让模型生成更精准的多模态推理路径,提升跨模态匹配效果。
Embed-RL: Reinforcement Learning for Reasoning-Driven Multimodal Embeddings
- 设计嵌入引导的强化学习框架,让推理过程与嵌入任务对齐。
- 提出追踪性思维链,聚焦检索相关多模态线索,提升匹配精度。
- 在资源受限下超越基准模型,在两个评测集上表现更优。
利用多模态大语言模型(MLLM)已成为推动通用多模态嵌入(UME)发展的关键,以应对多样化的跨模态任务。近期研究表明,引入生成式思维链(CoT)推理相比判别式方法能显著增强任务特定表征。然而,现有生成式嵌入方法的推理路径仅限于查询的文本分析,与目标检索无关。为此,我们提出一种推理驱动的UME框架,通过嵌入引导强化学习(EG-RL)优化推理器,生成具有证据追溯性的思维链(T-CoT)。主要贡献包括:(1)设计EG-RL框架,由嵌入器为推理器提供显式监督,确保生成的思维链与嵌入任务一致;(2)引入T-CoT,提取关键多模态线索,聚焦检索相关元素,并为嵌入器提供多模态输入;(3)在计算资源有限的情况下,本框架在MMEB-V2和UVRB两个基准测试中均优于开创性嵌入模型。结构化推理中融合多模态证据并面向检索对齐,有效增强了跨模态语义一致性,提升了细粒度匹配能力与复杂场景下的泛化性能。结果表明,针对性推理优化可显著提升多模态嵌入质量,为推理驱动的UME发展提供高效实用的解决方案。
原文摘要 · Abstract (English)
Leveraging Multimodal Large Language Models (MLLMs) has become pivotal for advancing Universal Multimodal Embeddings (UME) in addressing diverse cross-modal tasks. Recent studies demonstrate that incorporating generative Chain-of-Thought (CoT) reasoning can substantially enhance task-specific representations compared to discriminative methods. However, the generated reasoning CoTs of existing generative embedding methods are limited to the textual analysis of queries and are irrelevant to the retrieval of the targets. To address these limitations, we propose a reasoning-driven UME framework that integrates Embedder-Guided Reinforcement Learning (EG-RL) to optimize the Reasoner to produce evidential Traceability CoT (T-CoT). Our key contributions are threefold: (1) We design an EG-RL framework where the Embedder provides explicit supervision to the Reasoner, ensuring the generated CoT traces are aligned with embedding tasks. (2) We introduce T-CoT, which extracts critical multimodal cues to focus on retrieval-relevant elements and provides multimodal inputs for the Embedder. (3) With limited computational resources, our framework outperforms the pioneering embedding model on both MMEB-V2 and UVRB benchmarks. The integration of multimodal evidence in structured reasoning, paired with retrieval-oriented alignment, effectively strengthens cross-modal semantic consistency and boosts the fine-grained matching capability of the model as well as the generalization across complex scenarios. Our work demonstrates that targeted reasoning optimization can significantly improve multimodal embedding quality, providing a practical and efficient solution for reasoning-driven UME development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。