提出从底层融合视觉与文本信息的检索框架,提升复杂多模态查询理解能力。
Joint Fusion and Encoding: Advancing Multimodal Retrieval from the Ground Up
- 采用单塔架构,在编码阶段即实现图文早期融合。
- 两阶段训练使模型在多模态检索任务中全面超越传统方法。
- 特别适合需要深度跨模态交互的复杂查询场景。
信息检索对当今互联网应用至关重要,但传统语义匹配技术难以捕捉复杂查询所需的细粒度跨模态交互。尽管晚期融合的双塔架构通过独立编码视觉与文本数据并在高层合并来弥补这一差距,却常忽略对整体理解至关重要的细微交互。本文系统评估了这些局限,并提出一种统一的检索框架,从底层融合视觉与文本线索,实现早期跨模态交互以增强上下文理解。通过后训练适应与指令微调的两阶段训练过程,我们使用简单的单塔架构将多模态大语言模型(MLLM)适配为检索器。该方法在多种检索场景中表现优于传统方法,尤其在处理需模态融合的复杂多模态输入时优势更明显。结果显示,联合融合编码器在需模态融合的任务上提升更大,凸显早期集成策略的变革潜力,为情境感知且高效的检索指明了新方向。
原文摘要 · Abstract (English)
Information retrieval is indispensable for today's Internet applications, yet traditional semantic matching techniques often fall short in capturing the fine-grained cross-modal interactions required for complex queries. Although late-fusion two-tower architectures attempt to bridge this gap by independently encoding visual and textual data before merging them at a high level, they frequently overlook the subtle interplay essential for comprehensive understanding. In this work, we rigorously assess these limitations and introduce a unified retrieval framework that fuses visual and textual cues from the ground up, enabling early cross-modal interactions for enhancing context interpretation. Through a two-stage training process--comprising post-training adaptation followed by instruction tuning--we adapt MLLMs as retrievers using a simple one-tower architecture. Our approach outperforms conventional methods across diverse retrieval scenarios, particularly when processing complex multi-modal inputs. Notably, the joint fusion encoder yields greater improvements on tasks that require modality fusion compared to those that do not, underscoring the transformative potential of early integration strategies and pointing toward a promising direction for contextually aware and effective information retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。