arXiv:2505.19312cs.IR2025-05EMNLP被引 2

构建跨文档多模态检索框架,统一处理多种文本格式。

DocMMIR: A Framework for Document Multi-modal Information Retrieval

  • 提出统一多模态文档检索的框架,支持不同格式与领域。
  • 在45万样本数据集上,模型性能提升31% MRR@10。
  • 开源数据与代码,适合文档智能与跨模态研究者。

无监督表示学习与大规模预训练视觉语言模型的快速发展显著提升了跨模态检索性能。然而,现有多模态信息检索(MMIR)研究缺乏对文档级检索的全面探索,且缺少该粒度下的跨域数据集。为此,我们提出DocMMIR,一种专为统一多种文档格式与领域(如维基百科、arXiv论文、演示文稿)设计的多模态文档检索框架。构建了一个包含45万样本的大规模跨域多模态基准数据集,系统整合文本与视觉信息。实验表明,当前主流多模态大模型(CLIP、BLIP2、SigLIP-2、ALIGN)在该任务中表现有限,仅CLIP在零样本下具备合理性能。通过系统研究训练策略,包括跨模态融合方法与损失函数,我们针对该数据集定制训练方案,使CLIP在MRR@10上较零样本基线提升31%。所有数据与代码已公开于https://github.com/J1mL1/DocMMIR。

原文摘要 · Abstract (English)

The rapid advancement of unsupervised representation learning and large-scale pre-trained vision-language models has significantly improved cross-modal retrieval tasks. However, existing multi-modal information retrieval (MMIR) studies lack a comprehensive exploration of document-level retrieval and suffer from the absence of cross-domain datasets at this granularity. To address this limitation, we introduce DocMMIR, a novel multi-modal document retrieval framework designed explicitly to unify diverse document formats and domains, including Wikipedia articles, scientific papers (arXiv), and presentation slides, within a comprehensive retrieval scenario. We construct a large-scale cross-domain multimodal benchmark, comprising 450K samples, which systematically integrates textual and visual information. Our comprehensive experimental analysis reveals substantial limitations in current state-of-the-art MLLMs (CLIP, BLIP2, SigLIP-2, ALIGN) when applied to our tasks, with only CLIP demonstrating reasonable zero-shot performance. Furthermore, we conduct a systematic investigation of training strategies, including cross-modal fusion methods and loss functions, and develop a tailored approach to train CLIP on our benchmark. This results in a +31% improvement in MRR@10 compared to the zero-shot baseline. All our data and code are released in https://github.com/J1mL1/DocMMIR.

文档检索多模态CLIParXiv

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。