arXiv:2511.00903cs.CL2025-11EMNLP被引 6

ColMate提升多模态文档检索效果,专为图文结构设计。

ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval

  • 用OCR预训练+掩码对比学习,更好理解图文内容
  • 在ViDoRe V2上比现有模型提升3.61%准确率
  • 适合需要处理复杂图文文档的检索任务

检索增强生成在需要专业知识或最新数据时表现良好。然而,现有多模态文档检索方法常直接套用纯文本检索技术,无论在文档编码、训练目标还是相似度计算上。为此,我们提出ColMate,一种连接多模态表征学习与文档检索的模型。ColMate采用基于OCR的预训练目标、自监督掩码对比学习目标,以及更符合多模态文档结构和视觉特征的晚期交互打分机制。在ViDoRe V2基准上,ColMate相比现有检索模型提升3.61%,展现出更强的跨领域泛化能力。

原文摘要 · Abstract (English)

Retrieval-augmented generation has proven practical when models require specialized knowledge or access to the latest data. However, existing methods for multimodal document retrieval often replicate techniques developed for text-only retrieval, whether in how they encode documents, define training objectives, or compute similarity scores. To address these limitations, we present ColMate, a document retrieval model that bridges the gap between multimodal representation learning and document retrieval. ColMate utilizes a novel OCR-based pretraining objective, a self-supervised masked contrastive learning objective, and a late interaction scoring mechanism more relevant to multimodal document structures and visual characteristics. ColMate obtains 3.61% improvements over existing retrieval models on the ViDoRe V2 benchmark, demonstrating stronger generalization to out-of-domain benchmarks.

多模态检索文档理解对比学习OCR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。