arXiv:2505.07730cs.IR2025-05被引 6

晚交互机制显著提升文档图像检索效果,但带来计算开销

Reproducibility, Replicability, and Insights into Visual Document Retrieval with Late Interaction

  • 采用晚交互机制,在视觉文档检索中提升匹配精度
  • 晚交互使检索准确率明显提高,但推理速度下降
  • 适合关注检索效率与精度平衡的研究者

视觉文档检索(VDR)是一个新兴研究领域,旨在直接编码和检索文档图像,避免依赖光学字符识别(OCR)进行文档搜索。近期,ColPali通过引入晚交互机制显著提升了检索效果。本文系统评估了具有和不具晚交互机制的VDR方法在多个预训练视觉-语言模型上的表现,验证了晚交互在提升检索有效性方面的显著优势;同时发现该机制在推理阶段引入了较大的计算开销。此外,我们评估了VDR模型对文本输入的适应性,并在所提出的基准数据集上测试其在文本密集型场景下的鲁棒性,尤其是在扩展索引机制时的表现。进一步分析表明,尽管查询词元无法像文本检索那样直接匹配图像块,但它们倾向于与视觉相似的词元或其邻近区域匹配。

原文摘要 · Abstract (English)

Visual Document Retrieval (VDR) is an emerging research area that focuses on encoding and retrieving document images directly, bypassing the dependence on Optical Character Recognition (OCR) for document search. A recent advance in VDR was introduced by ColPali, which significantly improved retrieval effectiveness through a late interaction mechanism. ColPali's approach demonstrated substantial performance gains over existing baselines that do not use late interaction on an established benchmark. In this study, we investigate the reproducibility and replicability of VDR methods with and without late interaction mechanisms by systematically evaluating their performance across multiple pre-trained vision-language models. Our findings confirm that late interaction yields considerable improvements in retrieval effectiveness; however, it also introduces computational inefficiencies during inference. Additionally, we examine the adaptability of VDR models to textual inputs and assess their robustness across text-intensive datasets within the proposed benchmark, particularly when scaling the indexing mechanism. Furthermore, our research investigates the specific contributions of late interaction by looking into query-patch matching in the context of visual document retrieval. We find that although query tokens cannot explicitly match image patches as in the text retrieval scenario, they tend to match the patch contains visually similar tokens or their surrounding patches.

视觉文档检索晚交互检索效率多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。