arXiv:2505.01457cs.IRcs.CV2025-05被引 5

提出多粒度视觉文档检索框架,提升图文混排文档的精准查找能力。

A Multi-Granularity Retrieval Framework for Visually-Rich Documents

  • 分层编码+模态感知检索,融合文本与图像信息
  • 无需微调,布局感知搜索使准确率达65.56
  • 适合需要高效处理图文混合文档的研究者

检索增强生成(RAG)系统主要聚焦于文本检索,难以有效处理包含文本、图像、表格和图表的视觉丰富文档。为此,我们提出一个统一的多粒度多模态检索框架,针对MMDocIR和M2KR两个基准任务。该方法结合分层编码策略、模态感知检索机制以及基于视觉语言模型(VLM)的候选过滤,有效捕捉文本与视觉模态间的复杂依赖关系。通过利用现成的视觉语言模型并采用无训练的混合检索策略,框架在无需任务特定微调的情况下实现稳健性能。实验表明,引入布局感知搜索和VLM候选验证显著提升检索准确率,最高达到65.56。本工作凸显了可扩展、可复现方案在推进多模态文档检索中的潜力。

原文摘要 · Abstract (English)

Retrieval-augmented generation (RAG) systems have predominantly focused on text-based retrieval, limiting their effectiveness in handling visually-rich documents that encompass text, images, tables, and charts. To bridge this gap, we propose a unified multi-granularity multimodal retrieval framework tailored for two benchmark tasks: MMDocIR and M2KR. Our approach integrates hierarchical encoding strategies, modality-aware retrieval mechanisms, and vision-language model (VLM)-based candidate filtering to effectively capture and utilize the complex interdependencies between textual and visual modalities. By leveraging off-the-shelf vision-language models and implementing a training-free hybrid retrieval strategy, our framework demonstrates robust performance without the need for task-specific fine-tuning. Experimental evaluations reveal that incorporating layout-aware search and VLM-based candidate verification significantly enhances retrieval accuracy, achieving a top performance score of 65.56. This work underscores the potential of scalable and reproducible solutions in advancing multimodal document retrieval systems.

多模态检索视觉文档VLMRAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。