arXiv:2501.08828cs.IRcs.AI2025-01EMNLP被引 54

构建首个针对长文档多模态检索的综合基准,支持页面与版面级精细定位。

MMDocIR: Benchmarking Multimodal Retrieval for Long Documents

  • 设计页面级与版面级双任务评估体系,实现细粒度检索能力测试。
  • 包含1685个专家标注与17.3万条自动生成问题,覆盖丰富多模态内容。
  • 验证视觉模型优于纯文本模型,且基于VLM的文本检索更有效。

多模态文档检索旨在从长文档中识别并提取各类多模态内容,如图表、表格、公式和版面信息。尽管该任务日益重要,但缺乏全面且可靠的评估基准。为此,本文提出MMDocIR基准,涵盖两个任务:页面级检索与版面级检索。前者评估系统在长文档中定位相关页面的能力,后者则考察对特定版面元素(如段落、公式、图表、表格等)的检测精度,提供比整页分析更细粒度的评估方式。MMDocIR数据集包含1,685个专家标注问题和173,843个自举标签问题,是训练与评估多模态文档检索系统的宝贵资源。通过实验验证:(i) 视觉检索器显著优于纯文本模型;(ii) 使用MMDocIR训练集可有效提升多模态检索性能;(iii) 依赖视觉语言模型生成文本的检索器优于仅依赖OCR提取文本的模型。数据集已公开于https://mmdocrag.github.io/MMDocIR/。

原文摘要 · Abstract (English)

Multimodal document retrieval aims to identify and retrieve various forms of multimodal content, such as figures, tables, charts, and layout information from extensive documents. Despite its increasing popularity, there is a notable lack of a comprehensive and robust benchmark to effectively evaluate the performance of systems in such tasks. To address this gap, this work introduces a new benchmark, named MMDocIR, that encompasses two distinct tasks: page-level and layout-level retrieval. The former evaluates the performance of identifying the most relevant pages within a long document, while the later assesses the ability of detecting specific layouts, providing a more fine-grained measure than whole-page analysis. A layout refers to a variety of elements, including textual paragraphs, equations, figures, tables, or charts. The MMDocIR benchmark comprises a rich dataset featuring 1,685 questions annotated by experts and 173,843 questions with bootstrapped labels, making it a valuable resource in multimodal document retrieval for both training and evaluation. Through rigorous experiments, we demonstrate that (i) visual retrievers significantly outperform their text counterparts, (ii) MMDocIR training set effectively enhances the performance of multimodal document retrieval and (iii) text retrievers leveraging VLM-text significantly outperforms retrievers relying on OCR-text. Our dataset is available at https://mmdocrag.github.io/MMDocIR/.

多模态检索文档理解基准测试版面分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。