arXiv:2505.16470cs.IRcs.CL2025-05NeurIPS被引 39

构建首个多模态文档问答基准,评估图文混合证据整合能力

Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering

  • 设计跨页多模态证据链,支持图文联合推理
  • 60个模型测试显示闭源大模型显著优于开源模型
  • 图像描述微调可大幅提升开源模型表现

文档视觉问答(DocVQA)面临处理长篇多模态文档(文本、图像、表格)和跨模态推理的双重挑战。现有文档检索增强生成(DocRAG)方法仍受限于以文本为中心的模式,常忽略关键视觉信息。该领域也缺乏对多模态证据选择与整合的可靠评估基准。我们提出MMDocRAG,包含4,055个专家标注的问答对,涵盖跨页、跨模态证据链。框架引入创新指标评估多模态引用选择,支持答案中交织文本与相关视觉元素。通过对60个视觉语言模型(VLM)/大语言模型(LLM)和14个检索系统的大规模实验,发现多模态证据检索、选择与整合仍存持久挑战。关键发现:先进闭源视觉语言模型性能显著优于开源模型;闭源模型使用多模态输入时有适度优势,而开源模型在多模态输入下表现显著下降;值得注意的是,微调后的LLM在使用详细图像描述时取得显著提升。MMDocRAG建立了严格的测试环境,为开发更鲁棒的多模态DocVQA系统提供可行洞察。基准与代码已公开于https://mmdocrag.github.io/MMDocRAG/。

原文摘要 · Abstract (English)

Document Visual Question Answering (DocVQA) faces dual challenges in processing lengthy multimodal documents (text, images, tables) and performing cross-modal reasoning. Current document retrieval-augmented generation (DocRAG) methods remain limited by their text-centric approaches, frequently missing critical visual information. The field also lacks robust benchmarks for assessing multimodal evidence selection and integration. We introduce MMDocRAG, a comprehensive benchmark featuring 4,055 expert-annotated QA pairs with multi-page, cross-modal evidence chains. Our framework introduces innovative metrics for evaluating multimodal quote selection and enables answers that interleave text with relevant visual elements. Through large-scale experiments with 60 VLM/LLM models and 14 retrieval systems, we identify persistent challenges in multimodal evidence retrieval, selection, and integration.Key findings reveal advanced proprietary LVMs show superior performance than open-sourced alternatives. Also, they show moderate advantages using multimodal inputs over text-only inputs, while open-source alternatives show significant performance degradation. Notably, fine-tuned LLMs achieve substantial improvements when using detailed image descriptions. MMDocRAG establishes a rigorous testing ground and provides actionable insights for developing more robust multimodal DocVQA systems. Our benchmark and code are available at https://mmdocrag.github.io/MMDocRAG/.

多模态文档问答视觉语言模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。