arXiv:2502.12342cs.IRcs.CV2025-02ACL被引 49

构建真实场景多模态检索基准,揭示模型短板并提供训练数据。

REAL-MM-RAG: A Real-World Multi-Modal Retrieval Benchmark

  • 自动生成包含图文的复杂文档与真实查询,模拟现实挑战
  • 模型在表格密集文档和同义重述查询上表现明显下降
  • 提供重写训练集与金融领域新数据集,提升模型鲁棒性

准确的多模态文档检索对检索增强生成(RAG)至关重要,但现有基准未能充分反映真实世界中的挑战。我们提出REAL-MM-RAG,一个自动构建的基准,涵盖四个关键特性:(i) 多模态文档,(ii) 增强难度,(iii) 真实的RAG查询,(iv) 准确标注。此外,我们基于查询重述设计了多难度层级方案,以评估模型超越关键词匹配的语义理解能力。实验表明,现有模型在处理表格密集文档及查询重述时存在显著弱点。为此,我们构建了一个重述训练集,并引入一个聚焦金融领域的表格密集数据集。在这些数据上微调后,模型在REAL-MM-RAG基准上达到当前最优性能。本工作为多模态RAG系统的评估与优化提供了新路径,并提供了可复用的训练数据与模型。

原文摘要 · Abstract (English)

Accurate multi-modal document retrieval is crucial for Retrieval-Augmented Generation (RAG), yet existing benchmarks do not fully capture real-world challenges with their current design. We introduce REAL-MM-RAG, an automatically generated benchmark designed to address four key properties essential for real-world retrieval: (i) multi-modal documents, (ii) enhanced difficulty, (iii) Realistic-RAG queries and (iv) accurate labeling. Additionally, we propose a multi-difficulty-level scheme based on query rephrasing to evaluate models' semantic understanding beyond keyword matching. Our benchmark reveals significant model weaknesses, particularly in handling table-heavy documents and robustness to query rephrasing. To mitigate these shortcomings, we curate a rephrased training set and introduce a new finance-focused, table-heavy dataset. Fine-tuning on these datasets enables models to achieve state-of-the-art retrieval performance on REAL-MM-RAG benchmark. Our work offers a better way to evaluate and improve retrieval in multi-modal RAG systems while also providing training data and models that address current limitations.

多模态检索RAG数据集金融文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。