新评测基准Double-Bench揭示文档RAG系统真实短板,推动更严谨评估。
Are We on the Right Way for Assessing Document Retrieval-Augmented Generation?
- 构建多语言多模态细粒度评测体系,覆盖6语言4文档类型
- 3276份文档含72880页,5168个单/多跳查询,人工验证证据完备性
- 发现当前RAG框架盲目自信、视觉与文本嵌入差距缩小的深层问题
使用多模态大语言模型(MLLMs)的检索增强生成(RAG)系统在复杂文档理解中展现出巨大潜力,但其发展受限于评估不足。现有基准常聚焦于文档RAG系统的局部环节,且依赖带不完整标注的合成数据,无法反映真实世界瓶颈。为此,我们提出Double-Bench:一个大规模、多语言、多模态的评估系统,可对文档RAG各组件进行细粒度测评。该系统包含3,276份文档(共72,880页)和5,168个单/多跳查询,覆盖6种语言与4类文档类型,并支持动态更新以应对数据污染风险。所有查询均基于全面扫描的证据页,经专家人工验证以确保质量与完整性。我们在9种先进嵌入模型、4种MLLMs及4种端到端文档RAG框架上开展综合实验,发现文本与视觉嵌入模型差距正在缩小,凸显构建更强检索模型的必要性。研究还揭示当前文档RAG框架存在过度自信问题——即使缺乏证据也倾向于生成答案。我们希望完全开源的Double-Bench能为未来高级文档RAG研究提供严谨基础,并计划每年定期更新语料库并发布新基准。
原文摘要 · Abstract (English)
Retrieval-Augmented Generation (RAG) systems using Multimodal Large Language Models (MLLMs) show great promise for complex document understanding, yet their development is critically hampered by inadequate evaluation. Current benchmarks often focus on specific part of document RAG system and use synthetic data with incomplete ground truth and evidence labels, therefore failing to reflect real-world bottlenecks and challenges. To overcome these limitations, we introduce Double-Bench: a new large-scale, multilingual, and multimodal evaluation system that is able to produce fine-grained assessment to each component within document RAG systems. It comprises 3,276 documents (72,880 pages) and 5,168 single- and multi-hop queries across 6 languages and 4 document types with streamlined dynamic update support for potential data contamination issues. Queries are grounded in exhaustively scanned evidence pages and verified by human experts to ensure maximum quality and completeness. Our comprehensive experiments across 9 state-of-the-art embedding models, 4 MLLMs and 4 end-to-end document RAG frameworks demonstrate the gap between text and visual embedding models is narrowing, highlighting the need in building stronger document retrieval models. Our findings also reveal the over-confidence dilemma within current document RAG frameworks that tend to provide answer even without evidence support. We hope our fully open-source Double-Bench provide a rigorous foundation for future research in advanced document RAG systems. We plan to retrieve timely corpus and release new benchmarks on an annual basis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。