arXiv:2601.08620cs.AIcs.CV2026-01ACL被引 21

构建多模态RAG新基准,评估真实场景下图文混合信息的检索生成能力。

ViDoRe V3: A Comprehensive Evaluation of Retrieval Augmented Generation in Complex Real-World Scenarios

  • 设计跨10个领域的多模态文档库,支持图文混合查询与多语言输入。
  • 实测表明视觉检索优于纯文本,后期交互模型显著提升性能。
  • 适合研究多模态信息融合、视觉定位与复杂问答系统的开发者使用。

检索增强生成(RAG)系统需应对超越单文档检索的挑战,如解析表格、图表等视觉元素,跨文档信息整合及准确溯源。现有评测基准难以反映此类复杂性,多聚焦于纯文本、单文档或孤立评估检索与生成。本文提出ViDoRe v3,一个涵盖多类型查询的综合性多模态RAG评测基准,覆盖10个专业领域,包含约26,000页文档与3,099条人工验证查询,每条查询提供6种语言版本。通过12,000小时的人工标注,提供了高精度的检索相关性、边界框定位及验证答案。对主流RAG系统的评估显示:视觉检索器优于文本检索器,晚期交互模型与文本重排序显著提升效果,混合或纯视觉上下文有助于生成更高质量答案。但当前模型仍难以处理非文本元素、开放性问题及细粒度视觉定位。为推动该方向发展,本基准以商业友好许可发布于https://hf.co/vidore。

原文摘要 · Abstract (English)

Retrieval-Augmented Generation (RAG) pipelines must address challenges beyond simple single-document retrieval, such as interpreting visual elements (tables, charts, images), synthesizing information across documents, and providing accurate source grounding. Existing benchmarks fail to capture this complexity, often focusing on textual data, single-document comprehension, or evaluating retrieval and generation in isolation. We introduce ViDoRe v3, a comprehensive multimodal RAG benchmark featuring multi-type queries over visually rich document corpora. It covers 10 datasets across diverse professional domains, comprising ~26,000 document pages paired with 3,099 human-verified queries, each available in 6 languages. Through 12,000 hours of human annotation effort, we provide high-quality annotations for retrieval relevance, bounding box localization, and verified reference answers. Our evaluation of state-of-the-art RAG pipelines reveals that visual retrievers outperform textual ones, late-interaction models and textual reranking substantially improve performance, and hybrid or purely visual contexts enhance answer generation quality. However, current models still struggle with non-textual elements, open-ended queries, and fine-grained visual grounding. To encourage progress in addressing these challenges, the benchmark is released under a commercially permissive license at https://hf.co/vidore.

多模态RAG图文理解评测基准信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。