构建视觉主导的多模态检索增强评测基准,验证图像比文本更优的场景。
MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal Models
- 设计9类视觉优势场景,用1.6万张图+1353道题系统评估模型
- 所有大模型在图像增强下表现提升,最高达33.16%(人类)
- 揭示主流模型(如GPT-4o)仍难有效利用检索图像,仅提5.82%
现有多模态检索评测主要关注模型能否检索并使用外部文本知识回答问题。然而,在某些场景下,视觉信息比文本更易获取或更具优势,例如不同视角的图像。本文提出多模态检索增强生成评测基准MRAG-Bench,系统识别并分类了视觉知识优于文本知识的场景。该基准包含16,130张图像和1,353道人工标注的多选题,覆盖9个不同场景。我们对10个开源和4个专有大型视觉语言模型(LVLMs)进行了评估。结果表明,所有LVLM在图像增强下均表现出更大提升,证实了MRAG-Bench的视觉中心特性。此外,深入分析显示,表现最佳的GPT-4o在使用真实检索信息时仅提升5.82%,远低于人类参与者33.16%的提升,凸显当前模型在利用检索视觉知识方面的不足。这一发现强调了推动社区改进视觉知识利用能力的重要性。
原文摘要 · Abstract (English)
Existing multimodal retrieval benchmarks primarily focus on evaluating whether models can retrieve and utilize external textual knowledge for question answering. However, there are scenarios where retrieving visual information is either more beneficial or easier to access than textual data. In this paper, we introduce a multimodal retrieval-augmented generation benchmark, MRAG-Bench, in which we systematically identify and categorize scenarios where visually augmented knowledge is better than textual knowledge, for instance, more images from varying viewpoints. MRAG-Bench consists of 16,130 images and 1,353 human-annotated multiple-choice questions across 9 distinct scenarios. With MRAG-Bench, we conduct an evaluation of 10 open-source and 4 proprietary large vision-language models (LVLMs). Our results show that all LVLMs exhibit greater improvements when augmented with images compared to textual knowledge, confirming that MRAG-Bench is vision-centric. Additionally, we conduct extensive analysis with MRAG-Bench, which offers valuable insights into retrieval-augmented LVLMs. Notably, the top-performing model, GPT-4o, faces challenges in effectively leveraging retrieved knowledge, achieving only a 5.82% improvement with ground-truth information, in contrast to a 33.16% improvement observed in human participants. These findings highlight the importance of MRAG-Bench in encouraging the community to enhance LVLMs' ability to utilize retrieved visual knowledge more effectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。