评测图文检索增强生成在视觉知识问答中的效果
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
- 设计新基准,要求从文本检索图像并利用图像生成答案
- 发现图像能提供强证据,但现有模型利用率低
- 适合关注多模态RAG、视觉推理的研究者
检索增强生成(RAG)通过引入外部知识来提升大语言模型在知识密集型问答中的表现。尽管已有多个基准评估多模态大模型在多模态RAG场景下的能力,但多数仅依赖文本检索,未显式考察模型如何利用视觉证据进行生成。因此,尚无基准能独立衡量检索图像对生成的贡献。本文提出Visual-RAG,一个面向视觉语境、知识密集型问题的问答基准。与以往工作不同,Visual-RAG要求进行文本到图像的检索,并将检索到的线索图像整合以提取视觉证据用于答案生成。我们使用该基准评估了5个开源和3个专有多模态大模型,结果显示图像提供了强有力的信息支持;然而,即使是顶尖模型也难以高效提取和利用视觉知识。结果凸显了改进多模态RAG系统中视觉检索、定位与归因的必要性。
原文摘要 · Abstract (English)
Retrieval-augmented generation (RAG) is a paradigm that augments large language models (LLMs) with external knowledge to tackle knowledge-intensive question answering. While several benchmarks evaluate Multimodal LLMs (MLLMs) under Multimodal RAG settings, they predominantly retrieve from textual corpora and do not explicitly assess how models exploit visual evidence during generation. Consequently, there still lacks benchmark that isolates and measures the contribution of retrieved images in RAG. We introduce Visual-RAG, a question-answering benchmark that targets visually grounded, knowledge-intensive questions. Unlike prior work, Visual-RAG requires text-to-image retrieval and the integration of retrieved clue images to extract visual evidence for answer generation. With Visual-RAG, we evaluate 5 open-source and 3 proprietary MLLMs, showcasing that images provide strong evidence in augmented generation. However, even state-of-the-art models struggle to efficiently extract and utilize visual knowledge. Our results highlight the need for improved visual retrieval, grounding, and attribution in multimodal RAG systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。