arXiv:2511.12142cs.CV2025-11AAAI被引 2

首个评估多模态来源溯源的基准,提升视觉问答可信度

MAVIS: A Benchmark for Multimodal Source Attribution in Long-form Visual Question Answering

  • 构建多模态来源溯源系统,结合视觉与文本证据生成带引用的答案
  • 15.7万条数据集验证,多模态RAG模型在流畅性上优于单模态但图像溯源弱于文本
  • 发现信息量与准确性存在权衡,图像理解中的上下文偏差需重点改进

来源溯源旨在通过为每个陈述添加引用,提升AI生成答案的可靠性,帮助用户验证结果。然而现有研究主要聚焦于纯文本场景,忽视了多模态的重要性。我们提出MAVIS,首个专为评估多模态来源溯源系统设计的基准,该系统需理解视觉问题背后用户意图,检索多模态证据,并生成带引用的长篇回答。数据集包含15.7万条视觉问答实例,每条答案均标注事实级引用,指向多模态文档。我们开发了三维度细粒度自动评价指标(信息量、依据性、流畅性),其与人工评估高度相关。关键发现包括:(1) 多模态RAG的视觉语言模型生成的答案更信息丰富且流畅,但在图像文档上的依据性弱于文本文档,且多模态环境下差距加剧;(2) 在相同多模态文档下,不同提示方法在信息量与依据性间存在权衡;(3) 提出的方法强调缓解图像文档解释中的上下文偏差,是未来研究的关键方向。

原文摘要 · Abstract (English)

Source attribution aims to enhance the reliability of AI-generated answers by including references for each statement, helping users validate the provided answers. However, existing work has primarily focused on text-only scenario and largely overlooked the role of multimodality. We introduce MAVIS, the first benchmark designed to evaluate multimodal source attribution systems that understand user intent behind visual questions, retrieve multimodal evidence, and generate long-form answers with citations. Our dataset comprises 157K visual QA instances, where each answer is annotated with fact-level citations referring to multimodal documents. We develop fine-grained automatic metrics along three dimensions of informativeness, groundedness, and fluency, and demonstrate their strong correlation with human judgments. Our key findings are threefold: (1) LVLMs with multimodal RAG generate more informative and fluent answers than unimodal RAG, but they exhibit weaker groundedness for image documents than for text documents, a gap amplified in multimodal settings. (2) Given the same multimodal documents, there is a trade-off between informativeness and groundedness across different prompting methods. (3) Our proposed method highlights mitigating contextual bias in interpreting image documents as a crucial direction for future research.

多模态来源溯源视觉问答RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。