arXiv:2604.12033cs.CLcs.AI2026-04ACL被引 1

测试大模型在视觉语言任务中面对错误信息时是否能正确拒绝回答。

Benchmarking Deflection and Hallucination in Large Vision-Language Models

论文配图:Benchmarking Deflection and Hallucination in Large Vision-Language Models
图 1 · 摘自论文原文
  • 构建动态数据筛选流程,确保评测基准长期有效。
  • 创建2775个样本的基准,检验模型在证据冲突或缺失时的表现。
  • 提出细粒度评估方案,区分记忆与检索能力,适合模型鲁棒性研究者。

大型视觉语言模型(LVLMs)越来越多地依赖检索来回答知识密集型多模态问题。现有基准忽略了视觉与文本证据之间的矛盾,以及在检索知识不完整时生成拒答(如“抱歉,我无法回答”)的重要性。这些基准还因模型训练数据不断增长而迅速过时,导致许多问题无需检索即可作答。本文提出三项贡献:第一,设计动态数据筛选流程,通过过滤真正依赖检索的样本,保持基准难度随时间稳定;第二,引入VLM-DeflectionBench,包含2,775个样本,覆盖多种多模态检索场景,用于探测模型在证据冲突或不足时的行为;第三,定义细粒度评估协议,包含四个场景,分离参数化记忆与检索鲁棒性。对20个前沿LVLM的实验表明,模型在面对噪声或误导性证据时通常无法正确拒答。结果强调需评估模型‘不知道’时的行为,而非仅关注其已知知识,为可靠的知识增强多模态问答评估提供可复用、可扩展的基准。所有资源将在发表后公开。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) increasingly rely on retrieval to answer knowledge-intensive multimodal questions. Existing benchmarks overlook conflicts between visual and textual evidence and the importance of generating deflections (e.g., Sorry, I cannot answer...) when retrieved knowledge is incomplete. These benchmarks also suffer from rapid obsolescence, as growing LVLM training sets allow models to answer many questions without retrieval. We address these gaps with three contributions. First, we propose a dynamic data curation pipeline that preserves benchmark difficulty over time by filtering for genuinely retrieval-dependent samples. Second, we introduce VLM-DeflectionBench, a benchmark of 2,775 samples spanning diverse multimodal retrieval settings, designed to probe model behaviour under conflicting or insufficient evidence. Third, we define a fine-grained evaluation protocol with four scenarios that disentangle parametric memorization from retrieval robustness. Experiments across 20 state-of-the-art LVLMs indicate that models usually fail to deflect in the presence of noisy or misleading evidence. Our results highlight the need to evaluate not only what models know, but how they behave when they do not, and serve as a reusable and extensible benchmark for reliable KB-VQA evaluation. All resources will be publicly available upon publication.

视觉语言模型检索评估拒答行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。