arXiv:2604.14799cs.CLcs.CV2026-04被引 2

提出新基准评估多模态模型何时该不回答,发现现有模型很少主动放弃答题。

Knowing When Not to Answer: Evaluating Abstention in Multimodal Reasoning Systems

论文配图:Knowing When Not to Answer: Evaluating Abstention in Multimodal Reasoning Systems
图 1 · 摘自论文原文
  • 构建双维度无答案样本:视觉依赖性和证据充分性
  • 封闭与开源模型均极少主动拒答,信心基线表现更好
  • 需专门训练才能有效拒答,而非靠更多智能体或提示优化

有效拒答(Effective Abstention)——识别证据不足并拒绝回答——对可靠的多模态系统至关重要。然而,现有视觉语言模型(VLMs)和多智能体系统(MAS)的评估范式默认问题可答,迫使模型始终回应。拒答在纯文本领域已有研究,但在多模态场景仍被忽视;当前基准要么忽略不可答情况,要么使用粗粒度方法,难以捕捉真实失效模式。本文提出MM-AQA基准,通过视觉模态依赖性和证据充分性两个维度,从可答样本生成不可答实例。在2079个样本上评估三款前沿VLM(含闭源与开源)及两款MAS架构,发现:(1)标准提示下VLM几乎不拒答;即使简单置信度基线也优于该设置;(2)MAS虽提升拒答能力,但引入准确率-拒答率权衡;(3)序列设计性能匹配或超过迭代变体,表明瓶颈在于校准偏差而非推理深度;(4)当图像或文本证据缺失时模型会拒答,但面对劣化或矛盾证据仍试图调和。多模态有效拒答需基于拒答意识的训练,而非更优提示或更多智能体。

原文摘要 · Abstract (English)

Effective abstention (EA), recognizing evidence insufficiency and refraining from answering, is critical for reliable multimodal systems. Yet existing evaluation paradigms for vision-language models (VLMs) and multi-agent systems (MAS) assume answerability, pushing models to always respond. Abstention has been studied in text-only settings but remains underexplored multimodally; current benchmarks either ignore unanswerability or rely on coarse methods that miss realistic failure modes. We introduce MM-AQA, a benchmark that constructs unanswerable instances from answerable ones via transformations along two axes: visual modality dependency and evidence sufficiency. Evaluating three frontier VLMs spanning closed and open-source models and two MAS architectures across 2079 samples, we find: (1) under standard prompting, VLMs rarely abstain; even simple confidence baselines outperform this setup, (2) MAS improves abstention but introduces an accuracy-abstention trade-off, (3) sequential designs match or exceed iterative variants, suggesting the bottleneck is miscalibration rather than reasoning depth, and (4) models abstain when image or text evidence is absent, but attempt reconciliation with degraded or contradictory evidence. Effective multimodal abstention requires abstention-aware training rather than better prompting or more agents.

多模态拒答评估基准模型校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。