让视觉语言模型学会提问,主动澄清图像问答中的模糊问题。
Teaching Vision-Language Models to Ask: Resolving Ambiguity in Visual Questions
- 设计新基准ClearVQA,评估模型通过交互澄清歧义的能力。
- 发现现有模型倾向直接回答而非提问,难以主动求证。
- 适合研究人机交互、可解释AI与多模态理解的学者参考。
在视觉问答(VQA)场景中,用户因表达习惯差异常提出模糊问题。现有研究主要通过重述问题来应对,却忽视了用户与视觉语言模型(VLM)之间本应具备的互动性——模糊之处可通过用户反馈澄清。然而,交互式澄清研究面临两大挑战:(1)缺乏评估VLM通过交互解决歧义能力的基准;(2)当前训练范式使VLM更倾向于直接回答而非主动提问。为此,我们提出 extbf{ClearVQA} 基准,覆盖VQA中三种常见歧义类型,涵盖多种实际问答场景。
原文摘要 · Abstract (English)
In visual question answering (VQA) context, users often pose ambiguous questions to visual language models (VLMs) due to varying expression habits. Existing research addresses such ambiguities primarily by rephrasing questions. These approaches neglect the inherently interactive nature of user interactions with VLMs, where ambiguities can be clarified through user feedback. However, research on interactive clarification faces two major challenges: (1) Benchmarks are absent to assess VLMs' capacity for resolving ambiguities through interaction; (2) VLMs are trained to prefer answering rather than asking, preventing them from seeking clarification. To overcome these challenges, we introduce \textbf{ClearVQA} benchmark, which targets three common categories of ambiguity in VQA context, and encompasses various VQA scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。