arXiv:2606.31719cs.CLcs.AI2026-06中稿 · SIGDIAL 2026

视觉语言模型常误把共享信息当已达成共识,导致对话理解偏差。

Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue

论文配图:Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue
图 1 · 摘自论文原文
  • 通过1.3万条标注对话数据,测试模型对对话中真实共识的判断能力
  • 有地图图像时模型表现提升但过度预测共识,文本描述也引发同样偏差
  • 模型依赖静态地图线索而非对话过程追踪,适合评估模型对话可靠性

在协作对话中,共享感知不等于共享理解,必须通过互动建立共识。我们研究视觉语言模型(VLMs)能否区分潜在共享与实际共享的信息。基于HCRC MapTask对话中的13,077条标注参考表达,设计解释匹配任务,在系统控制对话上下文和地图信息访问条件下评估模型表现。结果表明:提供真实地图图像可提升整体性能,但使模型更倾向于高估一致性;相同内容的文本描述也诱发此偏差,而无关图像则完全抑制一致性预测,说明偏差源于任务相关地图内容,而非视觉通道本身。该提升以非一致情况下的准确率下降为代价。校准分析与参考链追踪显示,模型依赖地图上的静态指代线索,而非追踪对话历史中的认知演化过程。这一现象在Qwen3-VL-8B-Instruct中最为显著,并在另四个来自两种架构家族的模型中不同程度存在。具有该偏差的模型将地图内容(无论视觉或文本呈现)视为相互理解的证据,混淆了可能性与实际共识。

原文摘要 · Abstract (English)

In collaborative dialogue, shared perception does not guarantee shared interpretation. Mutual understanding must be established through interaction. We investigate whether vision-language models (VLMs) can distinguish what could be shared from what has been shared between dialogue participants through grounding. We formulate this as an interpretation-matching task on 13,077 annotated reference expressions from HCRC MapTask dialogues, and evaluate VLMs under systematically controlled manipulations of dialogue context and map-information access. Our results show that providing authentic map images improves overall performance but shifts models toward over-predicting alignment. Textual descriptions of the same map content reproduce this bias, while non-informative images suppress alignment predictions entirely, indicating that the bias is driven by task-relevant map content, not the visual channel. This improvement comes at the cost of degraded accuracy on non-aligned cases. Calibration analysis and reference-chain tracking further suggest that models rely on static referential cues on the maps rather than tracking how grounding unfolds through dialogue history. We observe these patterns most clearly in Qwen3-VL-8B-Instruct and, to varying degrees, in four additional models from two architecture families. In models that exhibit the bias, map content, whether presented visually or textually, is treated as evidence of mutual understanding, conflating potential with established common ground.

视觉语言模型对话理解共知判断认知偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。