评测视觉语言模型如何互动建共识,发现任务成功不等于真正理解。
Measuring How (Not Just Whether) VLMs Build Common Ground
- 设计四维指标,评估模型在对话中建立共同理解的能力
- 150场自对弈实验显示,模型普遍偏离人类互动模式
- 任务成功率高也不代表达成共识,需关注互动过程
大型视觉语言模型(VLMs)声称具备推理能力,但现有基准测试多采用单轮问答形式。然而,共识的建立是一个交互过程,人们通过持续沟通逐步形成共享理解。本文提出一套包含四项指标的评估体系:接地效率、内容一致性、词汇适应性与人类相似性,用于系统评估VLM在交互式共识构建中的表现。我们在三款私有VLM之间部署了150场自对弈参考游戏,并与人类双人组进行对比。结果显示,所有模型在至少三项指标上均偏离人类行为模式,其中GPT4o-mini表现最接近人类。研究发现:(i) 任务成功得分不能反映有效共识建立;(ii) 图像与话语高度对齐也不必然带来任务成功。该评估体系及发现为未来VLM共识研究提供了新框架。
原文摘要 · Abstract (English)
Large vision language models (VLMs) increasingly claim reasoning skills, yet current benchmarks evaluate them in single-turn or question answering settings. However, grounding is an interactive process in which people gradually develop shared understanding through ongoing communication. We introduce a four-metric suite (grounding efficiency, content alignment, lexical adaptation, and human-likeness) to systematically evaluate VLM performance in interactive grounding contexts. We deploy the suite on 150 self-play sessions of interactive referential games between three proprietary VLMs and compare them with human dyads. All three models diverge from human patterns on at least three metrics, while GPT4o-mini is the closest overall. We find that (i) task success scores do not indicate successful grounding and (ii) high image-utterance alignment does not necessarily predict task success. Our metric suite and findings offer a framework for future research on VLM grounding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。