arXiv:2608.29571cs.CL2026-08

测试视觉语言模型在多轮对话中理解语境指代的能力

Which one is banana man? Evaluating vision-language models in multi-turn pragmatic interpretation

论文配图:Which one is banana man? Evaluating vision-language models in multi-turn pragmatic interpretation
图 1 · 摘自论文原文
  • 用多轮指代游戏测试模型对上下文的语用理解
  • 模型能利用已有上下文,但难以构建有效语境
  • 适合研究人机协作与语言模型推理能力的学者

灵活适应语境和共享语用直觉是人类流畅对话的关键。反复进行的指代游戏——参与者通过语言多次识别新指代物——为评估智能体在多轮语言环境中进行上下文敏感语用推理的能力提供了测试场景。我们测试了人类和视觉-语言模型在迭代指代游戏中识别描述意图的能力,变化了上下文的数量、顺序和相关性。尽管人类表现稳定,所测试的模型虽能利用先前上下文解释人类的指称表达,但在构建有效上下文以准确理解这些表达方面存在困难。结果表明,所评估的模型缺乏高效语言协作所需的核心能力。

原文摘要 · Abstract (English)

Flexible adaptation to context and shared pragmatic intuitions contribute to smooth human conversation. Iterated reference games---in which players repeatedly pick out novel referents using language---present a test case for agents' ability to perform context-sensitive pragmatic reasoning in multi-turn linguistic environments. We tested humans and vision--language models on their ability to identify the intended meaning of descriptions produced in iterated reference games, varying the provided context in terms of amount, order, and relevance. While humans performed well consistently, the models we evaluated could make use of prior context to interpret humans' referring expressions, but they struggled to build up the relevant context to interpret those expressions effectively. Our results suggest that the models we evaluated lack core skills needed for efficient linguistic collaboration.

视觉语言模型多轮对话语用推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。