arXiv:2601.19792cs.CLcs.AI2026-01ACL被引 7

对比人类与视觉语言模型在指代沟通中的差异,发现模型难建共同理解。

LVLMs and Humans Ground Differently in Referential Communication

  • 设计四类互动对(人-人、人-AI、AI-人、AI-AI)进行多轮指代任务
  • 模型在生成与解析指代表达时准确率远低于人类,沟通效率低
  • 适用于研究人机协作、自然语言理解及对话系统优化的学者

为使生成式AI代理能与人类用户高效协作,准确预测人类意图至关重要。但当前协作能力受限于缺乏对共同基础的建模。本文通过因子设计的指代沟通实验,考察了四类互动对(人类-人类、人类-AI、AI-人类、AI-AI)在多轮重复交互中匹配无明显词汇标签物体图片的表现。结果显示,视觉语言模型(LVLMs)无法动态生成并解决指代表达,这一人类语言使用的核心能力严重缺失。研究发布包含356条对话(89组,每组4轮)的语料库,以及在线数据收集管道和分析工具,支持对准确性、效率与词汇重叠度的评估。

原文摘要 · Abstract (English)

For generative AI agents to partner effectively with human users, the ability to accurately predict human intent is critical. But this ability to collaborate remains limited by a critical deficit: an inability to model common ground. We present a referential communication experiment with a factorial design involving director-matcher pairs (human-human, human-AI, AI-human, and AI-AI) that interact with multiple turns in repeated rounds to match pictures of objects not associated with any obvious lexicalized labels. We show that LVLMs cannot interactively generate and resolve referring expressions in a way that enables smooth communication, a crucial skill that underlies human language use. We release our corpus of 356 dialogues (89 pairs over 4 rounds each) along with the online pipeline for data collection and the tools for analyzing accuracy, efficiency, and lexical overlap.

人机协作指代理解视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。