多模态大模型能对齐标签却无法形成对话默契,靠冗长描述而非约定俗成。
Aligned but Not Partner-Specific: Distinguishing How Multimodal LLM Agents Succeed in Reference Games Without Human-Like Conventions

- 构建伪搭档基线,切断互动历史以检验标签对齐是否依赖特定伙伴。
- 模型从第一轮起保持冗长描述,标签重合度接近天花板且与真人无异。
- 人类通过共演压缩表达,模型则始终冗余,缺乏动态协商机制。
重复参照游戏测试对话者是否能将初始长描述转化为基于共享交互史的简短、专属约定。以往研究发现,多模态大模型虽在标签使用上对齐,但效率未随回合提升。我们通过对比来自KTH Tangrams语料库的人类搭档与能力相当的多模态智能体搭档,回答这一问题。提出一种受控伪搭档基线,保留原参照任务结构但切断搭档历史,从而检验标签对齐是否依赖特定互动伙伴。在任务能力、描述策略与对齐动态三个层面分析均发现显著差异:人类通过共演压缩描述,提高与搭档的标签一致性;而智能体维持固定努力水平,从第一轮起即输出冗长描述,标签重合度接近上限,且真实搭档与伪搭档间无统计差异。因此,多模态大模型实现协调但无约定,成功依赖冗长描述,而非人类对话中典型的、依赖历史的简洁指称表达。
原文摘要 · Abstract (English)
Repeated reference games test whether interlocutors replace their initially long descriptions with shorter, partner-specific conventions grounded in shared interaction history. Prior work shows that multimodal LLMs fail to become more efficient across rounds, although they align on the labels they use. How can we determine whether this alignment reflects partner-specific grounding rather than a shared task vocabulary? We address this question by comparing capable multimodal agent dyads with human dyads from the KTH Tangrams corpus. Our novel methodological contribution is a constrained pseudo-dyad baseline that matches the original referential task structure, but breaks partner history. This baseline enables us to test whether the observed label alignment depends on interaction with a specific partner. Across three analytic layers (task competence, description strategy, alignment dynamics), we find clear differences. Humans reduce effort through entrainment, compressing descriptions and increasing label alignment with partners. Agents instead maintain fixed effort levels, producing verbose descriptions from round one, with near-ceiling label overlap that is statistically indistinguishable between real and pseudo dyads. MLLMs thus achieve coordination without convention, succeeding by verbose description rather than by forming the compact, history-dependent referring expressions characteristic of human dialogue.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。