arXiv:2601.20641cs.AI2026-01被引 1

视觉语言模型会自发形成高效隐秘的协作沟通方式。

Investigating the Development of Task-Oriented Communication in Vision-Language Models

  • 用参照游戏框架测试模型间协作沟通模式。
  • 模型能生成比自然语言更简洁且难懂的通信协议。
  • 适合研究AI协作透明性与可控性的研究人员参考。

我们探究基于大语言模型的智能体在协作推理任务中,是否能发展出不同于标准自然语言的任务导向沟通协议。重点关注两类核心特性:效率——以更简洁的方式传递任务相关信息;隐蔽性——对外部观察者难以理解,引发透明性与控制力的担忧。为此,我们采用参照游戏框架,让视觉语言模型(VLM)智能体进行通信,提供一个受控且可测量的评估环境。实验表明,VLM能够发展出有效的、任务适配的沟通模式,同时也能形成对人类和外部智能体而言难以解读的隐蔽协议。我们还观察到,相似模型间可在无显式共享协议的情况下自发实现协调。这些发现凸显了任务导向沟通的潜力与风险,并将参照游戏定位为该领域未来研究的重要测试平台。

原文摘要 · Abstract (English)

We investigate whether \emph{LLM-based agents} can develop task-oriented communication protocols that differ from standard natural language in collaborative reasoning tasks. Our focus is on two core properties such task-oriented protocols may exhibit: Efficiency -- conveying task-relevant information more concisely than natural language, and Covertness -- becoming difficult for external observers to interpret, raising concerns about transparency and control. To investigate these aspects, we use a referential-game framework in which vision-language model (VLM) agents communicate, providing a controlled, measurable setting for evaluating language variants. Experiments show that VLMs can develop effective, task-adapted communication patterns. At the same time, they can develop covert protocols that are difficult for humans and external agents to interpret. We also observe spontaneous coordination between similar models without explicitly shared protocols. These findings highlight both the potential and the risks of task-oriented communication, and position referential games as a valuable testbed for future work in this area.

多智能体沟通协议视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。