arXiv:2605.11695cs.CVcs.AI2026-05

不同视觉模型的智能体通过本地感知自发形成共享符号语言。

Emergent Communication between Heterogeneous Visual Agents through Decentralized Learning

论文配图:Emergent Communication between Heterogeneous Visual Agents through Decentralized Learning
图 1 · 摘自论文原文
  • 智能体仅交换离散词元序列,依靠自身视觉特征评估沟通效果。
  • 在MS-COCO上,通信组比无通信基线提升跨智能体对齐与图文检索性能。
  • 视觉编码器差异越大,共享词元越少且越不均衡,但保留局部语义特异性。

符号可共享,但感知私有。本文研究异构视觉智能体在去中心化学习下自发通信的可能性,探讨当智能体拥有不同视觉表示时,哪些视觉信息能成为可共享内容。不同于依赖外部共同通信目标优化消息的方法,我们的智能体仅交换离散词元序列,并基于本地感知证据更新自身模型。该设置聚焦于自发通信中尚未充分探索的方面:在无共享感知访问的情况下,共用符号能否产生,以及私有视觉空间间的相似性如何限制语言的内容与对称性。我们在马尔可夫链蒙特卡洛图像描述游戏(MHCG)中实现此设定,两个智能体通过交替提出词元序列,听者依据其自身视觉特征采用MH准则判断是否接受。实验使用三个冻结的视觉编码器组合,在MS-COCO数据集上验证,结果表明MHCG生成的共享词元序列具有视觉信息量,显著优于无通信基线,在跨智能体对齐、视觉特征预测和图像-文本检索任务中表现更佳;所有跨智能体指标随编码器差异增大而下降。适度的编码器异质性减少了共享序列数量,但保持了每个序列的视觉特异性;更强的异质性则导致更少、更粗糙且更不对称的序列。消融实验证明,听者侧的MH接受机制对防止退化词元生成至关重要。结果表明,仅凭本地感知评估即可促成共享符号的涌现,而编码器间视觉表征相似性塑造了语言的内容与对称性。

原文摘要 · Abstract (English)

Symbols are shared, but perception is private. We study emergent communication between heterogeneous visual agents through decentralized learning, asking what visual information can become shareable when agents have different visual representations. Instead of optimizing messages through a shared external communicative objective, our agents exchange only discrete token sequences and update their own models using local perceptual evidence. This setting focuses on an underexplored aspect of emergent communication, examining whether common symbols can arise without shared perceptual access, and how the similarity between private visual spaces constrains the content and symmetry of the resulting language. We instantiate this setting in the Metropolis-Hastings Captioning Game (MHCG), where two agents collaboratively form shared captions by exchanging proposed token sequences that a listener accepts or rejects using an MH-style criterion evaluated against its own visual features. We compare three pairings of frozen visual encoders, with agents starting from randomly initialized text modules. Experiments on MS-COCO show that MHCG produces visually informative shared token sequences that outperform a no-communication baseline in cross-agent alignment, visual-feature prediction, and image-text retrieval; all cross-agent metrics decline as encoder mismatch increases. Moderate encoder heterogeneity reduces the number of shared sequences while preserving per-sequence visual specificity, whereas stronger encoder heterogeneity yields fewer, coarser, and more asymmetric sequences. Ablations show that listener-side MH acceptance is critical for avoiding degenerate token formation. These results suggest that shared symbols can arise from local perceptual evaluation alone, with visual representational similarity across encoders shaping both the content and symmetry of the resulting language.

多智能体自发通信视觉语言去中心化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。