通过模拟语言演化,让神经网络自动生成结构清晰、高效可组合的通信协议。
A Compressive-Expressive Communication Framework for Compositional Representations
- 设计发送-接收双模型,用重建任务驱动语言自发形成。
- 在Shapes3D和MPI3D上实现更高重建精度与拓扑相似性。
- 适合研究语言涌现、符号系统生成的学者参考。
知识与语言的组合性——将复杂概念表示为简单成分的组合——是人类认知与交流的核心特征。尽管近年取得进展,深度神经网络仍难以可靠获得此能力。神经语言涌现模型通过模拟人类语言形成的压力来赋予人工代理组合性语言。本文提出CELEBI(基于离散瓶颈与迭代学习的压缩-表达式语言涌现),一种新的自监督框架,通过发送者与接收者之间的重建型通信游戏诱导组合表示。基于语言涌现理论与迭代学习框架,整合三种机制:渐进解码要求接收者每解析一个符号就输出部分重构;最终状态模仿使后续代际代理模仿重构而非原始消息,强化通信瓶颈;成对距离最大化通过鼓励消息间高距离来正则化多样性,与熵最大化存在理论关联。该方法在Shapes3D和MPI3D数据集上显著提升所学消息的效率与组合性,优于以往离散通信框架,在重建准确率与拓扑相似性上均表现更优。本工作为从基于简单的归纳偏置中涌现出结构化、可泛化通信协议提供了新的理论与实证支持。
原文摘要 · Abstract (English)
Compositionality in knowledge and language--the ability to represent complex concepts as a combination of simpler ones--is a hallmark of human cognition and communication. Despite recent advances, deep neural networks still struggle to acquire this property reliably. Neural models for emergent communication look to endow artificial agents with compositional language by simulating the pressures that form human language. In this work, we introduce CELEBI (Compressive-Expressive Language Emergence through a discrete Bottleneck and Iterated learning), a novel self-supervised framework for inducing compositional representations through a reconstruction-based communication game between a sender and a receiver. Building on theories of language emergence and the iterated learning framework, we integrate three mechanisms that jointly promote compressibility, expressivity, and efficiency in the emergent language. First, Progressive Decoding incentivizes intermediate reasoning by requiring the receiver to produce partial reconstructions after each symbol. Second, Final-State Imitation trains successive generations of agents to imitate reconstructions rather than messages, enforcing a tighter communication bottleneck. Third, Pairwise Distance Maximization regularizes message diversity by encouraging high distances between messages, with formal links to entropy maximization. Our method significantly improves both the efficiency and compositionality of the learned messages on the Shapes3D and MPI3D datasets, surpassing prior discrete communication frameworks in both reconstruction accuracy and topographic similarity. This work provides new theoretical and empirical evidence for the emergence of structured, generalizable communication protocols from simplicity-based inductive biases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。