无需人为设计,智能体自动生成语言符号实现高效协作。
AI Mother Tongue: Self-Emergent Communication in MARL via Endogenous Symbol Systems
- 用内生符号系统让智能体自发形成通信编码。
- 在无外部偏置下实现语义收敛与高效协作。
- 适合研究符号生成、多智能体协作的学者。
在去中心化多智能体强化学习(MARL)中,涌现通信长期受限于“联合探索困境”,导致智能体陷入“通信真空均衡”。传统方法通过引入归纳偏置促进通信出现。本文质疑这种人工偏置是否为过度工程化。基于向量量化变分自编码器(VQ-VAE)的“AI母语”(AIM)框架实验表明,当智能体具备内生符号系统时,其神经表征自然产生语义压缩与纳什均衡驱动的语义收敛,无需外部归纳偏置即可实现有效符号通信。该现象与近年神经科学发现一致——人类大脑内部思维并不直接使用人类语言;同时呼应大模型中“软思考”能力的研究。相比传统显式通信方法,AIM展现出更强泛化性与效率。可解释性分析工具揭示符号使用呈显著幂律分布,提出三大理论洞见:‘神经通信假说’、‘工具优先原则’、‘语义可解释范式’。未来将探索分层量化变分自编码器(HQ-VAE)以增强表达力,并研究‘强化学习低层级预训练’潜力。此发现为连接主义与符号主义的融合开辟新路径。
原文摘要 · Abstract (English)
In Decentralized Multi-Agent Reinforcement Learning (MARL), the development of Emergent Communication has long been constrained by the ``Joint Exploration Dilemma'', leading agents to fall into a ``Communication Vacuum Equilibrium'' . Traditional methods address this by introducing inductive biases to facilitate communication emergence . This study fundamentally questions whether such artificial inductive biases are, in fact, over-engineering. Through experiments with the ``AI Mother Tongue'' (AIM) framework, based on a Vector Quantized Variational Autoencoder (VQ-VAE), we demonstrate that when agents possess an endogenous symbol system, their neural representations naturally exhibit spontaneous semantic compression and Nash equilibrium-driven semantic convergence, achieving effective symbolic communication without external inductive biases. This aligns with recent neuroscience findings suggesting that the human brain does not directly use human language for internal thought , and resonates with research on ``soft thinking'' capabilities in Large Language Models (LLMs) . Compared to traditional explicit communication methods, AIM demonstrates stronger generality and efficiency. The interpretable analysis toolkit developed in this study confirms that symbol usage exhibits a significant power-law distribution, leading to three major theoretical insights: the ``Neural Communication Hypothesis'', the ``Tool-First Principle'', and the ``Semantic Interpretability Paradigm''. Future research will explore the integration of Hierarchical Quantized Variational Autoencoders (HQ-VAE) to enhance AIM's complex expressive capabilities and investigate the potential for ``Reinforcement Learning (RL) Low-Level Pre-training''. This discovery offers new avenues for bridging symbolism and connectionism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。