用大模型模拟语言演化,发现人工语言会自发形成结构。
Searching for Structure: Investigating Emergent Communication with Large Language Models
- 让大模型在参照游戏中自建语言,观察其演化过程
- 初始无结构的语言逐渐出现可成功沟通的结构性特征
- 适合对语言演化、人机交互感兴趣的读者
人类语言通过反复学习与使用而演化出结构,这些过程引入了影响语言习得的隐含偏见,推动语言系统朝高效传播发展。本文研究若将人工语言优化为适应大语言模型(LLMs)的隐含偏见,是否也会产生类似现象。为此,我们模拟经典参照游戏,让LLM在其中学习和使用人工语言。结果显示,初始无结构的全息语言确实被塑造成具有某些结构特征,使两个LLM代理能成功沟通。与人类实验观察一致,代际传递提升了语言可学性,但同时可能导致非人类特征的退化词汇。本工作扩展了实验发现,表明LLMs可用于语言演化模拟,并为未来人机语言演化实验开辟可能。
原文摘要 · Abstract (English)
Human languages have evolved to be structured through repeated language learning and use. These processes introduce biases that operate during language acquisition and shape linguistic systems toward communicative efficiency. In this paper, we investigate whether the same happens if artificial languages are optimised for implicit biases of Large Language Models (LLMs). To this end, we simulate a classical referential game in which LLMs learn and use artificial languages. Our results show that initially unstructured holistic languages are indeed shaped to have some structural properties that allow two LLM agents to communicate successfully. Similar to observations in human experiments, generational transmission increases the learnability of languages, but can at the same time result in non-humanlike degenerate vocabularies. Taken together, this work extends experimental findings, shows that LLMs can be used as tools in simulations of language evolution, and opens possibilities for future human-machine experiments in this field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。