用迭代学习模拟语言演化,让机器学会清晰表达复杂图像。
Image, Word and Thought: A More Challenging Language Task for the Iterated Learning Model
- 构建自编码器框架融合有监督与无监督学习
- 成功实现128个字符的无歧义、可组合语言生成
- 适合研究语言演化机制或人机通信系统设计
迭代学习模型通过模拟语言代际传递,探索语言传输约束如何促成语言结构的涌现。尽管每代学习者初始为白板,但受限于可接触语句数量的瓶颈,仍能催生出无歧义、受语法规则支配且跨代稳定的语言。最近提出的半监督迭代学习模型结合了自编码器架构,具备更强计算效率和生态合理性,首次应用于更复杂的语义-信号空间:七段数码管图像(共128种符号)。实验表明,该模型中的智能体能够学习并传递一种表达性强的语言:所有128个图形均有唯一编码;具有组合性:信号成分稳定对应语义成分;具备稳定性:语言在代际间保持不变。
原文摘要 · Abstract (English)
The iterated learning model simulates the transmission of language from generation to generation in order to explore how the constraints imposed by language transmission facilitate the emergence of language structure. Despite each modelled language learner starting from a blank slate, the presence of a bottleneck limiting the number of utterances to which the learner is exposed can lead to the emergence of language that lacks ambiguity, is governed by grammatical rules, and is consistent over successive generations, that is, one that is expressive, compositional and stable. The recent introduction of a more computationally tractable and ecologically valid semi supervised iterated learning model, combining supervised and unsupervised learning within an autoencoder architecture, has enabled exploration of language transmission dynamics for much larger meaning-signal spaces. Here, for the first time, the model has been successfully applied to a language learning task involving the communication of much more complex meanings: seven-segment display images. Agents in this model are able to learn and transmit a language that is expressive: distinct codes are employed for all 128 glyphs; compositional: signal components consistently map to meaning components, and stable: the language does not change from generation to generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。