arXiv:2409.07170cs.CL2024-09中稿 · CogSci 2025被引 1

用强化学习模拟语言演化,让智能体自动生成高效数制系统。

Learning Efficient Recursive Numeral Systems via Reinforcement Learning

  • 通过元语法引导双智能体交互,逐步演化出递归数制。
  • 智能体在通信效率压力下达成帕累托最优的词汇配置。
  • 结果与人类数制相似,适合研究语言起源与演化机制。

先前研究表明,强化学习(RL)可使智能体推导出类似人类的简单近似或精确受限数制(Carlsson, 2021)。然而,如何通过如强化学习这样简单的学习机制产生更复杂的递归数制(如英语数制)仍是重大挑战。本文提出一种方法,旨在为高效递归数制的出现提供机制解释。我们设计一对智能体,通过可逐步修改的元语法进行数值通信学习。采用稍作修改的Hurford(1975)元语法,实验表明,在高效沟通压力驱动下,我们的RL智能体能有效调整其词汇体系,形成与人类数制在效率上相当的帕累托最优配置。

原文摘要 · Abstract (English)

It has previously been shown that by using reinforcement learning (RL), agents can derive simple approximate and exact-restricted numeral systems that are similar to human ones (Carlsson, 2021). However, it is a major challenge to show how more complex recursive numeral systems, similar to for example English, could arise via a simple learning mechanism such as RL. Here, we introduce an approach towards deriving a mechanistic explanation of the emergence of efficient recursive number systems. We consider pairs of agents learning how to communicate about numerical quantities through a meta-grammar that can be gradually modified throughout the interactions. Utilising a slightly modified version of the meta-grammar of Hurford (1975), we demonstrate that our RL agents, shaped by the pressures for efficient communication, can effectively modify their lexicon towards Pareto-optimal configurations which are comparable to those observed within human numeral systems in terms of their efficiency.

强化学习语言演化数制系统元语法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。