arXiv:2410.17466cs.LGcs.GT2024-10

用强化学习模拟大规模异构智能体社会演化,揭示合作与竞争的进化机制。

Evolution of Societies via Reinforcement Learning

  • 设计可并行的快速策略梯度与对手学习感知算法,适配大规模演化仿真
  • 10万级智能体在三种经典博弈中演化,揭示先进学习策略对社会结构的影响
  • 适用于研究复杂社会行为演化,尤其适合关注合作进化的研究者

宇宙中存在大量独立共学习的智能体,是可观测环境的动态组成部分。然而,现有多智能体强化学习(MARL)应用通常局限于小规模同质群体,且计算成本高。本文提出一种方法,实现强化学习智能体在进化尺度上的模拟。具体而言,我们推导出适用于无状态标准形式博弈中随机成对交互场景的快速、可并行的策略梯度(PG)与对手学习感知(LOLA)实现方案。通过实验,我们在鹰鸽博弈、猎鹿博弈和石头剪刀布三种经典游戏中,模拟了20万智能体的大规模异构共学习群体演化,对比了基础与高级学习策略下的演化结果。结果显示,对手学习感知显著影响社会演化路径,揭示了不同学习机制如何塑造集体行为模式。

原文摘要 · Abstract (English)

The universe involves many independent co-learning agents as an ever-evolving part of our observed environment. Yet, in practice, Multi-Agent Reinforcement Learning (MARL) applications are typically constrained to small, homogeneous populations and remain computationally intensive. We propose a methodology that enables simulating populations of Reinforcement Learning agents at evolutionary scale. More specifically, we derive a fast, parallelizable implementation of Policy Gradient (PG) and Opponent-Learning Awareness (LOLA), tailored for evolutionary simulations where agents undergo random pairwise interactions in stateless normal-form games. We demonstrate our approach by simulating the evolution of very large populations made of heterogeneous co-learning agents, under both naive and advanced learning strategies. In our experiments, 200,000 PG or LOLA agents evolve in the classic games of Hawk-Dove, Stag-Hunt, and Rock-Paper-Scissors. Each game provides distinct insights into how populations evolve under both naive and advanced MARL rules, including compelling ways in which Opponent-Learning Awareness affects social evolution.

多智能体强化学习社会演化博弈论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。