用强化学习研究博弈中稀释与移动的影响,发现合作可自发形成。
Dilution, Diffusion and Symbiosis in Spatial Prisoner's Dilemma with Reinforcement Learning
- 采用独立多智能体Q-learning模拟博弈演化
- 固定规则与学习规则的博弈结果趋同
- 多行动设定下出现种群共生现象,适合群体智能研究
近期关于强化学习在空间囚徒困境中的研究显示,静态智能体可通过噪声注入、不同学习算法及邻居收益信息等机制学会合作。本文采用独立多智能体Q-learning算法,研究空间囚徒困境中稀释与移动的影响。算法定义了多种可能行为,关联经典非强化学习空间博弈结果,展现了算法在建模多种博弈场景中的灵活性及基准测试潜力。结果显示,具有固定更新规则的游戏与学习规则的游戏在定性上可等价,且当定义多种行为时,种群间出现互利共生效应。
原文摘要 · Abstract (English)
Recent studies in the spatial prisoner's dilemma games with reinforcement learning have shown that static agents can learn to cooperate through a diverse sort of mechanisms, including noise injection, different types of learning algorithms and neighbours' payoff knowledge. In this work, using an independent multi-agent Q-learning algorithm, we study the effects of dilution and mobility in the spatial version of the prisoner's dilemma. Within this setting, different possible actions for the algorithm are defined, connecting with previous results on the classical, non-reinforcement learning spatial prisoner's dilemma, showcasing the versatility of the algorithm in modeling different game-theoretical scenarios and the benchmarking potential of this approach. As a result, a range of effects is observed, including evidence that games with fixed update rules can be qualitatively equivalent to those with learned ones, as well as the emergence of a symbiotic mutualistic effect between populations that forms when multiple actions are defined.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。