通过加速更新与降学习率,将多智能体学习动态简化为可解的常微分方程。
Homogenization of Multi-agent Learning Dynamics in Finite-state Markov Games
- 用快慢变量分离思想,将学习过程重标度为慢变参数演化。
- 在状态遍历和更新连续条件下,证明了收敛到确定性常微分方程。
- 适用于分析有限状态马尔可夫博弈中多智能体学习行为,适合理论研究者。
本文提出一种新方法,用于近似多个强化学习智能体在有限状态马尔可夫博弈中的学习动态。核心思想是同时降低学习率并提高更新频率,将智能体参数视为受快速混合博弈状态影响的慢变变量。在状态过程遍历性和更新连续性的温和假设下,证明该重标度过程收敛至一个常微分方程(ODE)。该ODE提供了智能体学习动态的可处理、确定性近似。框架实现已开源:https://github.com/yannKerzreho/MarkovGameApproximation。
原文摘要 · Abstract (English)
This paper introduces a new approach for approximating the learning dynamics of multiple reinforcement learning (RL) agents interacting in a finite-state Markov game. The idea is to rescale the learning process by simultaneously reducing the learning rate and increasing the update frequency, effectively treating the agent's parameters as a slow-evolving variable influenced by the fast-mixing game state. Under mild assumptions-ergodicity of the state process and continuity of the updates-we prove the convergence of this rescaled process to an ordinary differential equation (ODE). This ODE provides a tractable, deterministic approximation of the agent's learning dynamics. An implementation of the framework is available at\,: https://github.com/yannKerzreho/MarkovGameApproximation
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。