提出新算法,让多智能体强化学习更稳定收敛。
Provably Convergent Actor-Critic for MARL through Risk-aversion
- 用风险规避机制设计新型演员-评论家算法
- 证明算法在有限样本下全局收敛
- 适合追求稳定性的多智能体系统研究者
在无限时域的非零和马尔可夫博弈(MGs)中学习平稳策略仍是多智能体强化学习(MARL)中的基础难题。尽管平稳策略因实用性被青睐,但计算经典博弈论均衡的平稳形式在计算上是不可行的——这与单智能体强化学习或零和博弈的相对易解形成鲜明对比。为弥合这一差距,我们研究了基于行为博弈论的风险规避量化响应均衡(RQE),该概念融合了风险规避与有限理性。我们证明RQE具备强正则性,使其特别适合在马尔可夫博弈中进行学习。我们提出一种新的单时间尺度演员-评论家算法,其演员更新更快、评论家更新更慢。借助RQE的正则性,我们证明该方法在有限样本下实现全局收敛。我们在多个环境中验证了该算法,结果表明其收敛性能优于风险中性基线。
原文摘要 · Abstract (English)
Learning stationary policies in infinite-horizon general-sum Markov games (MGs) remains a fundamental open problem in Multi-Agent Reinforcement Learning (MARL). While stationary strategies are preferred for their practicality, computing stationary forms of classic game-theoretic equilibria is computationally intractable -- a stark contrast to the comparative ease of solving single-agent RL or zero-sum games. To bridge this gap, we study Risk-averse Quantal response Equilibria (RQE), a solution concept rooted in behavioral game theory that incorporates risk aversion and bounded rationality. We demonstrate that RQE possesses strong regularity conditions that make it uniquely amenable to learning in MGs. We propose a novel single-timescale Actor-Critic algorithm characterized by a faster actor and a slower critic. Leveraging the regularity of RQE, we prove that this approach achieves global convergence with finite-sample guarantees. We empirically validate our algorithm in several environments to demonstrate superior convergence properties compared to risk-neutral baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。