提出MF-TRPO算法,为多智能体系统提供理论保证的稳定优化方法。
Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games
- 将TRPO拓展至平均场博弈框架,利用其稳定优化特性
- 给出算法在有限样本下的高概率收敛保证和复杂度分析
- 适合研究多智能体博弈与强化学习理论的学者参考
我们提出均值场信任域策略优化(MF-TRPO),一种用于在有限状态-动作空间中计算遍历性均值场博弈(MFG)近似纳什均衡的新算法。基于强化学习(RL)中TRPO的良好性能,我们将该方法扩展到MFG框架,利用其在策略优化中的稳定性与鲁棒性。在均值场文献的标准假设下,我们对MF-TRPO进行了严格分析,建立了其收敛性的理论保证。结果涵盖算法的精确形式及其基于样本的版本,后者给出了高概率收敛保证和有限样本复杂度。本工作通过连接强化学习技术与均值场决策机制,推进了均值场优化的发展,为解决复杂的多智能体问题提供了理论可靠的方案。
原文摘要 · Abstract (English)
We introduce Mean-Field Trust Region Policy Optimization (MF-TRPO), a novel algorithm designed to compute approximate Nash equilibria for ergodic Mean-Field Games (MFG) in finite state-action spaces. Building on the well-established performance of TRPO in the reinforcement learning (RL) setting, we extend its methodology to the MFG framework, leveraging its stability and robustness in policy optimization. Under standard assumptions in the MFG literature, we provide a rigorous analysis of MF-TRPO, establishing theoretical guarantees on its convergence. Our results cover both the exact formulation of the algorithm and its sample-based counterpart, where we derive high-probability guarantees and finite sample complexity. This work advances MFG optimization by bridging RL techniques with mean-field decision-making, offering a theoretically grounded approach to solving complex multi-agent problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。