arXiv:2503.22779cs.MAcs.GT2025-03

解决多智能体均值-方差博弈的非平稳与非可加性难题,提出收敛算法。

Policy Optimization and Multi-agent Reinforcement Learning for Mean-variance Team Stochastic Games

  • 基于敏感性优化,推导联合策略的性能差异与导数公式
  • 算法在微电网能源管理中实现收敛,性能提升显著
  • 适合研究多智能体金融或能源系统优化的读者

我们研究长期均值-方差团队随机博弈(MV-TSG),其中各智能体共享系统级均值-方差目标并独立决策以最大化该目标。该问题面临两大挑战:一是方差指标在动态设置下既非可加也非马尔可夫;二是所有智能体同时更新导致每个个体面对非平稳环境,使动态规划失效。本文从敏感性优化视角出发,推导出联合策略的性能差异与性能导数公式,为优化提供信息。证明了存在确定性纳什策略。随后提出顺序更新的均值-方差多智能体策略迭代(MV-MAPI)算法,证明其收敛至目标函数的一阶驻点。通过分析驻点局部几何结构,给出驻点为(局部)纳什均衡及严格局部最优的充分条件。针对未知环境参数的大规模场景,将信任域思想扩展至MV-MAPI,提出多智能体强化学习算法均值-方差多智能体信任域策略优化(MV-MATRPO),并推导每次联合策略更新的性能下界。在多个微电网能源管理系统上进行数值实验,验证方法有效性。

原文摘要 · Abstract (English)

We study a long-run mean-variance team stochastic game (MV-TSG), where each agent shares a common mean-variance objective for the system and takes actions independently to maximize it. MV-TSG has two main challenges. First, the variance metric is neither additive nor Markovian in a dynamic setting. Second, simultaneous policy updates of all agents lead to a non-stationary environment for each individual agent. Both challenges make dynamic programming inapplicable. In this paper, we study MV-TSGs from the perspective of sensitivity-based optimization. The performance difference and performance derivative formulas for joint policies are derived, which provide optimization information for MV-TSGs. We prove the existence of a deterministic Nash policy for this problem. Subsequently, we propose a Mean-Variance Multi-Agent Policy Iteration (MV-MAPI) algorithm with a sequential update scheme, where individual agent policies are updated one by one in a given order. We prove that the MV-MAPI algorithm converges to a first-order stationary point of the objective function. By analyzing the local geometry of stationary points, we derive specific conditions for stationary points to be (local) Nash equilibria, and further, strict local optima. To solve large-scale MV-TSGs in scenarios with unknown environmental parameters, we extend the idea of trust region methods to MV-MAPI and develop a multi-agent reinforcement learning algorithm named Mean-Variance Multi-Agent Trust Region Policy Optimization (MV-MATRPO). We derive a performance lower bound for each update of joint policies. Finally, numerical experiments on energy management in multiple microgrid systems are conducted.

多智能体强化学习均值方差博弈论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。