解决多智能体强化学习中因模型不匹配导致的性能下降问题
Stationary Robust Mean-Field Games under Model Mismatches

- 构建基于分布鲁棒性的无限时域均值场博弈框架
- 证明了平稳鲁棒均值场均衡的存在性并给出首个带收敛保证的算法
- 理论揭示大群体下均值场解可近似有限群体博弈的均衡行为
在真实环境中部署多智能体强化学习(MARL)常受限于训练模拟器与真实环境之间的模型不匹配,且战略交互会放大这一问题,导致部署后性能严重退化。分布鲁棒性通过在不确定集内优化最坏情况下的转移模型来应对该问题,但标准鲁棒MARL框架随智能体数量增长变得愈发不可行。本文提出一种无限时域、平稳的均值场博弈框架,将分布模型不确定性直接融入群体耦合动力学中。建立了具有压缩性贝尔曼算子的鲁棒动态规划原理,并通过不动点论证证明了平稳鲁棒均值场均衡的存在性。进一步提出了首个具有收敛性保证的明确算法。我们还将均值场解与依赖经验分布的有限群体鲁棒博弈相连接,表明当群体规模增大时,均值场均衡策略可诱导近似均衡行为。在压缩性鲁棒动力学条件下,我们还获得了明确的非渐近误差界。数值实验进一步验证了多种不确定性模型下鲁棒性对性能的定性和定量影响,支持了理论结论。
原文摘要 · Abstract (English)
Deploying multi-agent reinforcement learning (MARL) in the real world is often limited by model mismatches between the training simulators and the true environment, which could be further amplified through strategic interactions and result in severe performance degradation upon deployment. Distributional robustness offers a principled response by optimizing policies against worst-case transition models drawn from an uncertainty set, but standard robust MARL frameworks become increasingly intractable as the number of agents grows. This paper develops an infinite-horizon, stationary mean-field game framework that incorporates distributional model uncertainty directly into the population-coupled dynamics. We establish a robust dynamic programming principle with a contractive Bellman operator and prove the existence of a stationary robust mean-field equilibrium via a fixed-point argument. We further develop the first concrete algorithm with convergence guarantees. We then connect the mean-field solution to a finite-population robust game whose ambiguity sets depend on the empirical distribution, showing that the mean-field equilibrium policy induces approximate equilibrium behavior as the population size increases. Under a contractive robust-dynamics regime, we further obtain explicit non-asymptotic error bounds. Numerical experiments further illustrate the qualitative and quantitative impact of robustness under multiple uncertainty models, validating our theoretical findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。