arXiv:2603.09208cs.LGcs.GT2026-03

提出一种更鲁棒的多智能体博弈求解方法,兼顾性能与稳定性。

Strategically Robust Multi-Agent Reinforcement Learning with Linear Function Approximation

  • 基于风险敏感的量化响应均衡,设计优化值迭代算法
  • 实证显示跨智能体对战中表现更稳健,优于纳什均衡方法
  • 可调节参数平衡性能与鲁棒性,适合复杂多智能体系统

在一般和博弈的马尔可夫游戏中,实现可证明高效且鲁棒的均衡计算仍是多智能体强化学习的核心挑战。纳什均衡在一般情况下计算不可行,且因均衡多重性和对近似误差敏感而脆弱。本文研究风险敏感的量化响应均衡(RQRE),该均衡在有限理性与风险敏感下具有唯一且光滑的解。我们提出 exttt{RQRE-OVI},一种结合线性函数逼近的乐观值迭代算法,适用于大规模或连续状态空间。通过有限样本后悔分析,我们建立了收敛性,并明确刻画了样本复杂度随理性和风险敏感参数的变化规律。后悔界揭示出量化权衡:提高理性降低后悔,风险敏感则引入正则化,提升稳定性和鲁棒性。这暴露了期望性能与鲁棒性之间的帕累托前沿,纳什均衡为完美理性与风险中性极限情形。进一步证明,RQRE策略映射在估计收益上具有Lipschitz连续性,而纳什不满足;且RQRE具备分布鲁棒优化解释。实验表明, exttt{RQRE-OVI} 在自对弈中表现优异,跨对弈中行为显著更鲁棒,优于基于纳什的方法。结果表明, exttt{RQRE-OVI} 提供了一条原理清晰、可扩展且可调的均衡学习路径,提升了鲁棒性与泛化能力。

原文摘要 · Abstract (English)

Provably efficient and robust equilibrium computation in general-sum Markov games remains a core challenge in multi-agent reinforcement learning. Nash equilibrium is computationally intractable in general and brittle due to equilibrium multiplicity and sensitivity to approximation error. We study Risk-Sensitive Quantal Response Equilibrium (RQRE), which yields a unique, smooth solution under bounded rationality and risk sensitivity. We propose \texttt{RQRE-OVI}, an optimistic value iteration algorithm for computing RQRE with linear function approximation in large or continuous state spaces. Through finite-sample regret analysis, we establish convergence and explicitly characterize how sample complexity scales with rationality and risk-sensitivity parameters. The regret bounds reveal a quantitative tradeoff: increasing rationality tightens regret, while risk sensitivity induces regularization that enhances stability and robustness. This exposes a Pareto frontier between expected performance and robustness, with Nash recovered in the limit of perfect rationality and risk neutrality. We further show that the RQRE policy map is Lipschitz continuous in estimated payoffs, unlike Nash, and RQRE admits a distributionally robust optimization interpretation. Empirically, we demonstrate that \texttt{RQRE-OVI} achieves competitive performance under self-play while producing substantially more robust behavior under cross-play compared to Nash-based approaches. These results suggest \texttt{RQRE-OVI} offers a principled, scalable, and tunable path for equilibrium learning with improved robustness and generalization.

多智能体博弈论强化学习鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。