提出新算法提升非平稳多智能体协作稳定性,解决学习噪声导致的合作崩溃问题。
The Price of Paranoia: Robust Risk-Sensitive Cooperation in Non-Stationary Multi-Agent Reinforcement Learning

- 通过在线度量伙伴不可预测性,调节策略梯度更新以增强协作鲁棒性
- 在对称协调游戏中,合作稳定区扩大,且可理论证明其有效性
- 引入'焦虑代价'概念,精确衡量算法在稳定与采样效率间的最优平衡
合作均衡极为脆弱。当智能体在动态交互中学习时,每次梯度更新都会改变对方动作分布,使原本合作的伙伴变成高敏感决策阶段的随机噪声源。我们研究这种共学习噪声如何传播于协调博弈结构中,发现即使强帕累托占优的合作均衡,在标准风险中性学习下也呈指数级不稳定,一旦伙伴噪声超过临界阈值即不可逆崩溃。试图用分布鲁棒性应对不确定性反而恶化问题:风险规避的目标会惩罚高方差合作行为,加剧不稳定性。我们揭示根源在于鲁棒性应针对由伙伴不确定性引发的策略梯度方差,而非回报分布。由此提出新算法,其梯度更新受伙伴不可预测性在线度量调节,在对称协调游戏中可严格扩展合作基域。为统一稳定性、样本复杂度与福利后果,我们引入‘焦虑代价’作为‘无政府代价’的对偶结构,并结合新型‘合作窗口’,精准刻画学习算法在伙伴噪声下能恢复的福利上限,给出鲁棒性程度的闭式最优平衡。
原文摘要 · Abstract (English)
Cooperative equilibria are fragile. When agents learn alongside each other rather than in a fixed environment, the process of learning destabilizes the cooperation they are trying to sustain: every gradient step an agent takes shifts the distribution of actions its partner will play, turning a cooperative partner into a source of stochastic noise precisely where the cooperation decision is most sensitive. We study how this co-learning noise propagates through the structure of coordination games, and find that the cooperative equilibrium, even when strongly Pareto-dominant, is exponentially unstable under standard risk-neutral learning, collapsing irreversibly once partner noise crosses the game's critical cooperation threshold. The natural response to apply distributional robustness to hedge against partner uncertainty makes things strictly worse: risk-averse return objectives penalize the high-variance cooperative action relative to defection, widening the instability region rather than shrinking it, a paradox that reveals a fundamental mismatch between the domains where robustness is applied and instability originates. We resolve this by showing that robustness should target the policy gradient update variance induced by partner uncertainty, not the return distribution. This distinction yields an algorithm whose gradient updates are modulated by an online measure of partner unpredictability, provably expanding the cooperation basin in symmetric coordination games. To unify stability, sample complexity, and welfare consequences of this approach, we introduce the Price of Paranoia as the structural dual of the Price of Anarchy. Together with a novel Cooperation Window, it precisely characterizes how much welfare learning algorithms can recover under partner noise, pinning down the optimal degree of robustness as a closed-form balance between equilibrium stability and sample efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。