arXiv:2509.24047cs.LGcs.SY2025-09被引 1

用风险追求统一乐观机制,提升多智能体协作效率

Optimism as Risk-Seeking in Multi-Agent Reinforcement Learning

  • 将乐观性形式化为偏离惩罚的风险追求评估
  • 在合作基准上优于风险中立与启发式乐观方法
  • 适合追求理论与实践结合的多智能体强化学习研究者

风险敏感性已成为强化学习的核心议题,凸风险度量和鲁棒性公式提供了超越期望回报建模偏好的原则性方法。尽管多智能体强化学习(MARL)的近期扩展大多关注风险规避设置以增强对不确定性的鲁棒性,但在合作型MARL中,这种保守性常导致次优均衡。已有研究表明,乐观性可促进合作。然而,现有乐观方法虽在实践中有效,通常缺乏理论基础。基于凸风险度量的对偶表示,我们提出一个原则性框架,将风险追求目标解释为乐观性。我们引入乐观价值函数,形式化乐观性为发散惩罚的风险追求评估。在此基础上,推导出乐观价值函数的策略梯度定理,包含熵风险/KL惩罚设定下的显式公式,并开发了实现这些更新的去中心化乐观演员-评论家算法。在合作基准上的实证结果表明,风险追求型乐观性在协调性上持续优于风险中立基线和启发式乐观方法。因此,该框架统一了风险敏感学习与乐观性,为多智能体强化学习中的合作提供了一个理论坚实且实用有效的途径。

原文摘要 · Abstract (English)

Risk sensitivity has become a central theme in reinforcement learning (RL), where convex risk measures and robust formulations provide principled ways to model preferences beyond expected return. Recent extensions to multi-agent RL (MARL) have largely emphasized the risk-averse setting, prioritizing robustness to uncertainty. In cooperative MARL, however, such conservatism often leads to suboptimal equilibria, and a parallel line of work has shown that optimism can promote cooperation. Existing optimistic methods, though effective in practice, are typically heuristic and lack theoretical grounding. Building on the dual representation for convex risk measures, we propose a principled framework that interprets risk-seeking objectives as optimism. We introduce optimistic value functions, which formalize optimism as divergence-penalized risk-seeking evaluations. Building on this foundation, we derive a policy-gradient theorem for optimistic value functions, including explicit formulas for the entropic risk/KL-penalty setting, and develop decentralized optimistic actor-critic algorithms that implement these updates. Empirical results on cooperative benchmarks demonstrate that risk-seeking optimism consistently improves coordination over both risk-neutral baselines and heuristic optimistic methods. Our framework thus unifies risk-sensitive learning and optimism, offering a theoretically grounded and practically effective approach to cooperation in MARL.

多智能体强化学习风险敏感乐观性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。