arXiv:2410.16600cs.GTcs.AI2024-10ICML被引 5

提出凸马尔可夫博弈框架,解决多智能体决策中的非加性偏好问题。

Convex Markov Games: A New Frontier for Multi-Agent Reinforcement Learning

  • 基于占据测度的凸偏好建模,突破传统马尔可夫博弈限制
  • 在无限时域下仍存在纯策略纳什均衡,可梯度下降逼近
  • 实验验证公平协作、安全长程行为及人类模仿等新解法

行为多样性、专家模仿、公平性、安全性等序列决策偏好无法在时间上加性分解。本文引入凸马尔可夫博弈,支持对占据测度的一般凸偏好建模。尽管具有无限时域且比马尔可夫博弈更一般,纯策略纳什均衡依然存在。通过梯度下降优化可利用性的上界,可实证逼近均衡。实验发现经典重复型对策的新解法,在不对称协调博弈中找到公平解,并在机器人仓库环境中优先保障长期安全行为。在囚徒困境中,算法利用瞬时模仿,得到与人类行为仅微小偏差的策略组合,但每位玩家收益更高,且可被利用性降低三个数量级。

原文摘要 · Abstract (English)

Behavioral diversity, expert imitation, fairness, safety goals and others give rise to preferences in sequential decision making domains that do not decompose additively across time. We introduce the class of convex Markov games that allow general convex preferences over occupancy measures. Despite infinite time horizon and strictly higher generality than Markov games, pure strategy Nash equilibria exist. Furthermore, equilibria can be approximated empirically by performing gradient descent on an upper bound of exploitability. Our experiments reveal novel solutions to classic repeated normal-form games, find fair solutions in a repeated asymmetric coordination game, and prioritize safe long-term behavior in a robot warehouse environment. In the prisoner's dilemma, our algorithm leverages transient imitation to find a policy profile that deviates from observed human play only slightly, yet achieves higher per-player utility while also being three orders of magnitude less exploitable.

多智能体强化学习博弈论安全决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。