提出可随时约束的均衡新概念,实现高效计算。
Anytime-Constrained Equilibria in Polynomial Time
- 将即时约束引入马尔可夫博弈,定义新均衡概念
- 给出近似求解的多项式时间算法,逼近精度最优
- 适合研究博弈计算与约束决策的学者参考
我们将即时约束扩展至马尔可夫博弈场景,提出即时约束均衡(ACE)的概念。建立了完整的理论体系,包括:(1) 可行策略的计算表征;(2) 计算ACE的固定参数可追踪算法;(3) 近似计算ACE的多项式时间算法。由于在双人零和博弈中求可行策略已是NP难问题,因此我们的近似保证在P≠NP的前提下是最佳的。我们还首次构建了动作约束马尔可夫博弈的高效计算理论,可能具有独立研究价值。
原文摘要 · Abstract (English)
We extend anytime constraints to the Markov game setting and the corresponding solution concept of an anytime-constrained equilibrium (ACE). Then, we present a comprehensive theory of anytime-constrained equilibria that includes (1) a computational characterization of feasible policies, (2) a fixed-parameter tractable algorithm for computing ACE, and (3) a polynomial-time algorithm for approximately computing ACE. Since computing a feasible policy is NP-hard even for two-player zero-sum games, our approximation guarantees are optimal so long as $P \neq NP$. We also develop the first theory of efficient computation for action-constrained Markov games, which may be of independent interest.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。