提出新博弈均衡,让团队在信息不对称下更抗欺骗。
Beyond Bayesian Nash: Learning Minimax-Regret Equilibria for Adversarial Team Games under Asymmetric Information

- 用概率鲁棒的极小极大后悔机制,结合先验类型分布
- 在图结构对抗任务中,对隐藏对手类型表现更稳健
- 适合需要抗策略欺骗的多人对抗场景
具有非对称信息的对抗团队游戏(ATGs),如路径对抗、目标搜索与可达性游戏,要求策略能抵御隐藏对手类型(如隐藏目标旗)和欺骗。欺骗表现为对手类型分布的策略性调整,使全知对手可与自然协同,基于观测类型制定策略。现有风险中立解法(如贝叶斯纳什均衡)对分布偏移敏感,而分布鲁棒方法仅在预设模糊集内提供保证。为此,本文提出概率鲁棒极小极大后悔均衡(PR-MRE),融合极小极大后悔的分布无关鲁棒性与名义类型分布的概率信息。PR-MRE 在高置信度类型子空间上最小化最坏情况后悔,既抵御概率质量的战略重分配,又避免完全分布无关方法的过度保守。我们证明,在标准形式贝叶斯博弈中,PR-MRE 可表示为鲁棒双线性规划,并推导出可计算的半定松弛。进一步将该松弛嵌入鲁棒双轨奥尔策略优化(PRMRE-PSRO)框架,实现基于深度强化学习最佳响应的群体学习,获得近似 PR-MRE 策略。图结构对抗团队游戏实验表明,相较于风险中立均衡,PR-MRE 策略在隐藏类型下的最坏情况性能显著提升,展现出更强的对抗战略分布偏移能力。
原文摘要 · Abstract (English)
Adversarial team games (ATGs) with asymmetric information, such as adversarial path-finding, goal search, and reachability games on graphs, require strategies that are robust to hidden opponent types, such as a hidden goal flag, and to deception. Under asymmetric information, deception is seen as strategic shifts in the type distribution such that the omniscient opponent can collude with Nature and condition its play on the observed type. Existing risk-neutral solution concepts, such as Bayesian Nash equilibrium (BNE), are sensitive to distribution shifts, while distributionally robust approaches provide guarantees only within a prescribed ambiguity set. To address these limitations, we introduce Probabilistically Robust Minimax-Regret Equilibrium (PR-MRE), a novel equilibrium concept that combines the distribution-free robustness of minimax-regret reasoning with probabilistic information from a nominal type distribution. PR-MRE minimizes worst-case regret over a high-confidence subset of the type space, providing protection against strategic redistribution of probability mass while avoiding the conservatism of fully distribution-free approaches. We show that, for normal-form Bayesian games, PR-MRE can be formulated as a robust bilinear program and derive a tractable semidefinite relaxation. We then adapt this relaxation into a novel meta-solver within a robust double-oracle framework, PRMRE-PSRO, enabling population-based learning of approximate PR-MRE strategies via deep reinforcement learning best responses. Experiments on graph-structured adversarial team games demonstrate that PR-MRE discovers strategies with substantially improved worst-case performance across hidden types compared to risk-neutral equilibrium solutions, resulting in more robust behavior under strategic distribution shifts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。