让多智能体学会选对合作均衡,靠的是感知对手的更新策略。
Equilibrium Selection in Multi-Agent Policy Gradients via Opponent-Aware Basin Entry

- 引入对手感知的修正项,改变智能体进入合作均衡的概率。
- 在猎鹿、囚徒困境等场景中,合作均衡的进入概率显著提升。
- 适合研究多智能体协作与博弈均衡选择的学者参考。
多智能体策略梯度方法虽能局部收敛至稳定的纳什均衡,但无法确定具体收敛到哪个均衡。本文通过目标均衡集的吸引域进入概率来研究这一问题。对于有限回溯的Meta-MAPG,其更新可分解为普通策略梯度加上自身学习和同伴学习修正项,具有可控采样噪声与有限回溯偏差。我们识别出同伴学习修正项是主要的均衡选择机制:在局部对齐条件下,进入目标稳定纳什集吸引域的概率相对普通策略梯度提高。由于持续修正可能改变原博弈的零更新点,因此在进入吸引域后退火修正项,可恢复普通策略梯度动态并保留局部稳定纳什收敛性。在猎鹿博弈、迭代囚徒困境及初步神经策略协调环境中的实验支持该吸引域进入视角,显示同伴感知更新下合作吸引域的进入概率增加。
原文摘要 · Abstract (English)
Multi-agent policy-gradient methods have been shown to converge locally near stable Nash equilibria. Local convergence, however, does not determine which equilibrium is reached. We study this question through basin-entry probability with respect to a target set of equilibria selected by an external criterion, such as payoff dominance. For finite-unroll Meta-MAPG, we show that the update decomposes into ordinary policy gradient plus own-learning and peer-learning corrections, with controlled sampling noise and finite-unroll bias. We identify the peer-learning correction as the main equilibrium-selection mechanism: under a local alignment condition, the probability of entering the certified attraction region of the target stable-Nash set increases, relative to ordinary policy gradient. Because persistent correction may shift zero-update points of the original game, annealing the correction after entering the basin recovers ordinary policy-gradient dynamics and inherits local stable-Nash convergence guarantees. Experiments in Stag Hunt, iterated Prisoner's Dilemma, and preliminary neural-policy coordination environments support this basin-entry view, showing increased entry into cooperative basins under peer-aware updates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。