用扩散模型提升自对弈策略学习,更快更稳收敛到最优解
DiffFP: Learning Behaviors from Scratch via Diffusion-based Fictitious Play
- 基于扩散模型生成多样策略,动态逼近对手最佳应对
- 在连续空间零和博弈中实现ε-纳什均衡,收敛速度提升3倍
- 适合复杂对抗环境,对未知对手策略鲁棒性强
自对弈强化学习在竞争性多智能体游戏中展现出显著成效,但在连续决策空间中仍面临挑战。确保自对弈场景下的适应性和泛化能力对于在动态环境中取得竞争力至关重要。现有方法常出现收敛缓慢或无法收敛至纳什均衡的问题,使智能体易被未见对手策略利用。为此,我们提出DiffFP,一种基于扩散模型的虚构对弈(Fictitious Play)框架,可在学习过程中估计对未知对手的最佳响应,并生成稳健、多模态的行为策略。具体而言,采用扩散策略近似最佳响应,利用生成建模学习自适应且多样化的策略。实验表明,该框架在连续空间零和博弈中可收敛至ε-纳什均衡。在赛车与多粒子零和游戏等复杂多智能体环境中验证,所学策略对多样化对手具有鲁棒性,平均性能优于基于强化学习的基线方法,收敛速度提升最多达3倍,成功率高出30倍,展现出对对手策略的强鲁棒性与训练过程稳定性。
原文摘要 · Abstract (English)
Self-play reinforcement learning has demonstrated significant success in learning complex strategic and interactive behaviors in competitive multi-agent games. However, achieving such behaviors in continuous decision spaces remains challenging. Ensuring adaptability and generalization in self-play settings is critical for achieving competitive performance in dynamic multi-agent environments. These challenges often cause methods to converge slowly or fail to converge at all to a Nash equilibrium, making agents vulnerable to strategic exploitation by unseen opponents. To address these challenges, we propose DiffFP, a fictitious play (FP) framework that estimates the best response to unseen opponents while learning a robust and multimodal behavioral policy. Specifically, we approximate the best response using a diffusion policy that leverages generative modeling to learn adaptive and diverse strategies. Through empirical evaluation, we demonstrate that the proposed FP framework converges towards $ε$-Nash equilibria in continuous- space zero-sum games. We validate our method on complex multi-agent environments, including racing and multi-particle zero-sum games. Simulation results show that the learned policies are robust against diverse opponents and outperform baseline reinforcement learning policies. Our approach achieves up to 3$\times$ faster convergence and 30$\times$ higher success rates on average against RL-based baselines, demonstrating its robustness to opponent strategies and stability across training iterations
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。