arXiv:2503.08883cs.AI2025-03

提出新方法学习博弈中领导者与追随者的协同策略,提升多智能体交互建模精度。

Imitation Learning of Correlated Policies in Stackelberg Games

  • 设计专用于斯塔克尔伯格博弈的联合策略分布,捕捉领导者与追随者间的依赖关系
  • 引入隐空间差分网络(LSDN)实现高维环境下的稳定策略学习
  • 无需对抗训练,适用于复杂交互场景,如体育博弈和安全策略优化

斯塔克尔伯格博弈广泛应用于经济与安全等领域,其特点是领导者先行决策,追随者据此响应。在多智能体系统中,智能体行为相互依赖,传统多智能体模仿学习(MAIL)方法难以捕捉此类复杂交互。相关策略需考虑对手策略,但在斯塔克尔伯格博弈中,因决策不对称,领导者与追随者无法同时响应对方,导致现有方法难以生成相关策略。此外,基于占据度匹配或对抗训练(如GAIL、逆强化学习)的方法在高维环境中存在可扩展性差与训练不稳定问题。为此,本文提出专为斯塔克尔伯格博弈设计的相关策略占据度,并引入隐空间斯塔克尔伯格差分网络(LSDN),将双智能体交互建模为共享隐状态轨迹,采用多输出几何布朗运动(MO-GBM)有效捕捉联合策略。通过分离环境影响与智能体驱动转移,LSDN 实现了互依策略的同步学习,无需对抗训练,简化学习流程。在迭代矩阵博弈与多智能体粒子环境上的实验表明,LSDN 在再现复杂交互动态方面优于现有MAIL方法。

原文摘要 · Abstract (English)

Stackelberg games, widely applied in domains like economics and security, involve asymmetric interactions where a leader's strategy drives follower responses. Accurately modeling these dynamics allows domain experts to optimize strategies in interactive scenarios, such as turn-based sports like badminton. In multi-agent systems, agent behaviors are interdependent, and traditional Multi-Agent Imitation Learning (MAIL) methods often fail to capture these complex interactions. Correlated policies, which account for opponents' strategies, are essential for accurately modeling such dynamics. However, even methods designed for learning correlated policies, like CoDAIL, struggle in Stackelberg games due to their asymmetric decision-making, where leaders and followers cannot simultaneously account for each other's actions, often leading to non-correlated policies. Furthermore, existing MAIL methods that match occupancy measures or use adversarial techniques like GAIL or Inverse RL face scalability challenges, particularly in high-dimensional environments, and suffer from unstable training. To address these challenges, we propose a correlated policy occupancy measure specifically designed for Stackelberg games and introduce the Latent Stackelberg Differential Network (LSDN) to match it. LSDN models two-agent interactions as shared latent state trajectories and uses multi-output Geometric Brownian Motion (MO-GBM) to effectively capture joint policies. By leveraging MO-GBM, LSDN disentangles environmental influences from agent-driven transitions in latent space, enabling the simultaneous learning of interdependent policies. This design eliminates the need for adversarial training and simplifies the learning process. Extensive experiments on Iterative Matrix Games and multi-agent particle environments demonstrate that LSDN can better reproduce complex interaction dynamics than existing MAIL methods.

博弈学习多智能体策略建模模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。