arXiv:2509.18371eess.SYcs.MA2025-09中稿 · and will be presen…

用自注意力设计分布式策略,解决多智能体非线性博弈的通信与非平稳难题。

Policy Gradient with Self-Attention for Model-Free Distributed Nonlinear Multi-Agent Games

  • 基于自注意力建模时变通信拓扑下的非线性反馈策略
  • 在仿真与真实机器人追逃任务中表现优异
  • 适用于多团队协作、通信受限的分布式博弈场景

动态非线性多智能体博弈因智能体间交互随时间变化及(潜在)纳什均衡非平稳而极具挑战。本文研究模型无关博弈,即智能体的转移与代价可观测,但其生成函数未知。提出一种符合多团队博弈通信约束的新型分布式策略结构,每队含多个智能体,并通过策略梯度学习。该方法受线性二次博弈中分布式策略结构启发,后者表现为时变线性反馈增益。在非线性情况下,将策略建模为非线性反馈增益,由自注意力层参数化以适应时变的多智能体通信拓扑。实验表明,该方法在分布式线性和非线性调节任务,以及模拟和真实的多机器人追逃游戏中均取得优异性能。

原文摘要 · Abstract (English)

Multi-agent games in dynamic nonlinear settings are challenging due to the time-varying interactions among the agents and the non-stationarity of the (potential) Nash equilibria. In this paper we consider model-free games, where agent transitions and costs are observed without knowledge of the transition and cost functions that generate them. We propose a novel distributed policy structure that follows the communication constraints in multi-team games, with multiple agents per team, and learned through policy gradients. Our formulation is inspired by the structure of distributed policies in linear quadratic games, which take the form of time-varying linear feedback gains. In the nonlinear case, we model the policies as nonlinear feedback gains, parameterized by self-attention layers to account for the time-varying multi-agent communication topology. We demonstrate that our approach achieves strong performance in several settings, including distributed linear and nonlinear regulation, and simulated and real multi-robot pursuit-and-evasion games.

多智能体策略梯度自注意力分布式控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。