arXiv:2604.00717cs.MAcs.AI2026-04

提出主动共享感知机制,解决多智能体协同优化中的非平稳性问题。

GRASP: Gradient Realignment via Active Shared Perception for Multi-Agent Collaborative Optimization

  • 用独立梯度生成共识梯度,实现主动感知其他智能体策略更新
  • 理论证明共识方向可确保稳定均衡存在且可达
  • 在SMAC和GRF上验证了框架的可扩展性和高效性

非平稳性源于智能体并行策略更新,导致环境持续波动。现有方法如集中训练分散执行(CTDE)和顺序更新虽缓解此问题,但因对其他智能体的感知仍依赖环境交互采样,智能体处于被动感知状态,不可避免引发平衡震荡,显著降低系统收敛速度。为此,本文提出梯度重对齐主动共享感知(GRASP)框架,将广义贝尔曼均衡定义为策略演化的稳定目标。核心机制是利用各智能体的独立梯度推导出共识梯度,使智能体能主动感知策略更新,优化团队协作。理论上,借助Kakutani不动点定理,证明共识方向$u^*$保证该均衡的存在性与可达性。在星海争霸II多智能体挑战(SMAC)与谷歌研究足球(GRF)上的大量实验表明,该框架具备良好可扩展性与优异性能。

原文摘要 · Abstract (English)

Non-stationarity arises from concurrent policy updates and leads to persistent environmental fluctuations. Existing approaches like Centralized Training with Decentralized Execution (CTDE) and sequential update schemes mitigate this issue. However, since the perception of the policies of other agents remains dependent on sampling environmental interaction data, the agent essentially operates in a passive perception state. This inevitably triggers equilibrium oscillations and significantly slows the convergence speed of the system. To address this issue, we propose Gradient Realignment via Active Shared Perception (GRASP), a novel framework that defines generalized Bellman equilibrium as a stable objective for policy evolution. The core mechanism of GRASP involves utilizing the independent gradients of agents to derive a defined consensus gradient, enabling agents to actively perceive policy updates and optimize team collaboration. Theoretically, we leverage the Kakutani Fixed-Point Theorem to prove that the consensus direction $u^*$ guarantees the existence and attainability of this equilibrium. Extensive experiments on StarCraft II Multi-Agent Challenge (SMAC) and Google Research Football (GRF) demonstrate the scalability and promising performance of the framework.

多智能体协同优化梯度对齐强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。