arXiv:2505.06706cs.AI2025-05

提出双层平均场方法,动态分组智能体以减少学习噪声。

Bi-level Mean Field: Dynamic Grouping for Large-Scale MARL

  • 用变分自编码器动态分组智能体,捕捉个体差异。
  • 双层交互机制同时建模组间与组内关系,提升聚合精度。
  • 在多任务实验中优于现有最优方法,适合大规模多智能体场景。

大规模多智能体强化学习常因智能体间交互呈指数级增长而面临维度灾难,显著增加计算复杂度并降低学习效率。现有基于平均场(MF)的方法通过将邻近智能体近似为单一均值智能体来简化交互结构,使整体复杂度降至成对交互。然而,这些方法无法反映个体差异,导致平均场学习中迭代更新不准确,产生聚合噪声。本文提出双层平均场(BMF)方法,在大规模MARL中通过动态分组捕捉智能体多样性,以缓解聚合噪声。具体而言,BMF引入动态分组模块,利用变分自编码器(VAE)学习智能体表示,实现随时间动态分组;进一步设计双层交互模块,建模组间与组内交互,实现高效邻近聚合。在多个任务上的实验表明,所提BMF性能优于当前最先进方法。

原文摘要 · Abstract (English)

Large-scale Multi-Agent Reinforcement Learning (MARL) often suffers from the curse of dimensionality, as the exponential growth in agent interactions significantly increases computational complexity and impedes learning efficiency. To mitigate this, existing efforts that rely on Mean Field (MF) simplify the interaction landscape by approximating neighboring agents as a single mean agent, thus reducing overall complexity to pairwise interactions. However, these MF methods inevitably fail to account for individual differences, leading to aggregation noise caused by inaccurate iterative updates during MF learning. In this paper, we propose a Bi-level Mean Field (BMF) method to capture agent diversity with dynamic grouping in large-scale MARL, which can alleviate aggregation noise via bi-level interaction. Specifically, BMF introduces a dynamic group assignment module, which employs a Variational AutoEncoder (VAE) to learn the representations of agents, facilitating their dynamic grouping over time. Furthermore, we propose a bi-level interaction module to model both inter- and intra-group interactions for effective neighboring aggregation. Experiments across various tasks demonstrate that the proposed BMF yields results superior to the state-of-the-art methods.

多智能体平均场动态分组强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。