arXiv:2505.09756cs.LGcs.MA2025-05

提出基于社区的多智能体强化学习框架,支持动态协作与迁移探索。

Community-based Multi-Agent Reinforcement Learning with Transfer and Active Exploration

  • 智能体可隶属多个重叠社区,通过成员权重聚合共享策略与价值函数。
  • 理论证明在线性函数近似下,策略与价值更新均能收敛。
  • 兼具迁移学习与主动探索能力,适合动态多智能体系统应用。

我们提出一种新的多智能体强化学习(MARL)框架,其中智能体在具有潜在社区结构和混合成员关系的时间演化网络中协作。不同于传统的邻接或固定交互图,该社区框架通过允许每个智能体属于多个重叠社区,捕捉灵活且抽象的协调模式。每个社区维护共享的策略与价值函数,由个体智能体根据个性化成员权重进行聚合。我们设计了利用此结构的演员-评论家算法:智能体继承社区级估计用于策略更新与价值学习,实现结构化信息共享,且无需访问其他智能体的策略。重要的是,该方法支持通过成员估计适应新智能体或任务的迁移学习,以及通过优先探索不确定社区的主动学习。理论上,我们在线性函数近似下为演员与评论家更新建立了收敛性保证。据我们所知,这是首个将社区结构、可迁移性与主动学习结合并具备可证明保证的MARL框架。

原文摘要 · Abstract (English)

We propose a new framework for multi-agent reinforcement learning (MARL), where the agents cooperate in a time-evolving network with latent community structures and mixed memberships. Unlike traditional neighbor-based or fixed interaction graphs, our community-based framework captures flexible and abstract coordination patterns by allowing each agent to belong to multiple overlapping communities. Each community maintains shared policy and value functions, which are aggregated by individual agents according to personalized membership weights. We also design actor-critic algorithms that exploit this structure: agents inherit community-level estimates for policy updates and value learning, enabling structured information sharing without requiring access to other agents' policies. Importantly, our approach supports both transfer learning by adapting to new agents or tasks via membership estimation, and active learning by prioritizing uncertain communities during exploration. Theoretically, we establish convergence guarantees under linear function approximation for both actor and critic updates. To our knowledge, this is the first MARL framework that integrates community structure, transferability, and active learning with provable guarantees.

多智能体强化学习社区结构迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。