arXiv:2601.08210cs.LG2026-01

用集体影响估计让多智能体协作更高效,不依赖通信

Scalable Multiagent Reinforcement Learning with Collective Influence Estimation

  • 通过建模其他智能体对任务目标的集体影响,仅凭局部观测实现协作
  • 在通信受限环境下仍保持稳定高效,且支持新智能体无缝加入
  • 已在真实机器人平台验证,显著降低对通信的依赖

多智能体强化学习(MARL)在解决复杂协作任务方面潜力巨大,但现有方法通常依赖频繁的动作或状态信息交换,难以在实际机器人系统中实现。常见做法是引入估计器网络预测其他智能体行为,但其规模和计算开销随智能体数量迅速增长,限制了大规模应用。为此,本文提出一种增强型多智能体学习框架,引入集体影响估计网络(CIEN)。通过显式建模其他智能体对任务目标的集体影响,每个智能体仅需本地观测和任务目标状态即可推断关键交互信息,实现无需显式动作交换的高效协作。该框架有效避免了网络随团队规模扩展,新智能体可无缝集成而无需修改已有结构,展现出强可扩展性。基于Soft Actor-Critic(SAC)算法的多智能体协作任务实验表明,该方法在通信受限环境下仍能实现稳定高效的协调。此外,基于集体影响建模训练的策略已在真实机器人平台上部署,实验结果表明其具备显著提升的鲁棒性和部署可行性,同时大幅减少对通信基础设施的依赖。

原文摘要 · Abstract (English)

Multiagent reinforcement learning (MARL) has attracted considerable attention due to its potential in addressing complex cooperative tasks. However, existing MARL approaches often rely on frequent exchanges of action or state information among agents to achieve effective coordination, which is difficult to satisfy in practical robotic systems. A common solution is to introduce estimator networks to model the behaviors of other agents and predict their actions; nevertheless, such designs cause the size and computational cost of the estimator networks to grow rapidly with the number of agents, thereby limiting scalability in large-scale systems. To address these challenges, this paper proposes a multiagent learning framework augmented with a Collective Influence Estimation Network (CIEN). By explicitly modeling the collective influence of other agents on the task object, each agent can infer critical interaction information solely from its local observations and the task object's states, enabling efficient collaboration without explicit action information exchange. The proposed framework effectively avoids network expansion as the team size increases; moreover, new agents can be incorporated without modifying the network structures of existing agents, demonstrating strong scalability. Experimental results on multiagent cooperative tasks based on the Soft Actor-Critic (SAC) algorithm show that the proposed method achieves stable and efficient coordination under communication-limited environments. Furthermore, policies trained with collective influence modeling are deployed on a real robotic platform, where experimental results indicate significantly improved robustness and deployment feasibility, along with reduced dependence on communication infrastructure.

多智能体强化学习协作可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。