arXiv:2510.03592cs.LGcs.AI2025-10被引 9

用虚拟信息素让多机器人在狭窄空间自组织协作,无需通信

Deep Reinforcement Learning for Multi-Agent Coordination

  • 借鉴昆虫群落的痕迹通讯,用虚拟信息素建模局部与社会互动
  • 八机器人协作时实现非对称任务分配,显著降低拥堵并提升性能
  • 适合解决通信受限的密集场景多智能体协同问题

针对狭窄封闭环境中多机器人因拥堵和干扰导致协同效率低下的问题,受昆虫群体通过痕迹信息(stigmergy)实现鲁棒协作的启发,提出一种基于虚拟信息素的分布式多智能体强化学习框架(S-MADRL),无需显式通信即可实现去中心化的涌现式协调。为克服现有算法(如MADQN、MADDPG、MAPPO)在收敛性和可扩展性上的瓶颈,引入课程学习策略,将复杂任务逐步分解为更难的子问题。仿真结果表明,该框架在最多八名机器人协作时表现出最优协调能力,机器人自发形成非对称工作负载分布,有效缓解拥堵并调节群体表现。这种类自然演化的行为模式,为通信受限的拥挤环境中多智能体系统的可扩展协同提供了新方案。

原文摘要 · Abstract (English)

We address the challenge of coordinating multiple robots in narrow and confined environments, where congestion and interference often hinder collective task performance. Drawing inspiration from insect colonies, which achieve robust coordination through stigmergy -- modifying and interpreting environmental traces -- we propose a Stigmergic Multi-Agent Deep Reinforcement Learning (S-MADRL) framework that leverages virtual pheromones to model local and social interactions, enabling decentralized emergent coordination without explicit communication. To overcome the convergence and scalability limitations of existing algorithms such as MADQN, MADDPG, and MAPPO, we leverage curriculum learning, which decomposes complex tasks into progressively harder sub-problems. Simulation results show that our framework achieves the most effective coordination of up to eight agents, where robots self-organize into asymmetric workload distributions that reduce congestion and modulate group performance. This emergent behavior, analogous to strategies observed in nature, demonstrates a scalable solution for decentralized multi-agent coordination in crowded environments with communication constraints.

多智能体强化学习协同控制分布式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。