提出MFC-EQ框架,让多智能体在不通信下自动保持队形并快速到达目标。
MFC-EQ: Mean-Field Control with Envelope Q-Learning for Moving Decentralized Agents in Formation
- 用平均场理论简化多智能体状态,降低计算复杂度
- 通过包络Q学习实现单一策略适配不同任务偏好
- 能动态调整队形,适合复杂动态场景
我们研究了移动队形中多智能体路径规划(MAiF)的去中心化版本,目标是在部分观测和有限通信条件下,使多个智能体快速抵达目标的同时保持期望队形。由于队形依赖所有智能体的联合状态,其维度随智能体数量指数增长,导致学习过程难以处理。同时,设计一个能适应不同线性偏好组合的统一策略也极具挑战。本文提出均值场控制与包络Q学习结合的MFC-EQ框架,利用平均场理论近似全局动态,并通过包络Q学习训练出无需依赖具体偏好的通用策略。在多种实例上的实证评估表明,MFC-EQ优于现有最先进集中式MAiF基线。此外,该方法能有效应对目标队形动态变化的情况——这是现有MAiF规划器无法解决的难题。
原文摘要 · Abstract (English)
We study a decentralized version of Moving Agents in Formation (MAiF), a variant of Multi-Agent Path Finding aiming to plan collision-free paths for multiple agents with the dual objectives of reaching their goals quickly while maintaining a desired formation. The agents must balance these objectives under conditions of partial observation and limited communication. The formation maintenance depends on the joint state of all agents, whose dimensionality increases exponentially with the number of agents, rendering the learning process intractable. Additionally, learning a single policy that can accommodate different linear preferences for these two objectives presents a significant challenge. In this paper, we propose Mean-Field Control with Envelop $Q$-learning (MFC-EQ), a scalable and adaptable learning framework for this bi-objective multi-agent problem. We approximate the dynamics of all agents using mean-field theory while learning a universal preference-agnostic policy through envelop $Q$-learning. Our empirical evaluation of MFC-EQ across numerous instances shows that it outperforms state-of-the-art centralized MAiF baselines. Furthermore, MFC-EQ effectively handles more complex scenarios where the desired formation changes dynamically -- a challenge that existing MAiF planners cannot address.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。