arXiv:2508.12087cs.AIcs.MA2025-08被引 2

用自回归世界模型提升多智能体路径规划的全局决策能力

MAPF-World: Action World Model for Multi-Agent Path Finding

  • 构建可预测未来状态与动作的自回归世界模型
  • 在复杂场景中实现更优协同与零样本泛化性能
  • 模型更小、数据需求更低,适合真实场景部署

多智能体路径规划(MAPF)旨在为多个智能体规划无冲突的路径。该问题广泛应用于多机器人协作、物流辅助和社交导航等任务。近期基于基础模型与大规模数据的分布式可学习求解器在大规模MAPF中表现优异,但其作为反应式策略模型,难以建模环境的时间动态与智能体间依赖关系,在长期复杂规划中性能下降。为此,我们提出MAPF-World,一种用于MAPF的自回归动作世界模型,统一情境理解与动作生成,使决策超越即时局部观察。通过显式建模环境动态,包括空间特征与时间依赖,基于对未来状态与动作的预测,提升情境感知能力。结合预测未来信息,实现更知情、协调且具有前瞻性的决策,尤其在复杂多智能体环境中表现突出。此外,我们引入基于真实场景的自动地图生成器,构建符合实际布局的训练与评估基准。大量实验表明,MAPF-World优于现有最优可学习求解器,在分布外情形下展现出卓越零样本泛化能力。值得注意的是,该模型仅需96.5%更小的模型规模与92%更少的数据即可训练。

原文摘要 · Abstract (English)

Multi-agent path finding (MAPF) is the problem of planning conflict-free paths from the designated start locations to goal positions for multiple agents. It underlies a variety of real-world tasks, including multi-robot coordination, robot-assisted logistics, and social navigation. Recent decentralized learnable solvers have shown great promise for large-scale MAPF, especially when leveraging foundation models and large datasets. However, these agents are reactive policy models and exhibit limited modeling of environmental temporal dynamics and inter-agent dependencies, resulting in performance degradation in complex, long-term planning scenarios. To address these limitations, we propose MAPF-World, an autoregressive action world model for MAPF that unifies situation understanding and action generation, guiding decisions beyond immediate local observations. It improves situational awareness by explicitly modeling environmental dynamics, including spatial features and temporal dependencies, through future state and actions prediction. By incorporating these predicted futures, MAPF-World enables more informed, coordinated, and far-sighted decision-making, especially in complex multi-agent settings. Furthermore, we augment MAPF benchmarks by introducing an automatic map generator grounded in real-world scenarios, capturing practical map layouts for training and evaluating MAPF solvers. Extensive experiments demonstrate that MAPF-World outperforms state-of-the-art learnable solvers, showcasing superior zero-shot generalization to out-of-distribution cases. Notably, MAPF-World is trained with a 96.5% smaller model size and 92% reduced data.

多智能体路径规划世界模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。