arXiv:2506.23793cs.AIcs.LG2025-06被引 6

用主动微调提升多智能体路径规划模型,支持百万级机器人协同

Advancing Learnable Multi-Agent Pathfinding Solvers with Active Fine-Tuning

  • 基于预训练模型,通过中心化专家数据主动微调
  • 在多种场景下性能超越现有学习型求解器,最高支持百万智能体
  • 适合大规模机器人调度、物流等需要高并发协同的场景

多智能体路径规划(MAPF)是多机器人轨迹规划问题的常见抽象,多个同质机器人需在共享环境中同时移动。尽管最优求解已被证明为NP难问题,但高效可扩展的求解器对物流、搜救等实际应用至关重要。为此,基于机器学习的分布式次优求解器逐渐兴起。在近期提出的纯模仿学习求解器MAPF-GPT基础上,本文提出MAPF-GPT-DDG,该方法利用中心化专家数据对预训练的MAPF模型进行有效微调。通过一种新颖的增量数据生成机制,显著加速训练并大幅提升测试性能。实验表明,MAPF-GPT-DDG在多种测试场景中均优于所有现有学习型求解器,包括原始的MAPF-GPT。尤为突出的是,其可处理单个环境中多达100万智能体的实例,创下MAPF领域可扩展性新纪录。

原文摘要 · Abstract (English)

Multi-agent pathfinding (MAPF) is a common abstraction of multi-robot trajectory planning problems, where multiple homogeneous robots simultaneously move in the shared environment. While solving MAPF optimally has been proven to be NP-hard, scalable, and efficient, solvers are vital for real-world applications like logistics, search-and-rescue, etc. To this end, decentralized suboptimal MAPF solvers that leverage machine learning have come on stage. Building on the success of the recently introduced MAPF-GPT, a pure imitation learning solver, we introduce MAPF-GPT-DDG. This novel approach effectively fine-tunes the pre-trained MAPF model using centralized expert data. Leveraging a novel delta-data generation mechanism, MAPF-GPT-DDG accelerates training while significantly improving performance at test time. Our experiments demonstrate that MAPF-GPT-DDG surpasses all existing learning-based MAPF solvers, including the original MAPF-GPT, regarding solution quality across many testing scenarios. Remarkably, it can work with MAPF instances involving up to 1 million agents in a single environment, setting a new milestone for scalability in MAPF domains.

多智能体路径规划强化学习大规模系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。