用离线强化学习与大模型协作,让多智能体路径规划更快更稳。
Multi-Agent Path Finding via Offline RL and LLM Collaboration
- 基于决策变换器的离线强化学习,训练时间从周级缩短至小时级。
- 在稀疏奖励场景下表现优异,有效解决长时序信用分配难题。
- 融合GPT-4o动态引导策略,适应环境变化能力显著提升。
多智能体路径规划(MAPF)是机器人与物流领域中的关键挑战,因其组合复杂性及现实环境中固有的部分可观测性而难以求解。传统去中心化强化学习常因个体自利行为导致频繁碰撞,且依赖复杂通信模块使训练时间长达数周。为此,本文提出一种基于决策变换器(Decision Transformer, DT)的高效去中心化规划框架,利用离线强化学习将训练时间从数周压缩至数小时。该方法有效处理长时序信用分配问题,在稀疏与延迟奖励场景中性能显著提升。为进一步克服标准强化学习在动态环境中的适应性局限,我们引入GPT-4o以动态引导代理策略。大量实验表明,结合短时GPT-4o干预的DT框架,在静态与动态环境中均展现出更强的适应性与性能优势。
原文摘要 · Abstract (English)
Multi-Agent Path Finding (MAPF) poses a significant and challenging problem critical for applications in robotics and logistics, particularly due to its combinatorial complexity and the partial observability inherent in realistic environments. Decentralized reinforcement learning methods commonly encounter two substantial difficulties: first, they often yield self-centered behaviors among agents, resulting in frequent collisions, and second, their reliance on complex communication modules leads to prolonged training times, sometimes spanning weeks. To address these challenges, we propose an efficient decentralized planning framework based on the Decision Transformer (DT), uniquely leveraging offline reinforcement learning to substantially reduce training durations from weeks to mere hours. Crucially, our approach effectively handles long-horizon credit assignment and significantly improves performance in scenarios with sparse and delayed rewards. Furthermore, to overcome adaptability limitations inherent in standard RL methods under dynamic environmental changes, we integrate a large language model (GPT-4o) to dynamically guide agent policies. Extensive experiments in both static and dynamically changing environments demonstrate that our DT-based approach, augmented briefly by GPT-4o, significantly enhances adaptability and performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。