arXiv:2409.18435cs.LGcs.AI2024-09被引 2

用多智能体强化学习优化物料搬运系统调度,提升吞吐量最多7.4%。

Multi-agent Reinforcement Learning for Dynamic Dispatching in Material Handling Systems

  • 设计多智能体强化学习框架,融合现有调度规则提升探索效率。
  • 实验显示中位吞吐量比传统规则提升最高7.4%。
  • 可迭代训练,首轮结果作为启发式指导第二轮学习,性能持续优化。

本文提出一种多智能体强化学习(MARL)方法,用于学习动态调度策略,对优化跨多个行业的物料搬运系统吞吐量至关重要。为评估该方法,我们构建了一个反映实际系统复杂性的物料搬运环境,包含不同位置的多种活动、物理约束及内在不确定性。为增强学习过程中的探索能力,我们提出将现有动态调度启发式规则作为领域知识融入算法。实验结果表明,该方法在中位吞吐量上可比启发式规则最高提升7.4%。此外,我们分析了不同架构对多智能体训练性能的影响,并证明可通过将首轮MARL代理的结果作为启发式,训练第二轮MARL代理,进一步提升性能。本工作展示了将MARL应用于学习高效动态调度策略的潜力,可部署于真实系统以改善业务成果。

原文摘要 · Abstract (English)

This paper proposes a multi-agent reinforcement learning (MARL) approach to learn dynamic dispatching strategies, which is crucial for optimizing throughput in material handling systems across diverse industries. To benchmark our method, we developed a material handling environment that reflects the complexities of an actual system, such as various activities at different locations, physical constraints, and inherent uncertainties. To enhance exploration during learning, we propose a method to integrate domain knowledge in the form of existing dynamic dispatching heuristics. Our experimental results show that our method can outperform heuristics by up to 7.4 percent in terms of median throughput. Additionally, we analyze the effect of different architectures on MARL performance when training multiple agents with different functions. We also demonstrate that the MARL agents performance can be further improved by using the first iteration of MARL agents as heuristics to train a second iteration of MARL agents. This work demonstrates the potential of applying MARL to learn effective dynamic dispatching strategies that may be deployed in real-world systems to improve business outcomes.

多智能体强化学习调度优化物流系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。