arXiv:2502.20845cs.LGcs.AI2025-02被引 1

用引导策略提升采矿卡车调度的强化学习效率

Reinforcement Learning with Curriculum-inspired Adaptive Direct Policy Guidance for Truck Dispatching

  • 设计自适应直接策略引导机制,优化政策型强化学习
  • 在稀疏与密集奖励下均实现10%性能提升和更快收敛
  • 适合研究智能调度与强化学习应用的开发者

在露天矿场中,基于强化学习(RL)的高效卡车调度常受限于复杂的奖励设计和基于价值的方法。本文提出一种受课程学习启发的自适应直接策略引导方法,用于解决上述问题。通过引入时间差分与广义优势估计中的时间间隔,改进近端策略优化(PPO)以适应矿山调度中不均衡的决策周期,并采用最短处理时间教师策略,通过策略正则化与自适应引导实现有效探索。在OpenMines平台上的评估表明,该方法在稀疏与密集奖励设置下相比标准PPO均提升10%性能并加速收敛,展现出对奖励设计更强的鲁棒性。该直接策略引导方法为基于RL的卡车调度提供了一种通用且高效的课程学习方案,可支持后续先进架构研究。

原文摘要 · Abstract (English)

Efficient truck dispatching via Reinforcement Learning (RL) in open-pit mining is often hindered by reliance on complex reward engineering and value-based methods. This paper introduces Curriculum-inspired Adaptive Direct Policy Guidance, a novel curriculum learning strategy for policy-based RL to address these issues. We adapt Proximal Policy Optimization (PPO) for mine dispatching's uneven decision intervals using time deltas in Temporal Difference and Generalized Advantage Estimation, and employ a Shortest Processing Time teacher policy for guided exploration via policy regularization and adaptive guidance. Evaluations in OpenMines demonstrate our approach yields a 10% performance gain and faster convergence over standard PPO across sparse and dense reward settings, showcasing improved robustness to reward design. This direct policy guidance method provides a general and effective curriculum learning technique for RL-based truck dispatching, enabling future work on advanced architectures.

强化学习卡车调度课程学习PPO

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。