打造可配置的采矿卡车调度强化学习评测环境
Mining--Gym: A Configurable RL Benchmarking Environment for Truck Dispatch Scheduling
- 基于离散事件仿真构建可定制调度环境,支持真实场景模拟
- 在六种工况下验证RL策略优于传统启发式方法
- 适合研究采矿优化与强化学习落地的开发者和工程师
优化采矿过程——特别是卡车调度——是露天开采作业效率的关键。然而,设备故障、车辆维护和运输周期变化等动态随机性,给传统优化带来挑战。尽管强化学习(RL)在采矿物流自适应决策方面展现出强大潜力,但实际部署需要在真实、可配置的仿真环境中评估。当前缺乏标准化基准,阻碍了算法公平比较、可复现性及实际应用。为此,我们提出Mining-Gym——一个基于Salabim的离散事件仿真(DES)并集成Gymnasium的开源基准环境,用于训练、测试和评估采矿过程优化中的RL算法。其事件驱动的决策架构能捕捉采矿特有的不确定性,提供图形化界面用于参数配置、数据记录和实时可视化,支持可复现的RL策略与启发式基线对比评估。通过在六种场景(从正常运行到严重设备故障)中对比经典启发式与基于RL的调度策略,结果表明该环境是有效且可复现的测试平台,能够公平评估自适应决策能力,并展现RL训练调度器的强大性能潜力。
原文摘要 · Abstract (English)
Optimizing the mining process -- particularly truck dispatch scheduling -- is a key driver of efficiency in open-pit operations. However, the dynamic and stochastic nature of these environments, with uncertainties such as equipment failures, truck maintenance, and variable haul cycle times, challenges traditional optimization. While Reinforcement Learning (RL) shows strong potential for adaptive decision-making in mining logistics, practical deployment requires evaluation in realistic, customizable simulation environments. The lack of standardized benchmarking hampers fair algorithm comparison, reproducibility, and real-world applicability of RL solutions. To address this, we present Mining-Gym -- a configurable, open-source benchmarking environment for training, testing, and evaluating RL algorithms in mining process optimization. Built on Salabim-based Discrete Event Simulation (DES) and integrated with Gymnasium, Mining-Gym captures mining-specific uncertainties through an event-driven decision-point architecture. It offers a GUI for parameter configuration, data logging, and real-time visualization, supporting reproducible evaluation of RL strategies and heuristic baselines. We validate Mining-Gym by comparing classical heuristics with RL-based scheduling across six scenarios from normal operation to severe equipment failures. Results show it is an effective, reproducible testbed, enabling fair evaluation of adaptive decision-making and demonstrating the strong performance potential of RL-trained schedulers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。