对比五种强化学习算法在集装箱配载中的表现,提供可复用的仿真环境。
A Benchmark Study of Deep Reinforcement Learning Algorithms for the Container Stowage Planning Problem
- 构建包含起重机调度的CSPP强化学习评估环境。
- PPO和TRPO在复杂场景下表现更优,策略稳定性更强。
- 适合物流优化与强化学习交叉研究者参考。
集装箱配载规划(CSPP)是海运与码头运营的关键环节,直接影响供应链效率。由于问题复杂,传统上依赖人工经验。尽管强化学习(RL)近年被应用于CSPP,但不同算法间的系统性对比仍有限。为此,我们构建了一个捕捉CSPP核心特征的Gym环境,并扩展支持多智能体与单智能体形式的起重机调度。在此框架下,评估了五种RL算法:DQN、QR-DQN、A2C、PPO和TRPO,在多种复杂度场景下的表现。结果揭示随着问题复杂度增加,算法间性能差距显著,凸显算法选择与问题建模的重要性。本文系统评测了多种RL方法在CSPP中的应用,同时提供可复用的带起重机调度的Gym环境,为未来研究与实际部署奠定基础。
原文摘要 · Abstract (English)
Container stowage planning (CSPP) is a critical component of maritime transportation and terminal operations, directly affecting supply chain efficiency. Owing to its complexity, CSPP has traditionally relied on human expertise. While reinforcement learning (RL) has recently been applied to CSPP, systematic benchmark comparisons across different algorithms remain limited. To address this gap, we develop a Gym environment that captures the fundamental features of CSPP and extend it to include crane scheduling in both multi-agent and single-agent formulations. Within this framework, we evaluate five RL algorithms: DQN, QR-DQN, A2C, PPO, and TRPO under multiple scenarios of varying complexity. The results reveal distinct performance gaps with increasing complexity, underscoring the importance of algorithm choice and problem formulation for CSPP. Overall, this paper benchmarks multiple RL methods for CSPP while providing a reusable Gym environment with crane scheduling, thus offering a foundation for future research and practical deployment in maritime logistics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。