用强化学习优化多目标供应链,兼顾经济环保社会三方面。
Reinforcement Learning for Multi-Objective Multi-Echelon Supply Chain Optimisation
- 基于马尔可夫决策过程建模,用多目标强化学习求解复杂供应链。
- 在复杂场景下,比进化算法高75%的超体积,解集密度是单目标方法的11倍。
- 适合需要平衡多方目标、追求鲁棒性的供应链优化研究者。
本研究基于马尔可夫决策过程,构建了一个通用的多目标、多层级供应链优化模型,适用于非平稳市场环境,并综合考虑经济、环境与社会因素。采用多目标强化学习方法进行评估,对比了改进的加权求和单目标强化学习算法以及基于多目标进化算法(MOEA)的方法。通过可定制化模拟器,在不同网络复杂度下进行实验,模型能确定各供应链路径上的生产与配送量,实现多个目标间的近优权衡,逼近帕累托前沿。结果表明,该方法在最优性、多样性与密度之间取得最佳平衡,尤其引入共享经验回放缓冲区后,促进策略间知识迁移。在复杂场景中,其超体积较MOEA方法高出75%,解集密度约为单目标强化学习方法的11倍,且稳定维持生产与库存水平,有效降低需求损失。
原文摘要 · Abstract (English)
This study develops a generalised multi-objective, multi-echelon supply chain optimisation model with non-stationary markets based on a Markov decision process, incorporating economic, environmental, and social considerations. The model is evaluated using a multi-objective reinforcement learning (RL) method, benchmarked against an originally single-objective RL algorithm modified with weighted sum using predefined weights, and a multi-objective evolutionary algorithm (MOEA)-based approach. We conduct experiments on varying network complexities, mimicking typical real-world challenges using a customisable simulator. The model determines production and delivery quantities across supply chain routes to achieve near-optimal trade-offs between competing objectives, approximating Pareto front sets. The results demonstrate that the primary approach provides the most balanced trade-off between optimality, diversity, and density, further enhanced with a shared experience buffer that allows knowledge transfer among policies. In complex settings, it achieves up to 75\% higher hypervolume than the MOEA-based method and generates solutions that are approximately eleven times denser, signifying better robustness, than those produced by the modified single-objective RL method. Moreover, it ensures stable production and inventory levels while minimising demand loss.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。