对比经典与深度强化学习库存策略在药品供应链中的表现
Classical and Deep Reinforcement Learning Inventory Control Policies for Pharmaceutical Supply Chains with Perishability and Non-Stationarity
- 采用订单补货、投影库存和PPO强化学习三种方法优化药品库存
- 相比人工策略,三种方法平均成本更低但波动更大,其中投影库存最稳定
- 适用于医药供应链管理研究者及工业界决策者参考
我们研究药品供应链中的库存控制策略,解决易腐性、产量不确定性、非平稳需求等挑战,并考虑批量约束、提前期和缺货损失。与百时美施贵宝(BMS)合作,构建包含上述因素的真实案例,将三种策略——补货至目标(OUT)、投影库存水平(PIL)以及基于近端策略优化(PPO)的深度强化学习(DRL)——与基于人类经验的BMS基线进行比较。我们推导并验证了用于优化OUT和PIL参数的基于边界的方法,提出一种估计投影库存水平的方案,并将其与需求预测结合以提升在非平稳环境下的决策能力。相较于依赖高持有成本避免缺货的人工策略,三种实现策略均实现更低的平均成本,但成本波动更大。其中,PIL表现出稳健一致的性能;OUT在缺货成本高时表现不佳;而PPO在复杂多变场景中表现优异,但计算开销大。结果表明,尽管DRL具有潜力,但在所有数值实验中并未始终优于经典策略,凸显:1)需整合多种策略以有效应对药品供应链挑战;2)当前领域尚无单一策略能普遍适用。
原文摘要 · Abstract (English)
We study inventory control policies for pharmaceutical supply chains, addressing challenges such as perishability, yield uncertainty, and non-stationary demand, combined with batching constraints, lead times, and lost sales. Collaborating with Bristol-Myers Squibb (BMS), we develop a realistic case study incorporating these factors and benchmark three policies--order-up-to (OUT), projected inventory level (PIL), and deep reinforcement learning (DRL) using the proximal policy optimization (PPO) algorithm--against a BMS baseline based on human expertise. We derive and validate bounds-based procedures for optimizing OUT and PIL policy parameters and propose a methodology for estimating projected inventory levels, which are also integrated into the DRL policy with demand forecasts to improve decision-making under non-stationarity. Compared to a human-driven policy, which avoids lost sales through higher holding costs, all three implemented policies achieve lower average costs but exhibit greater cost variability. While PIL demonstrates robust and consistent performance, OUT struggles under high lost sales costs, and PPO excels in complex and variable scenarios but requires significant computational effort. The findings suggest that while DRL shows potential, it does not outperform classical policies in all numerical experiments, highlighting 1) the need to integrate diverse policies to manage pharmaceutical challenges effectively, based on the current state-of-the-art, and 2) that practical problems in this domain seem to lack a single policy class that yields universally acceptable performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。