用神经协调器提升库存管理中的资源约束适应能力。
Neural Coordination and Capacity Control for Inventory Management
- 设计神经协调器预测容量价格,替代传统控制器。
- 在亚马逊真实数据上验证,策略提升50%的容量合规性。
- 适合做智能供应链与强化学习结合的研究者。
本文研究有限共享资源下的多产品周期审查库存控制问题,聚焦零售商在存储或人力等资源受限场景下的决策。由于仅有一个历史容量样本路径,论文提出从可能的约束路径分布中采样,以更鲁棒地评估库存策略。将Madeka等人(2022)的外生决策过程(exo-IDP)扩展至有容量约束的问题,并证明某些容量控制问题复杂度不高于监督学习。提出“神经协调器”,通过预测容量价格引导系统遵守目标约束,替代传统模型预测控制器。采用改进的DirectBackprop算法联合训练深度强化学习采购策略与神经协调器。大规模回测显示,带神经协调器的强化学习策略在累积折扣收益和容量遵循方面均优于经典基线,部分情况下提升达50%。
原文摘要 · Abstract (English)
This paper addresses the capacitated periodic review inventory control problem, focusing on a retailer managing multiple products with limited shared resources, such as storage or inbound labor at a facility. Specifically, this paper is motivated by the questions of (1) what does it mean to backtest a capacity control mechanism, (2) can we devise and backtest a capacity control mechanism that is compatible with recent advances in deep reinforcement learning for inventory management? First, because we only have a single historic sample path of Amazon's capacity limits, we propose a method that samples from a distribution of possible constraint paths covering a space of real-world scenarios. This novel approach allows for more robust and realistic testing of inventory management strategies. Second, we extend the exo-IDP (Exogenous Decision Process) formulation of Madeka et al. 2022 to capacitated periodic review inventory control problems and show that certain capacitated control problems are no harder than supervised learning. Third, we introduce a `neural coordinator', designed to produce forecasts of capacity prices, guiding the system to adhere to target constraints in place of a traditional model predictive controller. Finally, we apply a modified DirectBackprop algorithm for learning a deep RL buying policy and a training the neural coordinator. Our methodology is evaluated through large-scale backtests, demonstrating RL buying policies with a neural coordinator outperforms classic baselines both in terms of cumulative discounted reward and capacity adherence (we see improvements of up to 50% in some cases).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。