用深度学习优化易腐品库存,融合人类经验提升决策效率。
Deep Learning for Perishable Inventory Systems with Human Knowledge
- 基于边际成本设计统一损失函数,实现端到端学习。
- 结构引导方法比纯黑箱模型降低20%以上风险,性能更优。
- 嵌入库存理论结构可减少学习复杂度,适合实际业务部署。
管理具有有限保质期的易腐产品是库存管理中的核心挑战,不当订货易导致缺货或过度浪费。本文研究随机提前期下的易腐品库存系统,其中需求过程和提前期分布均未知。在仅有有限历史数据、结合可观测协变量与系统状态的实际场景中,为提升小样本下的学习效率,采用边际成本核算机制,为每次订货分配单一生命周期成本,构建统一损失函数以支持端到端训练。提出两种端到端变体:纯黑箱方法(E2E-BB)直接输出订货量;结构引导方法(E2E-PIL)嵌入预测库存水平(PIL)策略,通过显式计算捕捉库存效应而非额外学习。进一步证明E2E-PIL目标函数具有一阶齐次性,可引入运营数据分析(ODA)中的增强技术,得到改进策略(E2E-BPIL)。合成与真实数据实验表明性能排序为:E2E-BB < E2E-PIL < E2E-BPIL。通过过量风险分解分析,嵌入启发式策略结构可降低有效模型复杂度,显著提升学习效率,仅牺牲轻微灵活性。结果表明,融合人类知识的深度学习决策工具更具有效性与鲁棒性,凸显先进分析与库存理论结合的价值。
原文摘要 · Abstract (English)
Managing perishable products with limited lifetimes is a fundamental challenge in inventory management, as poor ordering decisions can quickly lead to stockouts or excessive waste. We study a perishable inventory system with random lead times in which both the demand process and the lead time distribution are unknown. We consider a practical setting where orders are placed using limited historical data together with observed covariates and current system states. To improve learning efficiency under limited data, we adopt a marginal cost accounting scheme that assigns each order a single lifetime cost and yields a unified loss function for end-to-end learning. This enables training a deep learning-based policy that maps observed covariates and system states directly to order quantities. We develop two end-to-end variants: a purely black-box approach that outputs order quantities directly (E2E-BB), and a structure-guided approach that embeds the projected inventory level (PIL) policy, capturing inventory effects through explicit computation rather than additional learning (E2E-PIL). We further show that the objective induced by E2E-PIL is homogeneous of degree one, enabling a boosting technique from operational data analytics (ODA) that yields an enhanced policy (E2E-BPIL). Experiments on synthetic and real data establish a robust performance ordering: E2E-BB is dominated by E2E-PIL, which is further improved by E2E-BPIL. Using an excess-risk decomposition, we show that embedding heuristic policy structure reduces effective model complexity and improves learning efficiency with only a modest loss of flexibility. More broadly, our results suggest that deep learning-based decision tools are more effective and robust when guided by human knowledge, highlighting the value of integrating advanced analytics with inventory theory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。