arXiv:2507.22040cs.LGmath.OC2025-07被引 8

用深度强化学习解决库存管理问题,无需假设需求分布。

Structure-Informed Deep Reinforcement Learning for Inventory Management

  • 基于直接反向传播的DRL算法,仅用历史数据学习库存策略。
  • 在多种场景下表现优于或媲美传统方法,参数调优少。
  • 引入结构信息提升策略可解释性与泛化能力,适合工业应用。

本文研究深度强化学习(DRL)在经典库存管理问题中的应用,重点关注实际部署考量。采用基于DirectBackprop的DRL算法,处理多周期系统中缺货(含/不含提前期)、易腐品管理、双源采购以及联合采购与移除等典型场景。DRL仅依赖实际可获得的历史数据,避免对需求分布或参数的不切实际假设。实验表明,该通用DRL实现跨多种场景性能优于或媲美现有基准与启发式方法,且所需超参数调优极少。通过分析学习到的策略,发现其自然捕获了传统运筹学方法推导出的最优策略诸多结构性特征。为进一步提升策略性能与可解释性,提出结构感知策略网络(Structure-Informed Policy Network),将解析得出的最优策略特性显式融入学习过程。该方法增强了策略的可解释性与样本外鲁棒性,已在真实需求数据案例中验证。最后,展示了一种非平稳环境下的示范应用。本工作在保持实用性的同时,弥合了数据驱动学习与解析洞察之间的鸿沟。

原文摘要 · Abstract (English)

This paper investigates the application of Deep Reinforcement Learning (DRL) to classical inventory management problems, with a focus on practical implementation considerations. We apply a DRL algorithm based on DirectBackprop to several fundamental inventory management scenarios including multi-period systems with lost sales (with and without lead times), perishable inventory management, dual sourcing, and joint inventory procurement and removal. The DRL approach learns policies across products using only historical information that would be available in practice, avoiding unrealistic assumptions about demand distributions or access to distribution parameters. We demonstrate that our generic DRL implementation performs competitively against or outperforms established benchmarks and heuristics across these diverse settings, while requiring minimal parameter tuning. Through examination of the learned policies, we show that the DRL approach naturally captures many known structural properties of optimal policies derived from traditional operations research methods. To further improve policy performance and interpretability, we propose a Structure-Informed Policy Network technique that explicitly incorporates analytically-derived characteristics of optimal policies into the learning process. This approach can help interpretability and add robustness to the policy in out-of-sample performance, as we demonstrate in an example with realistic demand data. Finally, we provide an illustrative application of DRL in a non-stationary setting. Our work bridges the gap between data-driven learning and analytical insights in inventory management while maintaining practical applicability.

强化学习库存管理结构感知可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。