新商品库存管理难题:用迁移学习优化强化学习,降本提速。
Data-driven inventory management for new products: An adjusted Dyna-$Q$ approach with transfer learning
- 基于修正版Dyna-Q框架,融合模型与无模型学习。
- 相比Q-learning降低23.7%日均成本,训练速度提升77.5%。
- 利用相似产品数据迁移,早期训练更稳定,适合新商品场景。
本文提出一种新型强化学习算法,用于无历史需求数据的新产品库存管理。该算法在经典Dyna-Q结构基础上,平衡模型无关与模型依赖方法,加速训练并缓解模型偏差。通过迁移学习思想,将已有相似产品的销售数据作为预热信息融入算法,进一步稳定初期训练过程,降低最优策略估计方差。在真实烘焙店库存管理案例中验证:调整后的Dyna-Q相比Q-learning平均每日成本降低23.7%,相同周期内训练时间减少77.5%;使用迁移学习后,在30天测试期内,其总成本最低、成本波动最小,缺货率也相对较低。
原文摘要 · Abstract (English)
In this paper, we propose a novel reinforcement learning algorithm for inventory management of newly launched products with no historical demand information. The algorithm follows the classic Dyna-$Q$ structure, balancing the model-free and model-based approaches, while accelerating the training process of Dyna-$Q$ and mitigating the model discrepancy generated by the model-based feedback. Based on the idea of transfer learning, warm-start information from the demand data of existing similar products can be incorporated into the algorithm to further stabilize the early-stage training and reduce the variance of the estimated optimal policy. Our approach is validated through a case study of bakery inventory management with real data. The adjusted Dyna-$Q$ shows up to a 23.7\% reduction in average daily cost compared with $Q$-learning, and up to a 77.5\% reduction in training time within the same horizon compared with classic Dyna-$Q$. By using transfer learning, it can be found that the adjusted Dyna-$Q$ has the lowest total cost, lowest variance in total cost, and relatively low shortage percentages among all the benchmarking algorithms under a 30-day testing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。