arXiv:2603.19621cs.LGcs.AI2026-03被引 4

用经典库存策略增强强化学习,让库存管理更稳定高效。

DeepStock: Reinforcement Learning with Policy Regularizations for Inventory Management

  • 引入基于'基础库存'的策略正则化,约束强化学习行为
  • 在天猫平台100%部署中显著提升调参效率与最终表现
  • 适合关注工业级强化学习落地的算法工程师

深度强化学习(DRL)为利用大数据和算力训练库存策略提供了通用方法。然而,现成的DRL实现效果参差不齐,常因训练超参数敏感而难以应用。本文表明,通过引入基于经典库存概念如'基础库存'的策略正则化,可显著加速超参数调优并提升多种DRL方法的最终性能。我们在阿里巴巴电商主平台天猫实现了DRL的100%部署,验证了该方法的有效性。此外,大量合成实验显示,策略正则化重塑了库存管理领域最佳DRL方法的判断标准。

原文摘要 · Abstract (English)

Deep Reinforcement Learning (DRL) provides a general-purpose methodology for training inventory policies that can leverage big data and compute. However, off-the-shelf implementations of DRL have seen mixed success, often plagued by high sensitivity to the hyperparameters used during training. In this paper, we show that by imposing policy regularizations, grounded in classical inventory concepts such as "Base Stock", we can significantly accelerate hyperparameter tuning and improve the final performance of several DRL methods. We report details from a 100% deployment of DRL with policy regularizations on Alibaba's e-commerce platform, Tmall. We also include extensive synthetic experiments, which show that policy regularizations reshape the narrative on what is the best DRL method for inventory management.

强化学习库存管理正则化工业落地

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。