用运筹优化指导深度学习,提升库存管理效率
ORPR: An OR-Guided Pretrain-then-Reinforce Learning Model for Inventory Management
- 先用运筹模型生成高质量决策标签,再训练深度学习模型
- 实测库存周转缩短5.27天,缺货率下降2.29%,持有成本减少29.95%
- 适合需要精准、可解释库存策略的供应链场景
随着人工智能与运筹学在复杂库存系统中的融合日益深入,如何有效协调AI的自适应感知与运筹学的结构严谨性成为关键挑战。为此,我们提出一种新型的运筹学引导‘预训练-强化’框架。通过模拟增强的运筹模型生成高质量参考决策,隐式捕捉复杂的业务约束与管理偏好。利用这些运筹衍生的决策作为基础训练标签,构建领域感知的深度学习基础模型以建立基础决策能力,随后进入强化学习微调阶段。独特之处在于,将强化学习定位为深度对齐机制,使AI代理内化运筹学最优性原则,同时借助探索实现策略泛化优化,并支持专家指导进行场景适配(如促销活动)。经大量数值实验及京东公司实地部署(结合双重差分分析验证),该模型显著优于现有工业实践,在真实场景中实现库存周转减少5.27天、在库率提升2.29%、持有成本下降29.95%。研究表明,基于结构化运筹逻辑引导的轻量级模型,可在不依赖大规模模型扩展的前提下,实现顶尖性能与强泛化能力,为智能供应链管理提供可扩展、低成本的新范式。
原文摘要 · Abstract (English)
As the pursuit of synergy between Artificial Intelligence (AI) and Operations Research (OR) gains momentum in handling complex inventory systems, a critical challenge persists: how to effectively reconcile AI's adaptive perception with OR's structural rigor. To bridge this gap, we propose a novel OR-Guided "Pretrain-then-Reinforce" framework. To provide structured guidance, we propose a simulation-augmented OR model that generates high-quality reference decisions, implicitly capturing complex business constraints and managerial preferences. Leveraging these OR-derived decisions as foundational training labels, we design a domain-informed deep learning foundation model to establish foundational decision-making capabilities, followed by a reinforcement learning (RL) fine-tuning stage. Uniquely, we position RL as a deep alignment mechanism that enables the AI agent to internalize the optimality principles of OR, while simultaneously leveraging exploration for general policy refinement and allowing expert guidance for scenario-specific adaptation (e.g., promotional events). Validated through extensive numerical experiments and a field deployment at JD.com augmented by a Difference-in-Differences (DiD) analysis, our model significantly outperforms incumbent industrial practices, delivering real-world gains of a 5.27-day reduction in turnover and a 2.29% increase in in-stock rates, alongside a 29.95% decrease in holding costs. Contrary to the prevailing trend of brute-force model scaling, our study demonstrates that a lightweight, domain-informed model can deliver state-of-the-art performance and robust transferability when guided by structured OR logic. This approach offers a scalable and cost-effective paradigm for intelligent supply chain management, highlighting the value of deeply aligning AI with OR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。